Improve adherence to attribute-priority hierarchies

Solve the open problem of making e-commerce rerankers reliably follow the intended, query-conditioned attribute hierarchy rather than systematically violating or inconsistently applying its priorities.

Background

The AHP diagnostic evaluates whether a reranker prefers the product satisfying the higher-priority attribute in a controlled pair. The hand-designed hierarchy conflicts with judged preference, so it is not used as a training target.

Nevertheless, the authors report that even the best ZooWork-ShopRanker model only barely exceeds chance on the hierarchy diagnostic, and that preference alignment redistributes performance across attributes rather than uniformly improving hierarchy adherence. They therefore explicitly retain attribute hierarchy as unresolved.

References

Judge-labeled preference training therefore recovers part of the intended priority through scale---but only part; even the best model barely clears a coin flip, so attribute hierarchy remains an open problem.

— ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker  (2609.31002 - Xue et al., 25 Sep 2026) in Appendix, Section 'Attribute hierarchy: full analysis' (Section label app:ahp)