Principled treatment of negative patterning weights

Determine a principled treatment for negative patterning weights in Bradley–Terry reward-model retraining, beyond swapping the chosen and rejected responses and assigning the absolute value of the weight.

Background

Patterning can assign negative weights to preference pairs, but the retraining implementation uses a heuristic: it swaps the chosen and rejected responses and applies the absolute value of the weight. The appendix derives that this operation does not exactly implement the intended signed objective, because the resulting gradient is rescaled by an input- and parameter-dependent factor. The authors explicitly leave unresolved what principled treatment should replace this heuristic.

References

The discussion is preliminary; we are still thinking about what the right principled treatment is.

Patterning in Practice: Debiasing Reward Models with Susceptibilities  (2609.00699 - Wang et al., 1 Sep 2026) in Appendix, Section “Negative weights and the label-swap convention”