Inference-Time Alignment Under Generalized Preferences

Resolve the open challenge of designing inference-time alignment methods that can handle generalized preference models rather than relying exclusively on scalar Bradley–Terry reward functions.

Background

Inference-time alignment methods commonly rely on scalar reward models based on the Bradley–Terry assumption. Such models cannot represent general preference relations, including aggregated preferences that may be non-transitive even when individual preferences are transitive. Generalized preference models instead directly specify the probability that one response is preferred to another. The paper identifies the adaptation of inference-time alignment to these generalized preferences as an unresolved challenge and then initiates its study through the Best-of-Nash and Nash Mirror Descent algorithms.

References

To the best of our knowledge, tackling generalized preferences remains an open challenge in inference-time methods.

Inference-Time Nash Alignment  (2609.08082 - Hosseini et al., 8 Sep 2026) in Section 1, Introduction