Inference-Time Alignment Under Generalized Preferences
Resolve the open challenge of designing inference-time alignment methods that can handle generalized preference models rather than relying exclusively on scalar Bradley–Terry reward functions.
References
To the best of our knowledge, tackling generalized preferences remains an open challenge in inference-time methods.
— Inference-Time Nash Alignment
(2609.08082 - Hosseini et al., 8 Sep 2026) in Section 1, Introduction