Prove quadratic growth for DPO losses

Prove that the Chebyshev scalarization associated with the DPO losses used for LLM alignment satisfies the local quadratic-growth condition near its minimizers.

Background

The parameter-space distance guarantees rely on a local quadratic-growth assumption for the Chebyshev scalarization. The paper notes that smoothness and LoRA parameterization alone do not imply this condition, although it may hold when the active bottleneck is stable and has positive curvature. Establishing the condition for the DPO losses used in the experiments would strengthen the parameter-space interpretation of the convergence results.

References

It is plausible on a region where the active bottleneck is stable and the active weighted loss has positive curvature at its minimizer, in which case the max of finitely many such functions inherits quadratic growth, but we do not prove this for the DPO losses used here.

— Collaborative Personalized Preference Alignment for LLMs under Data Deficiency  (2610.05898 - Yang et al., 5 Oct 2026) in Remark H.35, Appendix H.5, p. 38