Prove quadratic growth for DPO losses
Prove that the Chebyshev scalarization associated with the DPO losses used for LLM alignment satisfies the local quadratic-growth condition near its minimizers.
References
It is plausible on a region where the active bottleneck is stable and the active weighted loss has positive curvature at its minimizer, in which case the max of finitely many such functions inherits quadratic growth, but we do not prove this for the DPO losses used here.
— Collaborative Personalized Preference Alignment for LLMs under Data Deficiency
(2610.05898 - Yang et al., 5 Oct 2026) in Remark H.35, Appendix H.5, p. 38