Multiplicative learning dynamics for classification losses

Derive an explicit multiplicative formula for the representation-learning dynamics induced by general classification losses within the power-iteration analysis of stochastic task weighting and MOSEL.

Background

The theoretical analysis of stochastic preference weighting is based on a power-iteration characterization in which feature projections evolve through an explicit multiplicative product of objective-dependent eigenvalues. The paper explains that this analysis is most directly established under assumptions associated with squared-loss-like settings and shared spectral structure.

For general classification losses, the relevant spectral object is described through a transformed correlation operator, and existing high-dimensional results provide qualitative support for eigenvalue-ordered representation learning. Nevertheless, the authors explicitly state that an explicit multiplicative formula has not yet been obtained, leaving a concrete theoretical problem.

References

While an explicit multiplicative formula for classification losses remains an open challenge, the high-dimensional analysis of \citet{ben2024high} provides rigorous support: SGD trajectories concentrate layer by layer on the low-rank outlier eigenspaces of the empirical Hessian and gradient matrices, confirming that an eigenvalue-ordered selection process governs representations under classification losses and suggesting our analysis extends to classification regimes in future work.

— Learning Pareto Stationary Fronts via Single-Pass Backpropagation  (2610.06397 - Celik et al., 5 Oct 2026) in Appendix, Section “Stochastic weighting as implicit regularization,” Remark “Extension and Scope”