Unexplained variance gap in cross-entropy training
Explain why the higher variance component of cross-entropy training is not accounted for by the reported experiments.
References
We do note that the higher $$ of CE is not explained by our experiments, but the results still illustrate that CA displays significant benefits in variance reduction, reaching a floor with 7.2M training samples.
— Constraint-Aware Training
(2610.02909 - Kim, 2 Oct 2026) in Section 5, paragraph beginning “Results”