Unexplained variance gap in cross-entropy training

Explain why the higher variance component of cross-entropy training is not accounted for by the reported experiments.

Background

The data-efficiency study measures both the loss of the ensemble mean prediction and the additional loss attributable to variation across independently sampled training corpora. Cross-entropy training exhibits a substantially larger variance component than constraint-aware training, but the paper does not identify the mechanism responsible for this difference. Explaining the higher cross-entropy variance would test and refine the paper’s theoretical account of parameter changes and corpus-induced prediction variability.

References

We do note that the higher $$ of CE is not explained by our experiments, but the results still illustrate that CA displays significant benefits in variance reduction, reaching a floor with 7.2M training samples.

— Constraint-Aware Training  (2610.02909 - Kim, 2 Oct 2026) in Section 5, paragraph beginning “Results”