Cause of the residual loss floor in constraint-aware training

Determine the cause of the approximately 7 × 10^-5 excess-loss floor observed in constraint-aware training, including whether it results from training noise or nondeterminism in the training kernel.

Background

In the data-efficiency experiments, constraint-aware training reaches a substantially lower excess loss than cross-entropy training and appears to flatten near an excess loss of approximately 7 × 10-5 nats. The authors explicitly state that they are unsure why this floor occurs and conjecture that training noise or nondeterminism in the training kernel may dominate corpus variance. Establishing the source of this residual error would clarify whether the observed floor is an optimization artifact, an implementation effect, or an intrinsic limitation of the training setup.

References

(We are unsure of the cause of the $\approx 7 \times 10{-5}$ floor observed in CA training, but conjecture that it is noise related with training or nondeterminism of the training kernel dominating corpus variance.)

— Constraint-Aware Training  (2610.02909 - Kim, 2 Oct 2026) in Section 5, paragraph beginning “Results”