Tight Characterizations Beyond the Established Scaling Regimes

Establish tight characterizations for data-limited mixed-training regimes and for two-stage training, extending beyond the optimization-saturated mixed-training regime and the upper-bound-only result currently available for two-stage training.

Background

The paper derives a scaling law for mixed training that is tight only when the optimization is saturated, namely when the effective sample size and learning rate are sufficiently large relative to the sketch dimension. For two-stage training, the paper establishes only an upper-bound scaling characterization because the propagated first-stage error and the source-mismatch contribution can have cross terms that cancel, preventing a uniform termwise lower bound.

Consequently, the precise scaling behavior in data-limited mixed-training settings and the exact scaling law for two-stage training remain unresolved. Resolving these issues would complete the comparison between the two training protocols and clarify when the reported upper bounds are sharp.

References

In addition, the mixed-training scaling law is tight only in the optimization-saturated regime, while the two-stage result remains an upper bound; this disparity between the two protocols leaves tight characterizations for data-limited and two-stage settings open.

Learning with Synthetic Data via SGD in High-Dimensional Linear Regression  (2609.09572 - Li et al., 9 Sep 2026) in Section 6, Conclusion, paragraph “Limitations”