Unresolved Questions on Neural Scaling Laws

Investigate the unresolved questions regarding neural scaling laws, focusing on clarifying what aspects of the observed performance scaling with training time, dataset size, and model size remain unestablished across architectures and tasks to better understand their theoretical foundations and practical implications.

Background

Neural scaling laws describe how performance metrics such as test loss improve predictably with increases in compute, model size, and dataset size. These laws are central to model and dataset design and to compute-optimal training strategies. Despite extensive empirical study, the introduction notes that many aspects of these laws are still not fully understood, motivating the development of theoretical models to explain and predict these phenomena.

The paper presents a solvable random feature model capturing several observed scaling behaviors, including differing exponents in time versus model size, finite-dataset and finite-width corrections, and the non-equivalence of ensembling and width scaling. Nonetheless, the authors explicitly acknowledge unresolved questions broadly about neural scaling laws.

References

Yet, many questions about neural scaling laws remain open.

A Dynamical Model of Neural Scaling Laws  (2402.01092 - Bordelon et al., 2024) in Introduction (Section 1), page 2

Looking forward in time though, there is a question-mark over whether this trend/model will continue to be `true enough', and whether it will be so all the way to ASI.

Understanding as an Explicit and Assessable Component of Frontier AI Safety Decisions  (2608.19816 - Barrett et al., 20 Aug 2026) in Appendix G, Section G.2, Table G.1 (“Examples of candidate felicitous falsehoods”)

Developing theoretical scaling laws relating training-data size to AIWP model accuracy, beyond emerging empirical estimates , remains an important open research direction and can build on our findings.

Missing the Butterfly and Predicting the Past: Features or Bugs of Accurate AI Weather Models?  (2608.25835 - Hassanzadeh et al., 26 Aug 2026) in Discussion, paragraph beginning “We should clarify that the surprising accuracy…”

The open question is how the benefits of tensorization change as the model grows, specifically how compression ratio and quality behave at each scale, thereby providing an analogue of scaling laws () for tensorized LLMs.

Tensor Methods for Language Models: From Token Representation to Training, Adaptation, Inference, Compression, and Interpretability  (2608.30505 - Tarasov et al., 31 Aug 2026) in Section 7.2, paragraph “Scaling”