Determine the scaling benefit of twelve times more training data

Determine how far beyond the 42.2 five-task benchmark bar the Daedalus-150M architecture can progress when trained with twelve times more data than the 5-billion-token ablation run.

Background

The paper reports that a parameter-matched Daedalus-150M hybrid trained on only 5 billion tokens scores 44.7 on the five-task evaluation harness, exceeding the pre-registered bar of 42.2. However, the full model is trained on 59.9 billion tokens, and the authors do not establish how much additional downstream-task improvement results from this larger training budget. The unresolved issue is therefore the extent of quality scaling beyond the already-cleared benchmark bar when the same architecture receives substantially more data.

References

This is not the Daedalus result and nothing is projected from it; it moves the open question from whether the architecture can reach the bar to how far beyond it twelve times more data carries.

Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference  (2608.20210 - Koutsiaris, 20 Aug 2026) in Section 7, subsection “One datapoint against the bar”