Compute-optimal scaling of the two-layer memory architecture

Determine the compute-optimal regime of the two-layer Kathleen architecture with its fixed-key holographic notebook beyond approximately 0.5 million parameters.

Background

The paper evaluates the byte-level Kathleen trunk and attached holographic notebook at only 0.2–3.4 million parameters, using no more than 512 MB of data and a single GPU. Although the WikiText-103 data ladder shows a monotonic increase in the notebook’s repeat-recall benefit, the authors characterize the evidence as limited to four data points, with a two-seed error estimate only at the largest rung.

The unresolved issue is whether the observed behavior and compute trade-offs persist in a compute-optimal regime beyond the approximately 0.5-million-parameter scale emphasized by the main model configurations. Establishing this would test whether the notebook’s benefits remain competitive when the underlying recurrent trunk and training budget are substantially enlarged.

References

The compute-optimal regime beyond ${\sim}0.5$M parameters remains unexplored.

Kathleen Remembers: Length-Invariant One-Shot Recall Without Attention  (2608.30376 - Fountzoulas, 31 Aug 2026) in Section 8, Limitations

Word-level absorption is reported as a boundary of the method, not resolved.

Kathleen Remembers: Length-Invariant One-Shot Recall Without Attention  (2608.30376 - Fountzoulas, 31 Aug 2026) in Section 8, Limitations