Determine what causes the residual first-letter accuracy advantage

Determine whether the intensity of a model's attempted first-letter encoding, measured by the number of north/south-initial sentences in a completion, predicts answer accuracy within the first-letter encoding condition.

Background

The replication study finds a residual answer-accuracy advantage for the first-letter condition in two models, even though the purported traces do not decode as encoded reasoning. The authors hypothesize that the encoding instruction may act as a procedural scaffold that makes the model step through the task, rather than functioning as a communication channel.

This proposed mechanism remains untested. The concrete unresolved question is whether the amount of attempted first-letter encoding within individual completions predicts answer accuracy, which would provide evidence for or against the procedural-scaffold explanation.

References

That remains untested; the natural next experiment is whether the intensity of the attempt, the number of north/south-initial sentences a completion contains, predicts its accuracy within the first-letter arm.

— Learning Steganography Is Easy, Learning Steganographic Reasoning Is Hard  (2609.39838 - Schulz et al., 30 Sep 2026) in Supplementary Material, Appendix "Replicating the encoded-reasoning uplift of Zolkowski et al.", paragraph "Conclusion"