Determine whether the deployment–encoding distinction disappears at larger scale

Determine whether the deployment–encoding distinction identified by the three-level framework disappears at model scales larger than those evaluated, particularly beyond the Qwen3-14B and Llama-3.2-3B models.

Background

The paper measures a surplus between probe recoverability and behavioral deployment across several model scales and finds that instruction tuning narrows this surplus. The narrowing is especially pronounced for the Llama-3.2-3B model, whose surplus falls to +0.042, but the evaluated models do not establish whether the distinction between encoded information and deployed behavior ultimately vanishes at still larger scales.

Resolving this issue would clarify whether the behavior–probe gap is a transient small- or medium-scale phenomenon or a persistent property of how larger LLMs transform internal syntactic representations into outputs.

References

The Llama-3.2-3B Instruct surplus collapses to $+0.042$, the same value that marks the Gemma-4 architectural boundary (§\ref{sec:rq1_crossfamily}); the deployment--encoding distinction the framework draws at small scale is therefore not stationary, and where it disappears at still larger scale remains open.

— Encoded but Not Decoded: Layer-Localized Evidence for a Three-Level Gap in LLM Syntax  (2609.29848 - Lu et al., 24 Sep 2026) in Section 4.3, “Instruction tuning and the deployment gap”