Transfer of reasoning behaviors beyond training distributions
Ascertain whether and under what conditions reasoning behaviors learned by large language models transfer beyond their training distributions.
References
Overall, our analyses expose fundamental gaps: we cannot know which training produces which cognitive capabilities a priori, cannot ensure behaviors transfer beyond training distributions, and cannot validate whether observed patterns reflect genuine cognitive mechanisms or spurious reasoning shortcuts.
While we have demonstrated strong performance on deterministic state tracking tasks, it remains unclear whether MHA-CSP’s inductive biases transfer to other structured reasoning domains, such as code generation or semantic parsing.
These comparisons establish that the 3B gain is not confined to the training databases, but they do not isolate whether it comes from BIRD's difficulty, from self-play, or from their combination.
Therefore this paper explicitly marks instance universality as a falsifiable open proposition rather than an established conclusion: its falsification condition is ``after training on sufficiently diverse rule schemata and domains, the operator's zero-shot reasoning accuracy on entirely new domains is significantly lower than on the training domains.''
The training and evaluation distributions are isomorphic, so the 100\% deduction accuracy cannot be extrapolated to richer reasoning forms: multi-premise reasoning, negation and quantifiers, and defeasible (non-monotonic) reasoning are all uncovered. In other words, this paper validates operator correctness ``within the template distribution''; robustness under distribution shift is an untested open question.