Extent of LLM performance attributable to reasoning vs memorized knowledge

Determine the extent to which large language models’ task performance reflects genuine reasoning as opposed to recall of memorized parametric world knowledge, by explicitly separating and measuring these contributions in controlled evaluations.

Background

The paper motivates SynthWorlds by noting that many evaluations confound genuine reasoning with memorized factual recall from pretraining. Because training data are massive and often undisclosed, benchmark scores may reflect parametric knowledge rather than reasoning ability. The authors aim to disentangle these effects via parallel real-mapped and synthetic-mapped corpora and tasks to quantify the “knowledge advantage gap.”

This open problem frames the need for methodologies that can isolate reasoning from memorization, enabling clearer scientific conclusions about LLMs’ capabilities and more reliable deployment in novel environments where prior memorized knowledge is less useful.

References

Yet, as LMs continue to be trained on massive web corpora (often with undisclosed training data), it remains unclear to what extent their performance reflects genuine reasoning versus the reciting of memorized knowledge.

SynthWorlds: Controlled Parallel Worlds for Disentangling Reasoning and Knowledge in Language Models  (2510.24427 - Gu et al., 28 Oct 2025) in Introduction (Section 1)

Finally, model responses may partly reflect prior exposure to well-known works, and the extent to which such exposure influences the generated captions remains unknown.

Unifying Score and Performance for Fine-Grained Music Understanding in Audio-Language Models  (2609.10351 - Dujardin et al., 9 Sep 2026) in Section 8, Limitations and Future Work

If those appeared in Kimi's or Claude's training data, the models may have reproduced parts of a known solution rather than deriving each step from evidence. No public target can remove this threat. Subtask checklists and transcript review show how far each agent got, but they cannot tell us whether a successful action was reasoned or recalled.

Big Enough to Break Out: Tracking the Rising Capability of LLM Penetration-Testing Agents  (2609.10780 - Lovelace et al., 9 Sep 2026) in Section 8, Threats to Validity, paragraph “Benchmark contamination”

The mechanism underlying the reasoning influence on verbatim retrieval is not established. On LD50 the published value is a unit conversion away from the measured one that includes molecular weight calculation and a logarithm. One reading is that reasoning supplies the arithmetic rather than the value, and the captured traces do contain that conversion (Appendix~\ref{app:traces}). Those traces do not separate the conversion from the alternatives within LD50, and they were captured at high effort only, so they establish only that the conversion occurs at a high reasoning level. On the other benchmarks, which involve no unit conversion, the increase is unexplained.

Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models  (2609.05381 - Busch et al., 4 Sep 2026) in Section 3, “Retrieval is influenced by the reasoning level”; Discussion, “Dependence on the reasoning level”