Extent of LLM performance attributable to reasoning vs memorized knowledge
Determine the extent to which large language models’ task performance reflects genuine reasoning as opposed to recall of memorized parametric world knowledge, by explicitly separating and measuring these contributions in controlled evaluations.
References
Yet, as LMs continue to be trained on massive web corpora (often with undisclosed training data), it remains unclear to what extent their performance reflects genuine reasoning versus the reciting of memorized knowledge.
Finally, model responses may partly reflect prior exposure to well-known works, and the extent to which such exposure influences the generated captions remains unknown.
If those appeared in Kimi's or Claude's training data, the models may have reproduced parts of a known solution rather than deriving each step from evidence. No public target can remove this threat. Subtask checklists and transcript review show how far each agent got, but they cannot tell us whether a successful action was reasoned or recalled.
The mechanism underlying the reasoning influence on verbatim retrieval is not established. On LD50 the published value is a unit conversion away from the measured one that includes molecular weight calculation and a logarithm. One reading is that reasoning supplies the arithmetic rather than the value, and the captured traces do contain that conversion (Appendix~\ref{app:traces}). Those traces do not separate the conversion from the alternatives within LD50, and they were captured at high effort only, so they establish only that the conversion occurs at a high reasoning level. On the other benchmarks, which involve no unit conversion, the increase is unexplained.