Interpreting synthetic-task performance: reliance on prior knowledge vs reasoning
Ascertain how much large language models rely on prior parametric factual knowledge when solving synthetic unseen tasks and distinguish whether failures arise from insufficient reasoning complexity or from missing background knowledge that models typically exploit.
References
Fifth, activation versus in-context teaching: because both E1 and the TRUE hint make the rule available in context, we cannot fully separate “activating latent knowledge” from “supplying missing knowledge”; the near-ceiling basic buckets and the sufficiency of a light hint make the activation reading more parsimonious but not decisive, and directly probing whether a model can state the rule unprompted would better separate the two---future work.
Crucially, evaluations based only on synthetic unseen tasks still leave open questions about performance. Success demonstrates reasoning in isolation, but it does not reveal how much models typically rely on prior knowledge as a scaffold.
We do not have that contrast. Our no-context arm uses a third prompt, neither strict nor permissive but a plain instruction to answer to the best of the model's knowledge, while the wrong-context arm uses the permissive template that invites the model to combine supplied information with what it knows.