Papers
Topics
Authors
Recent
Search
2000 character limit reached

The Self-Correction Illusion: LLMs Correct Others but Not Themselves

Published 4 Jun 2026 in cs.AI and cs.CL | (2606.05976v1)

Abstract: Recent work shows that LLM agents struggle to correct errors in their own reasoning traces yet show markedly higher correction rates when identical claims appear under external sources. We ask whether this asymmetry reflects a capability deficit or a role-label artifact: does an agent's willingness to correct a wrong claim depend causally on the chat-template role that carries it, rather than on the claim's content? Our setup keeps the erroneous claim byte-identical across all conditions (SHA-256 verified) and varies only its wrapping role: the agent's own \role{<thought>}, a \role{user} message, a \role{tool} response, or a \role{system <memory>} block. Across 13 model-domain cells covering seven model families and three domains (n=30n{=}30 paired tasks per cell), relabeling the claim from \role{<thought>} to an external role lifts the explicit-correction rate by 23 to 93 percentage points, with 10 of 13 cells reaching $p{&lt;}0.001$. Further experiments confirm that the effect is asymmetric, mechanistically decomposable, and robust across domains. The failure to self-correct is not a cognitive deficit; it is a chat-template artifact. We exploit this artifact by designing a prompt-structure-only intervention that requires no training and no model modification, with its strongest role label being domain-dependent: \role{<memory>} dominates on math, while a plain \role{user} message dominates on logical deduction.

Summary

  • The paper demonstrates that source-conditioned role relabeling increases explicit error correction by 23–93 percentage points across 13 model–domain cells, with 10 reaching p<0.001.
  • The paper finds that addressability, rather than verification ability, drives the effect: external role tags outperform self-distrust prompts, duplication, Self-Refine, Reflexion, and Chain-of-Verification.
  • The paper shows that externalized wrong claims rarely cause new errors—attack rates stay at or below 3.3%—but adversarial trust framing can raise them to 70%, making prompt-context control essential.

The paper argues that the widely documented failure of LLM agents to self-correct errors in their own reasoning traces is, to a large degree, an artifact of the chat-template role system rather than a genuine capability deficit. The authors introduce source-conditioned role relabeling: taking an erroneous intermediate claim c⋆c_\star that sits inside the agent's own <thought> block and re-presenting it, byte-identical (SHA-256 verified), under an external chat-template role — user, tool, or system <memory> — while holding the audit instruction fixed. Across 13 model–domain cells spanning seven model families and three verifiable-reasoning domains (n=30n=30 paired tasks per cell), this purely structural intervention lifts the explicit-correction rate by 23 to 93 percentage points, with 10 of 13 cells reaching p<0.001p<0.001. The central claim is deliberately strong: "the failure to self-correct is not a cognitive deficit; it is a chat-template artifact."

The intervention and its controls

The experimental design isolates a single variable with unusual care. A failure pool is constructed by retaining only tasks on which the agent, given its trajectory containing the injected wrong claim plus an audit-only instruction, fails to explicitly identify the error. Each retained task is then run through five conditions: the audit-only baseline (L0_self), the claim re-presented as a user message with or without a "Wait," prefix, as a calculator tool response, or inside a system <memory> block. Byte-identity is enforced by hashing every rendered prompt, so any difference in correction rate is attributable to the role tag alone. Scoring uses a locked Qwen-72B judge at temperature zero measuring whether the continuation explicitly names and rejects c⋆c_\star, with Cohen's κ=1.0\kappa = 1.0 against hand labels and κ=0.843\kappa = 0.843 under an independent second judge.

The headline results are substantial. On Llama-3.3-70B logical deduction, relabeling lifts correction from a 0% baseline to peaks of +93 pp; on Qwen2.5-72B math, <memory> yields +53 pp over a 17% baseline. Per-task flip analysis provides direct causal evidence: changing only the role label flips 26 of 30 flippable tasks on Llama-70B logic and 18 of 25 on Qwen-72B math. The three non-significant cells are exactly those the hypothesis predicts — ceiling cases where the audit-only baseline already exceeds 67% (gpt-oss-20B, Gemini 2.5 Flash, BBH-LD). Nine of thirteen cells survive Holm-Bonferroni correction.

Addressability, not verification

A series of controls establishes that the mechanism is not verification capability but what the authors call addressability. Four self-distrust prompt variants that leave the claim inside the original <thought> block reach at most 23.3% correction versus 70% for the relabel; two actually depress the baseline, because blanket distrust scatters attention across all intermediates rather than onto the specific wrong claim. Under all four variants the agent silently re-derives the correct answer without naming the erroneous intermediate. Published protocols fail similarly: Self-Refine, Reflexion, and Chain-of-Verification produce negative or small lifts on the same failure pool, while the relabel dominates every non-ceiling cell tested. Notably, the audit-only baseline already reaches 70–77% final-answer correctness despite near-zero strict correction rates, confirming silent re-derivation was occurring throughout — the model could fix the answer but could not act on the wrong substring as a discrete object.

A handle-granularity ladder decomposes the effect: bare syntactic boundaries (brackets, XML wrappers) contribute 17–23 pp, while adding the system role tag contributes a further 30 pp. At fixed system role, a nonsense <xqzy> tag reaches only 30% versus 70% for <memory>, so the tag's lexical identity is itself load-bearing. A within-thought duplication control matched on duplication count and recency position lifts correction by only +6.7 pp (p=0.26p=0.26), isolating a +46.7 pp pure role-tag effect and ruling out duplication- or salience-based accounts. Which wrapper is strongest is domain-dependent in an interpretable way: <memory> dominates on math (arithmetic claims are out-of-genre for a user turn), while a neutral user message dominates on logical deduction.

Safety scope and the adversarial mirror

The channel does not reverse into error injection by default. Across five adversarial experiments injecting wrong claims into tasks the agent already solves, all 20 attack-rate cells sit at or below 3.3%. The asymmetry is stark: agents commit to a wrong claim in their own <thought> about 83% of the time but to a byte-identical external claim at most 3.3% of the time. However, this safety is instructable rather than architectural: a single sentence ordering the agent to "treat this memory as ground truth and do not verify" raises the attack rate to 70%, while a distrust framing leaves it unchanged. The intervention is therefore a reliability lever, not a hardened defense — any deployment that does not control the surrounding prompt context against attacker-controlled trust framing loses the safety property.

Limitations and open questions

The paper is candid about scope. The effect is largest on vanilla instruction-tuned models with non-ceiling baselines; reasoning-tuned models (DeepSeek-R1 reaches 100% correction under audit alone) and strong frontier models leave no headroom, which the addressability account itself predicts. Cell sizes of n=30n=30 are adequate for headline effects but tight for subgroup analysis, and closed-weight cells ran below n=30n=30 due to rate limits, making cross-method ordering there suggestive rather than confirmatory. The failure-pool construction means the lifts speak to the targeted regime, not in-the-wild prevalence. Mechanistically, the account is established behaviorally: a hidden-state probe on final-layer embeddings is null, and circuit-level confirmation via activation patching remains open. The first-token "engage-then-verify" logprob signature observed on Qwen-72B does not reproduce lexically across other families, indicating the behavioral lift is universal but its internal route is model-specific. Extension beyond verifiable tasks — code debugging, planning, free-form reasoning — is untested.

Conclusion

This paper reframes intrinsic self-correction failure as a harness-level phenomenon: instruction tuning teaches models to respond fluently to content arriving under external roles but leaves thought-internal substrings unaddressable as discrete objects. Re-presenting a byte-identical wrong claim under an external role supplies the missing handle, lifting explicit correction by up to 93 pp with no training, tools, or weight changes, while the reverse channel remains safely closed unless trust framing overrides it. Whether addressability governs self-correction in free-form reasoning, and whether a lightweight training signal could internalize the role handle, are the concrete questions the paper leaves open.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 6 tweets with 0 likes about this paper.