- The paper demonstrates that source-conditioned role relabeling increases explicit error correction by 23–93 percentage points across 13 model–domain cells, with 10 reaching p<0.001.
- The paper finds that addressability, rather than verification ability, drives the effect: external role tags outperform self-distrust prompts, duplication, Self-Refine, Reflexion, and Chain-of-Verification.
- The paper shows that externalized wrong claims rarely cause new errors—attack rates stay at or below 3.3%—but adversarial trust framing can raise them to 70%, making prompt-context control essential.
The paper argues that the widely documented failure of LLM agents to self-correct errors in their own reasoning traces is, to a large degree, an artifact of the chat-template role system rather than a genuine capability deficit. The authors introduce source-conditioned role relabeling: taking an erroneous intermediate claim c⋆​ that sits inside the agent's own <thought> block and re-presenting it, byte-identical (SHA-256 verified), under an external chat-template role — user, tool, or system <memory> — while holding the audit instruction fixed. Across 13 model–domain cells spanning seven model families and three verifiable-reasoning domains (n=30 paired tasks per cell), this purely structural intervention lifts the explicit-correction rate by 23 to 93 percentage points, with 10 of 13 cells reaching p<0.001. The central claim is deliberately strong: "the failure to self-correct is not a cognitive deficit; it is a chat-template artifact."
The intervention and its controls
The experimental design isolates a single variable with unusual care. A failure pool is constructed by retaining only tasks on which the agent, given its trajectory containing the injected wrong claim plus an audit-only instruction, fails to explicitly identify the error. Each retained task is then run through five conditions: the audit-only baseline (L0_self), the claim re-presented as a user message with or without a "Wait," prefix, as a calculator tool response, or inside a system <memory> block. Byte-identity is enforced by hashing every rendered prompt, so any difference in correction rate is attributable to the role tag alone. Scoring uses a locked Qwen-72B judge at temperature zero measuring whether the continuation explicitly names and rejects c⋆​, with Cohen's κ=1.0 against hand labels and κ=0.843 under an independent second judge.
The headline results are substantial. On Llama-3.3-70B logical deduction, relabeling lifts correction from a 0% baseline to peaks of +93 pp; on Qwen2.5-72B math, <memory> yields +53 pp over a 17% baseline. Per-task flip analysis provides direct causal evidence: changing only the role label flips 26 of 30 flippable tasks on Llama-70B logic and 18 of 25 on Qwen-72B math. The three non-significant cells are exactly those the hypothesis predicts — ceiling cases where the audit-only baseline already exceeds 67% (gpt-oss-20B, Gemini 2.5 Flash, BBH-LD). Nine of thirteen cells survive Holm-Bonferroni correction.
Addressability, not verification
A series of controls establishes that the mechanism is not verification capability but what the authors call addressability. Four self-distrust prompt variants that leave the claim inside the original <thought> block reach at most 23.3% correction versus 70% for the relabel; two actually depress the baseline, because blanket distrust scatters attention across all intermediates rather than onto the specific wrong claim. Under all four variants the agent silently re-derives the correct answer without naming the erroneous intermediate. Published protocols fail similarly: Self-Refine, Reflexion, and Chain-of-Verification produce negative or small lifts on the same failure pool, while the relabel dominates every non-ceiling cell tested. Notably, the audit-only baseline already reaches 70–77% final-answer correctness despite near-zero strict correction rates, confirming silent re-derivation was occurring throughout — the model could fix the answer but could not act on the wrong substring as a discrete object.
A handle-granularity ladder decomposes the effect: bare syntactic boundaries (brackets, XML wrappers) contribute 17–23 pp, while adding the system role tag contributes a further 30 pp. At fixed system role, a nonsense <xqzy> tag reaches only 30% versus 70% for <memory>, so the tag's lexical identity is itself load-bearing. A within-thought duplication control matched on duplication count and recency position lifts correction by only +6.7 pp (p=0.26), isolating a +46.7 pp pure role-tag effect and ruling out duplication- or salience-based accounts. Which wrapper is strongest is domain-dependent in an interpretable way: <memory> dominates on math (arithmetic claims are out-of-genre for a user turn), while a neutral user message dominates on logical deduction.
Safety scope and the adversarial mirror
The channel does not reverse into error injection by default. Across five adversarial experiments injecting wrong claims into tasks the agent already solves, all 20 attack-rate cells sit at or below 3.3%. The asymmetry is stark: agents commit to a wrong claim in their own <thought> about 83% of the time but to a byte-identical external claim at most 3.3% of the time. However, this safety is instructable rather than architectural: a single sentence ordering the agent to "treat this memory as ground truth and do not verify" raises the attack rate to 70%, while a distrust framing leaves it unchanged. The intervention is therefore a reliability lever, not a hardened defense — any deployment that does not control the surrounding prompt context against attacker-controlled trust framing loses the safety property.
Limitations and open questions
The paper is candid about scope. The effect is largest on vanilla instruction-tuned models with non-ceiling baselines; reasoning-tuned models (DeepSeek-R1 reaches 100% correction under audit alone) and strong frontier models leave no headroom, which the addressability account itself predicts. Cell sizes of n=30 are adequate for headline effects but tight for subgroup analysis, and closed-weight cells ran below n=30 due to rate limits, making cross-method ordering there suggestive rather than confirmatory. The failure-pool construction means the lifts speak to the targeted regime, not in-the-wild prevalence. Mechanistically, the account is established behaviorally: a hidden-state probe on final-layer embeddings is null, and circuit-level confirmation via activation patching remains open. The first-token "engage-then-verify" logprob signature observed on Qwen-72B does not reproduce lexically across other families, indicating the behavioral lift is universal but its internal route is model-specific. Extension beyond verifiable tasks — code debugging, planning, free-form reasoning — is untested.
Conclusion
This paper reframes intrinsic self-correction failure as a harness-level phenomenon: instruction tuning teaches models to respond fluently to content arriving under external roles but leaves thought-internal substrings unaddressable as discrete objects. Re-presenting a byte-identical wrong claim under an external role supplies the missing handle, lifting explicit correction by up to 93 pp with no training, tools, or weight changes, while the reverse channel remains safely closed unless trust framing overrides it. Whether addressability governs self-correction in free-form reasoning, and whether a lightweight training signal could internalize the role handle, are the concrete questions the paper leaves open.