- The paper introduces CECoR, a decomposition-and-injection framework that represents verified multi-hop claims as QUESTION, VERIFY, and PREDICT programs before generating controlled, evidence-contradictory errors.
- CECoR’s lightweight T5 model reaches SARI scores of 76.76 on HOVER and 79.35 on FEVEROUS, outperforming prompting and prior correction systems, while RL improves semantic quality despite lower lexical overlap.
- The study shows that reference-based metrics can reward unchanged claims, whereas LLM judging better captures correction quality; limitations include synthetic test distributions, GPT-4o-mini dependence, and single-step error injection.
Motivation and problem statement
Factual Error Correction (FEC) revises factually inconsistent claims into evidence-supported statements, and is a natural post-hoc safeguard against hallucination in LLM-generated text. The paper identifies a structural gap in existing FEC paradigms: both Mask-then-Correct pipelines (e.g., evidence-based correction (Gan et al., 2021)) and reversed Mask-then-Corrupt approaches such as LIFE (He et al., 2023) and PivotFEC (Roussel et al., 2022) implicitly rest on an "Atomic Fact Assumption" — claims are treated as flat sequences and errors are injected or repaired as localized span edits, typically single-entity substitutions. This assumption fails for multi-hop claims whose truth conditions are distributed across reasoning chains over multiple evidence pieces (as in HOVER (Jiang et al., 2020) and FEVEROUS (Aloor et al., 2021)), where inconsistencies may be coupled across steps, involve temporal or relational constraints, or arise from coordinated multi-entity substitutions that no independent masking operation can express.
The CECoR framework
CECoR (Compositional Error Correction via Reasoning-aware Synthesis) replaces claim-level masking with a decomposition-and-injection paradigm built on program-guided reasoning (Ling et al., 2023). A planner converts each verified multi-hop claim into an interpretable reasoning program of QUESTION, VERIFY, and PREDICT steps written in controlled natural language. Errors are then injected one step at a time: entity substitution for PREDICT steps, factual-unit replacement or negation for VERIFY steps, and answer-plus-question rewriting for QUESTION steps, after which the corrupted program is recomposed into a fluent natural-language claim. Injecting into a single step at a time avoids cascading inconsistencies while yielding diverse faulty variants from each correct claim.
A five-criterion filter retains only synthetic pairs that satisfy length bounds, differ from the source claim, preserve genuine multi-hop dependency (at least two reasoning steps), are fluent by perplexity scoring, and are verifiably contradictory against the evidence. Training proceeds in two stages: supervised fine-tuning on filtered pseudo-parallel pairs, followed by reinforcement learning on naturally incorrect claims drawn from REFUTES subsets of fact-verification datasets. The RL reward combines evidence-based correctness, semantic similarity to the input claim, and fluency, encouraging minimal but sufficient edits.
Evaluation protocol
The paper makes a pointed methodological observation: surface-form metrics reward conservatism. On HOVER and FEVEROUS, the Do-Nothing baseline — which outputs the input unchanged — achieves SARI Final scores of 62.95 and 63.93 respectively, exceeding several competitive systems including GPT-4o few-shot prompting on HOVER. This follows directly from SARI's Keep component dominating when models make few edits. To address this, the authors complement rule-based metrics with LLM-as-a-judge evaluation using three independent judges (GPT-4o-mini, DeepSeek-V3, Gemini-2.5-flash), which assigns appropriately low scores to the Do-Nothing baseline and shows consistent cross-judge trends.
Main results
On HOVER and FEVEROUS, CECoR substantially outperforms all baselines under rule-based metrics despite lightweight backbones:
| Model |
HOVER SARI |
FEVEROUS SARI |
| LIFE (T5) |
45.13 |
61.45 |
| VENCE (T5) |
52.77 |
49.33 |
| GPT-4o-mini (8-shot) |
58.55 |
68.96 |
| CECoR (T5-sft) |
76.76 |
79.35 |
The roughly 18-point SARI margin over the strongest few-shot baseline on HOVER indicates that structured synthetic supervision compensates for small model capacity. Under LLM-judge evaluation, the picture shifts: SFT-only CECoR variants score below strong prompted GPT-4o baselines, but the RL-enhanced CECoR-L3-3b-rl achieves the best judge scores (0.83/0.80/0.80 on HOVER; 0.87/0.87/0.92 on FEVEROUS), confirming that RL aligns outputs with semantic correctness even where lexical overlap with references drops. The paper concedes this trade-off explicitly, noting RL variants show slightly lower rule-based scores because valid corrections diverge lexically from references.
Generalization, ablations, and robustness
Three additional experiments support the framework's breadth. First, on the human-curated single-hop FECDATA benchmark, CECoR-L3-3b-rl attains the highest GPT-judge score (0.94), surpassing GPT-4o 8-shot (0.84), demonstrating in-domain effectiveness beyond multi-hop settings. Second, filtering ablations show consistent gains across model sizes, particularly in SARI-Add and judge scores; notably, even the RL stage benefits, indicating filtered synthetic data provides a stronger policy initialization. Third, under retrieved rather than gold evidence (BM25 top-3 over a 5.2M-article Wikipedia dump), CECoR maintains large advantages over LIFE — which collapses entirely here because none of its synthetic examples pass its own filter under noisy retrieval — while suffering only moderate degradation itself. Cross-domain transfer from HOVER-trained models to FECDATA also outperforms distantly supervised baselines without any single-hop training data, with the RL variant achieving the best judge score (0.51).
Case studies illustrate the qualitative distinction: on implicit comparative-reasoning errors, only the RL-optimized model produces correct revisions, while baselines either hallucinate corrections inconsistent with evidence or leave claims unmodified.
Limitations and open questions
Several constraints temper these results. The HOVER test set is constructed by applying the authors' own error injection to validation examples, since the official test set is unavailable — meaning evaluation error distributions match the synthesis distribution, potentially inflating multi-hop results relative to naturally occurring errors. The framework depends on GPT-4o-mini both as the decomposition/injection engine and as the primary reward signal and judge, raising circularity concerns between generation, optimization, and evaluation. Only one reasoning step is corrupted per example, so coupled multi-step errors of the kind motivating the work are synthesized individually rather than jointly. FEVEROUS evaluation restricts to sentence-level textual evidence, excluding the dataset's tabular component. Finally, the RL stage draws incorrect claims exclusively from REFUTES subsets, leaving open how the approach handles partially supported claims or NotEnoughInfo cases.
Conclusion
CECoR demonstrates that exposing the latent reasoning structure of multi-hop claims enables controllable, step-level error synthesis that scales supervision without paired annotations, and that a two-stage SFT+RL pipeline converts this synthetic data into correctors that outperform both distantly supervised methods and few-shot LLM prompting on multi-hop benchmarks while transferring to single-hop and noisy-evidence settings. Its central empirical contribution is equally the demonstration that reference-based metrics systematically misrank FEC systems, with LLM-based judging providing a more faithful alternative.