IMPARA: GEC Metric & Concurrent Verification
- IMPARA is a term denoting two distinct artifacts: a reference-free grammatical error correction metric and a model checker for concurrent C/C++ programs.
- In GEC, the original IMPARA uses a similarity filter with a quality estimator, while IMPARA-GED removes the filter to enhance error detection and adversarial robustness.
- In software verification, Impara applies lazy abstraction, interpolants, and dynamic partial-order reduction to efficiently generate safety proofs in multi-threaded programs.
IMPARA denotes two distinct research artifacts that appear under closely similar names in separate technical literatures. In grammatical error correction (GEC), IMPARA is a reference-free automatic evaluation metric that scores a source sentence and a hypothesized correction through a similarity-estimation filter and a learned quality estimator; later work both diagnosed a robustness failure in the original design and introduced IMPARA-GED as a redesigned successor (Goto et al., 30 Sep 2025, Sakai et al., 3 Jun 2025). In software verification, Impara is a proof-generating model checker for concurrent C/C++ programs based on lazy abstraction with interpolants, and AbPress extends it with source-set–based dynamic partial-order reduction and summarization of shared-variable accesses (Kroening et al., 2014).
1. Terminological scope
The name is used in two research contexts, and the distinction is substantive rather than stylistic.
| Name in the literature | Domain | Core characterization |
|---|---|---|
| IMPARA | GEC evaluation | Reference-free metric over a source sentence and a hypothesis |
| IMPARA-GED | GEC evaluation | GED-enhanced redesign of IMPARA that removes the similarity filter |
| Impara | Concurrent software verification | Lazy-abstraction model checker for multi-threaded C/C++ programs |
For GEC, IMPARA belongs to the family of reference-free metrics, in contrast to reference-based metrics such as M², GLEU, and ERRANT. It evaluates system outputs without requiring gold corrections at test time. For verification, Impara belongs to the family of abstraction-based software model checkers and targets safety proofs for shared-memory concurrent programs. The two usages therefore occupy different methodological lineages: neural evaluation of language generation on one side, and symbolic reachability analysis with interpolation and partial-order reduction on the other (Sakai et al., 3 Jun 2025, Kroening et al., 2014).
2. IMPARA as a reference-free metric for grammatical error correction
In the GEC literature, IMPARA takes a source sentence containing errors and a hypothesis sentence produced by a GEC system, and returns a real-valued quality score: Its scoring rule is decomposed into a similarity estimation model and a quality estimation model , combined through a hard filter: with in practice (Goto et al., 30 Sep 2025).
The similarity estimator is defined as cosine similarity between sentence embeddings of and 0. In the configuration discussed in the robustness study, the embeddings are obtained with bert-base-cased and mean pooling, and the filter assigns score 1 to outputs deemed too different from the source. The quality estimator is a separate neural model that predicts how good 2 is as a correction; the experiments use the pre-trained model gotutiyan/IMPARA-QE, and good sentences typically receive values near 3, for example 4 (Goto et al., 30 Sep 2025).
As a reference-free metric, IMPARA does not require human references 5 at evaluation time. The reported advantages are no reference cost, tolerance for valid corrections that diverge from any single reference, and the possibility of domain adaptation by retraining or fine-tuning QE on new domains using human ratings rather than new parallel references. A later methodological account further describes the original QE as an encoder-based PLM followed by a linear layer, trained with pairwise ranking supervision on partially corrected sentences generated from GEC parallel data; in that account, IMPARA uses the final layer’s first token embedding as its sentence representation (Sakai et al., 3 Jun 2025).
Under non-adversarial conditions, IMPARA was regarded as strongly human-aligned. The robustness paper cites prior meta-evaluation on SEEDA indicating that IMPARA and related reference-free metrics often achieve Pearson and Spearman correlations greater than 6 with human system rankings when only reasonable systems are considered. The IMPARA-GED study reports original IMPARA sentence-level results of Acc 7 and 8 on SEEDA-S, and Acc 9 and 0 on SEEDA-E, reinforcing its earlier status as a high-performing reference-free evaluator (Goto et al., 30 Sep 2025, Sakai et al., 3 Jun 2025).
3. Adversarial vulnerability and the reliability crisis
The robustness study frames IMPARA as one instance of a broader reliability problem for reference-free GEC evaluation: high correlation with human judgments on ordinary systems does not imply robustness against systems optimized to exploit the metric itself. For IMPARA, the attack objective is to produce, for each input 1, a hypothesis 2 that maximizes the metric without attempting faithful correction: 3 Because the metric returns 4 once the similarity threshold is passed, the practical attack becomes
5
The paper instantiates this as a nearest-neighbor retrieval attack over a 3,574,070-sentence corpus formed from BEA2019-train, Troy-1BW, and Troy-Blogs, using the same bert-base-cased mean-pooled encoder as the similarity filter, the semsis retrieval library, and 6 neighbors (Goto et al., 30 Sep 2025).
The resulting system, Adversarial-IMPARA, is evaluated on the BEA-2019 dev set of 4,384 sentences against ten normal state-of-the-art systems and three other adversarial systems. Under IMPARA, the attack system outperforms all genuine GEC systems in both absolute and relative evaluation.
| System | IMPARA Abs. | IMPARA Rel. |
|---|---|---|
| T5-11B | 0.763 | -0.008 |
| UL2-20B | 0.758 | -0.017 |
| Adversarial-IMPARA | 0.911 | 0.384 |
The contrast is more revealing when compared with attacks designed for other metrics. Adversarial-SOME receives IMPARA Abs. 7 and Rel. 8; Adversarial-Scribendi receives IMPARA Abs. 9 and Rel. 0; Adversarial-LLM receives IMPARA Abs. 1 and Rel. 2. The attack is therefore highly targeted: IMPARA’s similarity filter blocks attacks that do not reuse its encoder and threshold, but it fails badly once the attacker does so (Goto et al., 30 Sep 2025).
A concrete example illustrates the failure mode. For the input sentence “You will be interesting in this job ?”, the adversarial system outputs “I hope it will be a suitable job for me .” with 3 and 4. Since 5, the output passes the filter and receives a near-maximal score. From a human standpoint, however, it does not preserve the original meaning: the source asks about the listener’s interest in the job, whereas the output expresses the speaker’s hope that the job will be suitable. The study summarizes the pathology directly: IMPARA confuses topic similarity plus fluency with correct, faithful error correction (Goto et al., 30 Sep 2025).
The paper identifies several reasons. Meaning preservation is enforced only through a single cosine-similarity threshold on sentence embeddings; generic topical overlap can therefore suffice. QE does not use the source sentence explicitly and behaves as a generic fluency or quality regressor. The attacker’s retrieval procedure is aligned exactly with the filter because it uses the same encoder. This makes it easy to find fluent, high-QE, topically related sentences that cross the threshold while changing propositional content, speaker, or sentence type. The paper states that these vulnerabilities undermine the reliability of automatic GEC evaluation and that reference-free metrics such as IMPARA should not be used in isolation for high-stakes evaluations (Goto et al., 30 Sep 2025).
4. IMPARA-GED: redesign through grammatical error detection
IMPARA-GED is a redesigned reference-free evaluator that focuses on strengthening the quality estimator rather than preserving the original similarity filter. Its central claim is that the quality estimator should itself be a strong grammatical error detector, and that the similarity estimator in original IMPARA functions poorly enough to justify removal (Sakai et al., 3 Jun 2025).
The critique of the original similarity module is empirical and specific. The paper reports that “I like cats.” and “I dislike cats.” receive similarity 6 under BERT-Base, so a meaning-changing output would pass the original threshold. Conversely, a good correction can be rejected: for the input “I think the family will stay mentally healty as it is, without having emtional stress.” and the output “I think the family will stay mentally healthy without having emotional stress.”, BERT-Large-uncased gives similarity 7, below the 8 threshold, so IMPARA would assign a score of 9. Removing the “healthy” correction increases the similarity to 0, which the paper presents as further evidence that vanilla PLMs treat spelling errors poorly in this setting (Sakai et al., 3 Jun 2025).
Accordingly, IMPARA-GED replaces the original score
1
with the simpler scoring function
2
The entire burden of evaluation moves to a strengthened QE. The method first trains a PLM on a grammatical error detection task using token-level labels automatically derived from ERRANT alignments, with loss
3
It then fine-tunes that GED-trained encoder as IMPARA’s QE using the pairwise ranking loss
4
The paper explores four GED label schemes—2-class, 4-class, 25-class, and 55-class—and three backbone PLMs: BERT5, DeBERTa-v36, and ModernBERT7. It also changes the sentence representation from the first token embedding to mean pooling over all final-layer token embeddings, arguing that this better aggregates token-level error information into a sentence-level quality signal (Sakai et al., 3 Jun 2025).
Training uses FCE and CoNLL-2013. The GED stage uses FCE plus CoNLL-2013 train for 5 epochs, selecting the best checkpoint on FCE dev. The QE stage then generates training pairs from CoNLL-2013 train, fine-tunes for 10 epochs, selects on CoNLL-2013 dev, and trains with 5 different random seeds, choosing the best on CoNLL-2013 devtest. The implementation uses Hugging Face transformers, ged_baselines, and the public IMPARA reproduction gotutiyan/IMPARA and gotutiyan/IMPARA-QE, on a single NVIDIA GeForce RTX 3090 (Sakai et al., 3 Jun 2025).
On SEEDA, the standout configuration is ModernBERT8 with binary GED. Its sentence-level results are SEEDA-S Acc 9, 0, and SEEDA-E Acc 1, 2. At system level on SEEDA-S, it achieves Pearson’s 3 and Spearman’s 4. These numbers exceed original IMPARA’s sentence-level results and are competitive with GPT-4-based evaluators. In SEEDA-S sentence-level correlation, the paper states that IMPARA-GED achieves the highest reported correlation with human judgments among all automatic metrics, including GPT-4 variants. A recurring empirical pattern is that 2-class GED performs best or close to best, which the authors attribute to higher label reliability under automatic alignment noise (Sakai et al., 3 Jun 2025).
The redesign does not claim to have solved meaning preservation. Instead, it argues that modern GEC systems rarely produce wildly off-meaning outputs in realistic settings, that existing PLM similarity filters are often harmful or vacuous, and that better modeling of grammatical correctness and fluency yields larger improvements in human alignment. The authors further state that they release IMPARA-GED as the official version under naist-nlp/IMPARA-GED (Sakai et al., 3 Jun 2025).
5. Impara as a model checker for concurrent C/C++ programs
In software verification, Impara is a proof-generating software model checker for multi-threaded C/C++ programs. It works with a fixed, finite number of threads, models each thread as a finite control-flow graph, handles shared-memory concurrency with global variables, semaphores, and locks, checks safety properties expressed as reachability of error locations, and produces safety proofs in the form of inductive invariants when the program is safe (Kroening et al., 2014).
Its core algorithmic basis is McMillan’s Impact, that is, lazy abstraction with interpolants. The program is unwound into an Abstract Reachability Tree (ART)
5
where each node is labeled by a global control location and a logical formula over program variables. Initially, node annotations are True. When an error location is reached, Impara checks whether the corresponding path is feasible via SMT solving. If the path is feasible, the program is unsafe; if infeasible, an interpolant sequence refines node annotations along the path and blocks that spurious path in the abstraction. This is the lazy-abstraction principle: refinement occurs only where infeasible error paths are encountered (Kroening et al., 2014).
A further key concept is covering. If two ART nodes share the same global control location and one annotation implies the other, then the more specific node can be covered by the more general one, and its subtree can be pruned. This makes safety proofs finite when a safe, complete, well-labeled ART exists. The paper states the central correctness condition directly: if there is a safe, complete, well-labeled ART of program 6, then 7 is safe (Kroening et al., 2014).
Before AbPress, Impara combined Impact with peephole partial-order reduction (PPOR), described as a static or symbolic POR technique based on small local patterns of dependent actions. AbPress replaces that earlier POR with source-set–based dynamic POR and adds a summarization mechanism for shared-variable accesses so that DPOR and Impact’s covers can be combined efficiently and soundly. The motivation is the dual explosion problem in concurrent verification: lazy abstraction addresses data-state explosion, while POR addresses schedule explosion (Kroening et al., 2014).
6. AbPress, source-sets, and practical significance
AbPress fuses source-set–based DPOR with Impara’s ART construction. Source-sets are a refinement of persistent sets. For a node 8, a source-set 9 must intersect 0 for every execution 1 from 2, where 3 contains the threads that can start execution independently from 4. The paper emphasizes that every source-set is a valid persistent set but not every persistent set is a source-set, and that source-sets are often smaller and yield greater reduction (Kroening et al., 2014).
The integration is operationalized through expansion, refinement, closing under covers, and backtracking. Expansion chooses a thread guided by the source-set, refinement computes interpolants on infeasible error paths, closing establishes covering relations, and backtracking performs race analysis to update source-sets. AbPress defines races through a happens-before relation over nodes on a path and uses ComputeBT to add a backtracking thread whenever a race exposes an alternative schedule that must be explored for completeness (Kroening et al., 2014).
The paper’s main additional contribution is summarization of shared-variable accesses in subtrees rooted at covered nodes. Without such summaries, DPOR would need to enumerate many paths beneath covering nodes merely to detect races against earlier actions. AbPress instead computes summaries 5 for covered subtrees and uses summary signatures during backtracking. The authors report that, without summarization, the number of paths involved in dependency analysis increases by an order of magnitude, and moderate programs with loops such as qrcu_true and stack_true time out at 900 seconds (Kroening et al., 2014).
Empirically, AbPress improves Impara substantially. On queue_ok_true (159 LOC, 3 threads, safe), FMCAD’13 Impara times out above 600 seconds, whereas AbPress completes in 63.7 seconds with 6 and SMT time 14.7 seconds. On stack_true (120 LOC, 3 threads, safe), FMCAD’13 Impara requires 360 seconds, 7, and SMT time 336.5 seconds, while AbPress completes in 30.5 seconds with 8 and SMT time 17.2 seconds. On the weak-memory litmus benchmark mix000_tso_false, AbPress runs in 2.9 seconds with 9 and SMT time 0.2 seconds, compared with FMCAD’13 Impara’s 4.5 seconds, 0, and SMT time 2.5 seconds. The paper summarizes the overall result by stating that AbPress compares favorably to state-of-the-art tools such as Threader and CBMC across many concurrent benchmarks while retaining proof-generating, unbounded verification capabilities (Kroening et al., 2014).
Both meanings of IMPARA therefore illustrate a common pattern in technical evaluation: strong nominal performance is not sufficient without careful treatment of failure modes. In GEC, the original IMPARA achieved high human correlation but proved vulnerable to metric-targeted retrieval attacks, motivating IMPARA-GED and a broader emphasis on adversarial robustness and metric ensembles (Goto et al., 30 Sep 2025, Sakai et al., 3 Jun 2025). In software verification, Impara’s lazy abstraction was effective but required stronger POR and shared-access summarization to remain practical under heavy interleaving. The two lines of work are unrelated in domain, yet each centers on the same methodological tension between elegant formal structure and the robustness demands of realistic use (Goto et al., 30 Sep 2025, Kroening et al., 2014).