Papers
Topics
Authors
Recent
Search
2000 character limit reached

QFrBLiMP: Quebec Minimal-Pair Benchmark

Updated 14 July 2026
  • The paper introduces QFrBLiMP as the first benchmark of linguistic minimal pairs tailored for Quebec French using attested, normative data from the BDL.
  • It covers 20 linguistic phenomena with 1,761 minimal pairs to assess formal grammatical competence and scaling trends in LLM performance.
  • Human annotation reveals high inter-annotator agreement, showing that models excel at frequent formal rules but struggle with nuanced lexical semantics and rare constructions.

Searching arXiv for QFrBLiMP and closely related benchmark papers. Using the arXiv search tool to verify the benchmark paper and related minimal-pair resources. QFrBLiMP, the Quebec-French Benchmark of Linguistic Minimal Pairs, is a corpus designed to evaluate how well LLMs encode grammatical knowledge of the Quebec-French variety rather than only broad downstream task performance. It is introduced as the first benchmark of linguistic minimal pairs specifically for Quebec French and is explicitly modeled after BLiMP for English: one sentence in each pair is acceptable, the other differs minimally and is unacceptable, and a model succeeds if it assigns higher probability to the acceptable member of the pair (Beauchemin et al., 30 Sep 2025). Relative to MultiBLiMP-Fr, the French subset of the multilingual MultiBLiMP resource, QFrBLiMP is larger, is not synthetic, is built from a Quebec governmental linguistic resource, targets Quebec-French norms, and covers 20 linguistic phenomena rather than 6 in the French MultiBLiMP comparison subset.

1. Definition and position within minimal-pair evaluation

QFrBLiMP was motivated by a gap in the evaluation ecosystem for French-language and Quebec-French-specific grammatical benchmarking. Minimal-pair benchmarks had already been created for English, Japanese, Dutch, Russian, Chinese, Basque, and many languages through MultiBLiMP, but not for Quebec French. The benchmark is intended to isolate formal grammatical competence more directly than common NLU benchmarks such as GLUE, which the paper characterizes as mixing syntax with world knowledge and task adaptation (Beauchemin et al., 30 Sep 2025).

The benchmark is framed as a Quebec-French analogue of BLiMP. Its purpose is not general task evaluation, but targeted testing of whether a model prefers the grammatical sentence over the minimally contrasted ungrammatical alternative. The paper argues that this distinction matters because Quebec French differs in some lexical and normative respects from hexagonal French, even if much of the syntax is shared.

QFrBLiMP is also positioned against MultiBLiMP-Fr. The paper emphasizes three contrasts. First, QFrBLiMP contains 1,761 instances, whereas the French comparison subset of MultiBLiMP contains 255. Second, QFrBLiMP covers 20 linguistic phenomena, whereas MultiBLiMP provides only 6 for French in that comparison. Third, QFrBLiMP is derived from attested normative examples rather than from a synthetic template pipeline.

2. Corpus construction and linguistic coverage

The source material for QFrBLiMP is the Banque de dépannage linguistique (BDL), an official online linguistic resource maintained by the Office québécois de la langue française (OQLF), a public institution in Quebec (Beauchemin et al., 30 Sep 2025). The BDL contains 2,667 articles across eleven categories such as spelling and syntax. Each article explains what the institution considers normatively correct and incorrect Quebec/French usage and includes example sentences written by linguists.

The construction pipeline is manual. The authors inspected all 2,667 articles and extracted 25,153 linguistic acceptability judgment sentences. Using the BDL’s own green/red presentation of correct versus incorrect examples, they labeled sentences 0 (ungrammatical) or 1 (grammatical). Minimal pairs were then manually organized from these extracted examples according to the same correctness scheme and assigned to one of twenty linguistic categories. The resulting dataset is therefore a manually curated, grammar-driven resource based on attested normative examples rather than automatically synthesized data.

The benchmark’s phenomenon inventory is as follows:

No. Linguistic phenomenon Minimal pairs
1 Accords participes passés (Past participle agreements) 97
2 Flexion du verbe (Verb inflection) 95
3 ne ... que (only ... that) 97
4 Sélection morphologie fonctionnelle (Functional morphology selection) 96
5 Clitique dans la négation de l'infinitif (Clitics in infinitive negation) 99
6 Montée du clitique (Rising clitics) 97
7 Négation standard (Standard negation) 114
8 Déterminants (Determinants) 106
9 Sémantique lexicale (Lexical semantics) 113
10 Accord dans l'expression idiomatique (Agreement in idiomatic expression) 91
11 Accord des adjectifs (Adjective agreement) 100
12 -é / -er 100
13 Sélection lexicale du complément (Lexical selection of the complement) 96
14 Négation de l'infinitif (Infinitive negation) 97
15 Îlot sujet (Subject island) 60
16 Îlot ajout (Addition/Adjunct island) 60
17 Îlot qu- (Wh-island) 60
18 Îlot SN (complex noun phrase / NP island) 60
19 Dépendance parasitique avec dont (Parasitic dependence with dont) 60
20 Préposition orpheline (Orphan preposition) 63

The appendix descriptions reported in the paper clarify the linguistic intent of several categories. Category 1 tests agreement of past participles in gender and number depending on the auxiliary and direct object position; category 2 tests subject-verb agreement, especially in the subjunctive; category 3 tests placement of ne...que; categories 5 and 6 test clitic placement; category 7 tests the formal two-part negation ne...pas; category 8 tests the des → de alternation before prenominal adjectives; category 12 tests the French -é / -er contrast; categories 15–18 test island constraints; category 19 tests correct use of dont as a de-complement relative pronoun; and category 20 tests the ban on preposition stranding.

A notable methodological caveat concerns category 9. Although it is labeled lexical semantics, the paper states that it often operationalizes the prescriptively preferred standard French lexical choice over an anglicism or non-standard term. The limitations section therefore treats it as less purely grammatical than many BLiMP-style categories.

3. Human annotation and reliability

Human annotation is central to QFrBLiMP. The paper reports that twelve native Quebec-French-speaking undergraduate students served as annotators (Beauchemin et al., 30 Sep 2025). They were not authors, were recruited via a University job board, were interviewed before the experiment, and were paid according to the University hourly pay scale. Each annotator worked at most 20 hours.

The annotation workflow included a preparation phase. Before annotation, the authors held a meeting to explain the task, the interface, and the annotation guidelines. Annotators then completed 15 practice instances in a pilot phase; these items were used to familiarize annotators and refine the guidelines, and they were discarded from the final dataset. A second meeting provided feedback and warned annotators about phenomena requiring special caution. Annotators were instructed to spend at most 5 minutes per sentence pair.

The interface was a customized version of Prodigy, in French. On each item, annotators saw the two members of a minimal pair and had to select the sentence they judged “well written.” The reported instruction was:

“Sélectionnez la phrase que vous jugez bien écrite parmi les deux phrases disponibles.”

with the English gloss:

“Select the sentence you think is well written from the two available.”

Order was randomized so that the grammatical sentence did not always appear in the same position. Each of the 1,761 final minimal pairs was annotated. Final gold labels were determined by majority vote; in the case of ties, the authors randomly selected one of the two options, and this occurred only 2 occurrences.

Agreement is reported using Worker Agreement With Aggregate (WAWA), defined as “the average fraction of the annotators’ votes that agree with the aggregated vote for each pair.” The average WAWA across categories is 86.31\%. The highest-agreement categories are LP 12 (-é/-er) at 96.08\%, LP 14 (Négation de l'infinitif) at 95.79\%, and LP 5 (Clitics in infinitive negation) at 95.54\%; LP 4 and LP 3 are both above 93\%. The lowest-agreement categories are LP 20 (Orphan preposition) at 69.84\%, LP 19 (Parasitic dependence with dont) at 72.78\%, LP 15 (Subject island) at 75.28\%, and LP 18 (NP island) at 75.56\%. The paper interprets these lower values as indicating either greater subjectivity or the need for refined definitions, while still treating the overall agreement pattern as evidence of benchmark reliability.

4. Model evaluation protocol

The paper benchmarks 77 open-source LLMs on both QFrBLiMP and the French portion of MultiBLiMP (Beauchemin et al., 30 Sep 2025). Because the task requires token probabilities or perplexities over specific sentences, evaluation is restricted to open models whose decoder probabilities were accessible. The benchmark suite intentionally varies model size, includes models marketed for reasoning, includes instruction-tuned and base pairs, and includes French-specialized models.

The evaluated range extends from FLAN-T5-small (76.9M) and SmolLM2-135m to Qwen2.5-72b and Meta-Llama-3.1-70b. French-specialized examples named in the paper include CamemBERT, Lucie, Claire, Chocolatine, BERT-base-French-europeana, and French-Alpaca-Llama3. Reasoning-branded models include Gemma variants, Llama 3.x variants, DeepSeek-R1 distills, Deepthink, QwQ, Phi-4, Reka-flash-3, and S1.1-32b.

The evaluation metric follows BLiMP and is reported as accuracy: a model is correct on a minimal pair if it prefers the grammatical sentence. The comparison itself is based on perplexity. The paper gives the following formula:

$\text{PPL}(s) = \exp\left(-\frac{1}{|s|}\sum_{i=0}^{|s|} \log\left({P_{\Theta}\left(x_i|x_{<i}\right)\right)\right) \right)$

where s|s| is the sentence length in tokens, xix_i are the words, and Θ\Theta are the model parameters. Following BLiMP and RuBLiMP, the sentences in a minimal pair are ranked based on their perplexity, so the model succeeds when the grammatical sentence receives lower perplexity and therefore higher probability. The implementation uses the official BLiMP Hugging Face metrics repository code.

Two baselines are reported. Random chooses one sentence of the pair uniformly at random. Human is defined as the majority choice of the twelve annotators. The human baseline is reported phenomenon-wise rather than as a single per-table scalar. Its values vary substantially across categories: 98.97 on infinitive negation (14), 95.00 on wh-islands (17), 94.85 on ne...que (3), but only 73.33 on NP islands (18), 75.00 on subject islands (15), 76.67 on dont dependencies (19), and 84.13 on orphan prepositions (20).

5. Empirical profile of model performance

A major result is a clear scaling trend. Overall QFrBLiMP accuracy rises strongly with model size, approximately linearly in log-parameter space. The fitted curve is drawn against a random baseline around 49.40 and a human baseline around 88.87 (Beauchemin et al., 30 Sep 2025). The paper concludes that grammatical competence on QFrBLiMP is an emergent ability that scales with model size, while also reporting a plateau: once models reach about 101010^{10} parameters, many cluster in the 85–90\% accuracy range, suggesting diminishing returns.

The benchmark also reveals a hierarchy of difficulty across phenomena. The easiest, described as “mastered,” are frequent and regular rules of morphology and syntax. Category 12 (-é/-er), category 5 (clitics in infinitive negation), and category 2 (verb inflection) are among the most successfully handled. The paper gives representative examples: Bloom-1b7 scores 100.00 on category 12, 97.89 on category 2, and 96.97 on category 5; Lucie-7b scores 100.00, 98.95, and 98.99 on those same categories; Qwen2.5-72b scores 99.00, 98.95, and 98.99. In category 12, where the human baseline is 93.00, such models exceed human majority consistency.

A middle tier includes more complex syntactic phenomena requiring longer-distance dependencies or more intricate rules. Past participle agreement, clitic movement, and island effects are often learned well but not flawlessly. The paper cites Qwen2.5-7b at 94.85 on past participle agreement (1) as a strong but imperfect example.

The hardest categories are 9 (Lexical semantics) and 20 (Orphan preposition). The paper treats this as its most important substantive finding. For category 9, the human baseline is 81.42. Strong models remain below it: Meta-Llama-3.1-8b-it reaches 78.76, Chocolatine-14b-it reaches 76.11, and Meta-Llama-3.1-70b reaches 75.22. Many otherwise strong models are in the high 60s or low 70s. For category 20, the gap is sharper: the human baseline is 84.13, the random baseline is 55.56, and the strongest reported model, Qwen2.5-72b, reaches only 60.32. Other large models remain poor, including DeepSeek-R1-distill-Qwen-32b at 57.14, Qwen2.5-72b-it at 57.14, XLM-RoBERTa-large at 55.56, Deepthink-reasoning-7b at 53.97, and Meta-Llama-3.1-70b and Phi-4 at 49.21. Some French encoders are near zero, including CamemBERT-large at 0.00, CamemBERT-base at 1.59, and BERT-base-French-europeana at 1.59.

The paper’s interpretation is that current LLMs are much better at formal, frequent, surface-learnable rules than at phenomena requiring deeper semantic or structural understanding. It explicitly argues that lexical semantics “requires real-world knowledge and an understanding of nuanced word meanings that cannot be resolved by syntactic plausibility alone,” while orphaned prepositions involve “a subtle and abstract syntactic rule that is rare in most training corpora.”

The human comparison is correspondingly nuanced. Models can outperform human majority-vote consistency on highly regular formal tasks. The paper does not treat this as evidence that models are “more linguistic” in a broad sense. Instead, it argues that models can be extremely consistent at applying frequent abstract patterns once learned, whereas human raters may be influenced by semantics, naturalness, or alternative heuristics. On semantically loaded or pragmatically delicate tasks, however, humans remain much stronger.

The comparison with MultiBLiMP-Fr reinforces this interpretation. Across the 77 models, performance on QFrBLiMP and MultiBLiMP-Fr is strongly positively correlated. Nearly all points in the reported figure lie above the parity line, meaning models score higher on QFrBLiMP than on the French MultiBLiMP subset. The paper does not interpret this as proof that QFrBLiMP is easier; instead, it hypothesizes that QFrBLiMP may be a cleaner benchmark because it is larger (1,761 vs 255) and more carefully controlled, thereby reducing noise and lexical confounds.

The results on specialization and instruction tuning are similarly restrained. French-specialized models are competitive but do not form a clearly superior class, and reasoning-branded models do not show a systematic advantage once scale is accounted for. The paper also highlights cases where instruction tuning appears detrimental to formal grammatical knowledge. Its example is Llama-3.1-70b versus Llama-3.1-70b-it on standard negation (7): 96.49\% for the base model versus 85.09\% for the instruction-tuned model. The authors speculate that alignment for conversational ability may interfere with preexisting grammatical knowledge, or that English-centered instruction tuning may mismatch a French grammatical judgment task.

6. Limitations, interpretive issues, and prospective extensions

The limitations section is explicit on several points (Beauchemin et al., 30 Sep 2025). First, because QFrBLiMP is sourced entirely from the BDL, there may be a distributional shift between benchmark data and ordinary web text used to pretrain LLMs. BDL examples are pedagogical, formal, and carefully selected to illustrate clear rules, so they may not represent naturally occurring language. Second, because the BDL is publicly available online, there is a real risk of data leakage / test contamination: some benchmark sentences may have appeared verbatim in model pretraining corpora, which could inflate scores through memorization rather than generalization.

Third, category 9 is not purely grammatical in the same way as agreement or island constraints. Because it involves anglicisms and prescriptive lexical choice, it reflects normative lexical prescription and Quebec sociolinguistic history; the paper therefore treats it as more subjective and less directly comparable to canonical BLiMP categories. A plausible implication is that QFrBLiMP measures not only formal morphosyntax but also institutionally codified standards of acceptable usage.

The ethical discussion adds that the BDL itself embodies institutional and political norms of Quebec French, so QFrBLiMP necessarily reflects a particular standard. Some minimal pairs judged valid in Quebec may not align with other Francophone norms. The authors also note a dual-use issue: improving LLM fluency and grammaticality can facilitate harmful applications such as disinformation or harassment.

Within those limits, QFrBLiMP fills a specific gap in French-language evaluation. It provides a benchmark tailored to Quebec French, with human judgments and manually curated examples, and it shows that current open LLMs learn a great deal of Quebec/French formal grammar, that this knowledge scales strongly with size, and that models may already rival or exceed aggregated human consistency on tightly formal contrasts. At the same time, it shows that average grammatical accuracy can obscure serious blind spots: when a task requires lexical meaning, semantic plausibility, or rare abstract constructions, performance can collapse.

The future directions proposed in the paper follow from these findings. They include building training corpora specifically optimized for Quebec French, investigating alternative evaluation approaches such as probing or semantic-similarity-based methods for low-performing categories, and studying more closely why instruction tuning may degrade formal grammatical knowledge. More broadly, the benchmark motivates alignment techniques that improve conversational ability without sacrificing foundational linguistic competence.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to QFrBLiMP.