- The paper provides a decomposition of wrong-consensus agreement in LLM self-consistency, evaluating the agreement index $\Gamma$ using a per-case preference channel using public data from Ding et al.
- The study reveals there is substantial per-case preference sufficiency of 59-78% giving the suggestion that previous models attributing the amplification primarily to shared bias could be over-estimating.
- From findings by different benchmarking methods (AIME and GPQA-Diamond), we notice that the first method when damaged measurably parental prep for such large effectiveness difference.
Overview and motivation
Majority voting over multiple samples of a LLM — self-consistency (Wang et al., 2022) — is a standard accuracy-raising technique, but its gains are erratic: on hard questions it can reduce accuracy relative to a single sample. This paper by Zhang et al. offers a quantitative, descriptive account of one facet of that failure: when a run's plurality answer is wrong, how much do its K samples coalesce around that wrong consensus, and what explains the coalescence? The paper's contribution is a counterfactual decomposition of a pluralistic agreement index Γ, applied to public per-run data from Ding's audit of self-consistency (Ding, 9 Jul 2026) on the GPT-4.1 family. It proposes no new voting method and claims no new backfire phenomenon; its claim is the decomposition itself.
The central question is whether wrong-consensus agreement is a generic mechanical effect (independent voters concentrating on an attractive per-case answer) or evidence of correlated error across runs. The paper answers this cell by cell against a hierarchy of progressively richer i.i.d. nulls: uniform wrong votes → fixed per-case preference → per-case preference with run-level heterogeneity.
The agreement index and its counterfactual nulls
For each case, a run consists of K=50 sampled answers; α is the share of votes for the run's plurality label. Conditioning on runs whose plurality is wrong, the empirical index is
Γemp​=dE[α∣run wrong]​,d=C−11−p​,
where p is single-sample accuracy and C the option count (C=4 for GPQA-Diamond; mean distinct answers per run for AIME). The normalization d is a reference scale, not a chance-correction claim; crucially, the headline coverage ratio ϕ cancels Γ0 exactly, so cross-benchmark comparisons are made on the Γ1-invariant Γ2 rather than on Γ3 magnitudes. The authors verify empirically that fixing Γ4 to 9 or 20 leaves Γ5 bit-identical, and that switching to random tie-breaking shifts Γ6 by at most 0.009.
Two i.i.d. multinomial counterfactuals are simulated (Γ7 draws per cell), both difficulty-matched at the case level:
- Uniform null Γ8: incorrect votes spread uniformly over the Γ9 wrong options. Pooling over difficulty instead destroys the phenomenon entirely — a pooled-K=500 control predicts K=501 wrong-consensus runs on three of four GPQA cells where the observed share is 42–60% — which justifies difficulty matching as the informative baseline.
- Leak-free per-case preference null K=502: each case's option preference K=503 and accuracy K=504 are estimated leave-one-out from its other runs only, so no run predicts its own agreement. K=505 is the maximal agreement attributable to per-case mechanical attraction under independence.
The headline quantity is the mechanical coverage K=506 on held-out test runs, with complement K=507, the preference-unexplained residual. A strictly more conservative uniform coverage K=508 is reported alongside. A pooled-preference variant serves as a negative control and sits far below K=509 everywhere (e.g., 1.01 vs. 4.26 on gpt-4.1 AIME), confirming that per-case structure, not any global preference, carries the mechanical agreement.
An important identification caveat is stated plainly: because temperature sampling produces positively correlated draws within a case, correlation can raise α0 above the correlation-free reference but never lower it. Hence a high α1 is identified in direction (the true mechanical coverage is at least as large), while α2 is an upper bound on the correlation-free residual. Both references are benchmark comparisons rather than certified bounds.
Mechanical coverage: multiple-choice versus open-domain
Across all eight cells, α3 — several times the reference scale — confirming strong attraction of wrong-run samples to the consensus. The decomposition then splits cleanly by benchmark type:
| Benchmark |
α4 range |
Residual α5 |
Reading |
| GPQA-Diamond (multiple-choice) |
0.806–0.927 |
0.07–0.19 |
Per-case preference largely suffices |
| AIME (open-domain) |
0.586–0.781 |
0.22–0.41 |
Substantial unexplained residual |
On GPQA-Diamond, the leak-free per-case preference reproduces roughly 81–93% of the observed agreement index. The implication is direct: accounts attributing multiple-choice amplification primarily to shared bias over-claim on this data, since a popular-but-wrong option the whole cohort latches onto is captured by the per-case preference channel. Whether that preference itself originates in shared training bias is not identified by the design.
On open-domain AIME, only 59–78% is mechanically explained at α6, leaving a residual of 1.56–2.80 α7 units. Shrinkage analysis shows this headline residual is conservative: shrinking the preference estimate toward pure noise lowers AIME α8 monotonically (to 0.206–0.292 at α9) while GPQA stays above 0.5 even at Γemp​=dE[α∣run wrong]​,d=C−11−p​,0, so the multiple-choice result does not ride on the quality of the preference estimate.
The benchmark split is described honestly as direction-consistent but partially overlapping: the four GPQA cells are the four largest Γemp​=dE[α∣run wrong]​,d=C−11−p​,1 values (exact permutation over all 70 labelings gives Γemp​=dE[α∣run wrong]​,d=C−11−p​,2), the mean difference is 0.19 (coupled bootstrap CI Γemp​=dE[α∣run wrong]​,d=C−11−p​,3), yet the boundary cells' clustered CIs overlap. Within benchmarks, Γemp​=dE[α∣run wrong]​,d=C−11−p​,4 is not monotone in single-sample accuracy, ruling out a simple difficulty-only account; across the eight cells, Γemp​=dE[α∣run wrong]​,d=C−11−p​,5 correlates positively with accuracy (Spearman Γemp​=dE[α∣run wrong]​,d=C−11−p​,6) and negatively with answer-space size (Γemp​=dE[α∣run wrong]​,d=C−11−p​,7) and openness (Γemp​=dE[α∣run wrong]​,d=C−11−p​,8), though these drivers are confounded at Γemp​=dE[α∣run wrong]​,d=C−11−p​,9.
A calibrated Dirichlet-multinomial null that grants run-level preference heterogeneity more than absorbs the AIME residual: p0–p1 on AIME versus 0.85–1.01 on GPQA, with the dispersion parameter fitted on a held-out half of cases and evaluated on the other half. The underlying dispersion signature is robust regardless of the null's fit: within-case cross-run variance of the plurality share is 2–3× larger on AIME than GPQA. The reading is that on open-domain tasks, runs of the same case differ in which wrong answers attract them — a correlated-error structure consistent with, but not identified as, the pretraining-bias account of systematic error patterns [McCoy et al., PNAS 2024].
Backfire, ceiling, fragility, and soft aggregation
Four complementary results characterize what high agreement actually buys:
Backfire replication. Difficulty-binned voting gaps reproduce the backfire pattern of Bahuguna et al. (Bahuguna, 11 Aug 2026): the hardest bins show negative gaps, e.g., p2 (coupled CI p3) for gpt-4.1-nano GPQA-ZS at per-case p4. Two of three representative hardest-bin intervals exclude zero; the third crosses zero. The global gap is positive because easy cases dominate, so the binned view is the honest scoping.
Low empirical ceiling. Runs in the top quintile of self-consistency reach consensus accuracy of only 0.42–0.83, a 1.2–3.6× lift over base rate (e.g., 0.59 vs. p5 for nano on AIME). High agreement is therefore graded evidence, never certification — a point reinforced by the decomposition, since wrong-run samples are themselves attracted to the consensus.
Champion fragility. Pairing two sampling conditions of the same case (gpt-4.1-mini; 353 pairs), the plurality winner flips between conditions 40% of the time on GPQA and 82% on AIME — the latter conflating winner instability with the larger candidate set, and measured cross-axis rather than same-parameter, as the authors note. High agreement only weakly stabilizes the winner: flip rates fall to 0.17 (GPQA, p6) and remain 0.78 (AIME, p7).
Jensen gap null result. Soft confidence p8 adds nothing over hard confidence p9: pooled AUROC 0.769 vs. 0.770, within 0.02 on every cell (per-cell AUROC 0.61–0.85). Notably, the Jensen gap C0 is diagnostic — roughly twice as large on wrong-plurality runs (mean 0.055 vs. 0.027) — showing that wrong runs carry more distributional mass beyond their winner, exactly the mass that available scores fail to exploit.
Limitations and scope
The principal limitation is the i.i.d. assumption: within-case sampling correlation is a third channel the design cannot separate from C1, and for non-exchangeable correlation structures even the direction of the comparison is unidentified. External validity is bounded by the GPT-4.1 family; a small Qwen3.5-9B arm shows direction-consistent uniform coverage (C2 on MMLU, 0.293 on MMLU-Pro) but uses different C3, lacks a rival null (one run per question makes leave-one-out preference undefined), and conflates difficulty with answer-space openness. Other disclosed constraints include unreleased sampling settings inherited from Ding's pipeline, string-level (not semantic) clustering of AIME answers — leaving a semantic-clustering sensitivity as future work — release majority labels inherited as-is for 30 of 5,300 runs, and a model-identity resolution through C4 keys internal to the release, which restricts paired analyses to mini-ZS. The difficulty/openness confound remains unresolved throughout.
Conclusion
This paper decomposes wrong-consensus agreement in LLM self-consistency into a mechanical per-case-preference component and a preference-unexplained residual, via a leak-free, difficulty-matched counterfactual hierarchy. On constrained multiple-choice data, the mechanical channel reproduces C5 of the agreement index — a result identified in direction even under sampling correlation — so shared-bias-dominant accounts over-claim there. On open-domain data, a residual survives that a calibrated run-heterogeneity null more than absorbs, consistent with run-to-run variation in error preferences. Together with the low high-agreement ceiling (0.42–0.83), substantial champion flip rates, and the failure of soft aggregation to improve on hard voting, the results reframe self-consistency confidence as graded evidence of correctness rather than certification. The claims are explicitly falsifiable through C6 and C7, and the natural open questions left by the paper are extending the rival decomposition across model families, separating within-case sampling correlation from the residual, and converting the channel decomposition into explicit uncertainty estimates.