Papers
Topics
Authors
Recent
Search
2000 character limit reached

Decomposing Wrong-Consensus Agreement in LLM Self-Consistency: A GPT-4.1 Case Study

Published 19 Aug 2026 in cs.CL and cs.AI | (2608.18795v1)

Abstract: Majority voting over multiple LLM samples is widely used to raise answer accuracy, yet its gain varies erratically: on hard questions it can even backfire. This paper gives a quantitative account of this failure. A pluralistic agreement index Gamma is defined as the expected fraction of the samples of a wrong run that agree with the consensus, normalized by a reference scale d=(1-p)/(C-1), and is decomposed into a mechanical component (what a vote delivers given only a per-case answer preference) and a preference-unexplained residual. The mechanical null is difficulty-matched and leak-free: each case is resimulated at its own accuracy and option preference, estimated from the case's other runs, so no run predicts its own agreement. On GPT-4.1 the decomposition shows benchmark-associated direction (an observational ordering over n=4 cells per benchmark, not a significance claim). On multiple-choice GPQA-Diamond, the per-case answer preference explains 81-93% of the held-out test-run agreement index: the shared-bias-dominates account over-claims here, because a wrong but attractive option the whole cohort latches onto is captured by the per-case preference channel (whether that preference is induced by shared training bias is not identified). On open-domain AIME, the mechanical preference explains only 59-78% (21-29% if shrunk to pure noise), and a preference-unexplained residual of 1.56-2.80 Gamma units survives, which a run-level preference-heterogeneity reference more than absorbs (1.4-2.1). A self-consistency backfire on hard questions is reproduced (binned voting gap down to -0.09, coupled CI [-0.12,-0.07]), and the highest-agreement bin reaches an accuracy of only 0.42-0.83, a 1.2-3.6x lift over base rate: agreement is graded evidence, not certification. No new voting method is proposed; code and evidence are committed and reproducible.

Summary

  • The paper provides a decomposition of wrong-consensus agreement in LLM self-consistency, evaluating the agreement index $\Gamma$ using a per-case preference channel using public data from Ding et al.
  • The study reveals there is substantial per-case preference sufficiency of 59-78% giving the suggestion that previous models attributing the amplification primarily to shared bias could be over-estimating.
  • From findings by different benchmarking methods (AIME and GPQA-Diamond), we notice that the first method when damaged measurably parental prep for such large effectiveness difference.

Overview and motivation

Majority voting over multiple samples of a LLM — self-consistency (Wang et al., 2022) — is a standard accuracy-raising technique, but its gains are erratic: on hard questions it can reduce accuracy relative to a single sample. This paper by Zhang et al. offers a quantitative, descriptive account of one facet of that failure: when a run's plurality answer is wrong, how much do its KK samples coalesce around that wrong consensus, and what explains the coalescence? The paper's contribution is a counterfactual decomposition of a pluralistic agreement index Γ\Gamma, applied to public per-run data from Ding's audit of self-consistency (Ding, 9 Jul 2026) on the GPT-4.1 family. It proposes no new voting method and claims no new backfire phenomenon; its claim is the decomposition itself.

The central question is whether wrong-consensus agreement is a generic mechanical effect (independent voters concentrating on an attractive per-case answer) or evidence of correlated error across runs. The paper answers this cell by cell against a hierarchy of progressively richer i.i.d. nulls: uniform wrong votes → fixed per-case preference → per-case preference with run-level heterogeneity.

The agreement index and its counterfactual nulls

For each case, a run consists of K=50K=50 sampled answers; α\alpha is the share of votes for the run's plurality label. Conditioning on runs whose plurality is wrong, the empirical index is

Γemp=E[α∣run wrong]d,d=1−pC−1,\Gamma_{\mathrm{emp}} = \frac{\mathbb{E}[\alpha \mid \text{run wrong}]}{d}, \qquad d = \frac{1-p}{C-1},

where pp is single-sample accuracy and CC the option count (C=4C=4 for GPQA-Diamond; mean distinct answers per run for AIME). The normalization dd is a reference scale, not a chance-correction claim; crucially, the headline coverage ratio ϕ\phi cancels Γ\Gamma0 exactly, so cross-benchmark comparisons are made on the Γ\Gamma1-invariant Γ\Gamma2 rather than on Γ\Gamma3 magnitudes. The authors verify empirically that fixing Γ\Gamma4 to 9 or 20 leaves Γ\Gamma5 bit-identical, and that switching to random tie-breaking shifts Γ\Gamma6 by at most 0.009.

Two i.i.d. multinomial counterfactuals are simulated (Γ\Gamma7 draws per cell), both difficulty-matched at the case level:

  • Uniform null Γ\Gamma8: incorrect votes spread uniformly over the Γ\Gamma9 wrong options. Pooling over difficulty instead destroys the phenomenon entirely — a pooled-K=50K=500 control predicts K=50K=501 wrong-consensus runs on three of four GPQA cells where the observed share is 42–60% — which justifies difficulty matching as the informative baseline.
  • Leak-free per-case preference null K=50K=502: each case's option preference K=50K=503 and accuracy K=50K=504 are estimated leave-one-out from its other runs only, so no run predicts its own agreement. K=50K=505 is the maximal agreement attributable to per-case mechanical attraction under independence.

The headline quantity is the mechanical coverage K=50K=506 on held-out test runs, with complement K=50K=507, the preference-unexplained residual. A strictly more conservative uniform coverage K=50K=508 is reported alongside. A pooled-preference variant serves as a negative control and sits far below K=50K=509 everywhere (e.g., 1.01 vs. 4.26 on gpt-4.1 AIME), confirming that per-case structure, not any global preference, carries the mechanical agreement.

An important identification caveat is stated plainly: because temperature sampling produces positively correlated draws within a case, correlation can raise α\alpha0 above the correlation-free reference but never lower it. Hence a high α\alpha1 is identified in direction (the true mechanical coverage is at least as large), while α\alpha2 is an upper bound on the correlation-free residual. Both references are benchmark comparisons rather than certified bounds.

Mechanical coverage: multiple-choice versus open-domain

Across all eight cells, α\alpha3 — several times the reference scale — confirming strong attraction of wrong-run samples to the consensus. The decomposition then splits cleanly by benchmark type:

Benchmark α\alpha4 range Residual α\alpha5 Reading
GPQA-Diamond (multiple-choice) 0.806–0.927 0.07–0.19 Per-case preference largely suffices
AIME (open-domain) 0.586–0.781 0.22–0.41 Substantial unexplained residual

On GPQA-Diamond, the leak-free per-case preference reproduces roughly 81–93% of the observed agreement index. The implication is direct: accounts attributing multiple-choice amplification primarily to shared bias over-claim on this data, since a popular-but-wrong option the whole cohort latches onto is captured by the per-case preference channel. Whether that preference itself originates in shared training bias is not identified by the design.

On open-domain AIME, only 59–78% is mechanically explained at α\alpha6, leaving a residual of 1.56–2.80 α\alpha7 units. Shrinkage analysis shows this headline residual is conservative: shrinking the preference estimate toward pure noise lowers AIME α\alpha8 monotonically (to 0.206–0.292 at α\alpha9) while GPQA stays above 0.5 even at Γemp=E[α∣run wrong]d,d=1−pC−1,\Gamma_{\mathrm{emp}} = \frac{\mathbb{E}[\alpha \mid \text{run wrong}]}{d}, \qquad d = \frac{1-p}{C-1},0, so the multiple-choice result does not ride on the quality of the preference estimate.

The benchmark split is described honestly as direction-consistent but partially overlapping: the four GPQA cells are the four largest Γemp=E[α∣run wrong]d,d=1−pC−1,\Gamma_{\mathrm{emp}} = \frac{\mathbb{E}[\alpha \mid \text{run wrong}]}{d}, \qquad d = \frac{1-p}{C-1},1 values (exact permutation over all 70 labelings gives Γemp=E[α∣run wrong]d,d=1−pC−1,\Gamma_{\mathrm{emp}} = \frac{\mathbb{E}[\alpha \mid \text{run wrong}]}{d}, \qquad d = \frac{1-p}{C-1},2), the mean difference is 0.19 (coupled bootstrap CI Γemp=E[α∣run wrong]d,d=1−pC−1,\Gamma_{\mathrm{emp}} = \frac{\mathbb{E}[\alpha \mid \text{run wrong}]}{d}, \qquad d = \frac{1-p}{C-1},3), yet the boundary cells' clustered CIs overlap. Within benchmarks, Γemp=E[α∣run wrong]d,d=1−pC−1,\Gamma_{\mathrm{emp}} = \frac{\mathbb{E}[\alpha \mid \text{run wrong}]}{d}, \qquad d = \frac{1-p}{C-1},4 is not monotone in single-sample accuracy, ruling out a simple difficulty-only account; across the eight cells, Γemp=E[α∣run wrong]d,d=1−pC−1,\Gamma_{\mathrm{emp}} = \frac{\mathbb{E}[\alpha \mid \text{run wrong}]}{d}, \qquad d = \frac{1-p}{C-1},5 correlates positively with accuracy (Spearman Γemp=E[α∣run wrong]d,d=1−pC−1,\Gamma_{\mathrm{emp}} = \frac{\mathbb{E}[\alpha \mid \text{run wrong}]}{d}, \qquad d = \frac{1-p}{C-1},6) and negatively with answer-space size (Γemp=E[α∣run wrong]d,d=1−pC−1,\Gamma_{\mathrm{emp}} = \frac{\mathbb{E}[\alpha \mid \text{run wrong}]}{d}, \qquad d = \frac{1-p}{C-1},7) and openness (Γemp=E[α∣run wrong]d,d=1−pC−1,\Gamma_{\mathrm{emp}} = \frac{\mathbb{E}[\alpha \mid \text{run wrong}]}{d}, \qquad d = \frac{1-p}{C-1},8), though these drivers are confounded at Γemp=E[α∣run wrong]d,d=1−pC−1,\Gamma_{\mathrm{emp}} = \frac{\mathbb{E}[\alpha \mid \text{run wrong}]}{d}, \qquad d = \frac{1-p}{C-1},9.

A calibrated Dirichlet-multinomial null that grants run-level preference heterogeneity more than absorbs the AIME residual: pp0–pp1 on AIME versus 0.85–1.01 on GPQA, with the dispersion parameter fitted on a held-out half of cases and evaluated on the other half. The underlying dispersion signature is robust regardless of the null's fit: within-case cross-run variance of the plurality share is 2–3× larger on AIME than GPQA. The reading is that on open-domain tasks, runs of the same case differ in which wrong answers attract them — a correlated-error structure consistent with, but not identified as, the pretraining-bias account of systematic error patterns [McCoy et al., PNAS 2024].

Backfire, ceiling, fragility, and soft aggregation

Four complementary results characterize what high agreement actually buys:

Backfire replication. Difficulty-binned voting gaps reproduce the backfire pattern of Bahuguna et al. (Bahuguna, 11 Aug 2026): the hardest bins show negative gaps, e.g., pp2 (coupled CI pp3) for gpt-4.1-nano GPQA-ZS at per-case pp4. Two of three representative hardest-bin intervals exclude zero; the third crosses zero. The global gap is positive because easy cases dominate, so the binned view is the honest scoping.

Low empirical ceiling. Runs in the top quintile of self-consistency reach consensus accuracy of only 0.42–0.83, a 1.2–3.6× lift over base rate (e.g., 0.59 vs. pp5 for nano on AIME). High agreement is therefore graded evidence, never certification — a point reinforced by the decomposition, since wrong-run samples are themselves attracted to the consensus.

Champion fragility. Pairing two sampling conditions of the same case (gpt-4.1-mini; 353 pairs), the plurality winner flips between conditions 40% of the time on GPQA and 82% on AIME — the latter conflating winner instability with the larger candidate set, and measured cross-axis rather than same-parameter, as the authors note. High agreement only weakly stabilizes the winner: flip rates fall to 0.17 (GPQA, pp6) and remain 0.78 (AIME, pp7).

Jensen gap null result. Soft confidence pp8 adds nothing over hard confidence pp9: pooled AUROC 0.769 vs. 0.770, within 0.02 on every cell (per-cell AUROC 0.61–0.85). Notably, the Jensen gap CC0 is diagnostic — roughly twice as large on wrong-plurality runs (mean 0.055 vs. 0.027) — showing that wrong runs carry more distributional mass beyond their winner, exactly the mass that available scores fail to exploit.

Limitations and scope

The principal limitation is the i.i.d. assumption: within-case sampling correlation is a third channel the design cannot separate from CC1, and for non-exchangeable correlation structures even the direction of the comparison is unidentified. External validity is bounded by the GPT-4.1 family; a small Qwen3.5-9B arm shows direction-consistent uniform coverage (CC2 on MMLU, 0.293 on MMLU-Pro) but uses different CC3, lacks a rival null (one run per question makes leave-one-out preference undefined), and conflates difficulty with answer-space openness. Other disclosed constraints include unreleased sampling settings inherited from Ding's pipeline, string-level (not semantic) clustering of AIME answers — leaving a semantic-clustering sensitivity as future work — release majority labels inherited as-is for 30 of 5,300 runs, and a model-identity resolution through CC4 keys internal to the release, which restricts paired analyses to mini-ZS. The difficulty/openness confound remains unresolved throughout.

Conclusion

This paper decomposes wrong-consensus agreement in LLM self-consistency into a mechanical per-case-preference component and a preference-unexplained residual, via a leak-free, difficulty-matched counterfactual hierarchy. On constrained multiple-choice data, the mechanical channel reproduces CC5 of the agreement index — a result identified in direction even under sampling correlation — so shared-bias-dominant accounts over-claim there. On open-domain data, a residual survives that a calibrated run-heterogeneity null more than absorbs, consistent with run-to-run variation in error preferences. Together with the low high-agreement ceiling (0.42–0.83), substantial champion flip rates, and the failure of soft aggregation to improve on hard voting, the results reframe self-consistency confidence as graded evidence of correctness rather than certification. The claims are explicitly falsifiable through CC6 and CC7, and the natural open questions left by the paper are extending the rival decomposition across model families, separating within-case sampling correlation from the residual, and converting the channel decomposition into explicit uncertainty estimates.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.