---
title: Accuracy Paradox Overview
url: https://www.emergentmind.com/topics/accuracy-paradox
type: topic
---

# Accuracy Paradox Overview

Searching arXiv for the cited "accuracy paradox" papers to ground the article in current literature.
Accuracy paradox denotes a family of phenomena in which higher measured, optimized, or reported “accuracy” ceases to track the property that actually matters, and can even degrade it. Across recent work, the paradox appears in at least five distinct forms: aggregation effects, where corpus-level correlations reverse segment-level tradeoffs; proxy optimization effects, where a more accurate intermediate signal or reward worsens downstream performance; strategic response effects, where better measurement changes behavior in ways that lower screening quality; metric insufficiency effects, where scalar accuracy obscures subgroup, epistemic, or cumulative harms; and estimation effects, where variance-based precision improves while bias-inclusive accuracy deteriorates [2402.12690] [2410.06554] [2605.13039].

## 1. Conceptual scope and metric structure

In its classical machine-learning form, the accuracy paradox refers to the fact that overall classification accuracy can be misleading under class imbalance, because a model can predict the majority class and still appear strong. Several of the cited works explicitly use this classical case as a point of departure, then generalize it to settings in which “accuracy” is either the wrong objective, an incomplete proxy, or a statistic whose interpretation changes under aggregation, strategic adaptation, or prevalence shift [2509.13345] [2607.05450] [2007.13632].

A recurring formal pattern is that single-number accuracy suppresses dimensions that remain decision-relevant. In the LLM hallucination literature, standard definitions are stated as
$$
\text{Acc} = \frac{TP + TN}{TP + TN + FP + FN}, \quad
\text{Prec} = \frac{TP}{TP + FP}, \quad
\text{Rec} = \frac{TP}{TP + FN},
$$
with related quantities such as F1, Brier score, expected calibration error, and coverage–risk used to expose confidence, calibration, and abstention behavior that raw accuracy omits [2509.13345]. In medical screening, positive predictive value is explicitly prevalence-dependent:
$$
\rho(\phi)=\frac{a\phi}{a\phi+(1-b)(1-\phi)},
$$
so unchanged sensitivity and specificity do not imply unchanged practical accuracy when prevalence changes [2011.06032]. In forecasting, pointwise RMSE, MAE, and \(R^2\) can improve while cumulative planning error worsens, motivating total percentage forecast error:
$$
TPFE = \frac{\left| \sum_{t \in T} \hat{y}_t - \sum_{t \in T} y_t \right|}{\sum_{t \in T} y_t} \times 100\%.
$$
This is an explicit redefinition of what “accurate” means for a planning task [2607.05450].

The shared lesson is not that accuracy is useless, but that it is task-conditional. Accuracy is reliable only when the metric, the decision unit, the sampling regime, and the behavioral environment align with the actual objective. When any of those conditions fails, paradoxical reversals become possible.

## 2. Simpson’s paradox and the translation accuracy–fluency tradeoff

In machine translation, the paradox is formalized most explicitly as a Simpson’s paradox. Accuracy, also called fidelity or adequacy, is defined as faithfulness to the source text; fluency is conformity to target-language norms and ease of processing for the recipient. The paper shows that these dimensions trade off within source segments but are positively correlated when aggregated over a corpus, because segment identity, difficulty, and genre act as confounders [2402.12690].

The formal statement uses accuracy \(A\), fluency \(F\), and stratum \(S\). Within a segment \(S_i\),
$$
\mathrm{Corr}(A,F\mid S_i)<0,
$$
but after aggregation,
$$
\mathrm{Corr}(A,F)>0.
$$
The mechanism is made explicit by the law of total covariance:
$$
\mathrm{Cov}(A,F)=E[\mathrm{Cov}(A,F\mid S)] + \mathrm{Cov}(E[A\mid S],E[F\mid S]).
$$
When the first term is negative and the second term is sufficiently positive, the aggregate correlation flips sign. Under the probability-based operationalization, the segment-level statistic is
$$
\rho=\mathrm{corr}(\vec{p}_{x|y},\vec{p}_{y}),
$$
computed across candidate translations \(y\) of a fixed source segment \(x\); \(\rho<0\) indicates tradeoff and \(\rho>0\) alignment [2402.12690].

The empirical evidence spans CRITT TPR-DB, RLTC, and MTMQM. Under the probability-based operationalization, corpus-level correlations are positive in all three datasets: CRITT \(r=.625\), RLTC \(r=.685\), and MTMQM \(r=.675\), all with \(p<.001\); the M2M100 replication gives .689, .703, and .801, again with \(p<.001\). By contrast, segment-level human-rating correlations are mostly negative and differ significantly from permutations with \(p<.001\). For human ratings at corpus level, MTMQM yields \(r=.396, p<.001\), whereas RLTC yields \(r=-.085, p<.001\), a small negative result. The segment-level English→German case study around “I issued you a refund of the book” illustrates the mechanism qualitatively: renderings using *ausstellen* are more faithful but less natural, whereas more idiomatic alternatives move away from the exact source sense [2402.12690].

The evaluation consequence is direct. Translation choices are made per segment, so segment-level analysis is the recommended unit. The paper therefore recommends reporting \(\mathrm{Corr}(A,F\mid S_i)\) alongside \(\mathrm{Corr}(A,F)\), stratifying by segment properties, and using partial correlation \(\mathrm{Corr}(A,F\cdot S)\) or \(\mathrm{Corr}(A,F\mid S)\). This suggests that many claims of “alignment” between adequacy and fluency are artifacts of aggregation rather than properties of the decision problem itself [2402.12690].

## 3. Proxy optimization, reward models, and strategic signaling

In RLHF, the accuracy paradox appears as a non-monotonic relationship between reward-model accuracy and downstream language-model quality. On QA-FEEDBACK, Longformer-base-4096 reward models are trained separately for relevance, factuality, and completeness, and then used to train T5-small, T5-base, and T5-large with PPO, a separate T5-base value model, and KL regularization. The central finding is that language models trained with moderately accurate reward models outperform those trained with the most accurate reward models when evaluated by independent high-accuracy evaluators. The RLHF objective is written as
$$
J(\pi)=E_{x\sim D,\, y\sim \pi(\cdot|x)}[r_\phi(x,y)]-\beta KL(\pi(\cdot|x)\|\pi_0(\cdot|x)),
$$
with KL coefficient \(0.3\), KL threshold \(20.0\), total episodes \(80{,}000\), learning rate \(1e{-5}\), and reward shaping settings of \(+0.3/-0.3\) for relevance, \(+0.5/-0.5\) for factuality, and mean \(-0.4468\), std \(8.3012\), bias \(0.0\), scale \(0.3\) for completeness [2410.06554].

The non-monotonicity is reported across all three tasks and all three T5 sizes. Reward-model accuracy ranges are task-specific: factuality \(0.64\)–\(0.77\), relevance \(0.49\)–\(0.69\), and completeness \(0.44\)–\(0.70\). Independent evaluator reward models attain offline accuracies \(69.6\), \(77.8\), and \(70.9\) for relevance, factuality, and completeness, respectively. Surface plots show an inverted-U pattern: too little reward accuracy underconstrains PPO, while too much accuracy exacerbates overoptimization of proxy-specific features and brittleness. The paper does not report correlation coefficients or significance tests for this relationship, but interprets the phenomenon through Goodhart’s law, reward hacking, distribution shift, overfitting, and calibration–discrimination mismatches [2410.06554].

A structurally related paradox appears in screening with strategic signaling. There, a principal observes a noisy signal
$$
s=e+\epsilon/\pi=e+\sigma\epsilon,
$$
where \(e\) is costly effort and \(\pi\) is signal precision. Approval follows a cutoff rule \(A(s)=\mathbbm{1}\{s>\tau\}\). Although larger \(\pi\) means more accurate signals in the accuracy order, the paper proves a “pitfall of precision”: when precision is already high, further improvements reduce screening accuracy and lower the principal’s welfare because marginally bad types are newly induced to exert effort and mimic good types [2605.13039].

The welfare derivative is decomposed as
$$
\frac{dW}{d\pi}=A_1(\pi)+A_2(\pi)+A_3(\pi),
$$
where \(A_1\) and \(A_2\) are direct informational effects and \(A_3\) is the indirect strategic effect via the cutoff type. The main theorem states that there exists \(\tilde{\pi}>0\) such that precision helps near \(\tilde{\pi}\), but there also exists \(\bar{\pi}>\tilde{\pi}\) such that \(dW/d\pi<0\) for \(\pi\ge \bar{\pi}\). Type I error \(\alpha(\pi)\) decreases in \(\pi\), but Type II error \(\beta(\pi)\) increases once precision is sufficiently large. The paper also reports a reversal of discrimination: groups with noisier technologies face lower approval rates yet may be preferred ex ante, and shows that commitment to an approval standard restores monotonicity, with \(\overline{W}(\pi)\) increasing in \(\pi\) [2605.13039].

Taken together, these results suggest that “better proxy” and “more precise signal” are not behaviorally innocent interventions. Once optimization or signaling responds endogenously to the proxy, the proxy’s nominal accuracy can diverge sharply from the quality of the final outcome.

## 4. Reasoning reliability beyond answer correctness

In mathematical reasoning, the paradox takes the form of a mismatch between final-answer correctness and the quality of the underlying computation. On a 500-problem GSM8K subset, Qwen2.5-Math-7B achieves \(61.0\%\) accuracy, but only \(56\) cases, or \(11.2\%\) of all predictions and \(18.4\%\) of correct predictions, are classified as faithful and stable. By contrast, \(249\) cases, or \(49.8\%\) of all predictions and \(81.6\%\) of correct predictions, are “lucky guesses,” and \(44\) cases, or \(8.8\%\), are silent failures: incorrect answers with high stability \(S(q)\ge 0.65\) [2603.03475].

The paper defines a composite faithfulness score
$$
F(q)=0.35\,S(q)+0.35\,A(q)+0.30\,E(q),
$$
where \(S(q)\) is activation stability, \(A(q)\) reasoning-hop alignment, and \(E(q)\) depth efficiency. Faithfulness is thresholded by \(T_F=0.60\), \(T_S=0.65\), and \(T_E=0.60\). Despite continuous fidelity being predictive with \(AUROC=0.78\), the observed correlation between latent fidelity and correctness is weakly negative: Pearson \(r=-0.2087\), Spearman \(\rho=-0.2054\), \(p=0.0020\). Reasoning steps also correlate negatively with correctness (\(r=-0.2699\), \(p<0.001\)), while reasoning depth \(D\) is not significantly correlated with correctness (\(r=0.0317\), \(p=0.4800\)). The interpretation offered is a binary-threshold artifact combined with a shallow-win phenomenon: correct answers often arise from surface heuristics rather than faithful multi-hop computation [2603.03475].

Self-correction work identifies another reversal. On GSM8K-Complex, DeepSeek-Chat has \(94.0\%\) baseline accuracy but an intrinsic correction rate of only \(16.7\%\) (\(5/30\)), whereas GPT-3.5-Turbo has \(66.4\%\) baseline accuracy and \(26.8\%\) intrinsic correction (\(45/168\)), and Claude-3-Haiku has \(70.4\%\) baseline accuracy and \(29.1\%\) intrinsic correction (\(43/148\)). Detection rates vary widely—GPT-3.5 \(81.5\%\), DeepSeek \(56.7\%\), Claude \(10.1\%\)—yet detection capability does not predict correction success. Model-generated localization hints reduce correction for GPT-3.5 from \(26.8\%\) to \(15.5\%\) and for Claude from \(29.1\%\) to \(12.8\%\), though DeepSeek increases from \(16.7\%\) to \(26.7\%\) on its smaller error set [2601.00828].

The proposed Error Depth Hypothesis explains the reversal by the composition of errors rather than answer accuracy alone. DeepSeek’s errors are \(44\%\) setup/interpretation, \(33\%\) logic, and \(22\%\) calculation, whereas GPT-3.5’s are \(25\%\), \(13\%\), and \(62\%\), respectively. Iterative reflection therefore helps weaker models more: cumulative correction over three rounds reaches \(67.9\%\) for GPT-3.5 and \(60.8\%\) for Claude, but only \(20.0\%\) for DeepSeek. This suggests that stronger models may fail less often, but when they fail, they fail in ways that are structurally harder to repair [2601.00828].

## 5. Fairness, hallucination, and screening

In fairness research, the paradox is that optimizing overall accuracy on imbalanced or spuriously correlated data can increase unfairness by concentrating errors on protected groups. The debiasing paper formalizes bias as an equal-opportunity-style TPR gap:
$$
\mathrm{bias}(\theta,t)=\left|P(\hat{t}=t\mid b=0,t^*=t)-P(\hat{t}=t\mid b=1,t^*=t)\right|,
$$
with overall model bias given by \(\sum_t \mathrm{bias}(\theta,t)\), and evaluates balanced average accuracy across target-by-bias subgroups. The proposed adversarial example–based data augmentation generates \(x_i^{adv}\) by solving
$$
x_i^{adv}=\arg\min_{x'}\big[\lambda L_b(x',a_{attack})+(1-\lambda)L_t(x',y_i)\big],
$$
thereby flipping the bias attribute while preserving the target label. On C-MNIST, AEDA\(_{robust}\) reaches bACC \(91.80\%\) and model bias \(0.53\), compared with Original \(55.62\%\) and \(7.84\); on CelebA, AEDA\(_{robust}\) reaches bACC \(74.30\%\) and bias \(3.27\), compared with Original \(73.57\%\) and \(5.48\) [2007.13632].

A conceptually broader LLM governance literature argues that overreliance on accuracy misdiagnoses hallucination risk. It distinguishes factuality hallucination, consistency hallucination, reference hallucination, sycophancy hallucination, consensus illusion, oversimplified hallucination, and prompt-sensitivity hallucination. The central claim is that optimizing for surface correctness and fluency fosters passive user trust in “accurate-looking” outputs that are epistemically ungrounded, manipulative, or socially distorting. Accuracy is therefore presented as a superficial proxy for reliability, because it does not guarantee epistemic validity, uncertainty communication, or interpretability. The paper extends the critique from outputs to individuals and society, emphasizing over-trust, rhetorical illusion, social sorting, privacy harms, equity harms, epistemic convergence, and social deskilling, and argues that the EU AI Act, GDPR, and DSA are not structurally equipped to address these harms when governance remains accuracy-centric [2509.13345].

In medical screening, the paradox is formalized through Bayes’ theorem. With sensitivity \(a\), specificity \(b\), and prevalence \(\phi\), positive predictive value is
$$
\rho(\phi)=\frac{a\phi}{(1-b)+\phi J}, \quad J=a+b-1,
$$
and
$$
\frac{d\rho}{d\phi}=\frac{a(1-b)}{\big[(1-b)+\phi J\big]^2}>0.
$$
Successful screening followed by treatment lowers prevalence, which therefore lowers PPV even if sensitivity and specificity remain unchanged. The ratio of later to initial PPV is
$$
\zeta(\phi_{0},k)=\frac{\phi_k(1-b)+J\phi_0\phi_k}{\phi_0(1-b)+J\phi_0\phi_k}.
$$
The paper also defines the number of positive test iterations needed to reverse the paradox:
$$
n_{i\phi_e}=\left\lceil\frac{\ln\left[\frac{\omega\phi_e\phi_k-\omega\phi_e}{\omega\phi_e\phi_k-\phi_k}\right]}{2\ln\omega}\right\rceil,
$$
where \(\omega=\sqrt{LR+}\). The result is a prevalence-dependent limit on practical screening accuracy that intensifies as false positives dominate at low prevalence [2011.06032].

These three literatures make a common point from different directions. Accuracy can hide subgroup error concentration, can overstate epistemic trustworthiness, and can decline dynamically even when the underlying test is unchanged. This suggests that fairness, calibration, provenance, and prevalence are not secondary auditing variables; they are constitutive of what counts as accurate performance in the first place.

## 6. Forecasting, tournaments, and metrology

In time-series forecasting, temporal disaggregation creates a granularity paradox. Finer grains increase sample size \(N(g)\) but also increase the recursive horizon \(H(g)\) needed to cover a fixed planning interval. Under recursive forecasting, small errors compound over the longer horizon. The paper formalizes this with, for example, an AR(1) recursion
$$
\hat{y}_{T+h}=\hat{\phi}^h y_T,
$$
so that with parameter bias \(\hat{\phi}=\phi+\delta\),
$$
e_{T+h}\approx -\, h\phi^{h-1}\delta y_T.
$$
Empirically, on a 13-year public procurement series resampled to six grains, Holt–Winters collapses at the Daily level with Test \(R^2=-151\) and \(TPFE=425.85\%\), while Linear Regression remains stable, with TPFE \(16.09\%\) at Monthly, \(16.96\%\) at Annual, and \(16.32\%\) at Daily. The LSTM exhibits a U-shaped curve: TPFE worsens from Monthly \(19.66\%\) to Bi-Weekly \(35.94\%\), then improves sharply at Daily to \(4.35\%\) with Test \(R^2=0.6627\). The proposed consensus–dissensus diagnostic compares the direction of change in pointwise metrics against TPFE to detect cases where standard diagnostics mask recursive error propagation [2607.05450].

In prediction tournaments, the paradox is an extreme-value phenomenon rather than a proxy-metric failure. With Brier score
$$
s_{ij}=(q_{ij}-X_i)^2,
$$
contestant \(j\)’s expected total score is
$$
E[S_j]=\sum_{i=1}^n p_i(1-p_i)+n\sigma_j^2,
$$
so lower RMS error \(\sigma_j\) implies better expected performance. In one-on-one play, this translates into strong win rates: a forecaster who is \(5\%\) more accurate wins about \(75\%\) of the time, and one who is \(10\%\) more accurate wins about \(90\%\) of the time. Yet in tournaments with \(300\) contestants and \(100\) questions, when \(\sigma\in[0,0.3]\), the winner is most likely to be around the 100th most accurate, and the top-ranked contestants never win in the reported simulations. The mechanism is a mean–variance tradeoff plus extreme-value selection: less accurate forecasters have worse expected scores but more variable scores, so one of them can realize an unusually favorable outcome trajectory and win the field [1903.02131].

Quantum metrology identifies an even sharper distinction between precision and accuracy. The paper defines accuracy as distributional separability,
$$
\alpha := \frac{|p(\theta)-p(\theta_0)|}{\Delta p(\theta)+\Delta p(\theta_0)},
$$
and precision as the smallest parameter shift that achieves a preset \(\alpha\). Using state distinguishability and the quantum Fisher information, it derives a corrected precision bound
$$
\Delta\theta \ge \frac{2\alpha}{\sqrt{\nu F_Q(\theta_0)}},
$$
which is a factor of \(2\) larger than the traditional \(1/\sqrt{\nu F_Q}\) resolution at \(\alpha=1\). The paper attributes the correction to the quadratic curvature of the probability–parameter relation near the working point. It further shows an inherent precision limit
$$
\Delta\theta_{inh}(\theta_0)=\theta_0-\arccos\left(\frac{2}{\nu}+\cos\theta_0\right),
$$
with \(\Delta\theta_{inh,\min}\approx 2/\nu\) at \(\theta_0=\pi/2\), achieving Heisenberg-like scaling even without entanglement. The cost is reduced accuracy, with \(\alpha\approx 1/\sqrt{\nu}\) under that strategy. The paper therefore concludes that accuracy may actually decrease with increasing sampling when one pursues excessive precision [2501.17427].

Across these domains, the paradox does not denote a single theorem but a recurring structural warning. Accuracy fails when it is aggregated over the wrong unit, optimized as a proxy, interpreted without strategic feedback, detached from subgroup or cumulative costs, or equated with variance while ignoring bias. The cited literature therefore replaces generic accuracy with task-aligned objects: segment-level conditional correlation in translation, independent evaluator performance and KL dynamics in RLHF, faithfulness and correction rates in reasoning, balanced subgroup metrics and provenance in LLM governance, PPV under changing prevalence in screening, TPFE under fixed planning horizons in forecasting, latent skill rather than winner-take-all rank in tournaments, and bias-inclusive distinguishability in metrology. This suggests that the central scientific question is not whether a system is accurate in the abstract, but which operational notion of accuracy remains invariant under the data-generating, optimization, and decision process of interest.

Source: https://www.emergentmind.com/topics/accuracy-paradox