---
title: No Evidence, No Score
url: https://www.emergentmind.com/topics/no-evidence-no-score
type: topic
---

# No Evidence, No Score

“No Evidence, No Score” denotes a methodological position according to which a numerical score, significance estimate, or benchmark outcome should not be treated as scientific evidence unless the evidential conditions that justify that number are themselves validated. In the works collected under this theme, the issue appears in several forms: evidence statistics that lack a legitimate measurement scale, benchmark scores that cannot show the evidence path that produced them, significance claims contaminated by construction artifacts or covariance mismatch, and empirical null-result studies in which improved instrumentation or stricter tests remove earlier signals. This suggests a common norm: a score is not self-authenticating merely because it is numeric [1805.11516][2605.10448].

## 1. Measurement theory and the problem of a zero point

Veronica Vieland’s “Absolutely Zero Evidence” treats the routine interpretation of \(P\)-values, maximum likelihood ratios, and Bayes factors as a measurement problem rather than merely a statistical one [1805.11516]. The paper reviews the standard scale types of representational measurement theory—ordinal, interval, ratio, and “absolute” in a stronger thermodynamic sense—and argues that familiar evidence statistics do not conform to any legitimate scale type. Scientists often use \(P\), \(-\log P\), \(\mathrm{MLR}\), \(\log \mathrm{MLR}\), \(\mathrm{BF}\), and \(\log \mathrm{BF}\) as if they preserved evidential meaning. Under measurement theory, that practice is already diagnostic: if logarithmic transformation is freely allowed, then the quantity is at best ordinal, because interval or ratio scales do not permit meaning-preserving nonlinear transformations. On that view, claims such as “twice as much evidence” or “the same evidential increment” are unjustified [1805.11516].

The paper’s central conceptual test is the existence of an “absolute 0 evidence.” Vieland argues that a genuine zero point for evidence cannot mean neutrality between hypotheses. Instead, it must correspond to the absence of relevant information, captured by the principle that evidence should approach its minimum as \(n \to 0\), or more generally as the amount of relevant information in the data approaches \(0\) [1805.11516]. This relocates the meaning of zero from “transition point” to “informational emptiness.”

That distinction is what invalidates the usual candidate evidence scores. In coin-toss examples, \(-\log P\), \(\log \mathrm{MLR}\), and \(|\log \mathrm{BF}|\) can all remain fixed at \(0\) while \(n\) increases and the evidential situation changes. Data at a Bayes-factor transition point may be highly informative and yet equally incompatible with both compared hypotheses. The score therefore conflates no relevant information with informative but hypothesis-balanced data. Vieland’s conclusion is accordingly stronger than a complaint about misuse: the usual minimum values “cannot be pressed into service as 0-points for a proper evidence measurement scale,” and without a defensible zero there is no scientifically validated evidence score in the strict representational sense [1805.11516].

## 2. Evidence-conditioned scoring in AI and agent evaluation

Recent work on agent evaluation makes the slogan operational by treating the evidence path as part of the object being scored. “GroundEval” defines a judge-free framework for stateful agents in which a plausible final answer does not receive credit unless the agent searched, fetched, cited, and was permitted to use the relevant artifacts [2606.22737]. The framework formalizes four inputs—event log, artifact corpus, access policy, and evaluation config—and evaluates both final answer and trajectory under grounded, time-bounded, and access-controlled constraints. In context mode, citation validity is defined by
\[
E_{\mathrm{valid}} = E \cap C \cap V_a \cap T_t,
\]
so a cited artifact counts only if it was injected into context, visible to actor \(a\), and available by the as-of time \(t\) [2606.22737]. The motivating Silence-track example is deliberately stark: two frontier LLM judges scored a plausible answer \(0.9\) and \(0.85\), but the agent had never fetched the decisive artifact, so GroundEval assigned an answer score of \(0.000\). Its compliance-adjusted score is
\[
S_{\text{adj}} = \left[w_a S_{\text{answer}} + w_t S_{\text{traj}}\right](1-v)^2,
\]
so repeated access, horizon, or subsystem violations cannot be washed out by a superficially correct answer [2606.22737].

A related diagnostic protocol for long-context and retrieval-augmented models introduces four matched evidence-availability conditions: no evidence, full context, retrieved evidence, and oracle-evidence reference [2606.06758]. Its central estimator, ONCU, is
\[
\mathrm{ONCU}_{\mathrm{raw}}(c)=\frac{S_c-S_{\mathrm{no}}}{S_{\mathrm{oracle}}-S_{\mathrm{no}}},
\]
and it is declared valid only when
\[
S_{\mathrm{oracle}} > S_{\mathrm{no}}.
\]
That denominator-validity requirement is the formal version of “no evidence, no score”: if oracle evidence does not improve over the no-evidence baseline, then there is no positive evidence-derived advantage to normalize, so a utilization ratio is uninterpretable [2606.06758].

The same logic appears in outcome verification for interactive agents. The outcome-evidence reporting layer assigns each completed record one of three labels—Evidence Pass, Evidence Fail, or Unknown—and refuses to collapse Unknown into a resolved success or failure [2605.10448]. With \(P\), \(F\), and \(U\) denoting those counts,
\[
N=P+F+U,
\qquad
\mathrm{CountedScore}(a,d)=\frac{P}{P+F},
\]
and the all-record performance is only partially identified:
\[
\mathrm{Perf}(a,d)\in\left[\frac{P}{N},\frac{P+U}{N}\right],
\qquad
\mathrm{Width}(a,d)=\frac{U}{N}.
\]
These are explicitly partial-identification bounds rather than confidence intervals. When \(U>0\), any point estimate resolves records without evidence [2605.10448].

Even in the absence of ground-truth labels, the same rule reappears in comparative safety scoring. “When No Benchmark Exists” argues that a score is deployment-relevant only under a fixed scenario pack, rubric, auditor, judge, sampling configuration, and rerun budget, and only after an instrumental-validity chain has been established: responsiveness to a safe-versus-abliterated contrast, dominance of target-driven variance, and rerun stability [2605.06652]. In the Norwegian validation of SimpleAudit, safe and abliterated targets separate with AUROC values between \(0.89\) and \(1.00\), target identity is the dominant variance component with \(\eta^2 \approx 0.52\), and severity profiles stabilize by ten reruns. The paper’s claim is narrow but exact: without that validation evidence, the score is merely output, not deployment evidence [2605.06652].

## 3. Contaminated scores and construction-dependent significance

A score may also fail because it is not measuring the intended object. The BOSS parity-violation analysis is a direct case in which the reported significance statistic was shown to be structurally contaminated [2407.03397]. Earlier work compressed the parity-odd four-point correlation function \(E_a\) into
\[
\chi^2 \equiv E_a (C_{\rm ana}^{-1})^{ab} E_b,
\]
but \(\chi^2\) is parity-even. Its expectation decomposes into a genuine parity-violation contribution and an additive contamination term,
\[
\chi^2_{\rm data}-\chi^2_{\rm mock}
=
\bar E_a (C_{\rm ana}^{-1})^{ab}\bar E_b
+
\mathrm{Tr}\!\left[(C_{\rm data}-C_{\rm mock})C_{\rm ana}^{-1}\right].
\]
The second term is an 8PCF mismatch term, so a large score can arise from data–mock covariance mismatch even when the parity-odd mean is zero [2407.03397]. The paper introduces \(\chi_\times^2\) to isolate the signal and \(\chi_{\rm null}^2\) to isolate the mismatch. Applied to BOSS, the clean parity-violation signal ranges from \(0\) to \(2.5\sigma\) depending on analysis choices, whereas the 8PCF bias term is \(\sim 6\sigma\); hence the conclusion that there is no compelling evidence for parity violation in BOSS [2407.03397].

An analogous problem appears in fact verification under the NEI (“Not Enough Information”) label. “Evidence Absence Is Not Evidence Insufficiency” argues that NEI is not construction-free: placeholder evidence, random irrelevant evidence, position-biased non-rationales, BM25 near-misses, cited non-rationales, same-document non-rationales, fixed-claim substitutions, and missing-hop constructions all instantiate different evidence conditions [2605.26663]. The paper’s decisive finding is that competence does not transfer reliably across those constructions. On SciFact, a DeBERTa-v3-base model trained on placeholder NEI achieves matched placeholder NEI-F1 of \(1.000\) but \(0.000\) on BM25 near-miss and cited hard NEI; on the held-out human-adjudicated hard-NEI set, placeholder and position-biased training both yield NEI recall \(0.000\), whereas BM25 near-miss and cited non-rationale training reach \(0.691\) and \(0.679\), respectively [2605.26663]. Fixed-claim diagnostics define a reference-label probability drop
\[
\Delta_i =
P_{\theta}(y_{\mathrm{ref}}\mid c_i,E_i^{\mathrm{ref}})
-
P_{\theta}(y_{\mathrm{ref}}\mid c_i,E_i^{\mathrm{hard}}),
\]
showing that changing only the evidence condition shifts confidence in the original Support/Refute label, not merely NEI recall [2605.26663]. The implication is precise: an aggregate NEI score can hide which evidential problem the model has actually solved.

## 4. Null tests in cosmology and galaxy evolution

In cosmology, “no evidence” often appears through explicit null-test parameterizations. A study of dynamical dark energy constructed two one-parameter deformations of flat late-time \(\Lambda\)CDM,
\[
\rho_{de1}=\rho_{de0}(1+z)^\alpha,
\qquad
\rho_{de2}=\rho_{de0}\left(1+\beta\frac{z}{1+z}\right),
\]
with the null conditions \(\alpha=0\) and \(\beta=0\) corresponding exactly to constant dark energy [1709.01074]. Using the combined “CBSLC” dataset—CMB, BAO, SNIa, lensing, and cosmic chronometers—the fitted values are \(\alpha=-0.14\pm0.22\) and \(\beta=0.15\pm0.13\). The latter is only about \(1.2\sigma\) from zero, so the paper concludes there is no evidence of dynamical dark energy at the \(1.2\sigma\) confidence level. The same analysis notes only slight alleviation of the \(H_0\) and \(\sigma_8\)-related tensions and a small one-parameter fit improvement of \(\Delta\chi^2=1.995\), which it does not interpret as strong support for dynamics [1709.01074].

A related Early Dark Energy analysis asks whether known Planck anomalies might be anti-correlated with EDE and therefore hide it in full Planck data [2203.12930]. The model uses the standard axion-like scalar field and the EDE fraction
\[
f_{\rm EDE}= \left.\frac{\rho_{\rm EDE}}{3M_{\rm Pl}^2 H^2}\right|_{z_c},
\]
then enlarges the parameter space with \(\Omega_K\), \(w\), and \(A_L\). Across all full-Planck-based fits, the posterior remains consistent with zero and yields only upper limits, typically \(f_{\rm EDE}\lesssim0.04\)–\(0.07\). Representative CMB-only bounds include \(f_{\rm EDE}<0.055\) for EDE+\(\Omega_K\) and \(f_{\rm EDE}<0.064\) for EDE+\(A_L\), with negligible correlation between \(f_{\rm EDE}\) and the anomaly parameters [2203.12930]. The paper’s conclusion is therefore not that EDE is impossible, but that full Planck contributes no positive evidence for it.

The same null-result logic extends to environmental effects in galaxy evolution. Using UNIONS \(u\)-band imaging as a resolved tracer of recent star formation, one study compared 5277 local satellites to 8360 matched field galaxies and quantified asymmetry with \(\delta_{\rm norm}\), \(\Delta(u-m)\), \(\Delta(u-r)\), \(A\), and \(A_s\) [2502.13123]. The full satellite and field distributions are described as “extremely similar” by eye; where formal AD-test differences appear, they often go in the opposite direction from the ram-pressure-enhancement picture, with slightly higher asymmetry in the field sample. Even the “RPS-likely” subset of 347 satellites in low-stellar-mass, high-halo-mass, early-infall conditions is indistinguishable from the low-mass field control. The conclusion is therefore that there is no strong statistical evidence for asymmetrically enhanced star formation in infalling galaxies, and any enhancement must be small, uncommon, or short-lived [2502.13123].

## 5. Experimental and observational null results in the physical sciences

A large class of “No Evidence, No Score” studies proceeds by remeasurement, contamination control, or higher-resolution follow-up. In helioseismology, a GONG-based reanalysis of 31 flares tested whether solar flares drive high-frequency global oscillations above the acoustic cutoff at \(5.3\) mHz [1206.6010]. The result is explicitly null: among the 31 flares, a decrease in acoustic power after the flare is just as likely as an increase; the Spearman rank correlation gives less than \(60\%\) confidence, and the Pearson correlation strongly indicates no correlation. The paper therefore finds no evidence that consistently supports flare-driven high-frequency waves [1206.6010].

In nuclear spectroscopy, a targeted remeasurement of the \(^{11}\mathrm{B}(^{3}\mathrm{He},d)^{12}\mathrm{C}\) reaction at \(44\) MeV revisited the long-cited \(11.16\) MeV \(2^+\) state in \(^{12}\mathrm{C}\) [1206.4217]. Modern low-background spectra at \(25^\circ\), \(30^\circ\), and \(35^\circ\) show “a deep valley where the 11.16 MeV state is expected,” and line-shape fitting reproduces the region with the known neighboring structures and background, without any additional resonance. The paper concludes that there is no evidence for the previously reported state and that the relevant \(2^+\) strength lies instead in the broad \(9.6\)–\(9.8\) MeV structure [1206.4217].

In star formation, higher-resolution \(1.2\) mm mapping of Barnard 59’s nuclear clump confirms the earlier extinction-based impression of a monolithic object [1201.2129]. After subtraction of known YSO contributions, the central clump remains smooth, with no significant evidence for prestellar fragmentation on scales below about \(1.5\times10^4\) AU. The system is not inert—the paper quotes a virial parameter \(\alpha_{\mathrm{vir}}=0.25\)—but it is not breaking into multiple dense subcores within the observed resolution and sensitivity [1201.2129].

In high-pressure hydrogen physics, a rebuttal to the claim of a new phase above \(325\) GPa at \(300\) K argues that the reported evidence does not meet ordinary phase-transition criteria [1605.05703]. The key points are all evidential. The \(L_2\) mode remains visible to at least \(350\) GPa; the \(L_3\) mode moves under the \(1332\ \mathrm{cm^{-1}}\) first-order diamond phonon near \(320\) GPa and is therefore masked rather than shown to disappear; and the linewidth change in \(L_1\) begins around \(360\) GPa rather than \(325\) GPa. The purported high-frequency vibron anomalies also overlap the second-order diamond Raman band between \(2300\) and \(2600\ \mathrm{cm^{-1}}\). The conclusion is that there is no evidence for a phase transition at \(325\) GPa [1605.05703].

In debris-disk astronomy, V488 Per provides a case in which the strongest results are again absences [2108.03700]. Herschel photometry confirms an outer dust population with a blackbody-fit temperature of \(\sim130\) K, while Subaru/COMICS spectroscopy of the \(\sim800\) K inner dust does not detect any obvious solid-state emission features. The combined radial-velocity and adaptive-optics campaign also finds no evidence for stellar or substellar companions within several hundred AU, with companions in the warm/cool belt gap constrained to \(\lesssim10\,M_{\rm Jup}\) under the adopted two-belt architecture [2108.03700]. The paper’s interpretation remains cautious: metallic iron or amorphous carbon are allowed because they do not produce strong \(8\)–\(13\,\mu\mathrm{m}\) features, but they are not directly detected [2108.03700].

## 6. Statistical, institutional, and practical consequences

The slogan also appears in classical significance testing and institutional scoring systems. In the critique of alleged planetary influence on solar activity, Cameron and Schüssler redo the significance calculations with corrected white- and red-noise generation, matched preprocessing, and unbiased frequency bins of width \(3.7\Delta P\) [1307.5988]. The resulting chance-coincidence probabilities rise to \(7.5\%\) for white noise and \(22\%\) for red noise, far above the values claimed in the original planetary-torque analysis. The conclusion is not that no planetary mechanism is possible in principle, but that the apparent period matches are statistically insignificant and provide no evidence for planetary influence on solar activity [1307.5988].

The same institutional lesson appears in consumer credit scoring. Albanesi and Vamossy argue that credit scores are central to U.S. debt allocation despite little public evidence on their performance [2409.00296]. Benchmarking a widely used score against a machine-learning default model built from legally permissible credit-report features, they find average AUC \(0.856\) for the score and \(0.906\) for the model, along with \(41\%\) misclassification across broad risk categories. The misclassification is especially severe among low-score borrowers: \(54.94\%\) for Deep Subprime and \(69.65\%\) for Near Prime [2409.00296]. The paper’s equity argument is performance-based rather than rhetorical: the alternative model improves predictive accuracy most for young, low-income, and minority groups because it performs better on thin-file and otherwise low-quality-data cases. This suggests that heavily relied-upon score categories can outrun the publicly available empirical case for their validity [2409.00296].

Taken together, these works suggest that “No Evidence, No Score” is not a single doctrine but a family of evidential constraints on quantification. A \(P\)-value, Bayes factor, AUROC, benchmark win rate, safety score, NEI-F1, or credit category may be computationally well defined and still fail as evidence if its zero point is equivocal, its denominator is invalid, its construction is shortcut-prone, its covariance is contaminated, its retained artifacts do not verify the claimed outcome, or its subgroup performance is unknown [1805.11516][2605.10448]. In that sense, the principle is not anti-measurement. It is a demand that measurement claims inherit the same rigor as the hypotheses they are used to support.

Source: https://www.emergentmind.com/topics/no-evidence-no-score