Papers
Topics
Authors
Recent
Search
2000 character limit reached

Quantifying Evidential Rigor in Meta-Analytic Corpora: A Simulation-Characterized, Bias-Robust Bayesian Workflow with a Nutrition Case Study

Published 31 May 2026 in stat.ME and stat.AP | (2606.01428v1)

Abstract: Conventional meta-analysis summarizes evidence through pooled estimates, intervals, and p-values, but these outputs do not directly measure evidence for an effect, evidence for no effect, or the degree to which conclusions depend on publication selection or small-study effects. We introduce a corpus-scale Bayesian evidential-audit workflow for meta-analytic corpora. The workflow reconstructs or accepts study-level effects and standard errors, harmonizes directions, fits a matched Bayesian random-effects baseline and a bias-aware model-averaged ensemble, and reports paired estimates with component and joint model-family evidence. The central estimand is rigor: a joint Bayes-factor summary combining resolved effect/no-effect evidence with absence of an explicit bias component in the fitted ensemble. Rigor is not a positive-finding score; no-effect evidence can score highly, whereas inconclusive or bias-dependent evidence scores poorly. We characterize the workflow using an ADEMP-framed simulation/resampling design with known-cell synthetic simulation, empirical registry resampling, and empirical fitted-profile-weighted synthetic sampling. A nutrition intervention corpus provides the worked case study, where bias-aware fitting often attenuates conventional estimates and many nominally meaningful effects lose clean evidential support. A public companion repository provides empirical inputs, generated artifacts, simulation source/design files, and documentation for reproducing and adapting the audit.

Authors (1)

Summary

  • The paper introduces a reproducible corpus-scale Bayesian audit workflow that defines rigor as a joint model-family Bayes factor for clean evidence supporting an effect or no effect.
  • Simulation and resampling analyses show that small effects are often difficult to resolve, while stronger modeled bias reduces clean support and increases evidence for selection effects.
  • In a 175-outcome nutrition corpus, 62.3% showed weak or inconclusive effect evidence, and 75.9% of conventionally moderate effects were attenuated by at least 50% under bias-aware modeling.

Overview and motivation

This paper introduces a corpus-scale Bayesian workflow for "evidential auditing" of meta-analytic corpora, built around a new estimand the author terms rigor (2606.01428). The motivation is a persistent gap between what conventional meta-analysis reports—pooled estimates, confidence intervals, and p-values—and what those summaries do not provide: a direct quantification of how strongly the data support an effect, support no effect, or remain inconclusive. Publication selection and small-study effects, though widely recognized, are typically handled as diagnostics or sensitivity analyses rather than being integrated into the primary inferential summary. At corpus scale, this gap becomes consequential: a field, journal, or research program can produce many nominally meaningful pooled estimates while generating little durable evidential resolution.

The workflow reconstructs or accepts study-level effect sizes and standard errors, harmonizes effect directions, fits each outcome twice—once with a matched Bayesian random-effects baseline and once with the bias-aware RoBMA-PSMA model-averaged ensemble—and reduces the fits to auditable sidecars from which component and joint model-family evidence are summarized. RoBMA-PSMA is used as a deliberately fixed audit engine: a common model space is required so that evidence measures remain comparable across outcomes, strata, and corpora. The paper's contribution is the corpus-level protocol wrapped around this engine—the rigor estimand, the sidecar/registry reporting contract, an ADEMP-framed simulation and resampling characterization, and a public companion repository—rather than a new bias-adjustment model.

The rigor estimand

The central methodological object is rigor, defined as a joint model-family Bayes-factor summary. The RoBMA-PSMA ensemble crosses three binary dimensions: effect presence/absence (μ+/μ0\mu_+/\mu_0), heterogeneity (τ+/τ0\tau_+/\tau_0), and explicit publication-selection or small-study-effect components (ω+/ω0\omega_+/\omega_0). Two pre-specified branches are extracted: the effect-supporting branch Rk+\mathcal{R}_k^+ (models with μ+\mu_+ and ω0\omega_0, marginalizing over heterogeneity) and the no-effect-supporting branch Rk0\mathcal{R}_k^0 (models with μ0\mu_0 and ω0\omega_0). The headline quantity is

log10BFkR=max{log10BFkR,+,log10BFkR,0},\log_{10} BF_k^R = \max\{\log_{10} BF_k^{R,+}, \log_{10} BF_k^{R,0}\},

with the selected branch retained as interpretive metadata (rigor_direction). Three properties of this definition deserve emphasis. First, rigor is not a positive-finding score: strong evidence for no effect can score as highly as strong evidence for an effect, which matters because ruling out ineffective interventions is itself successful evidence accumulation. Second, rigor is a joint estimand and is not generally equal to a product of marginal component Bayes factors, because the effect and no-bias families average over different partitions of the model space; exact extraction requires summing prior and posterior mass over the joint branches. Third, negative rigor indicates that even the better-supported clean branch lost support relative to its complement—a failure of clean evidential resolution, not evidence for no effect.

The paper is careful about the model-space meaning of "clean": absence of the specified τ+/τ0\tau_+/\tau_00 component does not certify that a literature is free of all bias, does not encode study-level risk-of-bias domains, and does not certify internal validity. Stratum- and corpus-level summaries are finite-corpus means of selected log rigor Bayes factors, reported both outcome-weighted and article-balanced, with the latter downweighting source articles that contribute many reconstructed outcomes.

Simulation and resampling characterization

The workflow is characterized through three ADEMP-framed exercises (Q1–Q3), with the audit engine and reducers held fixed and only the evidence-generation mechanism varying. Q1 (known-cell behavior) uses a 36-cell synthetic design crossing four effect anchors (τ+/τ0\tau_+/\tau_01), three heterogeneity anchors, and three bias-burden levels, with τ+/τ0\tau_+/\tau_02 fitted synthetic outcomes per cell at the publication tier. The rigor atlas shows that selected rigor is hardest to resolve for small-effect cells—clean but weakly informative evidence remains evidentially unresolved—and that stronger bias burden appropriately lowers clean resolved support while increasing bias-component evidence. This is presented as a feature rather than a defect: the estimand withholds clean credit exactly where a nonzero estimate might otherwise be mistaken for clean evidence. Stability is summarized by the 90% interval width (τ+/τ0\tau_+/\tau_03) of stratum-level rigor summaries across outcome counts, with a minimum-τ+/τ0\tau_+/\tau_04 viability display identifying which known-cell structures support stable summaries at practical corpus sizes.

Q2 (empirical-registry stability) resamples the observed nutrition registry in two modes—outcome-row and source-article cluster resampling—to quantify finite-registry stability of rigor, bias evidence, and attenuation summaries. The source-cluster mode distinguishes ordinary small-sample variability from dependence on influential multi-outcome source syntheses. Q3 (empirical fitted-profile-weighted synthetic sampling) projects empirical outcomes into fitted-profile cells defined by bands on the bias-aware effect estimate, heterogeneity estimate, and bias Bayes factor, then samples from the Q1 synthetic library under support-masked shrinkage weights (τ+/τ0\tau_+/\tau_05 pseudo-outcomes). The paper is explicit that Q3 is a portability check, not calibration: fitted-profile cells are deterministic coarsenings of fitted outputs, not latent truth labels, and the bridge comparison against empirical bootstrap intervals is interpreted as a contrast, not as a tuning target. The paper also concedes that the characterization is conditional on these mechanisms and provides no universal operating guarantee.

Nutrition case study

The empirical demonstration applies the workflow to 175 reconstructed outcome-level meta-analyses across eight intervention strata (caffeine, creatine, diet, fasting, fiber, omega-3, protein, vitamin D). The results are descriptive properties of the analyzed corpus, not population estimates for nutrition as a field, and the corpus is purposive and extractability-conditioned. The headline pattern is limited clean evidential resolution. The strongest numerical findings are:

Component pattern τ+/τ0\tau_+/\tau_06 %
Weak/inconclusive effect evidence (τ+/τ0\tau_+/\tau_07) 109/175 62.3%
Strong evidence for nonzero effect (τ+/τ0\tau_+/\tau_08) 5/175 2.9%
Strong evidence for null effect (τ+/τ0\tau_+/\tau_09) 16/175 9.1%
At least moderate bias-component evidence (ω+/ω0\omega_+/\omega_00) 68/175 38.9%
Strong bias-component evidence (ω+/ω0\omega_+/\omega_01) 33/175 18.9%
Strong heterogeneity evidence (ω+/ω0\omega_+/\omega_02) 73/175 41.7%
Bias evidence with weak effect evidence (joint pattern) 45/175 25.7%
Conventionally nontrivial baseline effects (ω+/ω0\omega_+/\omega_03) 108/175 61.7%
Of those, shrinkage ω+/ω0\omega_+/\omega_04 under RoBMA-PSMA 82/108 75.9%
Of those, effect evidence inconclusive 86/108 79.6%

The most consequential claim is in the conventional-magnitude subset: among outcomes whose matched-baseline estimates were of at least moderate standardized magnitude, 75.9% showed at least 50% attenuation under the bias-aware ensemble, and 79.6% remained evidentially inconclusive for the effect component. The implication is that many nutrition meta-analytic results that attract attention under conventional pooled-effect summaries do not translate into strong clean evidential support once model uncertainty and explicit bias-adjustment components are propagated. The joint pattern—modeled-bias evidence in 25.7% of outcomes despite weak effect evidence—illustrates the failure mode the workflow is designed to surface: a study record can favor selection structure without resolving toward an effect.

Limitations and open questions

The paper states its limitations plainly. The corpus is purposive, so corpus summaries cannot be generalized without a sampling design. Reconstruction from legacy meta-analyses involves conversions, standard-error derivation, and repeated-measures assumptions that are documented but not lossless. Direction harmonization is a coordinate-system choice requiring substantive judgment. Rigor is model-space dependent: it is evidence about the fitted RoBMA-PSMA ensemble, not a truth label, and RoBMA-PSMA does not encode study-level risk-of-bias domains. The fixed audit engine aids comparability but means findings may be specific to that ensemble's priors and component definitions; systematic prior-sensitivity analysis for the joint rigor Bayes factor is identified as a needed extension, since joint Bayes factors can be prior-sensitive in ways marginal summaries do not reveal. Outcome multiplicity and source clustering are addressed only by article-balanced summaries and cluster resampling, not by a full hierarchical dependence model. Open questions the paper leaves include: how rigor summaries behave under alternative bias-adjustment families and comparator ensembles, whether selected rigor functions as a comparative measure outside nutrition (e.g., in journal-window, funder-portfolio, or prospective evidence streams), and how audit outputs could be propagated into downstream evidence synthesis without double use of primary trials.

Conclusion

The paper defines rigor as a joint model-family Bayes factor for resolved effect/no-effect evidence in the bias-component-absent branch of a fixed bias-aware ensemble, embeds it in a reproducible corpus-scale audit workflow with a matched random-effects baseline, characterizes the workflow through a three-part ADEMP simulation and resampling design, and demonstrates it on a 175-outcome nutrition corpus in which bias-aware fitting frequently attenuated conventionally meaningful effects and most such effects lacked clean evidential support. The companion repository converts the proposal into an inspectable, reusable protocol for asking whether a declared corpus of meta-analytic evidence is actually accumulating resolved, bias-robust conclusions.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.