Papers
Topics
Authors
Recent
Search
2000 character limit reached

External anchors reveal target-population effects hidden by published clinical-trial evidence

Published 5 Jul 2026 in stat.ME | (2607.04327v1)

Abstract: Meta-analyses guide clinical and policy decisions, but they estimate the effect among published trials, not the effect in the population a decision concerns. We introduce reference-anchored meta-analysis, which uses an external untreated-outcome distribution for two purposes: to define the target population, and, through each trial's control-arm mean as an externally anchored instrument, to correct publication selection. In an antidepressant literature with recovered unpublished trials, a strict holdout -- removing the hidden trials from both fitting and anchoring -- recovers a held-out all-trials benchmark. Applied to insomnia total sleep time and the placebo-controlled HbA1c literature in type-2 diabetes, it returns target-population effects materially smaller than the published averages, while a first-stage diagnostic flags, in advance, where the correction should not be attempted. Evidence synthesis becomes a target-explicit estimate with auditable selection assumptions.

Summary

  • The paper introduces reference-anchored meta-analysis to explicitly separate publication bias from target-population mismatch in synthesizing clinical trial evidence.
  • It leverages external untreated-outcome distributions and control-arm means as shadow instruments to recalibrate effect estimates for specific target populations.
  • Empirical evaluations in insomnia and type-2 diabetes demonstrate that this method substantially adjusts naive estimates, enhancing reliability over conventional approaches.

Reference-Anchored Meta-Analysis: Identifying Target-Population Effects in the Presence of Publication Selection

Introduction

This work introduces "reference-anchored meta-analysis," a methodological innovation addressing two fundamental sources of discrepancy between meta-analytic summary effects and the effect sizes pertinent for clinical or policy decisions: publication selection bias and target mismatch. While prior correction approaches primarily focus on publication bias using only the geometry of the published data (effect size, precision, or p-value distributions), this framework employs external, untreated-outcome distributions to both anchor estimands to a declared patient population and to identify evidence distortion due to publication selection. This dual-use of external information enables separation of these two sources of bias and provides a pathway to effect estimation directly relevant to policy and practice.

Methodological Framework

The proposed estimator critically depends on the availability of an external reference distribution for the untreated outcome in the trial-eligible target population. Such distributional information is generally accessible for normed biomarkers or objectively measured outcomes (e.g., total sleep time (TST), HbA1c, systolic blood pressure (SBP)), enabling a credible definition of the population for which the treatment effect is to be reported.

The methodology is distinguished by two key components:

  1. Target Definition via External Reference: The external untreated-outcome distribution fixes the population at which the treatment effect is reported, transforming the classical meta-analytic estimand from an average over the published trials' enrolled populations to the effect in the declared target population.
  2. Selection Correction via Anchor Instrument: The control-arm mean of each published trial is used as a "shadow instrument," leveraging its external anchoring to identify deviations from the target distribution in the published sample. The resulting deviation structure, conditional on the reported effect and design covariates, identifies the relative publication selection profile.

Identification of the bias-corrected mean effect at the target reference is achieved under three assumptions: (i) exclusion (publication does not depend on the absolute control mean given reported effect and covariates), (ii) knowledge of the external reference distribution, and (iii) relevance (the control mean must be informative for the outcome). Absolute publication rates, which are not identifiable from the published record alone, are estimated via registry calibration using external sources such as ClinicalTrials.gov.

Empirical Evaluation and Numerical Results

The framework is validated in settings where ground truth is available, notably in the antidepressant literature with comprehensive recovery of both published and unpublished trials. Under a strict holdout design—with all recovered-unpublished trials excluded from both analysis and anchor—the method recovers the all-trials aggregate effect (4∗=0.224^* = 0.22, 95% CI [0.15,0.30][0.15, 0.30] versus all-trials benchmark $0.248$ on HAM-D standardized mean difference), while conventional correctors (precision-effect, funnel-based, or p-value based) diverge or fail to adjust meaningfully. This out-of-sample recovery underscores the estimator's ability to target the inverse selection-weighted mean.

In practical applications where the anchor is genuinely external:

  • Insomnia – Total Sleep Time (TST): Naïve pooling gave a standardized mean difference (SMD) of $0.34$ (≈\approx +17 min), but the reference-anchored effect at the target population was −0.07-0.07 (95% CI [−0.23,0.05][-0.23, 0.05], or −3.5-3.5 min), well below minimal clinically important difference. This substantial adjustment is attributable to both selection correction and a large negative target shift due to more severe cases participating in published trials than in the target population.
  • Type-2 Diabetes – HbA1c: Naïve effect was $0.71$ SMD (≈0.50\approx 0.50 percentage points), while the target-anchored estimate was [0.15,0.30][0.15, 0.30]0 (95% CI [0.15,0.30][0.15, 0.30]1, [0.15,0.30][0.15, 0.30]2 points), roughly 40% below the published mean. Sensitivity analyses verified stability across reasonable anchor specifications and drug class partitions.

In settings where the external anchor is weak or non-existent (e.g., self-rated sleep quality scales in insomnia, lacking objective referents and with a weak first-stage instrument [0.15,0.30][0.15, 0.30]3), the method self-abstains from reporting a corrected effect, addressing risks of spurious inference.

Theoretical and Practical Implications

Separation of Bias Mechanisms: The methodology enables explicit deconvolution of publication selection from target-population mismatch, a distinction unavailable to conventional correction approaches based solely on the published record. This is critical, as prior adjustment procedures cannot resolve the confounding of these two forms of bias—especially under selection mechanisms tied to unobservable effect size.

Calibration and Sensitivity: Bootstrap interval estimation, registry-calibrated publication rates, and sensitivity analyses (anchor-sensitivity and Copas-type direct selection) are used to ensure robustness. When the anchor is credible and the instrument strong, the method achieves near-unbiased and near-nominal frequentist coverage in simulation, in contrast to persistent biases manifested by published-only correction models in simulation scenarios involving target shift or effect-dependent selection.

Generalization and Limitations: The method's applicability is inherently limited to outcomes and clinical areas where credible external untreated-outcome distributions are available. Its validity further depends on the exclusion restriction, relevance of the anchoring instrument, and specification of the target population. It explicitly does not address selective outcome reporting within trials.

Connections to Causal Inference: The statistical structure leverages recent advances in proximal causal inference and nonresponse instrumental variables, but adapts these to situations in which the required instrument is unobserved for missing units, necessitating external information for identification.

Complementarity with Existing Methods: The framework is compatible with robust Bayesian model averaging approaches, which combine model uncertainty over selection mechanisms; the reference-anchored estimator provides an additional identification source that is unavailable to model-averaging over published-only data.

Prospective Directions

The integration of external outcome distributions for both targeting and identification of selection mechanisms marks a direction for transforming meta-analytic practice towards more policy-relevant estimands. Further development could extend into domains with rich registry or observational data streams, sophisticated exclusion-sensitivity modeling, and greater automation of anchor construction. The core principles invite generalization across biomedical and social-scientific meta-analyses, provided credible untreated-outcome references exist.

Conclusion

The reference-anchored meta-analysis paradigm elevates meta-analytic inference by explicitly defining the target population and correcting for publication selection using external untreated-outcome distributions. This framework not only corrects for bias but fundamentally shifts what evidence synthesis estimates—from the average effect among published trials to the effect in a declared patient population. Strong empirical and simulation evidence demonstrates accurate recalibration of effects—materially reducing estimated benefit in some literatures and appropriately abstaining in others where external anchoring is fragile. By making both targeting and selection assumptions explicit and auditable, this approach enhances the credibility and relevance of evidence synthesis for decision-making in clinical and policy contexts.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.