---
title: Copas Selection Model in Meta-Analysis
url: https://www.emergentmind.com/topics/copas-selection-model
type: topic
---

# Copas Selection Model in Meta-Analysis

Searching arXiv for recent and foundational papers on the Copas selection model and related extensions.
Publication bias in meta-analysis can be formalized through the Copas selection model, a Heckman-type selection framework that augments a random-effects outcome model with a latent publication mechanism. In its standard form, each study contributes an estimated effect size and a standard error, while publication is governed by an unobserved selection variable whose probability depends on study precision and may also depend on the study result through correlation with the outcome error. The model is used primarily for sensitivity analysis and bias adjustment, because the publication mechanism is only partially identified from published studies alone [2007.00836]. Subsequent work has extended the framework to rare-event generalized linear mixed models, diagnostic-test meta-analysis with exact likelihoods, registry-informed full-likelihood estimation, Bayesian robustification, and nonparametric worst-case bounds [2405.03603].

## 1. Core formulation

The standard Copas selection model starts from a random-effects meta-analysis for observed study effect sizes. One formulation is
\[
Y_i=\mu_i+\sigma_i\epsilon_i,\quad \epsilon_i \sim N(0,1), \quad \mu_i\sim N(\mu,\tau^2),
\]
or equivalently
\[
Y_i=\mu+\tau u_i+\sigma_i\epsilon_i,\quad u_i\sim N(0,1),\quad \epsilon_i\sim N(0,1)
\]
[2007.00836]. Here \(Y_i\) is the reported study effect size, \(\mu\) is the overall mean effect, \(\tau^2\) is the between-study heterogeneity, and \(\sigma_i^2\) is the within-study sampling variance. In practice, the observed \(s_i^2\) is typically used as a proxy for \(\sigma_i^2\) [2007.00836].

The publication process is represented by a latent variable \(Z_i\), with publication occurring iff
\[
Z_i>0.
\]
A standard precision-based selection equation is
\[
Z_i=\gamma_0+\gamma_1/s_i+\delta_i,
\]
with \(\delta_i\sim N(0,1)\) [2007.00836]. The corresponding marginal publication probability is
\[
\Pr(Z_i>0\mid s_i)=\Phi(\gamma_0+\gamma_1/s_i),
\]
so larger or more precise studies are more likely to be selected when \(\gamma_1>0\) [2007.00836].

The key structural feature is the dependence between the selection noise and the outcome noise:
\[
\left(\begin{array}{c} \epsilon_i \\ \delta_i \end{array}\right)\sim N\left\{\left(\begin{array}{c} 0 \\ 0 \end{array}\right),\left(\begin{array}{cc} 1 & \rho \\ \rho & 1 \end{array}\right)\right\}.
\]
The parameter \(\rho\) is the central publication-bias parameter. If \(\rho=0\), publication is independent of the observed effect size conditional on precision, so there is no publication bias in the Copas sense; nonzero \(\rho\) implies informative selection [2007.00836].

A closely related formulation, often called the Copas–Shi model, writes the selection variable as
\[
Y_i=\alpha_0+\alpha_1 / s_i+\delta_i,\qquad \delta_i \sim N(0,1),
\]
with \(\mathrm{corr}(\epsilon_i,\delta_i)=\rho\) [2005.14396]. The difference is not substantive; it is a notational variation of the same Heckman-type architecture.

## 2. Likelihood, conditional inference, and sensitivity analysis

Because only published studies are observed, inference proceeds from the likelihood conditional on selection. For the standard random-effects setup, the observed-data log-likelihood can be written as
\[
l \propto \sum_{i=1}^n \left[-\frac{1}{2} \log(\tau^2+\sigma_i^2)-\frac{(y_i-\mu)^2}{2(\tau^2+\sigma_i^2)}-\log\{\Phi(\gamma_0+\gamma_1/s_i)\}+\log\{\Phi(v_i)\}  \right],
\]
where
\[
v_i=\frac{\gamma_0+\gamma_1/s_i+\rho s_i(y_i-\mu)/(\tau^2+\sigma_i^2)}{\sqrt{1-\rho^2s_i^2/(\tau^2+\sigma_i^2)}}.
\]
In practice, \(\sigma_i^2\) is replaced by \(s_i^2\), leaving \(\mu,\tau,\rho,\gamma_0,\gamma_1\) as the unknown parameters [2007.00836].

This conditional-likelihood structure is the basis of classical Copas sensitivity analysis. The methodological difficulty is that the selection parameters are not fully identifiable from published studies alone. Accordingly, the traditional procedure fixes \((\alpha_0,\alpha_1)\) or \((\gamma_0,\gamma_1)\) as sensitivity parameters and estimates the remaining model parameters conditionally [2005.14396]. One interpretation aid is the expected number of unpublished studies,
\[
M=\sum_{i=1}^{N}\left\{\frac{1-P(Y_i>0\mid s_i)}{P(Y_i>0\mid s_i)}\right\},
\]
but choosing a plausible range for \(M\) is itself difficult [2005.14396].

A later GLMM extension for rare-event meta-analysis recasts this sensitivity analysis in terms of publication probabilities at the extremes of study size rather than directly through the latent coefficients. There, the analyst fixes \(min\) and \(max\), the probabilities of publishing the smallest and largest studies, and derives the selection parameters from
\[
min=\Phi(a_0+a_1\sqrt{n_{\min}}), \qquad max=\Phi(a_0+a_1\sqrt{n_{\max}}),
\]
after which \((\theta,\tau,\rho)\) are estimated by maximum likelihood [2405.03603]. This suggests an important conceptual point: the Copas framework is often best viewed not as a fully identified correction model, but as a structured sensitivity model for publication mechanisms.

## 3. Variants of the selection mechanism

The original and most widely used Copas specification links selection to study precision or size. In the standard aggregated-data version,
\[
Z_i = \alpha + \beta S_i^{-1} + \delta_i,
\]
and a study is published only if \(Z_i>0\) [2007.15955]. An alternative sample-size-based form is
\[
Y_i=\alpha_0+\alpha_1\sqrt{n_i}+\delta_i,
\]
which is especially useful when planned sample size is known from registries but standard errors are unavailable for unpublished studies [2005.14396]. A rare-event GLMM adaptation likewise uses
\[
Z_i = a_0 + a_1\sqrt{n_i} + \delta_i
\]
to reflect the usual small-study/publication-bias pattern without relying on normal approximations for sparse log-odds ratios [2405.03603].

A distinct line of development uses a Copas \(t\)-statistics selection model, in which publication depends monotonically on a study-specific \(t\)-statistic:
\[
P(\mathrm{select}\mid \gamma_0, \gamma_1, t^{(s)}) = a\left(t^{(s)}\right)=H(\gamma_0+\gamma_1 t^{(s)}),
\]
typically with \(H=\Phi\) [2406.04095]. In diagnostic-test meta-analysis, the key statistic is a linear combination of logit specificity and logit sensitivity, and special cases include lnDOR-based, sensitivity-based, and specificity-based selection [2406.04095]. This broadens the Copas idea from size-driven selection to significance-driven selection.

The distinction matters. One paper explicitly compares three publication-bias mechanisms: Copas’ original precision-based mechanism, significant-effect-size selection, and standardized-effect-size selection. It reports that Copas’ method performs well when data are generated under the mechanism it assumes, but can exhibit substantial bias and poor coverage under alternative realistic mechanisms, especially direct selection on standardized effect size [2007.15955]. A plausible implication is that the Copas model should not be interpreted as a universal structural law of publication, but as one parametric family within a larger class of selective-reporting mechanisms.

## 4. Methodological extensions

Several recent developments extend the Copas selection model beyond the conventional normal-normal random-effects setting.

For rare-event meta-analysis, the framework has been adapted to generalized linear mixed models with exact within-study likelihoods. In this setting, the within-study normal approximation for the log-odds ratio may be invalid when counts are sparse or zero, continuity corrections are ad hoc, and the estimated log-odds ratio and its standard error may not be independent. To avoid these issues, the extension combines a between-study normal random-effects model with either a hypergeometric-normal or binomial-normal GLMM, and then embeds a Copas-Heckman-type selection mechanism on top of that structure [2405.03603]. This extension also covers single-arm meta-analysis of proportions under a binomial GLMM [2405.03603].

For diagnostic-test accuracy meta-analysis, the Copas \(t\)-statistics selection model has been embedded in a bivariate binomial model that uses the exact within-study binomial likelihood instead of the bivariate normal approximation for empirical logits. Publication is modeled through a monotone function of a study-level statistic built from sensitivity and specificity, and the method yields publication-bias-adjusted estimates of the summary receiver operating characteristic curve and its area [2406.04095]. The authors identify this as the first application of the Copas \(t\)-statistics selection model to the bivariate binomial model [2406.04095].

A separate development replaces the conditional-likelihood-only approach with a full likelihood that combines the Copas-like conditional likelihood for published studies with a marginal semi-parametric empirical likelihood for the distribution of standard errors. In that framework, the model assumes
\[
\theta_i^* = \theta + \tau u_i + s_i^* \epsilon_i,
\qquad
Z_i = \gamma_1 + \gamma_2/s_i^* + \delta_i,
\]
with \((\epsilon_i,\delta_i)\) jointly normal and publication iff \(Z_i>0\) [2507.13615]. The paper proves identifiability when \(\rho \neq 0\), derives joint asymptotic normality for the MLEs, and shows that the full likelihood ratio follows an asymptotic central chisquare distribution [2507.13615]. Although the full and conditional methods are first-order asymptotically equivalent for several parameters, the full likelihood yields smaller mean squared errors and more accurate coverage probabilities in simulation [2507.13615].

## 5. Bayesian and robust formulations

The standard Copas model assumes normal study-specific random effects, but that assumption may be inadequate under heavy-tailed or outlier-prone between-study heterogeneity. A robust Bayesian Copas selection model addresses this by keeping the latent selection structure
\[
z_i = \gamma_0 + \gamma_1/s_i + \delta_i,\qquad y_i \text{ observed only if } z_i>0,
\]
while replacing the normal random effects with alternative heavy-tailed priors such as Laplace, Student’s \(t\), and slash distributions [2005.02930].

In this robust Bayesian Copas formulation,
\[
y_i = \theta + \tau u_i + s_i \epsilon_i,
\]
with \(u_i \sim p(u_i)\) and \(p(\cdot)\) chosen from normal, Laplace, Student’s \(t\), or slash families [2005.02930]. The model is estimated in a Bayesian hierarchical framework with weakly informative priors, and fitted using JAGS through a truncated bivariate-normal representation [2005.02930]. Model selection among competing random-effects distributions is based on DIC [2005.02930].

An additional contribution of that work is a quantitative measure of publication-bias magnitude based on Hellinger distance between the posterior distribution of \(\theta\) under the full robust Bayesian Copas model and the posterior obtained by fixing \(\rho=0\). The measure
\[
D= H\!\left( \widehat{\pi}_{\mathrm{rbc}}(\theta\mid \mathbf y), \widehat{\pi}_{\rho=0}(\theta\mid \mathbf y) \right)
\]
is interpreted on a scale from negligible to very high publication bias [2005.02930]. This shifts emphasis from testing whether \(\rho\) is nonzero to quantifying how much publication-bias correction changes inference on the overall effect.

## 6. Testing, identifiability, and model-robust bounds

The Copas model was originally used mainly for sensitivity analysis rather than formal hypothesis testing. A key reason is non-regularity under the null hypothesis of no publication bias. When testing
\[
H_0:\rho=0 \qquad \text{vs.} \qquad H_1:\rho \neq 0,
\]
the parameters governing the selection mechanism disappear from the relevant part of the likelihood under \(H_0\), creating non-identifiability and a singular Fisher information matrix [2007.00836]. Consequently, standard score, Wald, and likelihood ratio tests are invalid in their usual forms [2007.00836].

A score-based solution fixes \((\gamma_0,\gamma_1)\) at candidate values, constructs a score test for \(\rho=0\), and then maximizes over a grid:
\[
T=\sup_{\gamma_0,\gamma_1}\left[Z(\gamma_0,\gamma_1,\hat{\mu},\hat{\tau}) \right]^2.
\]
Its limiting distribution is the supremum of a squared mean-zero Gaussian process, and a parametric bootstrap is used for p-value computation in practice [2007.00836]. Simulation results reported in that work show well-controlled type I error and higher power than Egger’s test, Trim and Fill, and the earlier Copas naive test in most scenarios [2007.00836].

At the opposite end of the modeling spectrum, recent work develops worst-case bounds over classes of selection models related to the Copas-Jackson framework. The original Copas-Jackson bound assumes that the marginal publication probability conditional on total study standard deviation \(\sigma\),
\[
p(\sigma)=\Pr(\text{selected}\mid \sigma),
\]
is a non-increasing function of \(\sigma\) [2508.17716]. This assumption is compatible with Copas-Heckman-type selection but not with all \(t\)-statistics-based models, especially the 2-probit model [2508.17716]. To address this limitation, an extended bound is constructed over a broader class that includes both Copas-Heckman and common significance-based selection models, using simulation-based nonlinear optimization [2508.17716]. This suggests a broader methodological landscape: parametric Copas models provide interpretable sensitivity analyses, whereas Copas-Jackson-type bounds provide model-robust worst-case envelopes when the true selection mechanism is highly uncertain.

## 7. Empirical behavior, practical use, and controversies

Applied studies consistently portray the Copas selection model as a tool for structured sensitivity analysis rather than definitive bias correction. In a rare-event meta-analysis of catheter-related bloodstream infection, GLMM-based Copas-Heckman estimates of the log odds ratio remained fairly stable across increasingly severe publication-bias assumptions, whereas the normal-normal model showed more movement [2405.03603]. In a proportion meta-analysis from the same paper, the GLMM estimate again remained comparatively stable under sensitivity analysis while the normal-normal model changed more [2405.03603]. In diagnostic-test accuracy meta-analysis of CD64 for bacterial infection, SAUC remained above 0.5 under all assessed publication scenarios, although it changed noticeably as the assumed marginal publication probability decreased [2406.04095].

Registry-informed implementations are designed to reduce the arbitrariness of conventional Copas sensitivity grids by incorporating unpublished studies from clinical trial registries. In that approach, the full likelihood includes published studies through their effect sizes and standard errors, and unpublished studies through planned sample sizes only:
\[
\begin{aligned}
L_{full}(\theta,\tau,\rho,\alpha_0,\alpha_1) &=\sum_{i=1}^{N}\left[\log f(\hat{\theta_i}\mid Y_i>0;n_i)+\log P(Y_i>0;n_i)\right] \\
&\quad+\sum_{i=N+1}^{N+M}\left[\log P(Y_i<0;n_i)\right].
\end{aligned}
\]
By maximizing this likelihood, all unknown parameters can be estimated simultaneously [2005.14396]. Simulation results in that paper show smaller biases and more accurate confidence intervals than existing methods, and reanalyses of tiotropium and clopidogrel meta-analyses show that registry-informed adjustment can either preserve or materially weaken the original substantive conclusion, depending on the case [2005.14396].

The principal controversy concerns robustness to the assumed selection mechanism. One simulation study concludes that Copas’ method is not robust against realistic alternatives such as significance-based or standardized-effect-size-driven publication bias, and therefore questions its usefulness when the true mechanism is unknown [2007.15955]. Other work responds to the same concern by recommending multiple selection functions in sensitivity analysis [2406.04095], by developing broader worst-case bounds [2508.17716], or by using richer likelihoods and empirical likelihood components to stabilize inference [2507.13615]. This suggests that the main contemporary role of the Copas selection model is as a formal, likelihood-based language for publication-bias sensitivity analysis, whose conclusions gain credibility when triangulated across alternative selection mechanisms and modeling assumptions.

Source: https://www.emergentmind.com/topics/copas-selection-model