---
title: Marginal Sensitivity Model Overview
url: https://www.emergentmind.com/topics/marginal-sensitivity-model
type: topic
---

# Marginal Sensitivity Model Overview

Searching arXiv for the cited marginal sensitivity model papers to ground the article in current preprints.
The marginal sensitivity model (MSM) is a sensitivity-analysis framework for causal inference under unmeasured confounding that replaces no unmeasured confounding with an explicit bound on how strongly an unobserved variable may perturb treatment assignment relative to what is explained by measured covariates alone. In the canonical binary-treatment formulation, one observes i.i.d. copies of $(Y,T,X)$, introduces an unmeasured confounder $U$, and constrains the relationship between $e(x,u)=P(T=1\mid X=x,U=u)$ and $e(x)=P(T=1\mid X=x)$ through a bounded-odds-ratio or equivalent propensity-ratio condition parameterized by $\Gamma\ge 1$ [2505.13868]. This construction preserves partial identification, yields sharp bounds for common causal functionals such as $E[Y(1)]$, $E[Y(0)]$, and the average treatment effect (ATE), and has become a central template for extensions to weighted observational studies, continuous treatments, marginal structural models, alternative sensitivity scales, and joint treatment–outcome sensitivity models [1711.11286] [2307.00093] [2204.10022] [2210.04681] [2211.04697] [2504.08301] [2505.13868].

## 1. Formal definition and core assumptions

In the binary-treatment potential-outcomes setup, the observed outcome is
$$
Y = T Y(1) + (1-T)Y(0),
$$
with $T\in\{0,1\}$, measured covariates $X$, and unmeasured confounder $U$ [2505.13868]. The starting point is the latent-variable ignorability condition
$$
(Y(0),Y(1))\perp T \mid (X,U),
$$
which allows confounding through $U$ while preserving conditional independence after augmenting the covariate set by the latent variable [2505.13868] [2504.08301].

The defining feature of MSM is a restriction on how much $U$ may alter treatment assignment. One formulation imposes
$$
\frac{1}{\Gamma}\le \frac{e(x,u)}{e(x)}\le \Gamma \quad \text{for all }x,u,
$$
where $e(x,u)=P(T=1\mid X=x,U=u)$ and $e(x)=P(T=1\mid X=x)$ [2505.13868]. In weighted observational studies, the same model is written on the inverse-odds scale as
$$
\frac{1}{\Gamma}\le \frac{w^*(x,u)}{w(x)}\le \Gamma,
$$
where $w(x)=e(x)/(1-e(x))$ and $w^*(x,u)=e(x,u)/(1-e(x,u))$ [2307.00093]. A related formulation constrains the logit difference,
$$
\|h\|_\infty \le \lambda:=\log\Lambda,
$$
with $h(\mathbf x,y)=\logit e_0(\mathbf x)-\logit e(\mathbf x,y)$ and $\Lambda\ge 1$ [1711.11286]. These parameterizations are equivalent ways to encode a worst-case bound on latent confounding strength.

When $\Gamma=1$ or $\Lambda=1$, the model collapses to no unmeasured confounding [2505.13868] [1711.11286]. For $\Gamma>1$, the observed data no longer identify causal estimands pointwise, but MSM defines an identified interval over all complete-data laws compatible with the observed distribution and the sensitivity constraint [1711.11286] [2504.08301]. This is the source of the model’s interpretability: the sensitivity parameter directly quantifies the maximal multiplicative distortion of treatment assignment attributable to hidden confounding.

A common misconception is that MSM assumes a specific parametric form for the unmeasured confounder. The supplied formulations do not do so. They instead restrict the extent to which the full-data propensity may differ from the observable marginal propensity, leaving the latent structure otherwise unspecified [1711.11286] [2210.04681]. This suggests that MSM is best understood as a semiparametric sensitivity class rather than a structural confounding model.

## 2. Sharp bounds for potential outcomes and the ATE

Under MSM, sharp bounds are available for missing counterfactual regressions and for average potential outcomes. In the exposition of Zhang, Xu, and Tan, the sharp upper bound for
$$
v_1(x)=E[Y(1)\mid T=0,X=x]
$$
under MSM is
$$
v_{1,\mathrm{MSM}}^{+}(x)=E[Y\mid T=1,X=x]+(\Gamma-1)\,E[\rho_\Gamma(Y;x)\mid T=1,X=x],
$$
where $\rho_\Gamma(Y;x)$ is a suitable check-loss-based quantile-contrast [2505.13868]. The corresponding upper bound for
$$
p_1=E[Y(1)]
$$
is
$$
p_{1,\mathrm{MSM}}^{+}=E\!\left[T\,Y+(1-T)\,v_{1,\mathrm{MSM}}^{+}(X)\right].
$$
Analogous formulas hold for lower bounds, for $p_0=E[Y(0)]$, and hence for the ATE [2505.13868].

A more explicit quantile representation appears in the enhanced-MSM exposition. There, for each $x$,
\begin{alignat*}{2}
\nu^+(x)
&=E[Y\mid A=1,x] \\
&\quad +\bigl(\Lambda-1/\Lambda\bigr)\,
E\!\left[\rho_{\tau(x)}\bigl(Y,q_{\tau(x)}^*(x)\bigr)\mid A=1,x\right],\qquad
\tau(x)=\frac{\Lambda-1}{\Lambda-1/\Lambda},
\end{alignat*}
with
$$
\rho_\tau(y,q)=\tau (y-q)_+ + (1-\tau)(q-y)_+,
$$
and the sharp upper bound of $E[Y^1]$ given by
$$
\mu_{\rm MSM}^{1+}=E\bigl[A\,Y+(1-A)\,\nu^+(X)\bigr].
$$
The ATE interval is then
$$
\mathrm{ATE}\in \bigl[\mu_{\rm MSM}^{1-}-\mu_{\rm MSM}^{0+},\;\mu_{\rm MSM}^{1+}-\mu_{\rm MSM}^{0-}\bigr].
$$
These formulas show that the extremal law tilts mass toward outcome tails via quantile structure rather than through an arbitrary reweighting [2504.08301].

In weighted-estimator formulations, the same sharpness principle appears as an extremal linear program. Zhao et al. show that the extrema can be obtained by sorting control outcomes and reweighting the largest outcomes by the upper bound $\Gamma$ and the smallest by $1/\Gamma$ [2307.00093]. The percentile-bootstrap paper expresses the stabilized IPW estimate as a linear-fractional program in variables
$$
z_i=\exp(h(X_i,Y_i))\in[\Lambda^{-1},\Lambda],
$$
then applies a Charnes–Cooper transformation to obtain an ordinary linear program [1711.11286]. The optimal solution pushes each $z_i$ to one of the two endpoints in the same order as the outcomes, again exposing the thresholding character of sharp MSM bounds [1711.11286].

This quantile-threshold structure is significant because it explains why MSM bounds are computationally tractable despite being defined over infinite-dimensional sensitivity classes. The extremal confounding pattern is not arbitrary; it concentrates on tail outcomes in a way determined by the sensitivity parameter [1711.11286] [2307.00093] [2504.08301].

## 3. Weighted estimators, design sensitivity, and bootstrap inference

In inverse-probability weighting settings, MSM is used to assess how unmeasured confounding can shift weighted treatment-effect estimators away from the target causal estimand. Huang, Soriano, and Pimentel define the bias of a weighted ATT estimator through
$$
\mathrm{Bias}(\hat T_w/w^*) = T(w^*)-T(w),
$$
where
$$
T(w)=E[Y\mid Z=1]-E[w\,Y\mid Z=0],
$$
and $w^*$ denotes the full-data weight satisfying the marginal sensitivity constraint [2307.00093]. The identified sensitivity interval is written as
$$
[L_{\mathrm{msm}}(\Gamma),U_{\mathrm{msm}}(\Gamma)],
$$
formed by the infimum and supremum of $T(w^*)$ over admissible weights [2307.00093].

A notable development is the introduction of design sensitivity for MSM-based weighted studies. Design sensitivity is the critical $\Gamma^*$ at which asymptotic power drops from one to zero as sample size increases [2307.00093]. Huang et al. characterize $\Gamma^*$ through an estimating equation involving the control-potential-outcome quantile function and the indicator
$$
G_\theta(y)=1\{y\ge F^{-1}(1-\theta)\},
$$
with the solution arising from the worst-case allocation of weight to high and low $Y(0)$ values [2307.00093]. This places MSM not only in post hoc robustness analysis but also at the design stage of observational studies.

The same paper shows that trimming and augmentation affect design sensitivity differently under MSM. For trimming, one conditions on $w<m$ in the design-sensitivity equation and replaces $T$ by the trimmed ATT $T_{\rm trim}$ [2307.00093]. In simulations and the FARC example, trimming under the marginal model often brings little gain for small effects and can even harm design sensitivity for large effects, because one loses the units with the largest outcomes that a worst-case bias would exploit [2307.00093]. By contrast, augmentation replaces $Y(0)$ by residuals $e=Y-g(X)$ in the design-sensitivity equation and typically increases design sensitivity quite substantially [2307.00093].

For interval estimation, the percentile-bootstrap framework of Zhao, Small, and Bhattacharya constructs confidence intervals that cover the entire partially identified set under MSM [1711.11286]. For a fixed sensitivity function $h$, the percentile-bootstrap interval is
$$
[L^{(h)},U^{(h)}]
=
\left[
Q_{\alpha/2}\bigl(\hat{\hat\mu}^{(h)}_{\bullet}\bigr),
Q_{1-\alpha/2}\bigl(\hat{\hat\mu}^{(h)}_{\bullet}\bigr)
\right],
$$
and the final union interval is formed from the bootstrap distributions of
$$
\inf_{\|h\|\le \lambda}\hat{\hat\mu}^{(h)}_b
\quad\text{and}\quad
\sup_{\|h\|\le \lambda}\hat{\hat\mu}^{(h)}_b.
$$
A generalized minimax/maximin inequality yields
$$
L\le \inf_h L^{(h)},\qquad U\ge \sup_h U^{(h)},
$$
which supports asymptotic coverage of the partially identified region [1711.11286].

These developments clarify that MSM is not only a set of identifying restrictions. It also supports an inferential program: sharp optimization for point bounds, asymptotically valid bootstrap procedures for interval bounds, and asymptotic design criteria for comparing estimation strategies before analysis [1711.11286] [2307.00093].

## 4. Generalizations: continuous treatments and marginal structural models

MSM has been extended beyond binary static treatment. For continuous-valued interventions, the continuous marginal sensitivity model (CMSM) replaces treatment probabilities with conditional treatment densities. If $T\in\mathcal T\subset\mathbb R$ is continuous and $U$ is an unobserved confounder, CMSM assumes
$$
\Lambda^{-1}\le \frac{f(t\mid x,u)}{f(t\mid x)}\le \Lambda
\quad\forall (t,x,u),
$$
or equivalently introduces a reweighting function
$$
w(t,x,u)=\frac{f(t\mid x,u)}{f(t\mid x)}
$$
with $w(t,x,u)\in[\Lambda^{-1},\Lambda]$ and $E_U[w(t,x,U)\mid T=t,X=x]=1$ [2204.10022]. The estimand is the average dose–response curve
$$
\mu(t)=E[Y(t)],
$$
and sharp lower and upper bounds are given by
$$
\underline\mu(t)=\inf_{w\in\mathcal W}E_X[w(t,X)m(t,X)],
\qquad
\overline\mu(t)=\sup_{w\in\mathcal W}E_X[w(t,X)m(t,X)],
$$
where $m(t,x)=E[Y\mid T=t,X=x]$ [2204.10022].

These infinite-dimensional programs admit closed-form quantile-tilting solutions. Let $Q_{m\mid t}(p)$ denote the conditional quantile function of $m(T,X)$ given $T=t$, and let
$$
\alpha=\frac{\Lambda}{\Lambda^2+1}.
$$
Then
$$
\underline \mu(t)
=\int_{0}^{\alpha}Q_{m\mid t}(p)\Lambda\,dp
+\int_{\alpha}^{1}Q_{m\mid t}(p)\Lambda^{-1}\,dp,
$$
with an analogous expression for $\overline\mu(t)$ [2204.10022]. Computation proceeds by sorting predicted outcomes at each target dose and assigning weights $\Lambda$ or $\Lambda^{-1}$ to lower or upper tails, an $O(n\log n)$ procedure per dose point [2204.10022].

A separate line of work embeds MSM-style sensitivity analysis into marginal structural models. In that framework, one posits
$$
E[Y(a)]\equiv \psi(a)=g(a;\beta),
$$
permits discrete or continuous, static or time-varying treatment, and defines the propensity-sensitivity class
$$
\Pi(\gamma)=\{\tilde\pi(a\mid x,u):\gamma^{-1}\le \tilde\pi(a\mid x,u)/\pi(a\mid x)\le \gamma,\ \int \tilde\pi(a\mid x,u)\,da=1\ \forall x,u\}
$$
[2210.04681]. Equivalently one works with bias-weight functions
$$
v(z)=E[\pi(A\mid X)/\pi(A\mid X,U)\mid X,A,Y]\in[\gamma^{-1},\gamma],\qquad E[v(Z)\mid A,X]=1
$$
[2210.04681]. Bounds on the confounded outcome regression are then propagated through moment conditions for $\beta$, and the empirical problem is solved through U-statistic estimating equations [2210.04681].

This generality matters because it shows that MSM is not confined to simple ATE estimation with binary treatment. The same bounded-ratio idea can be adapted to continuous-dose causal curves, parametric MSM targets $g(a;\beta)$, and time-varying treatment regimes with stabilized weights [2204.10022] [2210.04681].

## 5. Alternative sensitivity scales and structural variants

Although the odds-ratio-based binary-treatment MSM remains the reference formulation, the supplied literature presents several alternative scales for sensitivity analysis. One variant is a risk-ratio-based marginal sensitivity model. In the binary-treatment case, Basit, Latif, and Wahed define
$$
e_0(x,y)=P(A=1\mid X=x,Y(1)=y),\qquad e_0(x)=P(A=1\mid X=x),
$$
and impose
$$
\frac{e_0(x)}{e_0(x,y)}\le \Gamma
\quad\text{and}\quad
\frac{e_0(x,y)}{e_0(x)}\le \Gamma
\quad\forall x,y.
$$
Equivalently,
$$
\sup_{y,y'}\frac{P(A=1\mid Y(1)=y,X=x)}{P(A=1\mid Y(1)=y',X=x)}\le \Gamma.
$$
The identification region for the ATE is obtained by optimizing over deviation functions $\ell_a(x,y)\in[-\log\Gamma,\log\Gamma]$, which yields a linear-fractional program reducible to a linear program by a Charnes–Cooper transformation [2309.15391]. The framework extends to multivalued treatments with generalized propensity scores $r_0(a\mid x)$ and shifted estimators $\tau^{(\ell)}(c)$ [2309.15391].

Another variant replaces the worst-case $L^\infty$ restriction by an $L^2$ bound. Zhang and Zhao retain the standard $L^\infty$-MSM class
$$
\mathcal H_\infty(\Gamma)=\{h\ge 0:\;W_-(X)\le h(X,Y)\le W_+(X),\;E[h\mid X,Z=1]=1\},
$$
with
$$
W_-(x)=(1-1/\Gamma)e(x)+1/\Gamma,\qquad
W_+(x)=(1-\Gamma)e(x)+\Gamma,
$$
and then propose an $L^2$-MSM that restricts the average squared deviation of the propensity score ratio:
$$
E[w(X,U)^2\mid Z=1]\le \gamma^2
$$
or equivalently
$$
E[h(X,Y)^2\mid Z=1]\le M^2,\qquad E[h\mid X,1]=1,\qquad h\ge e(X).
$$
The resulting sharp bounds admit closed-form solutions via Lagrangian decomposition with
$$
h_*(X,Y)=e(X)+\lambda g(X,Y),
$$
where
$$
g(X,Y)=(\xi_X-Y)\,1\{Y\le \xi_X\}
$$
and $\xi_X$ satisfies a quantile-like equation [2211.04697].

These alternatives reveal an important conceptual point. MSM is not a single immutable model but a family of sensitivity classes centered on a common principle: bounding the discrepancy between observed and full-data treatment assignment, or equivalent weighting functions, on a specified scale [2211.04697] [2309.15391]. A plausible implication is that model choice should be driven by the intended interpretability of the sensitivity parameter and by whether worst-case or average-strength confounding is the target of calibration.

## 6. Enhanced and distributionally enhanced models

A prominent recent direction strengthens MSM by constraining not only treatment assignment but also the way unmeasured confounders alter outcomes. The enhanced marginal sensitivity model (eMSM) adds an outcome-sensitivity bound alongside the conventional treatment-sensitivity bound [2504.08301]. In addition to the MSM assumptions, eMSM imposes for $a\in\{0,1\}$
$$
\Gamma^{-1}\le \frac{f(y\mid A=a,X,U)}{f(y\mid A=a,X)}\le \Gamma
\quad
(\text{or equivalently on means, if one prefers})\quad
\Gamma^{-1}\le \frac{E[Y\mid A=a,X,U]}{E[Y\mid A=a,X]}\le \Gamma.
$$
The sharp upper bound for the missing treated potential outcome becomes
$$
\nu_e^+(x)
=
E[Y\mid A=1,x]
+
(\Lambda-1/\Lambda)\,\psi_+(x)\,
E\!\left[\rho_{\tau(x)}(Y,q^*_{\tau(x)}(x))\mid A=1,x\right],
$$
where
$$
\psi_+(x)=\min\left\{1,\;\frac{\Gamma-1}{\Gamma+1}\,\frac{1/\Lambda}{\Lambda-1/\Lambda}\right\}.
$$
The resulting bounds satisfy
$$
\mu_{\rm MSM}^{1-}\le \mu_e^{1-}\le \mu_e^{1+}\le \mu_{\rm MSM}^{1+},
$$
and similarly for $E[Y^0]$, so eMSM always yields no-wider sensitivity intervals than MSM [2504.08301].

The distributionally enhanced marginal sensitivity model (deMSM) sharpens this idea by adding a density-ratio constraint directly on the full potential-outcome distribution [2505.13868]. For $t\in\{0,1\}$, it imposes outcome-sensitivity functions $T_1(x)\le 1\le T_2(x)$ such that
$$
w_t(y,x,u)
:=
\frac{dP(Y(t)=y\mid X=x,U=u)}
{dP(Y(t)=y\mid X=x,T=t)}
\in [1/T_1(x),T_2(x)]
\quad \text{for all }y,x,u.
$$
Thus deMSM combines: the latent-variable extension $(Y(0),Y(1))\perp T\mid (X,U)$, the treatment-sensitivity bound with $(A_1(x),A_2(x))=(1/\Gamma,\Gamma)$, and the outcome-sensitivity bound above [2505.13868].

Under deMSM, the sharp bounds for
$$
v_1(x)=E[Y(1)\mid T=0,X=x]
$$
admit closed-form quantile-type representations. The upper bound is summarized as
\begin{align*}
v_{\mathrm{deMSM}}^+(x)
&=E[Y\mid T=1,X=x] \\
&\quad +(\Gamma-1)(T_2-T_1)\cdot
\min\left\{
T\,E[(Y-q_+)_+\mid T=1,X],
(1-T)\,E[(q_+-Y)_+\mid T=1,X]
\right\},
\end{align*}
with $q_+$ chosen as the appropriate quantile of $Y\mid T=1,X=x$, and
$$
p_1^+=E[T\,Y+(1-T)v_{\mathrm{deMSM}}^+(X)],
\qquad
p_1^-=E[T\,Y+(1-T)v_{\mathrm{deMSM}}^-(X)].
$$
A key feature is symmetry: $v_{\mathrm{deMSM}}^+(x)$ depends only on the product $(\Gamma-1)(T_2-T_1)$ and a quantile-level function that treats treatment and outcome sensitivity symmetrically. Exchanging the treatment-sensitivity pair $(A_1,A_2)=(1/\Gamma,\Gamma)$ with the outcome pair $(T_1,T_2)$ leaves the bound unchanged [2505.13868].

The paper further notes a practical symmetric choice $\Gamma=T\ge 1$, under which the quantile level is $T/(1+T)$ and the total allowed shift is $(T-1)^2/T$ [2505.13868]. It also states that the same extremal joint distribution $Q$ attains both the upper bound for $p_1$ and the lower bound for $p_0$, so the ATE bounds are simultaneously sharp [2505.13868].

deMSM is compared directly with MSM, MSM-U, and eMSM. MSM and MSM-U use only the treatment-side constraint; eMSM adds a mean-shift constraint on $E[Y(1)\mid X,U]$ versus $E[Y\mid T=1,X]$ with difference-scale parameters $\Delta_1,\Delta_2$; deMSM replaces that by the density-ratio constraint and yields bounds that are always at least as tight as MSM’s and eMSM’s [2505.13868]. The new interpretation proposed there is that outcome sensitivity should live on the same scale as treatment sensitivity, namely a density-ratio scale [2505.13868]. This suggests a unifying sensitivity axis rather than separate ad hoc calibrations.

## 7. Interpretation, relation to other sensitivity frameworks, and empirical use

MSM is closely related to, but distinct from, Rosenbaum’s matched-study sensitivity model. Zhao, Small, and Bhattacharya show
$$
\mathcal E(\sqrt\Gamma)\subset \mathcal R(\Gamma)
\quad\text{and}\quad
\mathcal R(\Gamma)\cap \mathcal C\subset \mathcal E(\Gamma)\cap \mathcal C,
$$
where $\mathcal R(\Gamma)$ is Rosenbaum’s class and $\mathcal C$ is a compatibility condition [1711.11286]. Unlike Rosenbaum’s model, the marginal model does not require exact matching nor constant treatment effects, and it applies to missing-data inverse-probability weighting as well as observational-treatment settings [1711.11286]. This distinction is often consequential in practice: MSM operates on the relationship between full-data and marginal propensities, not on matched-pair treatment odds.

Applied examples illustrate how MSM-based procedures are used to study robustness rather than to deliver a single corrected estimate. In the fish-consumption and blood-mercury application, percentile-bootstrap intervals were computed for ATE and ATT under $\Lambda=e^0=1$, $e^{0.5}\approx 1.65$, $e^1\approx 2.72$, $e^2\approx 7.39$, and $e^3\approx 20.1$, and the 90% CI of the ATE excluded zero even at $\Lambda\approx 2.7$ [1711.11286]. In the weighted-study reanalysis of drivers of support for the 2016 Colombian peace agreement, MSM-based design sensitivity was used to compare standard IPW, trimmed IPW, and doubly augmented estimators, with augmentation improving robustness more consistently than trimming [2307.00093]. In the risk-ratio-based framework, maternal education and female fertility in Bangladesh were analyzed over $\Gamma\in\{1,1.25,1.5,1.75,2,2.25,2.5,2.75\}$, producing widening point-estimate intervals and bootstrap confidence intervals as sensitivity increased [2309.15391].

In more elaborate causal models, the same logic carries into continuous or time-varying exposures. The continuous-treatment CMSM was motivated by aerosol optical depth and cloud optical thickness, where domain expertise was used to choose $\Lambda$ in the range $[1,2]$ and derive uncertainty bands for the average dose–response curve under hidden meteorological confounding [2204.10022]. Sensitivity analysis for marginal structural models was illustrated with mothers’ smoking and birthweight, and with mobility and Covid-19 deaths, in each case tracing how the estimated structural parameter changes over $\gamma$ [2210.04681].

A recurring controversy concerns conservatism. The $L^\infty$-type marginal sensitivity model protects against the worst-case hidden-confounding configuration and may therefore produce wide bounds [2211.04697]. The $L^2$ alternative was proposed precisely because the worst-case bound can be overly conservative when extreme confounding events are rare [2211.04697]. Another controversy concerns calibration: the sensitivity parameter is interpretable, but not directly estimable from observed data. The papers therefore emphasize comparison to omitted measured covariates, design-sensitivity thresholds, or symmetric parameterizations such as $\Gamma=T$ in deMSM [2307.00093] [2211.04697] [2505.13868].

Taken together, these developments position MSM as a foundational sensitivity-analysis paradigm in modern causal inference. Its main contribution is not identification under weaker assumptions in the usual sense, but disciplined partial identification under transparent, quantitative restrictions on hidden confounding. Later variants—CMSM, risk-ratio MSM, $L^2$-MSM, eMSM, and deMSM—retain this core logic while altering the sensitivity scale, the target causal structure, or the balance between treatment-side and outcome-side restrictions [2204.10022] [2309.15391] [2211.04697] [2504.08301] [2505.13868].

Source: https://www.emergentmind.com/topics/marginal-sensitivity-model