Optimal Treatment Regime: Foundations & Methods
- Optimal Treatment Regime (OTR) is a decision rule that uses patient covariates to assign treatments aimed at maximizing expected outcomes.
- It encompasses static and dynamic settings, varying value functionals, and diverse estimation paradigms, balancing flexibility with interpretability.
- Methodological challenges include nonregular inference, identification issues, and trade-offs between rule complexity and clinical usability.
An optimal treatment regime (OTR)—also called an optimal treatment rule, individualized treatment rule, or policy—is a decision rule that assigns treatment on the basis of covariates so as to maximize a population-level criterion, typically an expected outcome. In the standard static binary-treatment setup, with covariates , treatment , outcome , and potential outcomes , a regime induces and value ; the OTR is any . Contemporary work extends this formulation to dynamic regimes, alternative value functionals such as conditional medians or survival probabilities, observational settings with nuisance estimation, and nonregular inference problems in which the optimal rule itself may be difficult to estimate or summarize (Cheng et al., 2024, Williams et al., 2024).
1. Formal definition and causal structure
In the static potential-outcomes formulation, the observed data are i.i.d. copies of , where , 0, and 1. A regime is a mapping
2
and its value is
3
Under unconfoundedness or ignorability,
4
consistency,
5
and positivity,
6
the value function is identified from observed data, and optimality is defined by maximizing 7 over a class of admissible rules (Cheng et al., 2024).
The same logic extends to dynamic settings. For a finite horizon 8, a dynamic treatment rule at stage 9 is a map
0
which conditions on realized histories of outcomes and treatments. A dynamic regime is the sequence 1, and its welfare is a linear functional of the induced joint distribution of counterfactual outcomes, with the main case
2
The optimal dynamic treatment regime is
3
which is the direct dynamic analogue of the static OTR criterion (Han, 2019).
This formalism fixes the central object of OTR analysis: a policy-valued estimand defined through counterfactual welfare. The main technical differences across the literature arise from the choice of value functional, the admissible regime class, and the identification conditions used to recover that value from observed data.
2. Objective functions and classes of regimes
The literature represented here does not treat optimality as synonymous with mean-outcome maximization. It instead develops several value criteria, each inducing a different notion of optimality.
| Setting | Value or objective | Representative source |
|---|---|---|
| Static mean-optimal regime | 4 | (Leqi et al., 2021) |
| Linear mean-optimal regime | 5, 6 | (Cheng et al., 2024) |
| Conditional-median regime | 7 | (Leqi et al., 2021) |
| 8-year survival regime | 9 | (Jiang et al., 2014) |
| Treatment-initiation regime | 0 with RMRL at 1 | (Chen et al., 2021) |
| Dynamic welfare | 2 | (Han, 2019) |
| Indefinite-horizon survival regime | 3 | (She et al., 30 Jan 2025) |
A major structured subclass is the linear regime
4
usually indexed over a compact or normalized parameter space such as
5
This class is attractive because signs and magnitudes of 6 directly encode which covariates push treatment toward 7 or 8, and because the rule can be communicated as a weighted score. The resulting trade-off is explicit: nonlinear classes can approximate the truly optimal rule more closely, but linear rules improve interpretability, bedside usability, and inferential tractability (Cheng et al., 2024).
Alternative criteria materially change the regime. In the median-based formulation, the rule is
9
where 0 is the conditional median under treatment 1, and the criterion is the Average Conditional Median Effect
2
This criterion is within-group and median-based, rather than mean-based or based on the marginal median of 3 (Leqi et al., 2021).
Survival settings also diversify the objective. One strand maximizes 4-year survival probability 5 over a linear class of rules (Jiang et al., 2014). Another treats treatment timing itself as the decision and evaluates regimes through restricted mean residual lifetime at a prespecified 6,
7
with value 8 (Chen et al., 2021). In dynamic IV-based settings, optimality may become set-valued: when welfare is only partially identified, the identified set of optimal regimes consists of the maximal elements in a sharp partial welfare order rather than a unique rule (Han, 2019).
These formulations imply that OTR is not a single estimator or algorithm but a family of decision problems. What changes across papers is the welfare functional being optimized, the structural assumptions under which it is identified, and the extent to which interpretability is imposed ex ante through the regime class.
3. Estimation paradigms
The broader OTR literature includes regression-based methods such as Q-learning and A-learning, direct policy-search methods based on IPW or AIPW value estimation, TMLE-based policy learning, outcome-weighted learning, and modern RL/OPE/OFL approaches (Cheng et al., 2024). Across these paradigms, the core estimation problem is to recover either the rule itself, the value of a candidate rule, or both.
For linear static regimes in observational data, a central direct-value strategy uses the augmented inverse probability weighted representation. With propensity score 9, outcome regressions 0, and
1
the value of 2 can be written as
3
where
4
Replacing 5 by estimators yields
6
and the optimal linear regime is estimated by
7
In the cited implementation, the optimization over the unit sphere is carried out by a genetic algorithm via rgenoud (Cheng et al., 2024).
Median-optimal regimes replace conditional means by conditional medians and estimate the conditional median treatment effect
8
The learned rule is
9
The associated ACME estimator is influence-function based and doubly robust-style: 0 while 1 is obtained by regressing a pseudo-outcome 2 on 3 through a linear smoother (Leqi et al., 2021).
When unmeasured confounding precludes standard ignorability, proximal causal learning replaces 4 by proxy-based identification. With 5, where 6 and 7 are outcome-inducing and treatment-inducing confounding proxies, the value of a 8-based regime can be identified through an outcome confounding bridge 9,
0
and the value of a 1-based regime through a treatment confounding bridge 2,
3
The cited paper further defines an expanded class
4
and estimates bridge functions with neural moment methods, followed by weighted classification steps for 5, 6, and a learned predilection function 7 (Shen et al., 2022).
Longitudinal EHR settings create an additional nuisance process: visit times. In that case, single-stage individualized treatment rules can be estimated via doubly weighted dWOLS using both an inverse probability of treatment weight and an inverse intensity of visits weight,
8
within weighted estimating equations for a blip model
9
The resulting rule is
0
and the method is designed to break the collider-stratification bias induced by covariate-driven observation times (Coulombe et al., 2022).
For genuinely dynamic regimes, one prominent modern strategy is backward recursion with doubly robust pseudo-outcomes. At a single time point, the DR transformation
1
satisfies
2
where 3 is the blip. Repeating this backwards over time yields a sequence of estimated rules
4
thereby learning an ODTR by dynamic programming with DR pseudo-outcomes rather than via a fully parametric Q-function (Williams et al., 2024).
Across these approaches, a common architecture emerges: identify a policy value; orthogonalize or augment it against nuisance estimation; restrict or regularize the policy class; and then optimize the estimated value. The divergence across papers lies mainly in how identification is achieved and how much structure is imposed on the decision rule.
4. Inference, nonregularity, and asymptotic phenomena
Inference for OTRs is not a routine by-product of policy estimation. The central obstacle is that the rule itself is often defined by a non-smooth sign or indicator operation, and the difficulty is more acute when the optimal regime is nonunique.
For optimal linear regimes estimated by AIPW maximization, the value estimator remains regular: 5 but the regime parameter does not. Under additional smoothness, density, and curvature assumptions,
6
and
7
that is, the argmax of a mean-zero Gaussian process with a quadratic drift. The nonregularity comes from the indicator 8, which creates a kink in the objective. Standard nonparametric bootstrap is therefore invalid for 9, and the paper adapts the Cattaneo–Jansson–Nagasawa bootstrap by centering and drift-adjusting the objective with an estimated Hessian 0 (Cheng et al., 2024).
A distinct but related nonregularity appears in inference for the mean outcome under the OTR itself. When
1
the regularity assumption that the OTR is unique fails, classical inference that treats the estimated OTR as if it were the true OTR can be biased, and pathwise differentiability of the target can fail except in a degenerate deterministic-outcome case (Xu et al., 11 Sep 2025). The proposed remedy is adaptive smoothing of the estimated rule: 2 with 3 chosen adaptively as the propensity score 4. The resulting estimator is valid regardless of whether regularity holds, achieves asymptotic normality, and attains a derived lower bound on the asymptotic variance for the class of robust asymptotically linear unbiased estimators. In that sense, the procedure is efficient not only in the classical regular case but also in the explicitly nonregular setting (Xu et al., 11 Sep 2025).
Survival-timing OTRs generate yet another asymptotic pattern. When the decision variable is a continuous initiation time and the value is a kernel-based estimator of restricted mean residual lifetime, the regime parameter is asymptotically normal at rate
5
whereas the optimized value has a nonstandard limit: 6 The weighted chi-squared limit arises because continuous treatment timing requires kernel smoothing in both 7 and 8, and the optimized value is dominated by quadratic fluctuation of the estimated parameter (Chen et al., 2021).
A recurrent misconception is that value estimation is always easier than regime estimation. The recent literature makes a sharper distinction: sometimes the value remains regular while the rule is nonregular; sometimes the rule can be estimated consistently but the optimized value has a nonstandard law; and when uniqueness of the OTR fails, even inference on the value can require specialized smoothing or resampling. This separation between inference on the policy and inference on the value is now a defining technical feature of OTR methodology.
5. Extensions beyond the standard static binary-treatment model
A substantial part of the modern OTR literature is devoted to departures from the canonical single-population, single-stage, no-unmeasured-confounding setting.
When sequential ignorability of treatment fails in dynamic settings, IV-based methods replace point identification of welfare by partial identification. Under a sequential IV assumption,
9
welfare differences between dynamic regimes can be bounded by linear programs over a latent-state distribution 00. The identified set of optimal regimes is then
01
the set of maximal elements of a sharp partial welfare order, rather than a single regime (Han, 2019).
When observation times are covariate-driven, as in EHR data, standard DTR assumptions fail because conditioning on visit occurrence opens collider paths. The proposed remedy is a repeated-measures ITR estimated by dWOLS with both IPT and inverse-intensity-of-visits weighting, under assumptions such as
02
This explicitly adjusts the observation process jointly with the treatment process (Coulombe et al., 2022).
When measured covariates are insufficient to remove confounding, proximal causal learning uses confounding proxies 03 and 04 together with bridge functions 05 and 06 to identify regime values and construct optimal regimes in classes indexed by 07, 08, or by adaptive combinations of both: 09 The resulting proximal regime dominates earlier proximal policies that globally selected either the 10-based or 11-based policy class (Shen et al., 2022).
Transportability problems arise when the source and target populations differ in covariate distribution and only target summary statistics are available. In that setting, calibrated AIPW estimators replace the usual source-sample average by weighted averages over source units, with calibration constraints
12
The learned regime
13
targets a pseudo-population that can coincide with the target population when the calibration weights recover the density ratio 14 (Chu et al., 2022).
Survival and longitudinal RL formulations further broaden the field. Indefinite-horizon DTRs for censored survival outcomes can be estimated by generalized survival random forests that repeatedly apply a single rule
15
over pooled visits, while optimizing truncated mean survival
16
The method uses summarized histories of fixed dimension, backward survival-curve augmentation, and iterative refitting, and it allows patients to have different numbers of decision points (She et al., 30 Jan 2025). Continuous-time timing decisions produce another generalization: treatment initiation itself becomes the action 17, and the regime is chosen to maximize expected restricted mean residual lifetime rather than a binary treatment contrast (Chen et al., 2021).
These extensions show that the notion of an OTR survives substantial relaxation of the standard setup, but the meaning of “optimal” often changes with identification strength. Under ignorability it is typically a maximizer of an identified value; under IV assumptions it may be an identified set; under transportability it may be optimal in a calibrated pseudo-population; and under nonregularity its value may require specialized inference even when the rule class is simple.
6. Interpretation, applications, and methodological tensions
A major motivation for restricting OTRs to interpretable classes is substantive interpretability. In the linear-rule setting,
18
the sign of 19 indicates whether larger 20 pushes toward treatment 21 or control 22, and the magnitude 23 reflects relative importance for the decision boundary. This is why the 2024 linear-regime paper places particular emphasis on inference for the regime parameter itself rather than only on the value (Cheng et al., 2024).
In the eICU vasopressor application, that interpretive emphasis yields clinically legible findings. With 24, admission temperature had a positive coefficient with 95% CI excluding zero, reported as Temp: estimate 25, CI 26; with 27, WBC had a positive coefficient with 95% CI excluding zero, reported as WBC: estimate 28, CI 29. The interpretation offered is that fever and leukocytosis tilt the optimal decision rule toward vasopressor use when optimizing fluid balance (Cheng et al., 2024).
Other applications illustrate how different value criteria change substantive recommendations. In the ACTG 175 HIV trial, the median-optimal policy had the highest estimated ACME, slightly outperforming treat-all and mean-optimal policies, while the motivating female subgroup histograms showed a large mean difference but almost no median difference between treated and control outcomes, exactly the situation in which median-based OTRs are intended to differ from mean-based rules (Leqi et al., 2021). In the oropharynx cancer study, a Bayesian loss on the bivariate potential outcomes was used to avoid unnecessary chemotherapy burden. Under the OTR.25 loss, the proposed regime reduced the frequency of CRT assignment by approximately 75% without reducing the average survival probability, thereby changing the rule not by improving survival per se but by reweighting survival against treatment burden (Klausch et al., 2018).
Dynamic applications make the same point in sequential settings. In the buprenorphine–naloxone analysis, the learned ODTR outperformed a clinically defined dose-escalation strategy when evaluated by the week-6 risk of return-to-regular-opioid-use, illustrating how DR backward recursion can yield a clinically interpretable dynamic rule that differs from standard practice (Williams et al., 2024). In the antidepressant BMI study using CPRD, the doubly weighted dWOLS estimator identified effect modification not detected by simpler estimators, and applying the estimated rule yielded about a 30–31 point increase in the BMI-related utility on a roughly 100-point scale (Coulombe et al., 2022). In pediatric Crohn’s disease, the indefinite-horizon survival-forest regime improved restricted mean hospital-free survival relative to observed practice, with the two-strata implementation increasing mean survival by about 30 days (She et al., 30 Jan 2025).
The principal methodological tension throughout the literature is between flexibility and interpretability. Flexible rules—trees, RKHS-based learners, deep nets, or pooled forest-based DTRs—can achieve better approximation of the optimal value, but often at the cost of harder inference, greater sensitivity to nuisance estimation, and weaker clinical transparency. Structured rules—linear scores, decision lists, or single-index blips—sacrifice approximation flexibility in exchange for direct substantive interpretation and, in some settings, tractable asymptotics. A second tension concerns identification: stronger assumptions such as sequential ignorability yield point-identified optimal policies, whereas weaker assumptions such as IV or proximal identification can produce only partial orders, pseudo-population optimality, or sensitivity-dependent recommendations. The contemporary OTR literature can therefore be read as a sequence of answers to the same question—what treatment rule is best?—posed under progressively different assumptions about outcomes, observability, transportability, and inferential regularity.