---
title: First-Best Optimal Policy Learning Explained
url: https://www.emergentmind.com/topics/first-best-optimal-policy-learning-opl
type: topic
---

# First-Best Optimal Policy Learning Explained

First-Best Optimal Policy Learning (OPL) is the problem of selecting a policy that maximizes a welfare criterion over a policy class, often the unrestricted class of measurable maps from observed covariates or states to actions. In the potential-outcome treatment setting, the first-best or unconstrained optimal policy assigns each covariate profile to the treatment with the larger conditional mean outcome; in multi-action treatment settings, the same optimization can be posed under risk neutrality or under risk-averse criteria that penalize conditional variance or second moments [2011.04993] [2509.06851]. Across contextual bandits, constrained allocation, censored survival analysis, episodic reinforcement learning, and partially identified structural models, the term denotes the welfare-maximizing benchmark relative to the admissible policy class and feasibility constraints, not a single algorithmic template [2605.12235] [2603.22900] [2112.10935] [2012.11046].

## 1. Formal welfare criteria and oracle characterizations

In the multi-action treatment formulation, observed covariates are denoted by $X\in\mathcal X\subseteq\mathbb R^p$, treatment by $A\in\mathcal A=\{0,1,\dots,M-1\}$, and observed outcome by $Y$. Using potential outcomes $Y(a)$, the maintained identifying conditions are conditional independence,
$$
\{Y(a):a\in\mathcal A\}\perp A\mid X,
$$
and overlap,
$$
e(a\mid x)=P[A=a\mid X=x]>0
$$
for all $a\in\mathcal A$ and $x$ in the support of $X$. A deterministic policy is a mapping $\pi:\mathcal X\to\mathcal A$, and first-best OPL chooses $\pi$ from the unrestricted class $\Pi$ of all measurable functions $X\mapsto A$ [2509.06851].

Under risk neutrality, the objective is
$$
\pi^*=\arg\max_{\pi\in\Pi}V(\pi),
\qquad
V(\pi)=E\bigl[Y\bigl(\pi(X)\bigr)\bigr].
$$
The same framework permits risk-sensitive criteria. With conditional mean $\mu(x,a)=E[Y\mid X=x,A=a]$ and conditional variance $\sigma^2(x,a)=\mathrm{Var}[Y\mid X=x,A=a]$, a linear risk-averse decision maker with aversion parameter $\theta\ge 0$ uses
$$
U_{\mathrm{lin}}(Y(a))=\mu(x,a)-\theta\,\sigma^2(x,a),
$$
and the induced first-best policy solves
$$
\pi^*_{\mathrm{lin}}
=\arg\max_{\pi\in\Pi}
E\bigl[\mu\bigl(X,\pi(X)\bigr)-\theta\,\sigma^2\bigl(X,\pi(X)\bigr)\bigr].
$$
A quadratic risk-averse criterion instead penalizes the second moment,
$$
U_{\mathrm{quad}}(Y(a))=E[Y(a)]-\theta\,E[Y(a)^2],
$$
so that
$$
\pi^*_{\mathrm{quad}}
=\arg\max_{\pi\in\Pi}
E\bigl[\mu\bigl(X,\pi(X)\bigr)-\theta\,E\bigl[Y(\pi(X))^2\mid X\bigr]\bigr],
$$
with
$$
E[Y^2\mid X=x,a]=\sigma^2(x,a)+\mu(x,a)^2
$$
in practice [2509.06851].

In the binary-treatment case, let $\mu_t(x)=E[Y(t)\mid X=x]$ for $t\in\{0,1\}$ and $\tau(x)=\mu_1(x)-\mu_0(x)$. The oracle first-best policy is
$$
\pi^*(x)\in\arg\max_{t\in\{0,1\}}\mu_t(x),
\qquad
\pi^*(x)=1\{\tau(x)>0\},
$$
and the maximized welfare is
$$
W^*=W(\pi^*)=E[\max\{\mu_1(X),\mu_0(X)\}].
$$
This is the omniscient benchmark obtained when the policymaker knows the conditional mean outcomes [2011.04993].

| Setting | First-best criterion | Source |
|---|---|---|
| Multi-action treatment | $\arg\max_{\pi\in\Pi}E[Y(\pi(X))]$ or risk-averse variants | [2509.06851] |
| Binary treatment | $\pi^*(x)=1\{\tau(x)>0\}$ | [2011.04993] |
| Budget and coverage | $\max_x \sum_i w_i x_i$ subject to budget and coverage | [2605.12235] |
| Censored survival | maximize $V^\tau(\pi)$ subject to $C(\pi)\le B$ | [2603.22900] |
| Episodic RL | policy minimizing $V_h^\pi(s)$ for all $(h,s)$ | [2112.10935] |
| Partial identification | $\arg\max_{\pi\in\Pi}\inf_{s\in\mathcal S}\tau(\pi,s)$ | [2012.11046] |

This taxonomy indicates that “first-best” is defined relative to the feasible policy space and welfare functional. In unrestricted treatment OPL it is the measurable-policy optimum; in constrained or robust settings it is the optimum after imposing budget, coverage, censoring, or ambiguity structure.

## 2. Value-function estimation and implementation in multi-action treatment OPL

With i.i.d. observations $\{(X_i,A_i,Y_i)\}_{i=1}^n$, first-best OPL in the Stata implementation `opl_ma_fb` proceeds by estimating the conditional mean and, when needed, the conditional variance. The algorithm fits a linear model of $Y$ on $X$ and $A\times X$ interactions, or separate regressions by action, to obtain $\hat\mu(x,a)$; computes squared residuals and regresses them on the same design to obtain $\hat\sigma^2(x,a)$; forms an individual-level predicted welfare score $\hat W_i(a)$; and assigns
$$
\hat\pi^*(X_i)=\arg\max_{a\in\mathcal A}\hat W_i(a),
$$
saving the result by default as `_opt_policy` [2509.06851].

The predicted welfare score depends on the chosen model:
$$
\hat W_i(a)=
\begin{cases}
\hat\mu(X_i,a), & \text{risk-neutral},\\[4pt]
\hat\mu(X_i,a)-\theta\,\hat\sigma^2(X_i,a), & \text{linear risk-averse},\\[4pt]
\hat\mu(X_i,a)-\theta\bigl[\hat\sigma^2(X_i,a)+\hat\mu(X_i,a)^2\bigr], & \text{quadratic risk-averse}.
\end{cases}
$$
The implementation can also compare $\hat\pi^*$ to a benchmark policy, save match indicators, and produce summary frequencies. It then computes and reports the estimated value-function under $\hat\pi^*$ using the regression-adjustment formula, reported as `e(V_opt_train)` and `e(V_opt_new)` [2509.06851].

The companion command `opl_ma_vf` evaluates the maximal welfare after policy construction:
```stata
. opl_ma_vf depvar varlist, policy_train(A) policy_new(_opt_policy)
```
It returns three estimators of $V(\hat\pi^*)=E[Y(\hat\pi^*(X))]$: regression adjustment (RA), inverse-probability weighting (IPW), and doubly robust (DR), stored in `e(RA)`, `e(IPW)`, and `e(DR)`. The paper gives the example
“RA = 124.19, IPW = 109.50, DR = 125.31” [2509.06851].

The RA and DR estimators are
$$
\hat V_{RA}(\pi)
=\frac{1}{n}\sum_{i=1}^n
\sum_{a\in\mathcal A}
\hat\mu(X_i,a)\,
1\{\pi(X_i)=a\},
$$
and
$$
\hat V_{DR}(\pi)
=\frac{1}{n}\sum_{i=1}^n
\Biggl[
\sum_{a\in\mathcal A}\hat\mu(X_i,a)\,1\{\pi(X_i)=a\}
+\frac{\bigl(Y_i-\hat\mu(X_i,A_i)\bigr)\,1\{\pi(X_i)=A_i\}}
{\hat e(A_i\mid X_i)}
\Biggr].
$$
The DR estimator remains consistent if either $\hat\mu$ or $\hat e$ is correctly specified [2509.06851].

The same implementation includes graphical output. `opl_ma_fb` can produce `gr_action_train`, a bar-chart comparing the frequency of actual vs. optimal actions in the training sample; `gr_reward_train`, a scatter or line-plot of observed vs. maximal expected reward; and `gr_reward_new`, a plot of predicted reward $\hat\mu(X_i,\hat\pi^*(X_i))$ across the new sample. Alternatively, `opl_plot_best` displays side-by-side a scatter of “observed” vs. “maximal” predicted rewards and a categorical plot of “observed treatment” vs. “optimal treatment” for each $i$ [2509.06851].

## 3. Restricted policy classes and empirical welfare maximization

Although first-best OPL is defined over unrestricted measurable policies, applied work often imposes a lower-dimensional structure. A prominent restriction is the threshold policy
$$
\pi_\tau(x)=1\{f(x)\ge \tau\},
$$
where $f:\mathbb R^d\to\mathbb R$ is a one-dimensional score such as an estimated conditional treatment effect $\hat\tau(x)$, a linear index, or a nonparametric score derived from domain knowledge [2011.04993].

The threshold-based formulation separates specification from policy selection. First, one estimates $f(x)$, often using regression adjustment or machine-learning methods such as causal forest, random forest, or boosted trees. Second, one chooses the scalar threshold $\tau$ that maximizes welfare within the restricted class. With estimated outcome models $\hat\mu_t(x)$, the empirical welfare for threshold $\tau$ is
$$
\hat W_n(\tau)
=
\frac{1}{n}\sum_{i=1}^n
\Bigl[
\hat\mu_0(X_i)
+
\bigl(\hat\mu_1(X_i)-\hat\mu_0(X_i)\bigr)\,1\{f(X_i)\ge\tau\}
\Bigr].
$$
The unrestricted first-best threshold is
$$
\tau^*\in\arg\max_{\tau\in\mathbb R}W(\tau),
$$
and the resulting policy $\pi_{\tau^*}$ converges to $\pi^*$ as $f$ tracks the true $\tau(x)$ [2011.04993].

The empirical implementation protocol is a five-step “empirical welfare-maximizing” procedure: input $(Y_i,T_i,X_i)$; estimate $\hat\mu_1(x)$ and $\hat\mu_0(x)$ and compute $\hat f(x)=\hat\mu_1(x)-\hat\mu_0(x)$; compute the unconstrained first-best welfare
$$
\hat W^*=\frac{1}{n}\sum_{i=1}^n \max\{\hat\mu_0(X_i),\hat\mu_1(X_i)\};
$$
evaluate $\hat W_n(\tau_k)$ on a fine grid; and select
$$
\hat\tau=\arg\max_k \hat W_n(\tau_k).
$$
To obtain valid out-of-sample welfare estimates and confidence intervals, the procedure recommends sample splitting into two folds, cross-fit welfare evaluation, and bootstrap over the entire procedure, including sample splitting [2011.04993].

Under unconfoundedness, overlap, bounded outcomes, uniform consistency of $\hat\mu_t$ and $\hat f$, and Lipschitz continuity of $W(\tau)$, the paper states a uniform regret consistency result:
$$
\sup_{\tau\in T}\bigl|\hat W_n(\tau)-W(\tau)\bigr|=O_p(n^{-1/2}),
$$
and any sequence $\hat\tau_n\in\arg\max_\tau \hat W_n(\tau)$ satisfies
$$
W(\tau^*)-W(\hat\tau_n)=O_p(n^{-1/2}).
$$
The formal statement is framed as a proposition using Assumptions A1–A3: unconfoundedness, overlap, bounded $Y$; uniform consistency of $\hat\mu_t(x)$ and $\hat f(x)$; and Lipschitz continuity of $W(\tau)$ [2011.04993].

This restricted-policy literature clarifies a central distinction within first-best OPL: the oracle target is the unrestricted welfare maximizer, whereas operational policies may be threshold rules chosen because they are easy to interpret, easy to implement, or compatible with treatment-share constraints.

## 4. Constraint-driven and outcome-specific extensions

A major extension studies first-best OPL under explicit resource constraints. In the budget-and-coverage setting, the planner chooses binary decisions $x_i\in\{0,1\}$, where $x_i=1$ if unit $i$ is treated, with expected gain $w_i$, cost $c_i>0$, total budget $B>0$, and minimum coverage requirement $C$. The optimization problem is
$$
\max_{x\in\{0,1\}^n}\sum_{i=1}^n w_i x_i
\quad\text{s.t.}\quad
\sum_{i=1}^n c_i x_i\le B,
\qquad
\sum_{i=1}^n x_i\ge C.
$$
The paper shows that this problem admits a knapsack-type structure and that the optimal policy can be characterized by an affine threshold rule involving budget and coverage shadow prices. With dual variables $\lambda\ge 0$ for the budget constraint and $\mu\ge 0$ for the coverage constraint, the LP relaxation yields
$$
x_i^*=\mathbf 1\{\,w_i-\lambda^* c_i+\mu^*\ge 0\},
$$
under a non-degeneracy assumption, equivalently treating unit $i$ iff
$$
w_i+\mu^*\ge \lambda^* c_i.
$$
The LP relaxation has at most two fractional components, and if $|w_i|\le W_{\max}$ for all $i$, then
$$
0\le \mathrm{OPT}^{\mathrm{LP}}_n-\mathrm{OPT}^{01}_n\le 2W_{\max},
$$
implying asymptotic equivalence in per-unit welfare. Two implementable algorithms are analyzed: Greedy–Lagrangian with Coverage (GLC), which is asymptotically exact in per-capita regret under boundedness and non-degeneracy, and rank-and-cut (RC), which is exact if costs are constant or the coverage multiplier is zero, but may misallocate when cost heterogeneity interacts with a binding coverage constraint [2605.12235].

A separate extension addresses right-censored survival outcomes. With logged data
$$
\mathcal D=\{(x_i,a_i,T_i,r_i)\}_{i=1}^n,
\qquad
T_i=\min\{L_i,C_i\},
\qquad
r_i=1\{L_i\le C_i\},
$$
the first-best OPL problem becomes
$$
\max_\pi V^\tau(\pi)
\quad\text{subject to}\quad
C(\pi)\le B,
$$
where
$$
V(\pi,t)=E_{x,a\sim \pi}[S(x,a,t)],
\qquad
V^\tau(\pi)=\int_0^\tau V(\pi,t)\,dt.
$$
To correct censoring bias, the framework introduces IPCW-IPS and IPCW-DR. Under independent censoring, with $\hat G(t\mid x,a)=\hat P(C>t\mid x,a)$ and importance weight $w(x_i,a_i)=\pi(a_i\mid x_i)/\pi_0(a_i\mid x_i)$,
$$
\hat V_{\rm IPCW\text{-}IPS}(\pi,t)
=
\frac1n\sum_{i=1}^n
w(x_i,a_i)\,
\frac{\mathbf 1\{T_i>t\}}{\hat G(t\mid x_i,a_i)}.
$$
With any estimator $\hat S(x,a,t)$ and residual
$$
\delta_i(t)=\frac{1\{T_i>t\}}{\hat G(t\mid x_i,a_i)}-\hat S(x_i,a_i,t),
$$
the doubly robust estimator is
$$
\hat V_{\rm IPCW\text{-}DR}(\pi,t)
=
\frac1n\sum_{i=1}^n
\Bigl[
w(x_i,a_i)\,\delta_i(t)
+
\sum_{a'}\pi(a'\mid x_i)\hat S(x_i,a',t)
\Bigr].
$$
The paper states that IPCW-IPS is unbiased if $\hat G$ is correct and $\pi_0$ is known, while IPCW-DR is unbiased if $\hat G$ is correct and either $\pi_0$ or $\hat S$ is correct; it also states a variance reduction result in which IPCW-DR variance is no larger than IPCW-IPS variance by a nonnegative gap [2603.22900].

These constrained and outcome-specific formulations preserve the first-best principle—maximize policy value over admissible policies—while altering either the feasible set or the statistical object being estimated.

## 5. Large action spaces, new actions, and optimization pathologies

In large-action contextual bandits, first-best OPL is often defined as learning a policy $\hat\pi_n$ that nearly maximizes
$$
V(\pi)=E_{x\sim\mu,\;a\sim\pi(\cdot\mid x)}[r(x,a)],
$$
relative to the oracle policy that concentrates on the action with maximal true reward at each context,
$$
\pi^*(a\mid x)=1[a=\arg\max_{a'} r(x,a')].
$$
A central recent finding is that optimization difficulties can dominate estimator quality. For OPE-based objectives such as IPS, cIPS, DR, MIPS, OffCEM, and POTEC, the paper proves that for any OPE estimator linear in $\pi$ and a standard softmax parameterization $\pi_\theta$, there exist problem instances where gradient descent is trapped in a suboptimal region for $\Theta(K)$ iterations, and that the nonconcave landscape can have exponentially many local maxima in $K$. By contrast, for weighted log-likelihood objectives
$$
\hat U_n^g(\pi)=\frac1n\sum_{i=1}^n g(r_i,\pi_0(a_i\mid x_i))\log \pi(a_i\mid x_i),
$$
with an $\ell_2$-regularized linear softmax policy, $\hat U_n^g(\pi_\theta)$ is strongly concave in $\theta$, so the loss admits a unique global minimum, no spurious stationary points, and can be optimized in $O(\mathrm{poly}(d,K)/\epsilon)$ gradient steps to $\epsilon$-optimality. Empirically, OPE-based methods are reported as highly fragile to batch size and learning-rate schedule, whereas PWLL methods remain robust across all configurations and often achieve higher validation reward and lower regret [2509.03456].

A different response to large action spaces is structural decomposition. POTEC defines a clustering map $c:A\to C$ with $|C|\ll |A|$ and factorizes the policy as
$$
\pi(a\mid x)=\sum_{c\in C}\pi^1(c\mid x;\theta)\,\pi^2(a\mid x,c;\psi).
$$
Stage 1 learns a cluster-selection policy by a low-variance policy-gradient estimator that only reweights cluster probabilities, while Stage 2 uses a regression-based within-cluster selector. Under Full-Cluster-Support and Local Correctness, the estimator is unbiased and the second-stage policy
$$
\pi^2(a\mid x,c)=1_{a=\arg\max_{b:c(b)=c}\hat h_\psi(x,b)}
$$
is optimal within each cluster. The framework strictly generalizes pure policy-based and pure regression-based OPL: if $|C|=1$, it degenerates to a pure regression-based policy; if $C=A$, it degenerates to IPS policy gradient [2402.06151].

The evolving-action setting exposes a further limitation of standard first-best OPL. Let
$$
\mathcal A_{\mathrm{existing}}=\{a\mid \exists x:\pi_0(a\mid x)>0\},
\qquad
\mathcal A_{\mathrm{new}}=\{a\mid \forall x:\pi_0(a\mid x)=0\}.
$$
Standard IPS and DR require $\pi_0(a\mid x)>0$ to produce unbiased estimates, and DM requires an action to appear in the logged data to train $\hat q(x,a)$. Hence no standard estimator can assign positive probability to any $a\in\mathcal A_{\mathrm{new}}$. To address this, LCPI uses discrete action features and a pseudo-inverse correction:
$$
G_{LCPI}(\theta)
=
\frac{1}{n}\sum_{i=1}^n
\Bigl[\sum_{a\in\mathcal A}
\pi_\theta(a\mid x_i)\,\nabla_\theta\ln\pi_\theta(a\mid x_i)\,\mathbb I_a^T\Bigr]
\Gamma_{\pi_0,x_i}^\dagger
\bigl(\mathbb I_{a_i}r_i\bigr),
$$
where
$$
\Gamma_{\pi_0,x}
=
E_{a\sim\pi_0(\cdot\mid x)}[\mathbb I_a\mathbb I_a^T].
$$
Under Local Combination Support and Local Linearity, LCPI is unbiased for the policy gradient and can extrapolate to new actions that share supported features. PONA then blends LCPI and DR:
$$
G_{PONA}(\theta;\lambda)
=
\lambda\,G_{LCPI}(\theta)
+
(1-\lambda)\,G_{DR}(\theta),
\qquad
\lambda\in[0,1],
$$
thereby trading off aggressiveness on new actions against conservatism on existing actions [2605.18509].

Taken together, these results identify a recurring misconception: better off-policy value estimation does not automatically imply better first-best policy learning. In large action spaces, trainability, decomposition, and feature-based extrapolation can become the binding considerations.

## 6. Offline reinforcement learning, uniform guarantees, and partially identified environments

In episodic offline reinforcement learning, first-best OPL is tied to uniform policy evaluation. For a finite-horizon MDP with horizon $H$, state space $\mathcal S$, action space $\mathcal A$, and offline data generated by a logging policy $\mu$, the model-based plug-in estimator forms empirical reward and transition estimates, computes estimated occupancies $\hat d_t^\pi$, and evaluates
$$
\hat v^\pi=\sum_{t,s,a}\hat d_t^\pi(s,a)\,\hat r_t(s,a).
$$
Uniform convergence means
$$
\sup_{\pi\in\Pi}|\hat v^\pi-v^\pi|\le \epsilon
$$
with high probability. Once this holds, the empirical maximizer
$$
\hat\pi^*\in\arg\max_{\pi\in\Pi}\hat v^\pi
$$
satisfies
$$
v^{\pi^*}-v^{\hat\pi^*}\le 2\sup_\pi |\hat v^\pi-v^\pi|\le 2\epsilon.
$$
For the local class
$$
\Pi_1=\{\pi:\sup_t\|\hat V_t^\pi-\hat V_t^{\hat\pi^*}\|_\infty\le \epsilon_{\mathrm{opt}}\},
$$
with $\epsilon_{\mathrm{opt}}\le O(\sqrt H/S)$, the sample complexity is
$$
n=\tilde O\bigl(H^3/(d_m\epsilon^2)\bigr),
$$
and the paper states that this local-class result is rate-optimal up to logarithmic factors [2007.03760].

Online policy optimization in episodic tabular MDPs yields a related first-best notion. In that setting, the optimal policy $\pi^*$ simultaneously minimizes $V_h^\pi(s)$ for all $(h,s)$, and learning quality is measured by regret
$$
\mathrm{Reg}(K)=\sum_{k=1}^K\bigl[V_1^*(s_1)-V_1^{\pi_k}(s_1)\bigr].
$$
RPO-SAT combines optimistic $Q$-estimates, an $\ell_2$-OMD policy-improvement step with decaying step size, and a reference-value mechanism satisfying the “Stable at Any Time” property. With high probability, it achieves
$$
\mathrm{Reg}(K)\le \tilde O\bigl(\sqrt{SAH^3K}+\sqrt{AH^4K}\bigr).
$$
When $S>H$, this matches the information-theoretic lower bound up to logarithmic factors, making it nearly minimax optimal in that regime [2112.10935].

First-best OPL becomes still broader in incomplete or partially identified models. Russell’s framework defines a policy transform
$$
\tau(\pi,P)=\int \varphi\!\bigl(V_\pi(v)\bigr)\,dP_{V_\pi}(v),
$$
where $V_\pi=(Y_\pi^\star,Y,Z,U)$ and $\varphi$ is a known bounded function. If the data are compatible with a family $\mathcal S$ of admissible states rather than a single distribution, the policymaker chooses
$$
\pi^*\in\arg\max_{\pi\in\Pi}\inf_{s\in\mathcal S}\tau(\pi,s).
$$
Writing
$$
I_{\ell b}(\pi)=\inf_{s\in\mathcal S}\tau(\pi,s),
$$
the first-best problem is
$$
\max_{\pi\in\Pi} I_{\ell b}(\pi).
$$
The paper characterizes ex-ante learnability through PAMPAC-learnability and gives ex-post guarantees for the maximin empirical rule
$$
\hat\pi\in
\Bigl\{\pi:\widehat I_{\ell b}(\pi)\ge \sup_{\pi'}\widehat I_{\ell b}(\pi')-\varepsilon\Bigr\},
$$
together with $\delta$-level set guarantees based on local Rademacher complexity [2012.11046].

This broader literature shows that first-best OPL is not confined to fully identified, unconstrained treatment assignment. Depending on the environment, it may denote an empirical welfare maximizer in a treatment study, a knapsack-optimal allocation under budget and coverage constraints, a censoring-aware optimizer of restricted mean survival time, a near-minimax policy in reinforcement learning, or a maximin rule under partial identification.

Source: https://www.emergentmind.com/topics/first-best-optimal-policy-learning-opl