---
title: 'Early-Cutoff Policy: Safe Learning in RD Designs'
url: https://www.emergentmind.com/topics/early-cutoff-policy
type: topic
---

# Early-Cutoff Policy: Safe Learning in RD Designs

An Early-Cutoff Policy (ECP) in a multi-cutoff regression discontinuity design is a threshold policy that replaces existing treatment cutoffs with earlier ones, typically by setting \(c(g) < c^{(0)}(g)\) when higher values of the running variable trigger treatment. In the sharp RD setting, the treatment rule is deterministic, so moving a cutoff changes treatment assignment in regions where the counterfactual policy value is not point identified from the status quo alone. The framework developed in "Safe Policy Learning under Regression Discontinuity Designs with Multiple Cutoffs" [2208.13323] addresses this problem by decomposing policy value into identifiable and unidentifiable components, bounding the latter through cross-group smoothness restrictions, and choosing new cutoffs by robust optimization so that the learned policy is safe relative to the baseline under the maintained assumptions.

## 1. Formal setup and policy class

The framework assumes a running variable \(X \in \mathbb{R}\) with density \(f_X > 0\) on support \(\mathcal{X}\), a group label \(G \in \mathcal{G}:=\{1,\ldots,Q\}\), and a known baseline cutoff \(c^{(0)}(g)\) for each group \(g\), sorted so that \(c^{(0)}(1) < \cdots < c^{(0)}(Q)\) [2208.13323]. Under the baseline sharp RD assignment rule,
\[
D^{(0)} = 1\{X \ge c^{(0)}(G)\},
\]
with the reverse inequality used for lower-is-treated designs. Potential outcomes are \(Y(1)\) and \(Y(0)\), the observed outcome is
\[
Y = D^{(0)}Y(1) + (1-D^{(0)})Y(0),
\]
and the sample \(\{(X_i,G_i,Y_i(1),Y_i(0))\}_{i=1}^n\) is i.i.d. from population \(\mathbb{P}\).

An ECP is indexed by a cutoff function \(c:\mathcal{G}\to\mathbb{R}\). Under \(c\), treatment becomes
\[
D_c := 1\{X \ge c(G)\}.
\]
The admissible class is restricted to
\[
\mathcal{C}:=\{c:\mathcal{G}\to[c^{(0)}(1),c^{(0)}(Q)]\},
\]
which avoids extrapolation beyond observed support. The induced threshold-policy class is
\[
\Pi := \{\pi_c(x,g)=1\{x\ge c(g)\}: c\in\mathcal{C}\},
\]
and the status quo policy is \(\pi^{(0)}(x,g)=1\{x\ge c^{(0)}(g)\}\), with \(\pi^{(0)}\in\Pi\).

Policy quality is measured by expected utility
\[
U(c):=\mathbb{E}[u(Y(D_c),X,G)].
\]
If not otherwise specified, the utility is \(u(y,w)=y\). A cost-adjusted specification is also allowed:
\[
u(y,w)=y-Cw,
\]
for a per-treatment cost \(C\ge 0\). For \(u(y,w)=y\), writing \(m_d(x,g):=\mathbb{E}[Y(d)\mid X=x,G=g]\) and \(m(w,x,g):=\mathbb{E}[Y(w)\mid X=x,G=g]\), one has
\[
U(c)=\mathbb{E}[Y(D_c)]=\mathbb{E}[m(D_c,X,G)].
\]

## 2. Identification through decomposition

The central identification difficulty is that a new cutoff policy disagrees with the baseline in regions where the realized data reveal only one potential outcome. The framework handles this by decomposing the value of any threshold policy \(\pi\equiv\pi_c\) into baseline-agreement terms and disagreement terms [2208.13323]:
\[
V(\pi):=\mathbb{E}[Y(\pi(X,G))]
\]
\[
= \mathbb{E}\big[\pi\pi^{(0)}Y + (1-\pi)(1-\pi^{(0)})Y\big]
+ \mathbb{E}\big[\pi(1-\pi^{(0)})m(1,X,G)\big]
+ \mathbb{E}\big[(1-\pi)\pi^{(0)}m(0,X,G)\big].
\]
The first expectation is point identified because the policy agrees with the realized treatment assignment. The remaining two terms are unidentifiable without extrapolation.

A useful observable object is
\[
\tilde m(x,g):=\mathbb{E}[Y\mid X=x,G=g]=m(\pi^{(0)}(x,g),x,g),
\]
the conditional mean under the baseline assignment rule. The framework also defines the cross-group difference function
\[
d(w,x,g,g'):=m(w,x,g)-m(w,x,g').
\]
Regression discontinuity identifies \(\tilde m(x,g)\) pointwise and identifies \(d(w,x,g,g')\) only on the side of the nearest cutoff where both potential outcomes are observed.

Using multiple cutoffs, the value is further decomposed as
\[
V(\pi)=\mathbb{I}\mathrm{den}+\Xi_1+\Xi_2.
\]
Here,
\[
\mathbb{I}\mathrm{den}:=\mathbb{E}\big[\pi\pi^{(0)}Y+(1-\pi)(1-\pi^{(0)})Y\big]
\]
is point identified. The term \(\Xi_1\) is identifiable by borrowing \(\tilde m(X,g')\) from the nearest available group interval, while \(\Xi_2\) contains the residual unknown difference-function terms \(d(w,X,g,g')\). The intuition is explicit: for group \(g\), when the candidate policy enters a region where \(\pi\) disagrees with \(\pi^{(0)}\), the nearest group \(g'\) with an observed cutoff supplies an observable surrogate \(\tilde m(X,g')\), and the remaining mismatch is recorded in \(d(w,X,g,g')\).

This decomposition is the analytic core of ECP learning. It separates the part of the policy value that can be estimated efficiently from the part that must be bounded rather than point estimated.

## 3. Smooth heterogeneity and partial identification

The bounding step relies on two assumptions. First, for each group \(g\) and treatment state \(w\), \(m(w,x,g)\) is continuous in \(x\) at \(c^{(0)}(g)\). Second, cross-group heterogeneity varies smoothly along the running variable. In Lipschitz form, the setup allows
\[
|m_d(x_1,g)-m_d(x_2,g)|\le L_d |x_1-x_2|.
\]

Operationally, the key restriction is imposed on the difference function:
\[
|d(w,x,g,g')-d(w,c^{(0)}_{g^*},g,g')|\le \lambda_{wgg'}|x-c^{(0)}_{g^*}|,
\]
with \(g^*=g\vee g'\) when \(w=1\) and \(g^*=g\wedge g'\) when \(w=0\). The radius \(\lambda_{wgg'}\ge 0\) parameterizes how quickly cross-group heterogeneity may vary away from the boundary [2208.13323]. This is the formal version of the paper’s “slowly varying heterogeneity across \(x\)” condition.

The identified cross-group contrast is
\[
\tilde d(x,g,g'):=\mathbb{E}[Y\mid X=x,G=g]-\mathbb{E}[Y\mid X=x,G=g'].
\]
Given \(\tilde d\) in the overlap region and the smoothness radius \(\lambda_{wgg'}\), the admissible set for the unidentifiable component is
\[
\mathcal{M}_d:=\{d: B_\ell(w,x,g,g')\le d(w,x,g,g')\le B_u(w,x,g,g')\},
\]
where
\[
B_\ell(w,x,g,g')=
\lim_{x'\to c^{(0)}_{g^*},\,x'\in\mathcal{X}_{gw}\cap\mathcal{X}_{g'w}}
\big[\tilde d(x',g,g')-\lambda_{wgg'}|x-x'|\big],
\]
\[
B_u(w,x,g,g')=
\lim_{x'\to c^{(0)}_{g^*},\,x'\in\mathcal{X}_{gw}\cap\mathcal{X}_{g'w}}
\big[\tilde d(x',g,g')+\lambda_{wgg'}|x-x'|\big].
\]
The set \(\mathcal{X}_{gw}=\{x\in\mathcal{X}:\pi^{(0)}(x,g)=w\}\) indexes the observed side for group \(g\).

This partial-identification step is what turns multiple cutoffs into a policy-learning device. Without multiple cutoffs, the framework states that one would need stronger smoothness directly on \(m(w,x)\) to construct analogous bounds.

## 4. Robust optimization and the safety guarantee

Safety is defined relative to the status quo policy. For a model \(m\), regret is
\[
R(c;m):=U(c^{(0)};m)-U(c;m).
\]
The robust objective chooses a cutoff function that minimizes the worst-case regret over the admissible model class:
\[
c^\star \in \arg\min_{c\in\mathcal{C}}\sup_{m\in\mathcal{M}}R(c;m),
\]
or equivalently,
\[
c^\star \in \arg\max_{c\in\mathcal{C}}\inf_{m\in\mathcal{M}}U(c;m).
\]

Because the baseline policy \(\pi^{(0)}\) belongs to the candidate policy class \(\Pi\), the resulting optimizer satisfies the safety guarantee
\[
\sup_{m\in\mathcal{M}}R(c^\star;m)\le 0.
\]
Under the maintained continuity and smooth-difference assumptions, this implies
\[
U(c^\star)\ge U(c^{(0)}).
\]
The guarantee is therefore relative, conservative, and model-class dependent: the learned policy is protected against the worst admissible extrapolation error rather than against an unrestricted counterfactual.

This robust formulation also clarifies a common misconception. The framework is not an RD generalization of local treatment-effect estimation to arbitrary thresholds by direct identification. It is a partially identified policy problem in which safe learning is obtained by combining identifiable components with worst-case bounds on the nonidentified remainder.

## 5. Estimation and implementable learning

The identifiable term \(\mathbb{I}\mathrm{den}\) is estimated by its sample analog,
\[
\widehat{\mathbb{I}\mathrm{den}}
=
n^{-1}\sum_{i=1}^n
\big[
\pi(X_i,G_i)\pi^{(0)}(X_i,G_i)Y_i
+
(1-\pi(X_i,G_i))(1-\pi^{(0)}(X_i,G_i))Y_i
\big].
\]
The identifiable extrapolation component \(\Xi_1\) is handled by a doubly robust representation using the group propensity
\[
e_g(x):=\mathbb{P}(G=g\mid X=x)
\]
and the observable outcome regression \(\tilde m(x,g)\). The nuisance functions \(\tilde m\) and \(e_g\) are estimated nonparametrically, with local polynomial RD estimation for \(\tilde m\) and, for example, a multinomial or logistic model for \(e_g\). Cross-fitting is then used to form \(\hat\Xi_{DR}\) from out-of-fold nuisance estimates [2208.13323].

The estimator is rate doubly robust: it achieves root-\(n\) consistency for \(\Xi_1\) if the product of the \(L_2\) errors of \(\tilde m\) and \(e_g\) converges faster than \(n^{-1/2}\). For the unidentifiable component \(\Xi_2\), the procedure estimates the bounds
\[
\hat B_\ell(w,x,g,g') \quad\text{and}\quad \hat B_u(w,x,g,g')
\]
using a two-stage DR learner \(\hat{\tilde d}^{DR}\) for \(\tilde d(x,g,g')\), and forms an empirical uncertainty set \(\widehat{\mathcal{M}}_d\). The in-sample worst case for \(\Xi_2\) is
\[
\inf_{d\in\widehat{\mathcal{M}}_d}\hat\Xi_2(c;d),
\]
which reduces to substituting \(d\) by \(\hat B_\ell\) pointwise wherever it enters with positive weight.

The resulting pessimistic value estimator is
\[
\hat U_{\mathrm{worst}}(c)
=
\widehat{\mathbb{I}\mathrm{den}}(c)
+
\hat\Xi_{DR}(c)
+
\inf_{d\in\widehat{\mathcal{M}}_d}\hat\Xi_2(c;d).
\]
Learning then proceeds by solving
\[
c^\star \in \arg\max_{c\in\mathcal{C}}\hat U_{\mathrm{worst}}(c).
\]
Because \(\pi_c(x,g)\) depends only on \(c(g)\), and because the objective activates only where \(\pi_c\) disagrees with \(\pi^{(0)}\), the optimization typically decomposes across groups. In practice, the paper recommends a groupwise grid search over
\[
c(g)\in[c^{(0)}(1),c^{(0)}(Q)].
\]

The implementable workflow therefore consists of estimating \(\tilde m\) and \(e_g\), selecting smoothness radii \(\lambda_{wgg'}\), constructing \(\widehat{\mathcal{M}}_d\), evaluating \(\hat U_{\mathrm{worst}}(c)\) on a grid, and selecting the maximizer. By construction, the fitted policy satisfies the finite-sample pessimistic inequality
\[
\hat U_{\mathrm{worst}}(c^\star)\ge \hat U_{\mathrm{worst}}(c^{(0)}),
\]
and asymptotically satisfies \(U(c^\star)\ge U(c^{(0)})\).

## 6. Asymptotic guarantees and diagnostics

The asymptotic analysis assumes bounded outcomes,
\[
|Y(w)|\le \Lambda \quad \text{a.s.},
\]
and group overlap on the relevant running-variable interval:
\[
\eta \le e_g(x)\le 1-\eta
\quad\text{for all } g \text{ and } x\in[c^{(0)}(1),c^{(0)}(Q)].
\]
For cross-fitted nuisance estimators, the required convergence conditions are
\[
\sup_{x,g}|\hat{\tilde m}^{(-k)}(x,g)-\tilde m(x,g)|\to_p 0,
\qquad
\sup_x|\hat e_g^{(-k)}(x)-e_g(x)|\to_p 0,
\]
together with
\[
\mathbb{E}\big[(\hat{\tilde m}-\tilde m)^2\big]\cdot
\mathbb{E}\big[(\hat e_g/\hat e_{g'}-e_g/e_{g'})^2\big]
=
o(n^{-1})
\]
for each \(g,g'\).

Under these conditions, the safety regret relative to baseline satisfies
\[
V(\pi^{(0)})-V(\hat\pi)
\le
O_p\!\left(
(\Lambda\sqrt{Q}+Q\sqrt{\max_g V_g^\ast})n^{-1/2}
\;\vee\;
\rho_n^{-1}
\right),
\]
where \(\rho_n^{-1}\) is the convergence rate for estimating the boundary limits of \(\tilde d\). If \(\tilde d\) is \(\gamma\)-smooth in one dimension, then typically
\[
\rho_n^{-1}=n^{-1/(2+1/\gamma)},
\]
which approaches \(n^{-1/2}\) near parallel trends [2208.13323].

The optimality gap relative to the oracle best-in-class threshold policy \(\pi^\ast\) is
\[
V(\pi^\ast)-V(\hat\pi)
\le
O_p\!\left(
(\Lambda\sqrt{Q}+Q\sqrt{\max_g V_g^\ast})n^{-1/2}
\;\vee\;
\rho_n^{-1}
\right)
+
\frac{2}{n}\sum_i \max_{w,g,g'} \lambda_{wgg'}|X_i-c^{(0)}_{g^\ast}|.
\]
The final term is the sample-average width of the identification region implied by the smoothness radii. The paper’s interpretation is direct: regret vanishes at the slower of \(n^{-1/2}\) and \(\rho_n^{-1}\), while the extra term measures the price of safety under partial identification.

The proposed diagnostics are likewise structured around the partial-identification logic. Smoothness sensitivity is assessed by varying \(\lambda_{wgg'}\) through multipliers \(M\in\{0,1,2,4,8\}\); \(M=0\) recovers a parallel-trend-like benchmark. The paper also recommends overlap checks for \(e_g(x)\), plots of fitted \(\tilde d(x,g,g')\) in overlap zones, RD bandwidth robustness, and uncertainty quantification for \(\widehat{\mathbb{I}\mathrm{den}}+\hat\Xi_{DR}\) using influence-function-based variance or bootstrap, combined with worst-case \(\Xi_2\) bounds to form a conservative lower confidence bound for \(U(\hat c)-U(c^{(0)})\).

## 7. Interpretation and empirical illustration

Geometrically, in a higher-is-treated sharp RD, the baseline decision boundary in \((x,g)\)-space is \(X\ge c^{(0)}(g)\). An ECP moves this boundary left by setting \(c(g)<c^{(0)}(g)\), thereby treating additional units with
\[
X\in[c(g),c^{(0)}(g)).
\]
Safety matters because these newly treated units lie farther from the original cutoff, so identification requires extrapolation. The framework’s answer is to borrow \(\tilde m(x,g')\) from the nearest observed group and hedge the remaining uncertainty with smooth-difference bounds [2208.13323]. For lower-is-treated designs, the paper states that one swaps inequalities and left/right language; the framework itself is unchanged.

The empirical illustration uses the Colombian ACCES program in 2010. The application contains 23 departments, each with its own cutoff for college loan eligibility based on a test score after sign normalization, and enrollment as the outcome. Baseline heterogeneity across groups is substantial. With parallel-trend-like \(\lambda=0\), many departments shift substantially, and some shifts are large and possibly implausible. Using data-driven \(\lambda\) from the maximum absolute first derivative of \(\tilde d\) in overlap regions, learned cutoffs move earlier in many departments but more conservatively than with \(\lambda=0\). Strong negative estimated RD effects lead to later cutoffs in a few departments. Under the cost-adjusted utility \(u(y,w)=y-Cw\), moderate \(C\) still tends to produce earlier cutoffs, while large \(C\) moves cutoffs later to avoid treatment costs.

The simulation results reinforce the same interpretation. In a two-group multi-cutoff RD design, when the status quo is optimal, the learned policy converges back to it as sample size grows and as \(\lambda\) increases, becoming more conservative. When a better policy exists, the learned ECP improves over baseline even at moderate \(n\). The optimality gap shrinks for smaller \(\lambda\) and larger \(n\), but cannot vanish when identification remains partial with \(\lambda>0\).

Taken together, these results place the ECP framework in a specific methodological niche. It is a safe policy-learning method for deterministic treatment-assignment environments in which multiple observed cutoffs provide the cross-group structure needed to turn RD extrapolation into a partially identified but learnable optimization problem.

Source: https://www.emergentmind.com/topics/early-cutoff-policy