---
title: 'CROMS: Decision-Aware Robust Optimization'
url: https://www.emergentmind.com/topics/croms
type: topic
---

# CROMS: Decision-Aware Robust Optimization

Searching arXiv for the exact topic and nearby acronym collisions to ground the article in the latest relevant paper.
CROMS denotes **Conformalized Robust Optimization with Model Selection**, a framework for **decision-making under uncertainty** in which prediction models are selected not by predictive accuracy or set width alone, but by the **downstream decision risk** induced by conformalized **contextual robust optimization (CRO)** decisions. In this setting, the unknown outcome \(Y\) is represented through a prediction set \(\mathcal U(X)\subseteq\mathcal Y\), and a decision \(z\in\mathcal Z\) is chosen by minimizing the worst-case loss over that set. The central contribution of CROMS is to make model selection **decision-aware**: among candidate score models \(\{S_\lambda:\lambda\in\Lambda\}\), it selects the model whose conformal set yields the most favorable robust decision, while preserving robustness guarantees in either approximate or exact finite-sample form depending on the algorithmic variant [2507.04716].

## 1. Definition and decision-theoretic setting

The underlying optimization problem is the basic **contextual robust optimization** problem
\[
z(X)=\arg\min_{z\in \mathcal Z}\max_{c\in \mathcal U(X)} \phi(c,z).
\]
Here, \(X\) is the observed covariate for the test instance, \(\mathcal U(X)\) is a prediction set for the unknown label \(Y\), and \(\phi(c,z)\) is the loss incurred if the true label is \(c\) and decision is \(z\). The inner maximization makes the decision robust to all labels in the set [2507.04716].

Standard conformal prediction supplies a prediction set with marginal coverage \(1-\alpha\), and this implies a corresponding **marginal robustness** guarantee for the CRO decision. The paper’s central insight is that conformal validity alone does not resolve a crucial design choice: **which predictive model should be used to construct the conformal set**. CROMS addresses precisely this issue by selecting models to approximately minimize the **average decision risk** in CRO solutions rather than optimizing prediction-set size or generic predictive performance [2507.04716].

For a candidate score model \(S_\lambda\), the conformal set is defined by
\[
\mathcal U_\lambda(X_{n+1}) = \{c\in \mathcal Y : S_\lambda(X_{n+1},c)\le \hat q_\lambda\}, \quad \hat q_\lambda = Q_{(1-\alpha)(1+n^{-1})}\big(\{S_\lambda(X_i,Y_i)\}_{i=1}^n\big),
\]
and the corresponding robust decision is
\[
z_\lambda(X_{n+1}) = \arg\min_{z\in \mathcal Z}\max_{c\in \mathcal U_\lambda(X_{n+1})}\phi(c,z).
\]
CROMS differs from standard CRO plus conformal prediction because it chooses among the candidate models by minimizing a **decision-risk proxy** rather than treating model choice as external to the robust decision problem [2507.04716].

The paper defines the population oracle model through the decision risk
\[
\lambda^* = \arg\min_{\lambda\in\Lambda} \mathbb E\!\left[\phi\big(Y,z_\lambda^{\mathsf o}(X)\big)\right],
\]
where \(z_\lambda^{\mathsf o}(X)\) is the decision using the oracle conformal set for model \(\lambda\). In this sense, CROMS is explicitly **decision-aware model selection**.

## 2. Oracle formulation, robustness, and optimality criteria

The oracle construction begins with the threshold
\[
q_\lambda^{\mathsf o} = \inf\{q:\Pr(S_\lambda(X,Y)\le q)\ge 1-\alpha\},
\]
which defines the oracle conformal set
\[
\mathcal U_\lambda^{\mathsf o}(X) = \{c\in\mathcal Y: S_\lambda(X,c)\le q_\lambda^{\mathsf o}\}.
\]
The corresponding oracle robust decision is
\[
z_\lambda^{\mathsf o}(X) = \arg\min_{z\in\mathcal Z}\max_{c\in\mathcal U_\lambda^{\mathsf o}(X)}\phi(c,z),
\]
and the oracle-optimal model is
\[
\lambda^*=\arg\min_{\lambda\in\Lambda}\mathbb E[\phi(Y,z_\lambda^{\mathsf o}(X))].
\]
The minimum expected loss is denoted
\[
v_\Lambda^*=\mathbb E[\phi(Y,z_{\lambda^*}^{\mathsf o}(X))].
\]

The paper defines **marginal robustness** as
\[
\Pr\!\left( \phi(Y,z(X)) \le \max_{c\in\mathcal U(X)} \phi(c,z(X)) \right)\ge 1-\alpha.
\]
This is guaranteed whenever the prediction set has marginal coverage \(1-\alpha\). The emphasis is not merely on set validity, but on a guarantee that the realized decision loss is controlled by the robust loss evaluated over the conformal set [2507.04716].

A data-driven decision \(\hat z(X)\) is said to be asymptotically optimal if
\[
\lim_{n\to\infty}\mathbb E[\phi(Y,\hat z(X))] = v_\Lambda^*.
\]
For individualized selection, the analogous conditional notion is
\[
\lim_{n\to\infty}\mathbb E[\phi(Y,\hat z(X))\mid X]=v_\Lambda^*(X) \quad\text{a.s.}
\]
These criteria make explicit that CROMS is formulated as a model-selection problem for **robust decision quality**, not only for predictive calibration [2507.04716].

## 3. E-CROMS: efficient empirical-risk-based selection

**E-CROMS** is the computationally simpler variant. For each \(\lambda\in\Lambda\), it first builds the conformal set
\[
\mathcal U_\lambda(x)
=
\{c:S_\lambda(x,c)\le \hat q_\lambda\},\qquad
\hat q_\lambda = Q_{(1-\alpha)(1+n^{-1})}(\{S_\lambda(X_i,Y_i)\}_{i=1}^n).
\]
It then computes, for each training point \(i\),
\[
z_\lambda(X_i)
=
\arg\min_{z\in\mathcal Z}\max_{c\in \mathcal U_\lambda(X_i)}\phi(c,z),
\]
selects the model by ERM,
\[
\hat\lambda_n
=
\arg\min_{\lambda\in\Lambda}
\frac{1}{n}\sum_{i=1}^n \phi(Y_i,z_\lambda(X_i)),
\]
and outputs
\[
\widehat{\mathcal U}(X_{n+1})=\mathcal U_{\hat\lambda_n}(X_{n+1}),
\qquad
\hat z(X_{n+1}) = z_{\hat\lambda_n}(X_{n+1}).
\]
Its defining feature is that model selection is performed directly against empirical decision loss [2507.04716].

The main tradeoff is explicit. The paper describes E-CROMS as **cheap** because it requires only \(n\) CRO solves per candidate model, but the model-selection step depends asymmetrically on the labeled sample. Consequently, the final selected model no longer has exact conformal symmetry, which creates a small coverage and robustness gap [2507.04716].

The finite-sample guarantee is correspondingly weakened. Under i.i.d. sampling,
\[
\Pr\big(Y_{n+1}\in \widehat{\mathcal U}(X_{n+1})\big) \ge (1-\alpha)(1+n^{-1}) - \frac{\sqrt{\log(2|\Lambda|)/2}+1/3}{\sqrt n}.
\]
The deviation from nominal is therefore \(O(\sqrt{\log|\Lambda|/n})\). The paper also gives an asymptotic decision-efficiency bound:
\[
\left| \mathbb E[\phi(Y_{n+1},\hat z(X_{n+1}))]-v_\Lambda^* \right| \le O\!\left( B+\frac{L}{\mu}\sqrt{\frac{\log(n\vee|\Lambda|)}{n}} \right),
\]
which implies asymptotic optimality [2507.04716].

These guarantees are derived under three assumptions. First, **quantile estimation regularity**, with score density lower bounded near the relevant quantile region:
\[
f_\lambda(s)\ge \mu>0.
\]
Second, **bounded loss**:
\[
\sup_{y,z} |\phi(y,z)|\le B.
\]
Third, the CRO solution is **locally Lipschitz in the threshold**:
\[
\sup_{x,y}
\left|
\phi\big(y,z_\lambda(x;q)\big)-\phi\big(y,z_\lambda(x;q')\big)
\right|
\le L|q-q'|.
\]
A common misconception would be to treat E-CROMS as exact conformal model selection. The paper explicitly does not make that claim; its finite-sample robustness is approximate rather than exact.

## 4. F-CROMS: symmetric full-conformal selection

**F-CROMS** is the full-conformal variant designed to restore exact conformal symmetry. Its central construction is hypothesis-based: for each candidate label \(y\) of the test point, \((X_{n+1},y)\) is treated as if it were part of the sample, and model selection is performed symmetrically on the augmented data [2507.04716].

For each \(\lambda\) and hypothesized \(y\), the paper defines
\[
\mathcal U_\lambda^{y}(x) = \{c\in\mathcal Y: S_\lambda(x,c)\le Q_{1-\alpha}( \{S_\lambda(X_i,Y_i)\}_{i=1}^n \cup \{S_\lambda(X_{n+1},y)\} )\}.
\]
For \(i\in[n+1]\), the auxiliary robust decision is
\[
z_\lambda^{y}(X_i) = \arg\min_{z\in\mathcal Z}\max_{c\in \mathcal U_\lambda^{y}(X_i)}\phi(c,z),
\]
and the corresponding hypothesized model selector is
\[
\hat\lambda^{y} = \arg\min_{\lambda\in\Lambda} \frac{1}{n+1} \left\{ \sum_{i=1}^n \phi(Y_i,z_\lambda^{y}(X_i)) + \phi(y,z_\lambda^{y}(X_{n+1})) \right\}.
\]
The final prediction set is
\[
\widehat{\mathcal U}(X_{n+1}) = \left\{ y\in\mathcal Y: S_{\hat\lambda^y}(X_{n+1},y) \le Q_{1-\alpha}( \{S_{\hat\lambda^y}(X_i,Y_i)\}_{i=1}^n \cup \{S_{\hat\lambda^y}(X_{n+1},y)\} ) \right\},
\]
followed by the final robust decision
\[
\hat z(X_{n+1}) = \arg\min_{z\in\mathcal Z}\max_{c\in \widehat{\mathcal U}(X_{n+1})}\phi(c,z).
\]

Because the construction is symmetric in the augmented sample, F-CROMS inherits the usual full-conformal validity logic. The paper states that if the data are i.i.d. or exchangeable, then F-CROMS satisfies **finite-sample marginal robustness at level \(1-\alpha\)**. It also provides an optimality result under the same quantile, Lipschitz, and boundedness assumptions used for E-CROMS, together with a nontrivial **efficiency gap**
\[
E[\phi(Y,z_\lambda^{\mathsf o}(X))] \ge E[\phi(Y,z_{\lambda^*}^{\mathsf o}(X))]+\beta_n \quad \forall \lambda\neq \lambda^*,
\]
where
\[
\beta_n \gtrsim Bn^{-1}+n^{-\gamma}\sqrt{\log(n\vee|\Lambda|)},\quad \gamma<1/2.
\]
Under these conditions,
\[
\left| \mathbb E[\phi(Y_{n+1},\hat z(X_{n+1}))]-v_\Lambda^* \right| \le O\!\left(\frac{L}{\mu}\sqrt{\frac{\log n}{n}}\right).
\]

The principal cost of this symmetry is computational. F-CROMS is described as **much more expensive** because it must evaluate all \(y\in\mathcal Y\) and, for each \(y\), re-run model selection. For regression, the paper therefore proposes a **discretized implementation** to make F-CROMS feasible [2507.04716].

## 5. CROiMS and F-CROiMS: individualized model selection

The individualized extension, **CROiMS**, is designed for settings in which the “best” model depends on the test covariate \(X_{n+1}\). Instead of minimizing average decision risk, it minimizes **conditional decision risk given \(X\)**. The conditional quantile for model \(\lambda\) is
\[
q_\lambda^{\mathsf{co}(x)} = \inf\{q:\Pr(S_\lambda(X,Y)\le q\mid X=x)\ge 1-\alpha\},
\]
which induces the conditional oracle conformal set
\[
\mathcal U_\lambda^{\mathsf{co}(x)} = \{y\in\mathcal Y:S_\lambda(x,y)\le q_\lambda^{\mathsf{co}(x)}\}.
\]
The corresponding oracle decision is
\[
z_\lambda^{\mathsf{co}(x)} = \arg\min_{z\in\mathcal Z}\max_{c\in \mathcal U_\lambda^{\mathsf{co}(x)}}\phi(c,z),
\]
and the individualized oracle model is
\[
\lambda^*(x)=\arg\min_{\lambda\in\Lambda} \mathbb E[\phi(Y,z_\lambda^{\mathsf{co}(x)})\mid X=x].
\]
This is the defining distinction from global CROMS: **the selected model is allowed to vary with \(x\)** [2507.04716].

CROiMS uses **localized conformal prediction** with kernel weights
\[
w_i(x)=\frac{H(X_i,x)}{\sum_{j=1}^n H(X_j,x)},
\]
and estimates the conditional threshold by
\[
\hat q_\lambda(x) = Q_{1-\alpha}\big(\{S_\lambda(X_i,Y_i)\}_{i=1}^n;\{w_i(x)\}_{i=1}^n\big).
\]
The localized conformal set is
\[
\mathcal U_\lambda^{\mathsf{LCP}(x)} = \{c\in\mathcal Y:S_\lambda(x,c)\le \hat q_\lambda(x)\}.
\]
For the test point \(X_{n+1}\), model selection is performed by weighted ERM:
\[
\hat\lambda(X_{n+1}) = \arg\min_{\lambda\in\Lambda} \sum_{i=1}^n w_i(X_{n+1}) \cdot \phi(Y_i,z_\lambda^{\mathsf{LCP}(X_i)}).
\]
The final set and decision are then
\[
\widehat{\mathcal U}^{\mathsf{CROiMS}(X_{n+1})} = \mathcal U_{\hat\lambda(X_{n+1})}^{\mathsf{LCP}(X_{n+1})},
\]
and
\[
\hat z^{\mathsf{CROiMS}(X_{n+1})} = \arg\min_{z\in\mathcal Z}\max_{c\in \widehat{\mathcal U}^{\mathsf{CROiMS}(X_{n+1})}}\phi(c,z).
\]

The paper is explicit that **exact distribution-free conditional coverage is impossible in general**, so CROiMS targets **asymptotic conditional robustness** rather than exact finite-sample conditional coverage. Under the assumptions
\[
p(x)\ge \rho>0,
\]
\[
\sup_s |F_\lambda(s|x)-F_\lambda(s|x')|\le \tau\|x-x'\|,
\]
\[
f_\lambda(s|x)\ge \rho_c>0,
\]
and
\[
\left|
\mathbb E[\Phi_\lambda(X,Y)\mid X=x]
-
\mathbb E[\Phi_\lambda(X,Y)\mid X=x']
\right|
\le L_c\|x-x'\|,
\]
together with the earlier bounded-loss and Lipschitz-in-threshold conditions, CROiMS satisfies
\[
\Pr\big(Y_{n+1}\in \widehat{\mathcal U}^{\mathsf{CROiMS}(X_{n+1})}\mid X_{n+1}\big) \ge 1-\alpha - O\!\left( \sqrt{\frac{\log(n\vee|\Lambda|)}{\rho n h_n^d} + \frac{\tau}{\rho}h_n\log(h_n^{-d})} \right)
\]
almost surely. It also satisfies a conditional optimality bound:
\[
\left| \mathbb E\!\left[\phi\!\left(Y_{n+1},\hat z^{\mathsf{CROiMS}(X_{n+1})}\right)\mid X_{n+1}\right] - v_\Lambda^*(X_{n+1}) \right|
\]
\[
\le O\!\left( B + \frac{L}{\rho_c} \sqrt{\frac{\log(n\vee|\Lambda|)}{\rho n h_n^d} + L_c + \frac{\tau L}{\rho_c h_n\log(h_n^{-d})}} \right).
\]
The paper notes that choosing \(h_n\asymp n^{-1/(d+2)}\) yields near-optimal nonparametric rates up to a \(\log|\Lambda|\) factor [2507.04716].

To restore exact finite-sample marginal validity for individualized selection, the appendix introduces **F-CROiMS**, based on a **swapping / leave-one-out full conformal** construction. Its prediction set is
\[
\widehat{\mathcal U}^{\textsf{F-CROiMS}(X_{n+1})} = \left\{ y\in\mathcal Y: S_{\hat\lambda(X_{n+1})}(X_{n+1},y) \le Q_{(1-\alpha)(1+n^{-1})} \big(\{S_{\hat\lambda^y(X_j)}(X_j,Y_j)\}_{j=1}^n\big) \right\},
\]
and it achieves
\[
\Pr\big(Y_{n+1}\in \widehat{\mathcal U}^{\textsf{F-CROiMS}(X_{n+1})}\big)\ge 1-\alpha.
\]

## 6. Empirical evaluation, metrics, and computational tradeoffs

The empirical study includes both synthetic and real-data settings and compares CROMS and CROiMS against several baselines: **Naive-CP / Naive-LCP**, which randomly choose a candidate model and then do conformal prediction, and **E2E**, a sample-splitting end-to-end method that first selects a model and then calibrates on held-out data [2507.04716].

For the averaged setting, the reported metrics are **marginal miscoverage**,
\[
\frac{1}{m}\sum_{j=n+1}^{n+m}\mathbf 1\{Y_j\notin \widehat{\mathcal U}(X_j)\},
\]
**marginal misrobustness**,
\[
\frac{1}{m}\sum_{j=n+1}^{n+m}
\mathbf 1\!\left\{
\phi(Y_j,\hat z(X_j))>
\max_{c\in\widehat{\mathcal U}(X_j)}\phi(c,\hat z(X_j))
\right\},
\]
and **average loss**,
\[
\frac{1}{m}\sum_{j=n+1}^{n+m}\phi(Y_j,\hat z(X_j)).
\]
For individualized settings, the paper additionally reports **worst-case conditional miscoverage** over local neighborhoods, **worst-case conditional misrobustness**, **group conditional loss** over covariate partitions, and **coverage gap** and **robustness gap** across subgroups [2507.04716].

| Method | Main property | Principal tradeoff |
|---|---|---|
| E-CROMS | Computationally efficient global model selection | Small finite-sample coverage gap |
| F-CROMS | Exact finite-sample marginal robustness | Much more expensive |
| CROiMS | Covariate-aware individualized model selection | Conditional guarantees are asymptotic |
| F-CROiMS | Finite-sample marginal validity for individualized selection | Uses a swapping/full-conformal correction |

In synthetic **classification**, the paper uses labels \(\mathcal Y=\{1,\dots,5\}\), decisions \(\mathcal Z=\{1,\dots,5\}\), and a loss matrix encoding ordinal or clinical severity. Candidate models vary by score penalty \(\lambda\) or by feature subsets. The reported result is that E-CROMS and F-CROMS both improve decision loss over baselines; F-CROMS maintains exact coverage, while E-CROMS has a small coverage gap for small \(n\) or large \(|\Lambda|\) [2507.04716].

In synthetic **regression**, the response is multivariate with \(Y\in\mathbb R^2\), and candidate models are trained on different datasets with covariate shift. The reported result is that CROiMS performs best in conditional and group-specific loss and better controls local coverage and robustness than averaged-selection methods [2507.04716].

The real-world applications are **COVID-19 radiography diagnosis**, with 4 classes—Normal, Pneumonia, COVID-19, Lung Opacity—and candidate CNN models trained on different label distributions, and **HAM10000 dermoscopic diagnosis**, with two decision categories described as “no action / additional test / disease” under a medical loss matrix and candidate models trained on age-stratified subsets. The paper reports that CROiMS gives the lowest average loss and best subgroup performance while maintaining conditional misrobustness near the nominal level in the radiography task, and substantially improves group-conditional loss in HAM10000 while maintaining good robustness control across age groups [2507.04716].

The broad methodological takeaway is that the paper shifts robust conformal decision-making from **constructing any valid set** to **selecting the model whose conformal set yields the best downstream decision**. A plausible implication is that, in heterogeneous decision environments, model selection for conformal prediction cannot be treated as a purely predictive pre-processing step.

## 7. Terminological scope and acronym collisions

Within the decision-focused conformal prediction literature, **CROMS** refers specifically to **Conformalized Robust Optimization with Model Selection** [2507.04716]. However, nearby acronyms in other fields can create confusion.

In astroparticle instrumentation, **CROME** denotes **Cosmic-Ray Observation via Microwave Emission**, an experiment built to search for microwave signals from extensive air showers [1108.0588]. In global cosmic-ray network science, **CREDO** denotes the **Cosmic Ray Extremely Distributed Observatory**, a collaboration dedicated to observing cosmic rays and cosmic ray ensembles [2010.08351]. In reionization simulation, **CROC** denotes **Cosmic Reionization On Computers**, a long-term numerical program for modeling cosmic reionization [1403.4245]. In computational mechanics, **CROMs** refers to **clustering-based reduced-order models**, and the adaptive extension **ACROM** generalizes that family for localized history-dependent phenomena [2109.11897].

This terminological overlap suggests that the acronym should be interpreted by disciplinary context. In the present usage, CROMS is a model-selection framework for conformalized robust decision-making rather than an observational program, a cosmological simulation suite, or a reduced-order modeling methodology.

Source: https://www.emergentmind.com/topics/croms