- The paper presents a novel covariate-assisted inferential framework for estimating support functions in under-identified linear systems.
- It leverages aggregation over continuous covariates to smooth non-regularities and achieve root-N inference using duality and influence function methods.
- The method is empirically validated on the Jobs First welfare reform experiment, demonstrating tighter bounds on treatment effects.
Adaptive Estimation of Aggregated Values in Conditional Linear Programs
The paper "Adaptive Estimation of Aggregated Values of Conditional Linear Programs" (2606.08359) introduces a covariate-assisted inferential framework for partially identified parameters derived from under-identified linear systems. The authors focus on economic settings—such as bounds on treatment effects, models with state dependence, and discrete choice models—where the number of identifying moment conditions fails to point-identify all parameters of interest. Their key innovation is to represent the boundary (support function) of the identified set as an average of intersections of regression functions, integrated over the distribution of covariates. They establish regularity, present asymptotic theory, and provide empirical illustration using data from the Jobs First welfare reform experiment.
The inferential target in many empirical problems is a projection of a partially identified parameter vector β0∈Rd, which solves a system of k<d linear equations, often constrained to the non-negative orthant: Aβ0=b0,β0≥0
where A is a known k×d matrix and b0 is an unknown mean of an observable vector. Such under-determined linear systems commonly emerge in econometrics, including IV models with imperfect compliance, principal stratification, random utility models, and state dependence models.
Practical inference in these models is hampered by two issues:
- Closed-form solutions for the identified set's bounds are rarely known.
- The value function of the linear program often exhibits non-regularity due to flat faces ('multiplicity of solutions')—this non-differentiability implies that standard root-N inference fails.
To address these, the authors propose to systematically incorporate covariates into the linear program, letting the right-hand side depend upon observed characteristics and modeling it as a nonparametric function.
Support Function Approach and Aggregation Over Covariates
The central parameter—the support function of the identified set along a direction q—is defined conditionally as: σ(q,x)=β0maxq⊤β0s.t.Aβ0=b0(x), β0≥0
A primary contribution is to aggregate this support function over the covariate distribution: σ(q)=E[σ(q,X)]
This procedure leverages the fact that aggregation over a continuous covariate space resolves the non-regularity: the set of covariate values corresponding to flat faces in the LP has measure zero under mild smoothness assumptions. As a result, k<d0 is regular and pathwise differentiable, enabling parametric-rate inference.
Duality and Identification Results
A pivotal technical result uses strong LP duality: k<d1
where k<d2 is the finite set of vertices of the dual feasible region (i.e., extreme points of k<d3). Aggregation then yields: k<d4
This is always weakly sharper than the bound from the unconditional linear program on aggregate data due to Jensen-type arguments.

Figure 1: Possible block expansions of the coarse design matrix, illustrating how granular partitions induce structured sparsity in the linear system.
Regularization and Influence Function Theory
The regularity of the aggregated support function is characterized precisely under a unique-dual-vertex assumption—the minimizer k<d5 is almost surely unique under continuously distributed covariates. This enables derivation of an explicit influence function for k<d6: k<d7
This representation is notable; it matches the efficient influence function in special cases (e.g., the semiparametric efficient bound for always-taker estimates as in [Luedtke & van der Laan, 2016]), and admits plug-in estimation. The cross-fitted plug-in estimator using modern machine learning is shown to be root-k<d8-consistent and asymptotically normal.
Estimation and Bootstrap Inference
Estimation proceeds as follows:
- First-stage estimation: Nonparametrically (or with sparsity-regularized methods, e.g., k<d9-penalized logistic regression) estimate Aβ0=b0,β0≥00.
- Dual value computation: For each Aβ0=b0,β0≥01, determine the minimizing dual vertex Aβ0=b0,β0≥02 and compute Aβ0=b0,β0≥03.
- Cross-fitting: Use sample splits to avoid overfitting in the first-stage estimation.
- Aggregation: Average Aβ0=b0,β0≥04 over the sample.
- Bootstrap: Implement a multiplier bootstrap procedure for valid confidence intervals.
Under sufficient first-stage rates (e.g., Aβ0=b0,β0≥05 uniform convergence of Aβ0=b0,β0≥06), the estimator achieves uniformly valid inference.
The empirical section applies the method to the Jobs First experiment, which randomized welfare applicants into different benefit regimes. The partially identified parameters are transition probabilities between welfare participation and earnings states. The baseline estimand is the share of women who, in response to the program, reduce their earnings to opt into welfare.
The approach facilitates:
- Use of rich covariate information: Baseline covariates are incorporated to shrink the width of identified sets.
- Flexible discretization: The linear programming representation can be extended to finer outcome grids. Specifically, above-FPL earnings bins are refined to test whether observed opt-in is due to trivial adjustments or substantive labor-supply changes.

Figure 2: Example granular specification designs, contrasting aggregate and highly refined partitioning in outcome space.
This empirically demonstrates the computational and inferential scalability of the method, allowing identification of meaningful behavioral patterns at granular subpopulations.
Theoretical and Practical Implications
Theoretical Implications
- Regularization via aggregation: Continuous covariate aggregation smooths away nonregularity, permitting root-Aβ0=b0,β0≥07 inference where standard support-function methods fail.
- Unified influence function: The dual-based influence function generalizes to a broad class of LP-based intersection bounds and facilitates the use of modern ML first-stage estimators.
- Generality and extensibility: The framework encompasses canonical econometric models (e.g., discrete choice, IV, sample selection) and is robust to high-dimensional covariates.
Practical Implications
- Sharper bounds: Conditioning on covariates systematically provides tighter identified sets than aggregate-only approaches.
- Modular computation: The block structure of the design matrix enables efficient implementation even as outcome granularity increases.
- Flexible ML integration: Regularized regression or cross-fitting allows for nonparametric estimation of nuisance components, preserving validity under modern modeling pipelines.
Future Directions
Potential extensions include:
- Relaxation of uniqueness: When the margin condition fails (ties among dual vertices on a positive-measure subset), robust inference may be obtained via non-sharp but valid moment-inequality or smoothing approaches [see also, e.g., (Ji et al., 2023)].
- High-dimensional asymptotics: Extending asymptotic theory to settings where the LP dimension (number of dual vertices) grows with sample size.
- Policy learning with partial identification: Integrating these conditional LP values as primitives in data-driven treatment assignment, policy evaluation, or semi-parametric causal inference.
Conclusion
This paper establishes a robust and general methodology for inference on aggregate values in conditional linear programs under partial identification. By leveraging covariate information and exploiting duality, the authors resolve the regularity barriers of traditional support-function inference and enable scalable, root-Aβ0=b0,β0≥08-valid inference with modern machine learning first stages. The empirical demonstration underscores both flexibility and computational feasibility, offering a new standard for partial identification analysis in complex, high-dimensional applied economic problems.