Papers
Topics
Authors
Recent
Search
2000 character limit reached

Adaptive Estimation of Aggregated Values of Conditional Linear Programs

Published 6 Jun 2026 in econ.EM | (2606.08359v1)

Abstract: We develop a covariate-assisted approach to partially identified parameters that are solutions to an under-identified system of linear equations with known coefficients. Examples include bounds on treatment effects, models of unemployment with state dependence, choice-theoretic models of IV, and random utility models. The boundary (i.e., support function) of the proposed identified set is represented as an average of intersections of regression functions, aggregated over the covariate distribution. We show that the boundary is a regular parameter, propose asymptotic theory, and demonstrate using an empirical application to Jobs First.

Summary

  • The paper presents a novel covariate-assisted inferential framework for estimating support functions in under-identified linear systems.
  • It leverages aggregation over continuous covariates to smooth non-regularities and achieve root-N inference using duality and influence function methods.
  • The method is empirically validated on the Jobs First welfare reform experiment, demonstrating tighter bounds on treatment effects.

Adaptive Estimation of Aggregated Values in Conditional Linear Programs

The paper "Adaptive Estimation of Aggregated Values of Conditional Linear Programs" (2606.08359) introduces a covariate-assisted inferential framework for partially identified parameters derived from under-identified linear systems. The authors focus on economic settings—such as bounds on treatment effects, models with state dependence, and discrete choice models—where the number of identifying moment conditions fails to point-identify all parameters of interest. Their key innovation is to represent the boundary (support function) of the identified set as an average of intersections of regression functions, integrated over the distribution of covariates. They establish regularity, present asymptotic theory, and provide empirical illustration using data from the Jobs First welfare reform experiment.


Problem Formulation and Motivation

The inferential target in many empirical problems is a projection of a partially identified parameter vector β0Rd\bm\beta_0 \in \mathbb{R}^d, which solves a system of k<dk < d linear equations, often constrained to the non-negative orthant: Aβ0=b0,β00A\bm\beta_0 = \bm{b}_0, \quad \bm\beta_0 \geq 0 where AA is a known k×dk \times d matrix and b0\bm{b}_0 is an unknown mean of an observable vector. Such under-determined linear systems commonly emerge in econometrics, including IV models with imperfect compliance, principal stratification, random utility models, and state dependence models.

Practical inference in these models is hampered by two issues:

  1. Closed-form solutions for the identified set's bounds are rarely known.
  2. The value function of the linear program often exhibits non-regularity due to flat faces ('multiplicity of solutions')—this non-differentiability implies that standard root-NN inference fails.

To address these, the authors propose to systematically incorporate covariates into the linear program, letting the right-hand side depend upon observed characteristics and modeling it as a nonparametric function.


Support Function Approach and Aggregation Over Covariates

The central parameter—the support function of the identified set along a direction qq—is defined conditionally as: σ(q,x)=maxβ0qβ0s.t.Aβ0=b0(x), β00\sigma(q, x) = \max_{\bm\beta_0} q^\top \bm\beta_0 \quad \text{s.t.} \quad A\bm\beta_0 = \bm{b}_0(x),\ \bm\beta_0 \geq 0 A primary contribution is to aggregate this support function over the covariate distribution: σ(q)=E[σ(q,X)]\sigma(q) = \mathbb{E}[\sigma(q,X)] This procedure leverages the fact that aggregation over a continuous covariate space resolves the non-regularity: the set of covariate values corresponding to flat faces in the LP has measure zero under mild smoothness assumptions. As a result, k<dk < d0 is regular and pathwise differentiable, enabling parametric-rate inference.


Duality and Identification Results

A pivotal technical result uses strong LP duality: k<dk < d1 where k<dk < d2 is the finite set of vertices of the dual feasible region (i.e., extreme points of k<dk < d3). Aggregation then yields: k<dk < d4 This is always weakly sharper than the bound from the unconditional linear program on aggregate data due to Jensen-type arguments.

Figure 1

Figure 1: Possible block expansions of the coarse design matrix, illustrating how granular partitions induce structured sparsity in the linear system.


Regularization and Influence Function Theory

The regularity of the aggregated support function is characterized precisely under a unique-dual-vertex assumption—the minimizer k<dk < d5 is almost surely unique under continuously distributed covariates. This enables derivation of an explicit influence function for k<dk < d6: k<dk < d7 This representation is notable; it matches the efficient influence function in special cases (e.g., the semiparametric efficient bound for always-taker estimates as in [Luedtke & van der Laan, 2016]), and admits plug-in estimation. The cross-fitted plug-in estimator using modern machine learning is shown to be root-k<dk < d8-consistent and asymptotically normal.


Estimation and Bootstrap Inference

Estimation proceeds as follows:

  1. First-stage estimation: Nonparametrically (or with sparsity-regularized methods, e.g., k<dk < d9-penalized logistic regression) estimate Aβ0=b0,β00A\bm\beta_0 = \bm{b}_0, \quad \bm\beta_0 \geq 00.
  2. Dual value computation: For each Aβ0=b0,β00A\bm\beta_0 = \bm{b}_0, \quad \bm\beta_0 \geq 01, determine the minimizing dual vertex Aβ0=b0,β00A\bm\beta_0 = \bm{b}_0, \quad \bm\beta_0 \geq 02 and compute Aβ0=b0,β00A\bm\beta_0 = \bm{b}_0, \quad \bm\beta_0 \geq 03.
  3. Cross-fitting: Use sample splits to avoid overfitting in the first-stage estimation.
  4. Aggregation: Average Aβ0=b0,β00A\bm\beta_0 = \bm{b}_0, \quad \bm\beta_0 \geq 04 over the sample.
  5. Bootstrap: Implement a multiplier bootstrap procedure for valid confidence intervals.

Under sufficient first-stage rates (e.g., Aβ0=b0,β00A\bm\beta_0 = \bm{b}_0, \quad \bm\beta_0 \geq 05 uniform convergence of Aβ0=b0,β00A\bm\beta_0 = \bm{b}_0, \quad \bm\beta_0 \geq 06), the estimator achieves uniformly valid inference.


Practical Illustration: Jobs First Welfare Reform

The empirical section applies the method to the Jobs First experiment, which randomized welfare applicants into different benefit regimes. The partially identified parameters are transition probabilities between welfare participation and earnings states. The baseline estimand is the share of women who, in response to the program, reduce their earnings to opt into welfare.

The approach facilitates:

  • Use of rich covariate information: Baseline covariates are incorporated to shrink the width of identified sets.
  • Flexible discretization: The linear programming representation can be extended to finer outcome grids. Specifically, above-FPL earnings bins are refined to test whether observed opt-in is due to trivial adjustments or substantive labor-supply changes.

Figure 2

Figure 2: Example granular specification designs, contrasting aggregate and highly refined partitioning in outcome space.

This empirically demonstrates the computational and inferential scalability of the method, allowing identification of meaningful behavioral patterns at granular subpopulations.


Theoretical and Practical Implications

Theoretical Implications

  • Regularization via aggregation: Continuous covariate aggregation smooths away nonregularity, permitting root-Aβ0=b0,β00A\bm\beta_0 = \bm{b}_0, \quad \bm\beta_0 \geq 07 inference where standard support-function methods fail.
  • Unified influence function: The dual-based influence function generalizes to a broad class of LP-based intersection bounds and facilitates the use of modern ML first-stage estimators.
  • Generality and extensibility: The framework encompasses canonical econometric models (e.g., discrete choice, IV, sample selection) and is robust to high-dimensional covariates.

Practical Implications

  • Sharper bounds: Conditioning on covariates systematically provides tighter identified sets than aggregate-only approaches.
  • Modular computation: The block structure of the design matrix enables efficient implementation even as outcome granularity increases.
  • Flexible ML integration: Regularized regression or cross-fitting allows for nonparametric estimation of nuisance components, preserving validity under modern modeling pipelines.

Future Directions

Potential extensions include:

  • Relaxation of uniqueness: When the margin condition fails (ties among dual vertices on a positive-measure subset), robust inference may be obtained via non-sharp but valid moment-inequality or smoothing approaches [see also, e.g., (Ji et al., 2023)].
  • High-dimensional asymptotics: Extending asymptotic theory to settings where the LP dimension (number of dual vertices) grows with sample size.
  • Policy learning with partial identification: Integrating these conditional LP values as primitives in data-driven treatment assignment, policy evaluation, or semi-parametric causal inference.

Conclusion

This paper establishes a robust and general methodology for inference on aggregate values in conditional linear programs under partial identification. By leveraging covariate information and exploiting duality, the authors resolve the regularity barriers of traditional support-function inference and enable scalable, root-Aβ0=b0,β00A\bm\beta_0 = \bm{b}_0, \quad \bm\beta_0 \geq 08-valid inference with modern machine learning first stages. The empirical demonstration underscores both flexibility and computational feasibility, offering a new standard for partial identification analysis in complex, high-dimensional applied economic problems.


Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.