Papers
Topics
Authors
Recent
Search
2000 character limit reached

Conditional Vector Quantile Regression

Updated 14 July 2026
  • Conditional Vector Quantile Regression is a framework that extends scalar quantile regression to multivariate settings using monotone transport maps derived from convex potentials.
  • It integrates optimal transport theory to construct a unique, joint quantile function that represents the full conditional distribution while ensuring global monotonicity.
  • CVQR has wide applications, including econometrics and manifold learning, and offers robust estimation even under misspecified models through convex-envelope approximations.

Searching arXiv for recent and foundational work on conditional vector quantile regression to ground the article. Conditional Vector Quantile Regression (CVQR) is a conditional multivariate quantile framework that extends scalar quantile regression to vector-valued outcomes by representing conditional distributions through monotone transport maps from a reference rank law to the law of the response given covariates. In the foundational formulation, the central object is the conditional vector quantile function (CVQF), a map QYZ(u,z)Q_{Y\mid Z}(u,z) that is monotone in the multivariate sense of being the gradient of a convex function in uu, pushes a non-atomic reference distribution FUF_U onto the conditional law of YZ=zY\mid Z=z, and yields the strong representation Y=QYZ(U,Z)Y = Q_{Y\mid Z}(U,Z) almost surely for an appropriate latent rank vector UU (Carlier et al., 2014). The term “CVQR” is not the authors’ formal label in the foundational paper; the formal regression model is “vector quantile regression” (VQR), which operationalizes the conditional framework through linear or sieve-type specifications of the CVQF (Carlier et al., 2014).

1. Definition and geometric structure

Let YRdY \in \mathbb{R}^d and ZRkZ \in \mathbb{R}^k be random vectors, with conditional distribution FYZ(,z)F_{Y\mid Z}(\cdot,z). Fix a reference, non-atomic distribution FUF_U on uu0 with density uu1 and convex support, for example uu2 or uu3. A conditional vector quantile function is a measurable map

uu4

such that for each uu5, the map uu6 is monotone in the sense of being the gradient of a convex function,

uu7

for some convex potential uu8. Equivalently, for all uu9 in the support of FUF_U0,

FUF_U1

With FUF_U2, the pushforward condition is

FUF_U3

and the strong representation is

FUF_U4

Under the non-atomicity and density condition on FUF_U5, the CVQF exists and is unique FUF_U6-almost everywhere for each FUF_U7, as a conditional Brenier map (Carlier et al., 2014).

When each conditional law FUF_U8 admits a density, there is also a conditional inverse or rank map,

FUF_U9

where YZ=zY\mid Z=z0 is the Legendre transform of YZ=zY\mid Z=z1. It satisfies

YZ=zY\mid Z=z2

and

YZ=zY\mid Z=z3

This gives CVQR a precise multivariate rank structure: YZ=zY\mid Z=z4 is simultaneously a transport coordinate, a conditional rank, and a latent factor with prescribed marginals (Carlier et al., 2014).

A common misconception is that multivariate quantiles can be defined componentwise without loss. The CVQR framework rejects that simplification: dependence across outcome components is encoded through the joint rank YZ=zY\mid Z=z5 and the convex potential YZ=zY\mid Z=z6, not through separate scalar quantiles. In the scalar case YZ=zY\mid Z=z7, the construction collapses to the usual conditional quantile function, but for YZ=zY\mid Z=z8 the monotone map is intrinsically joint rather than coordinatewise (Carlier et al., 2014).

2. Optimal transport formulation

A defining feature of CVQR is that it embeds Monge–Kantorovich optimal transport at its core. Under finite second moments, the CVQF arises as the solution to a conditional transport problem with quadratic cost. In primal form, the problem may be written as

YZ=zY\mid Z=z9

or equivalently,

Y=QYZ(U,Z)Y = Q_{Y\mid Z}(U,Z)0

The optimizer is Y=QYZ(U,Z)Y = Q_{Y\mid Z}(U,Z)1, and the transport map is the conditional Brenier map

Y=QYZ(U,Z)Y = Q_{Y\mid Z}(U,Z)2

which pushes Y=QYZ(U,Z)Y = Q_{Y\mid Z}(U,Z)3 to Y=QYZ(U,Z)Y = Q_{Y\mid Z}(U,Z)4 (Carlier et al., 2014).

The corresponding dual problem is

Y=QYZ(U,Z)Y = Q_{Y\mid Z}(U,Z)5

subject to

Y=QYZ(U,Z)Y = Q_{Y\mid Z}(U,Z)6

At the optimum, Y=QYZ(U,Z)Y = Q_{Y\mid Z}(U,Z)7 and Y=QYZ(U,Z)Y = Q_{Y\mid Z}(U,Z)8 are convex conjugates in Y=QYZ(U,Z)Y = Q_{Y\mid Z}(U,Z)9 for each UU0, and

UU1

with the gradients acting as mutual inverses almost everywhere (Carlier et al., 2014).

Under injectivity and differentiability, the density relation is governed by conditional Monge–Ampère equations: UU2 and symmetrically,

UU3

These equations tie conditional densities to Jacobians of the forward and inverse maps and formalize the change-of-variables structure induced by the CVQF (Carlier et al., 2014).

This transport-based definition clarifies why scalar quantile regression does not extend directly to multivariate targets. Scalar QR relies on one-dimensional order and pinball loss, whereas CVQR replaces order statistics by monotone transport maps defined as gradients of convex potentials. The monotonicity notion is therefore geometric rather than ordinal (Rosenberg et al., 2022).

3. Linear VQR as the canonical CVQR model

The canonical regression model is the linear CVQF specification, called vector quantile regression. Let UU4 collect known transformations of UU5, including an intercept. A linear CVQF posits

UU6

where UU7 is a UU8 matrix-valued function and, for each UU9, the map YRdY \in \mathbb{R}^d0 is the gradient of a convex function in YRdY \in \mathbb{R}^d1: YRdY \in \mathbb{R}^d2 Under correct specification,

YRdY \in \mathbb{R}^d3

Convexity of YRdY \in \mathbb{R}^d4 enforces monotonicity of YRdY \in \mathbb{R}^d5 in the multivariate sense (Carlier et al., 2014).

The interpretation of YRdY \in \mathbb{R}^d6 is analogous to scalar quantile regression, but generalized to multivariate ranks. As YRdY \in \mathbb{R}^d7 varies over the reference domain, the columns of YRdY \in \mathbb{R}^d8 trace conditional quantile surfaces of YRdY \in \mathbb{R}^d9 given ZRkZ \in \mathbb{R}^k0; variation of ZRkZ \in \mathbb{R}^k1 across ZRkZ \in \mathbb{R}^k2 reveals heterogeneity of covariate effects across the conditional distribution, while cross-component dependence is carried by the joint map ZRkZ \in \mathbb{R}^k3 (Carlier et al., 2014).

As ZRkZ \in \mathbb{R}^k4 becomes richer, the model becomes nonparametric in the sense of series modeling. A sieve specification approximates a smooth convex potential ZRkZ \in \mathbb{R}^k5 with a tensor-product basis,

ZRkZ \in \mathbb{R}^k6

Under smoothness, ZRkZ \in \mathbb{R}^k7 and its gradient uniformly approximate ZRkZ \in \mathbb{R}^k8 and ZRkZ \in \mathbb{R}^k9 as FYZ(,z)F_{Y\mid Z}(\cdot,z)0, providing a nonparametric pathway for CVQF estimation. Convexity can be enforced by construction, for example through positive-definite quadratic forms in FYZ(,z)F_{Y\mid Z}(\cdot,z)1 or convex basis constraints (Carlier et al., 2014).

A frequent misunderstanding is that CVQR estimates individual quantile points independently, as in separate scalar regressions. The framework is explicitly non-local in FYZ(,z)F_{Y\mid Z}(\cdot,z)2: the monotonicity requirement is a global shape restriction, so the map cannot be estimated pointwise without reference to the full transport structure (Carlier et al., 2014).

4. Identification, misspecification, and the univariate connection

Identification in the foundational framework relies on four classes of conditions: a reference rank law with density and convex support; conditional regularity ensuring existence of inverse ranks; finite second moments for the OT variational characterization; and correctness of the linear CVQF model together with convexity in FYZ(,z)F_{Y\mid Z}(\cdot,z)3 (Carlier et al., 2014). For feasible implementation and misspecification, the formulation can be relaxed from conditional independence to mean independence. The relaxed primal becomes

FYZ(,z)F_{Y\mid Z}(\cdot,z)4

When the linear model holds and FYZ(,z)F_{Y\mid Z}(\cdot,z)5 is full rank, this program identifies FYZ(,z)F_{Y\mid Z}(\cdot,z)6 and FYZ(,z)F_{Y\mid Z}(\cdot,z)7 (Carlier et al., 2014).

The mean-independence condition is weaker than conditional independence. With rich FYZ(,z)F_{Y\mid Z}(\cdot,z)8, or saturated FYZ(,z)F_{Y\mid Z}(\cdot,z)9, they coincide; otherwise the relaxed program targets a quasi-linear representation rather than literal conditional independence of ranks and covariates (Carlier et al., 2014). This distinction becomes central in the analysis beyond correct specification (Carlier et al., 2016).

The paper “Vector quantile regression beyond correct specification” shows that even under misspecification, the VQR problem still has a solution and still yields a global representation of dependence between random vectors (Carlier et al., 2016). In that setting, define

FUF_U0

Even when FUF_U1 is not convex, primal-dual optimality implies

FUF_U2

where FUF_U3 is the convex envelope of FUF_U4. The solution therefore replaces a misspecified primitive by its convex envelope and retains a cyclically monotone subgradient representation in FUF_U5 (Carlier et al., 2016).

This suggests a natural interpretation of misspecified CVQR as a best convex, monotone approximation generated by the variational system rather than as a failure of the framework. The statement is inferential in phrasing, but it follows the paper’s “best approximation / projection characterization,” in which the convex envelope delivers the closest convex specification induced by the dual constraints (Carlier et al., 2016).

In the univariate case FUF_U6, CVQR reduces to the Koenker–Bassett setting with an additional global monotonicity requirement. Under correct specification, it coincides with classical quantile regression. Beyond correct specification, the 2016 paper proves that CVQR is equivalent to Koenker–Bassett quantile regression with a global noncrossing constraint; more precisely, the global monotonicity-constrained Koenker–Bassett program has the same value as the CVQR mean-independence correlation maximization (Carlier et al., 2016). This establishes the scalar theory as a special case rather than an analogy.

5. Estimation and computation

At the population level, the relaxed linear VQR dual problem is an infinite-dimensional linear program: FUF_U7 subject to

FUF_U8

where

FUF_U9

At the optimum, uu00, where uu01 defines uu02 (Carlier et al., 2014).

In finite samples, one approximates uu03 by the empirical distribution over uu04 and uu05 by a grid uu06. The discretized primal linear program is

uu07

subject to

uu08

The last constraint enforces mean independence through moment matching. The discretized dual is

uu09

subject to

uu10

Efficient implementation leverages sparsity and standard solvers such as Gurobi; the linear program has uu11 variables and uu12 constraints, where uu13 (Carlier et al., 2014).

Subsequent work emphasizes that exact formulations become computationally prohibitive for moderate target dimension, quantile grid size, or feature dimension. “Fast Nonlinear Vector Quantile Regression” extends VQR beyond linear-in-uu14 parameterizations, introduces vector monotone rearrangement, and proposes fast, GPU-accelerated solvers for linear and nonlinear VQR with fixed memory footprint (Rosenberg et al., 2022). In that paper, the relaxed dual replaces max constraints by a log-sum-exp objective, yielding an unconstrained convex objective for the linear case: uu15 As uu16, the relaxed dual approaches the exact dual and is equivalent to an entropic-regularized primal (Rosenberg et al., 2022).

The same paper defines a nonlinear specification by introducing a learned feature map uu17: uu18 This allows lifting or compressing the covariates before fitting the quantile map (Rosenberg et al., 2022). Because relaxed optimization may slightly violate monotonicity, the paper proposes vector monotone rearrangement (VMR), which projects an estimated map onto the set of monotone maps via an OT problem between uu19 and the estimated quantile values. In one dimension, this reduces to sorting quantiles to remove crossings (Rosenberg et al., 2022).

The computational trade-off remains explicit throughout the literature. Complexity grows with the product of sample size and rank-grid size, and with target dimension through the uu20 quantile grid. GPU acceleration, double mini-batching, and entropic smoothing mitigate but do not remove the curse of dimensionality in uu21 (Rosenberg et al., 2022).

6. Applications, extensions, and scope

The foundational empirical application is multiple Engel curve estimation with household expenditure data and a bivariate response consisting of food and housing/heating expenditures. In the one-dimensional case, regressing each component separately via scalar VQR yields quantile curves very close to classical quantile regression, with minimal crossing issues. In the two-dimensional VQR specification,

uu22

with uu23 independent of uu24 under correct specification. The fitted surfaces reveal strong own-propensity effects and significant negative cross-covariation at median income, indicating local substitutability between food and housing in that region—an effect not available from separate scalar regressions (Carlier et al., 2014).

The framework has also been extended to non-Euclidean outcome spaces. “Vector Quantile Regression on Manifolds” defines manifold conditional vector quantile functions using Riemannian optimal transport with quadratic geodesic cost and uu25-concave potentials. In that setting, the conditional map is

uu26

and for each uu27, it pushes a manifold base law uu28 to the conditional law of uu29 (Pegoraro et al., 2023). The dual formulation averages the OT losses over uu30, uses bounded continuous uu31-concave functions in uu32, and supports conditional quantile estimation, confidence sets, and likelihood computation on manifolds such as uu33 and uu34 (Pegoraro et al., 2023).

The manifold paper also derives a conditional likelihood formula through the inverse map,

uu35

and defines conditional uu36-contours by pushing forward base contours centered at a Fréchet mean (Pegoraro et al., 2023). This widens the scope of CVQR from Euclidean multivariate responses to structured geometric domains.

The limitations stated across the literature are consistent. The reference law must be non-atomic with convex support in the Euclidean theory, or appropriate manifold regularity in the geometric theory. Conditional density assumptions are needed for inverse ranks and Monge–Ampère relations, though existence of the forward CVQF does not strictly require continuity of uu37. Monotonicity is a global constraint, and relaxed dual solvers may violate it unless corrective procedures such as VMR or involution regularization are used. Alternative reference distributions are possible, but they alter the geometry of the rank space [(Carlier et al., 2014); (Rosenberg et al., 2022); (Pegoraro et al., 2023)].

Taken together, these developments define CVQR as a transport-based theory of conditional multivariate quantiles. Its distinctive commitments are deterministic coupling through uu38, monotonicity via convex or uu39-concave potentials, and representation of the full conditional distribution rather than selected marginals. In the foundational Euclidean setting, the formal model is VQR; under misspecification, the theory survives through convex-envelope representations; and in recent extensions, the same logic supports nonlinear solvers, monotonicity repair, and regression on manifolds [(Carlier et al., 2014); (Carlier et al., 2016); (Rosenberg et al., 2022); (Pegoraro et al., 2023)].

Definition Search Book Streamline Icon: https://streamlinehq.com
References (4)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Conditional Vector Quantile Regression (CVQR).