---
title: Conditional Vector Quantile Regression
url: https://www.emergentmind.com/topics/conditional-vector-quantile-regression-cvqr
type: topic
---

# Conditional Vector Quantile Regression

Searching arXiv for recent and foundational work on conditional vector quantile regression to ground the article.
Conditional Vector Quantile Regression (CVQR) is a conditional multivariate quantile framework that extends scalar quantile regression to vector-valued outcomes by representing conditional distributions through monotone transport maps from a reference rank law to the law of the response given covariates. In the foundational formulation, the central object is the conditional vector quantile function (CVQF), a map \(Q_{Y\mid Z}(u,z)\) that is monotone in the multivariate sense of being the gradient of a convex function in \(u\), pushes a non-atomic reference distribution \(F_U\) onto the conditional law of \(Y\mid Z=z\), and yields the strong representation \(Y = Q_{Y\mid Z}(U,Z)\) almost surely for an appropriate latent rank vector \(U\) [1406.4643]. The term “CVQR” is not the authors’ formal label in the foundational paper; the formal regression model is “vector quantile regression” (VQR), which operationalizes the conditional framework through linear or sieve-type specifications of the CVQF [1406.4643].

## 1. Definition and geometric structure

Let \(Y \in \mathbb{R}^d\) and \(Z \in \mathbb{R}^k\) be random vectors, with conditional distribution \(F_{Y\mid Z}(\cdot,z)\). Fix a reference, non-atomic distribution \(F_U\) on \(\mathbb{R}^d\) with density \(f_U\) and convex support, for example \(U \sim U([0,1]^d)\) or \(U \sim N(0,I_d)\). A conditional vector quantile function is a measurable map
\[
(u,z) \mapsto Q_{Y\mid Z}(u,z)\in \mathbb{R}^d
\]
such that for each \(z\), the map \(u \mapsto Q_{Y\mid Z}(u,z)\) is monotone in the sense of being the gradient of a convex function,
\[
Q_{Y\mid Z}(u,z)=\nabla_u \phi(u,z),
\]
for some convex potential \(\phi(\cdot,z)\). Equivalently, for all \(u,\bar u\) in the support of \(F_U\),
\[
\big(Q_{Y\mid Z}(u,z)-Q_{Y\mid Z}(\bar u,z)\big)^\top (u-\bar u)\ge 0.
\]
With \(U\sim F_U\), the pushforward condition is
\[
Q_{Y\mid Z}(U,z)\overset{d}{=}Y\mid Z=z,
\]
and the strong representation is
\[
Y = Q_{Y\mid Z}(U,Z)\quad \text{almost surely}.
\]
Under the non-atomicity and density condition on \(F_U\), the CVQF exists and is unique \(F_U\)-almost everywhere for each \(z\), as a conditional Brenier map [1406.4643].

When each conditional law \(F_{Y\mid Z}(\cdot,z)\) admits a density, there is also a conditional inverse or rank map,
\[
Q^{-1}_{Y\mid Z}(y,z)=\nabla_y \phi^*(y,z),
\]
where \(\phi^*\) is the Legendre transform of \(\phi\). It satisfies
\[
Q^{-1}_{Y\mid Z}(Q_{Y\mid Z}(u,z),z)=u \quad \text{for }F_U\text{-a.e. }u,
\]
and
\[
U = Q^{-1}_{Y\mid Z}(Y,Z),\qquad U\mid Z \sim F_U.
\]
This gives CVQR a precise multivariate rank structure: \(U\) is simultaneously a transport coordinate, a conditional rank, and a latent factor with prescribed marginals [1406.4643].

A common misconception is that multivariate quantiles can be defined componentwise without loss. The CVQR framework rejects that simplification: dependence across outcome components is encoded through the joint rank \(u\) and the convex potential \(\phi\), not through separate scalar quantiles. In the scalar case \(d=1\), the construction collapses to the usual conditional quantile function, but for \(d>1\) the monotone map is intrinsically joint rather than coordinatewise [1406.4643].

## 2. Optimal transport formulation

A defining feature of CVQR is that it embeds Monge–Kantorovich optimal transport at its core. Under finite second moments, the CVQF arises as the solution to a conditional transport problem with quadratic cost. In primal form, the problem may be written as
\[
\min_V E\|Y-V\|^2 \quad \text{subject to } V\mid Z \sim F_U,
\]
or equivalently,
\[
\max_V E(V^\top Y)\quad \text{subject to } V\mid Z \sim F_U.
\]
The optimizer is \(V=U=Q^{-1}_{Y\mid Z}(Y,Z)\), and the transport map is the conditional Brenier map
\[
T(u,z):=\nabla_u \phi(u,z)=Q_{Y\mid Z}(u,z),
\]
which pushes \(F_U\) to \(F_{Y\mid Z}(\cdot,z)\) [1406.4643].

The corresponding dual problem is
\[
\min_{\phi,\phi^*} E[\phi(U,Z)] + E[\phi^*(Y,Z)]
\]
subject to
\[
\phi(u,z)+\phi^*(y,z)\ge u^\top y \quad \text{for all }(u,y,z).
\]
At the optimum, \(\phi\) and \(\phi^*\) are convex conjugates in \((u,y)\) for each \(z\), and
\[
Q_{Y\mid Z}(u,z)=\nabla_u \phi(u,z), \qquad Q^{-1}_{Y\mid Z}(y,z)=\nabla_y \phi^*(y,z),
\]
with the gradients acting as mutual inverses almost everywhere [1406.4643].

Under injectivity and differentiability, the density relation is governed by conditional Monge–Ampère equations:
\[
f_U(u)=f_{Y\mid Z}(Q_{Y\mid Z}(u,z),z)\det[D_uQ_{Y\mid Z}(u,z)],
\]
and symmetrically,
\[
f_{Y\mid Z}(y,z)=f_U(Q^{-1}_{Y\mid Z}(y,z))\det[D_yQ^{-1}_{Y\mid Z}(y,z)].
\]
These equations tie conditional densities to Jacobians of the forward and inverse maps and formalize the change-of-variables structure induced by the CVQF [1406.4643].

This transport-based definition clarifies why scalar quantile regression does not extend directly to multivariate targets. Scalar QR relies on one-dimensional order and pinball loss, whereas CVQR replaces order statistics by monotone transport maps defined as gradients of convex potentials. The monotonicity notion is therefore geometric rather than ordinal [2205.14977].

## 3. Linear VQR as the canonical CVQR model

The canonical regression model is the linear CVQF specification, called vector quantile regression. Let \(X=f(Z)\in \mathbb{R}^p\) collect known transformations of \(Z\), including an intercept. A linear CVQF posits
\[
Q_{Y\mid Z}(u,z)=Q_{Y\mid X}(u,x)=\beta(u)^\top x,
\]
where \(u\mapsto \beta(u)\) is a \(p\times d\) matrix-valued function and, for each \(x\), the map \(u\mapsto \beta(u)^\top x\) is the gradient of a convex function in \(u\):
\[
\beta(u)^\top x=\nabla_u \Phi_x(u), \qquad \Phi_x(u):=B(u)^\top x.
\]
Under correct specification,
\[
Y=\beta(U)^\top X \quad \text{almost surely}, \qquad U\mid X\sim F_U.
\]
Convexity of \(\Phi_x(\cdot)\) enforces monotonicity of \(u\mapsto \beta(u)^\top x\) in the multivariate sense [1406.4643].

The interpretation of \(\beta(u)\) is analogous to scalar quantile regression, but generalized to multivariate ranks. As \(u\) varies over the reference domain, the columns of \(\beta(u)\) trace conditional quantile surfaces of \(Y\) given \(Z\); variation of \(\beta(u)\) across \(u\) reveals heterogeneity of covariate effects across the conditional distribution, while cross-component dependence is carried by the joint map \(u\mapsto \beta(u)^\top x\) [1406.4643].

As \(f(Z)\) becomes richer, the model becomes nonparametric in the sense of series modeling. A sieve specification approximates a smooth convex potential \(\phi(u,z)\) with a tensor-product basis,
\[
\Phi^{J,L}(u,z)=\sum_{j=1}^J\sum_{l=1}^L \theta_{jl} q_j(u)f_l(z)
= B_0^L(u)^\top f^L(z).
\]
Under smoothness, \(\Phi^{J,L}\) and its gradient uniformly approximate \(\phi\) and \(Q_{Y\mid Z}\) as \(J,L\to\infty\), providing a nonparametric pathway for CVQF estimation. Convexity can be enforced by construction, for example through positive-definite quadratic forms in \(u\) or convex basis constraints [1406.4643].

A frequent misunderstanding is that CVQR estimates individual quantile points independently, as in separate scalar regressions. The framework is explicitly non-local in \(u\): the monotonicity requirement is a global shape restriction, so the map cannot be estimated pointwise without reference to the full transport structure [1406.4643].

## 4. Identification, misspecification, and the univariate connection

Identification in the foundational framework relies on four classes of conditions: a reference rank law with density and convex support; conditional regularity ensuring existence of inverse ranks; finite second moments for the OT variational characterization; and correctness of the linear CVQF model together with convexity in \(u\) [1406.4643]. For feasible implementation and misspecification, the formulation can be relaxed from conditional independence to mean independence. The relaxed primal becomes
\[
\max_V E(V^\top Y) \quad \text{subject to } V\sim F_U,\; E(X\mid V)=E(X).
\]
When the linear model holds and \(E(XX^\top)\) is full rank, this program identifies \(U\) and \(\beta(u)\) [1406.4643].

The mean-independence condition is weaker than conditional independence. With rich \(f(Z)\), or saturated \(X\), they coincide; otherwise the relaxed program targets a quasi-linear representation rather than literal conditional independence of ranks and covariates [1406.4643]. This distinction becomes central in the analysis beyond correct specification [1610.06833].

The paper “Vector quantile regression beyond correct specification” shows that even under misspecification, the VQR problem still has a solution and still yields a global representation of dependence between random vectors [1610.06833]. In that setting, define
\[
\phi_x(t):=\phi(t)+b(t)\cdot x.
\]
Even when \(\phi_x\) is not convex, primal-dual optimality implies
\[
\phi_x(U)=\phi_x^{**}(U), \qquad Y\in \partial (\phi_x^{**})(U)\quad \text{a.s.},
\]
where \(\phi_x^{**}\) is the convex envelope of \(\phi_x\). The solution therefore replaces a misspecified primitive by its convex envelope and retains a cyclically monotone subgradient representation in \(u\) [1610.06833].

This suggests a natural interpretation of misspecified CVQR as a best convex, monotone approximation generated by the variational system rather than as a failure of the framework. The statement is inferential in phrasing, but it follows the paper’s “best approximation / projection characterization,” in which the convex envelope delivers the closest convex specification induced by the dual constraints [1610.06833].

In the univariate case \(d=1\), CVQR reduces to the Koenker–Bassett setting with an additional global monotonicity requirement. Under correct specification, it coincides with classical quantile regression. Beyond correct specification, the 2016 paper proves that CVQR is equivalent to Koenker–Bassett quantile regression with a global noncrossing constraint; more precisely, the global monotonicity-constrained Koenker–Bassett program has the same value as the CVQR mean-independence correlation maximization [1610.06833]. This establishes the scalar theory as a special case rather than an analogy.

## 5. Estimation and computation

At the population level, the relaxed linear VQR dual problem is an infinite-dimensional linear program:
\[
\min_{\psi,b} E[\psi(X,Y)] + E[b(V)^\top E(X)]
\]
subject to
\[
\psi(x,y)+b(u)^\top x \ge u^\top y \quad \text{for all }(x,y,u),
\]
where
\[
\psi(x,y):=\sup_{u\in \mathcal U}\{u^\top y - b(u)^\top x\}.
\]
At the optimum, \(b(u)=B(u)\), where \(B(u)\) defines \(\Phi_x(u)=B(u)^\top x\) [1406.4643].

In finite samples, one approximates \(F_{YX}\) by the empirical distribution over \(\{(y_i,x_i)\}_{i=1}^n\) and \(F_U\) by a grid \(\{u_k\}_{k=1}^m\). The discretized primal linear program is
\[
\max_{\Pi\ge 0} \operatorname{Tr}(U^\top \Pi Y)
\]
subject to
\[
\Pi^\top 1_m = w,\qquad \Pi 1_n=\nu,\qquad \Pi X = \Pi^\top X.
\]
The last constraint enforces mean independence through moment matching. The discretized dual is
\[
\min_{\alpha,b} w^\top \alpha + (X^\top b)^\top \nu
\]
subject to
\[
1_m^\top \alpha + Xb^\top \ge YU^\top.
\]
Efficient implementation leverages sparsity and standard solvers such as Gurobi; the linear program has \(mn\) variables and \(mr+n\) constraints, where \(r=p\) [1406.4643].

Subsequent work emphasizes that exact formulations become computationally prohibitive for moderate target dimension, quantile grid size, or feature dimension. “Fast Nonlinear Vector Quantile Regression” extends VQR beyond linear-in-\(x\) parameterizations, introduces vector monotone rearrangement, and proposes fast, GPU-accelerated solvers for linear and nonlinear VQR with fixed memory footprint [2205.14977]. In that paper, the relaxed dual replaces max constraints by a log-sum-exp objective, yielding an unconstrained convex objective for the linear case:
\[
\min_{\psi,\beta}\; \psi^\top \nu + \mathrm{tr}(\beta^\top \bar X)
+ \varepsilon \sum_{i=1}^{T^d}\mu_i \log\!\left(\sum_{j=1}^N \exp\!\left[\frac{1}{\varepsilon}\left(\langle u_i,y_j\rangle - \beta_i^\top x_j - \psi_j\right)\right]\right).
\]
As \(\varepsilon\downarrow 0\), the relaxed dual approaches the exact dual and is equivalent to an entropic-regularized primal [2205.14977].

The same paper defines a nonlinear specification by introducing a learned feature map \(g_\theta\):
\[
\widehat Q^{NL}_{Y\mid X}(u;x)=\nabla_u\{\beta(u)^\top g_\theta(x)+\varphi(u)\}.
\]
This allows lifting or compressing the covariates before fitting the quantile map [2205.14977]. Because relaxed optimization may slightly violate monotonicity, the paper proposes vector monotone rearrangement (VMR), which projects an estimated map onto the set of monotone maps via an OT problem between \(U\) and the estimated quantile values. In one dimension, this reduces to sorting quantiles to remove crossings [2205.14977].

The computational trade-off remains explicit throughout the literature. Complexity grows with the product of sample size and rank-grid size, and with target dimension through the \(T^d\) quantile grid. GPU acceleration, double mini-batching, and entropic smoothing mitigate but do not remove the curse of dimensionality in \(d\) [2205.14977].

## 6. Applications, extensions, and scope

The foundational empirical application is multiple Engel curve estimation with household expenditure data and a bivariate response consisting of food and housing/heating expenditures. In the one-dimensional case, regressing each component separately via scalar VQR yields quantile curves very close to classical quantile regression, with minimal crossing issues. In the two-dimensional VQR specification,
\[
Y_1 = \frac{\partial \Phi_x(u)}{\partial u_1} = \beta_1(u)^\top x,\qquad
Y_2 = \frac{\partial \Phi_x(u)}{\partial u_2} = \beta_2(u)^\top x,
\]
with \(U=(U_1,U_2)\sim U([0,1]^2)\) independent of \(X\) under correct specification. The fitted surfaces reveal strong own-propensity effects and significant negative cross-covariation at median income, indicating local substitutability between food and housing in that region—an effect not available from separate scalar regressions [1406.4643].

The framework has also been extended to non-Euclidean outcome spaces. “Vector Quantile Regression on Manifolds” defines manifold conditional vector quantile functions using Riemannian optimal transport with quadratic geodesic cost and \(c\)-concave potentials. In that setting, the conditional map is
\[
Q_{Y\mid X}(u,x)=\operatorname{exp}_u[-\nabla_u \varphi(u,x)],
\]
and for each \(x\), it pushes a manifold base law \(\mu\) to the conditional law of \(Y\mid X=x\) [2307.01037]. The dual formulation averages the OT losses over \(x\), uses bounded continuous \(c\)-concave functions in \(u\), and supports conditional quantile estimation, confidence sets, and likelihood computation on manifolds such as \(S^2\) and \(\mathcal T^2\) [2307.01037].

The manifold paper also derives a conditional likelihood formula through the inverse map,
\[
p_{Y\mid X}(y;x)=\frac{1}{V(\mathcal M)}\cdot \left|\nabla_y Q^{-1}_{Y\mid X}(y;x)\right|,
\]
and defines conditional \(\tau\)-contours by pushing forward base contours centered at a Fréchet mean [2307.01037]. This widens the scope of CVQR from Euclidean multivariate responses to structured geometric domains.

The limitations stated across the literature are consistent. The reference law must be non-atomic with convex support in the Euclidean theory, or appropriate manifold regularity in the geometric theory. Conditional density assumptions are needed for inverse ranks and Monge–Ampère relations, though existence of the forward CVQF does not strictly require continuity of \(Y\mid Z\). Monotonicity is a global constraint, and relaxed dual solvers may violate it unless corrective procedures such as VMR or involution regularization are used. Alternative reference distributions are possible, but they alter the geometry of the rank space [1406.4643; 2205.14977; 2307.01037].

Taken together, these developments define CVQR as a transport-based theory of conditional multivariate quantiles. Its distinctive commitments are deterministic coupling through \(Y=Q_{Y\mid Z}(U,Z)\), monotonicity via convex or \(c\)-concave potentials, and representation of the full conditional distribution rather than selected marginals. In the foundational Euclidean setting, the formal model is VQR; under misspecification, the theory survives through convex-envelope representations; and in recent extensions, the same logic supports nonlinear solvers, monotonicity repair, and regression on manifolds [1406.4643; 1610.06833; 2205.14977; 2307.01037].

Source: https://www.emergentmind.com/topics/conditional-vector-quantile-regression-cvqr