---
title: Convolution-Smoothed Quantile Regression
url: https://www.emergentmind.com/topics/convolution-smoothed-quantile-regression
type: topic
---

# Convolution-Smoothed Quantile Regression

Convolution-smoothed quantile regression is a family of methods that replaces the non-differentiable quantile check loss, or an equivalent discontinuous quantile score or moment condition, by a kernel-convolution surrogate. In the linear quantile model \(Q(\tau\mid x)=x^\top\beta(\tau)\), a canonical construction writes the standard empirical objective as \(\widehat R(b;\tau)=n^{-1}\sum_{i=1}^n \rho_\tau(Y_i-X_i^\top b)\) and then replaces it by a smoothed criterion \(\widehat R_h(b;\tau)=\int \rho_\tau(t)\,d\widehat F_h(t;b)=n^{-1}\sum_{i=1}^n (\rho_\tau*k_h)(e_i(b))\), where \(e_i(b)=Y_i-X_i^\top b\), \(k_h(u)=k(u/h)/h\), and \(*\) denotes convolution. This construction appears as “smoothing the objective function, rather than only the indicator on the check function,” and it has been used to obtain differentiable quantile paths, bandwidth-uniform asymptotics, scalable penalized algorithms, quantile-density estimation, and a range of extensions in panel, functional, transfer-learning, and neural-network settings [1905.08535] [2205.02432] [2605.06265].

## 1. Formal construction and core variants

In standard linear quantile regression, the empirical criterion is
\[
\widehat R(b;\tau)=\frac1n\sum_{i=1}^n \rho_\tau(Y_i-X_i^\top b),
\qquad
\rho_\tau(u)=u\bigl[\tau-\mathbf 1(u<0)\bigr].
\]
The central difficulty is that \(\rho_\tau\) is not differentiable at zero, so the map \(b\mapsto \widehat R(b;\tau)\) is non-smooth, and the estimator path in \(\tau\) can have jumps [1905.08535] [2412.05736].

A defining convolution-type construction smooths the residual distribution before integrating the check loss. With residual cdf \(\widehat F(\cdot;b)\), smoothed density
\[
\widehat f_h(v;b):=\frac1n\sum_{i=1}^n k_h\bigl(v-e_i(b)\bigr),
\]
and corresponding smoothed cdf \(\widehat F_h(\cdot;b)\), the objective becomes
\[
\widehat R_h(b;\tau):=\int \rho_\tau(t)\,d\widehat F_h(t;b)
=\int \rho_\tau(t)\,\widehat f_h(t;b)\,dt
=\frac1n\sum_{i=1}^n (\rho_\tau*k_h)\bigl(e_i(b)\bigr),
\]
with estimator
\[
\widehat\beta_h(\tau):=\arg\min_{b\in\mathbb R^d}\widehat R_h(b;\tau).
\]
This is the formulation studied in “Smoothing quantile regressions” and used as the quantile-regression stage in “Convolution Mode Regression” [1905.08535] [2412.05736].

The resulting gradient and Hessian take explicit forms. Writing \(K(t)=\int_{-\infty}^t k(v)\,dv\), the smoothed QR derivatives are
\[
\widehat R_h^{(1)}(b;\tau)=\frac1n\sum_{i=1}^n X_i\left[K\!\left(\frac{X_i^\top b-Y_i}{h}\right)-\tau\right],
\]
and
\[
\widehat R_h^{(2)}(b;\tau)=\frac1n\sum_{i=1}^n X_iX_i^\top k_h(X_i^\top b-Y_i).
\]
Thus the discontinuous indicator in the ordinary QR score is replaced by the smooth cdf-like term \(K((X_i^\top b-Y_i)/h)\), and the Hessian becomes a kernel-weighted second-moment matrix [2412.05736].

A distinct but closely related line smooths the score or moment condition directly rather than starting from the full objective. In smoothed GMM for quantile models, the unsmoothed moment
\[
g_i^u(\beta,\tau)=Z_i\bigl[\mathbf 1\{\Lambda(Y_i,X_i,\beta)\le 0\}-\tau\bigr]
\]
is replaced by
\[
g_{ni}(\beta,\tau)=Z_i\left[\tilde I\!\left(-\frac{\Lambda(Y_i,X_i,\beta)}{h_n}\right)-\tau\right],
\]
with \(\tilde I'(\cdot)\) a kernel of order \(r\ge 2\) [1707.03436]. In panel quantile regression with fixed effects, the smoothed second-step loss replaces \(\mathbf 1\{u\le 0\}\) by \(K(u/h)\) and uses
\[
\varrho_\tau(u)=[\tau-K(u/h)]u
\]
as a smooth approximation to the check loss [1911.04729]. In limited dependent variable models, the same logic appears as the substitution
\[
I\{x_i^\top b\ge 0\}\leadsto K(x_i^\top b/h_n),
\]
with the cumulative distribution of a Gaussian kernel used in implementation [2112.06822].

These constructions are not treated as interchangeable in the literature. One recurring distinction is between smoothing the entire objective via a smoothed residual cdf and smoothing only the indicator inside the check loss. The former is explicitly presented as different from Horowitz-style smoothing and as asymptotically preferable in several second-order respects [1905.08535].

## 2. Differentiability, quantile-density estimation, and uniform asymptotics

A central technical consequence of convolution smoothing is differentiability of the fitted quantile path in the quantile index. Because \(\widehat R_h^{(1)}(\widehat\beta_h(\tau);\tau)=0\) and \(\widehat R_h\) is smooth, the Implicit Function Theorem yields
\[
\widehat\beta_h^{(1)}(\tau)
=
\frac{\partial \widehat\beta_h(\tau)}{\partial \tau}
=
\bigl[\widehat R_h^{(2)}(\widehat\beta_h(\tau);\tau)\bigr]^{-1}\bar X,
\]
with \(\bar X=n^{-1}\sum_i X_i\) [1905.08535] [2412.05736]. This contrasts with ordinary QR, whose estimator path is described as piecewise linear, non-differentiable, or step-like in \(\tau\).

This asymptotic differentiability is exploited to estimate the conditional quantile density
\[
q(\tau\mid x)=\frac{\partial Q(\tau\mid x)}{\partial \tau}
=x^\top\beta^{(1)}(\tau)
=\frac{1}{f(Q(\tau\mid x)\mid x)}.
\]
The smoothed estimator is
\[
\widehat q_h(\tau\mid x)=x^\top \widehat\beta_h^{(1)}(\tau),
\]
which, in the linear QR model, avoids the curse of dimensionality because it leverages the parametric structure rather than smoothing in the full covariate space [1905.08535]. This derivative-based construction is the mechanism behind the mode-regression estimator in “Convolution Mode Regression,” where the mode is recovered by minimizing the quantile density, or equivalently maximizing the negative quantile density (“sparsity”) [2412.05736].

The asymptotic theory developed for smoothed QR is explicitly uniform in both the quantile index and the bandwidth interval. In “Smoothing quantile regressions,” the smoothing bias satisfies
\[
\beta_h(\tau)=\beta(\tau)-h^{s+1}B(\tau)+o(h^{s+1})
\]
uniformly over \((\tau,h)\), and the paper states that the asymptotic theory holds uniformly with respect to the bandwidth and quantile level [1905.08535]. The same paper states that the smoothed estimator has a smaller Bahadur–Kiefer remainder and lower asymptotic mean squared error than the standard estimator, and that smoothing reduces variance by an \(O(h)\) correction:
\[
\Sigma_h(\tau)=\Sigma(\tau)-c_k h D(\tau)^{-1}+o(h),
\qquad
c_k=2\int_0^\infty K(y)\bigl[1-K(y)\bigr]\,dy>0
\]
[1905.08535].

In the mode-regression application, the smoothed Hessian estimator obeys the uniform bound
\[
\left\Vert \widehat D_h(\tau)-D(\tau)\right\Vert
=
o(h^j)+O_P\!\left(\sqrt{\log n/(nh)}\right),
\]
uniformly in \(\tau\in[\alpha,1-\alpha]\) and \(h\in[\underline h_n,\bar h_n]\). This yields
\[
\vert \widehat s_{x,h}(\tau)-s_x(\tau)\vert
=
o(h)+O_P\!\left(\sqrt{\log n/(nh)}\right)
\]
uniformly over \(\tau\), \(x\), and \(h\), and then
\[
\widehat{\tau}_{x,h}
=
\tau_x+o(h^{1/2})
+
O_P\!\left(\left(\frac{\log n}{nh}\right)^{1/4}\right),
\]
\[
\widehat m_h(x)
=
m(x)+o(h^{1/2})
+
O_P\!\left(\left(\frac{\log n}{nh}\right)^{1/4}\right)
\]
uniformly for \(x\in\mathcal X\) and \(h\in[\underline h_n,\bar h_n]\) [2412.05736]. The same paper emphasizes as an advantage that the estimator is uniform not only over design points \(x\) but also over the bandwidth interval.

## 3. Penalization, convex surrogates, and large-scale optimization

Convolution smoothing is especially prominent in high-dimensional penalized quantile regression because it turns the non-differentiable check loss into a smooth convex surrogate. In the generic penalized formulation,
\[
\min_{\beta\in\mathbb R^p}
\left\{
Q(\beta)+P(\beta)
\right\},
\qquad
Q(\beta)=\frac1n\sum_{i=1}^n \ell_{h,\tau}(y_i-x_i^\top\beta),
\]
the smoothed loss
\[
\ell_{h,\tau}(u)=\frac1h\int \rho_\tau(v)\,K\!\left(\frac{v-u}{h}\right)\,dv
\]
is convex, twice differentiable, and locally strongly convex [2205.02432]. The resulting objective supports first-order smooth optimization rather than linear programming or subgradient methods.

A unified algorithmic template for penalized convolution-smoothed QR is Local Adaptive Majorize-Minimization (LAMM). Given current iterate \(\hat\beta^{k-1}\), LAMM constructs the quadratic majorizer
\[
F(\beta\mid \phi_k,\hat\beta^{k-1})
=
Q(\hat\beta^{k-1})
+
\langle \nabla Q(\hat\beta^{k-1}),\beta-\hat\beta^{k-1}\rangle
+
\frac{\phi_k}{2}\|\beta-\hat\beta^{k-1}\|_2^2
\]
and updates
\[
\hat\beta^k
=
\arg\min_\beta\{F(\beta\mid \phi_k,\hat\beta^{k-1})+P(\beta)\}.
\]
The paper gives closed-form proximal updates for weighted lasso, elastic net, group lasso, and sparse group lasso, and states that each LAMM iteration is dominated by matrix-vector multiplication with per-iteration cost \(O(np)\) [2205.02432].

The same smooth-plus-penalty structure underlies concave-regularized high-dimensional QR. “High-Dimensional Quantile Regression: Convolution Smoothing and Concave Regularization” studies
\[
\hat Q_h(\beta)=\frac1n\sum_{i=1}^n \ell_h(y_i-x_i^\top\beta),
\qquad
\ell_h=\rho_\tau*K_h,
\]
with gradient and Hessian
\[
\nabla \hat Q_h(\beta)
=
\frac1n\sum_{i=1}^n
\left\{
\bar K\!\left(-\frac{r_i(\beta)}{h}\right)-\tau
\right\}x_i,
\]
\[
\nabla^2 \hat Q_h(\beta)
=
\frac1n\sum_{i=1}^n
K_h(-r_i(\beta))x_ix_i^\top.
\]
This paper proves that the smoothed empirical loss is twice continuously differentiable and locally strongly convex with high probability, and that iteratively reweighted \(\ell_1\)-penalized smoothed QR achieves the optimal rate of convergence, the oracle rate, and the strong oracle property under an almost necessary and sufficient minimum signal strength condition [2109.05640].

Composite quantile regression admits the same treatment. In high-dimensional CQR, the empirical composite loss
\[
\hat Q(\alpha,\beta)=\frac{1}{nq}\sum_{i=1}^n\sum_{k=1}^q \rho_{\tau_k}(y_i-\alpha_k-x_i^\top\beta)
\]
is replaced by
\[
\hat Q_h(\alpha,\beta)=\frac1{nq}\sum_{i=1}^n\sum_{k=1}^q \ell_{h,k}(y_i-\alpha_k-x_i^\top\beta),
\qquad
\ell_{h,k}=\rho_{\tau_k}*K_h.
\]
The smoothed composite loss is convex, twice continuously differentiable, and locally strongly convex with high probability, and a gradient-based majorize-minimization algorithm substantially improves computational efficiency over ADMM without compromising statistical performance [2208.09817].

For practical implementation, the package `conquer` is repeatedly used as the software realization of convolution-smoothed quantile regression. In penalized settings it is implemented for several convex penalties [2205.02432], and in the mode-regression application the estimation of \(\widehat\beta_h(\tau)\) “was done via the package \textsf{conquer}” [2412.05736].

## 4. Functional, panel, and limited-dependent-variable extensions

Convolution smoothing has been extended beyond finite-dimensional exogenous linear QR to settings where the unsmoothed problem is especially difficult.

In panel quantile regression with fixed effects, a two-step estimator replaces Canay’s unsmoothed second step by the smoothed objective
\[
\hat\beta(\tau)
=
\arg\min_\beta
\sum_{i=1}^N\sum_{t=1}^T
\left[
\tau-K\!\left(\frac{Y_{it}-\beta^\top W_{it}-\hat\alpha_i}{h}\right)
\right]
\left(Y_{it}-\beta^\top W_{it}-\hat\alpha_i\right).
\]
The paper derives explicit derivatives
\[
\varrho^{(1)}(u)=\tau-K(u/h)+k(u/h)\frac{u}{h},
\]
\[
\varrho^{(2)}(u)=2k(u/h)\frac1h+k^{(1)}(u/h)\frac{u}{h^2},
\]
and establishes
\[
\sqrt{NT}(\hat\beta(\tau)-\beta(\tau))
\overset d\longrightarrow
\mathcal N\!\left(\kappa b,\Sigma^{-1}\Omega\Sigma^{-1}\right),
\]
with an explicit incidental-parameter bias formula and both analytical and split-panel jackknife corrections [1911.04729]. Here smoothing is not merely computational; it is used to derive the higher-order expansion exposing the fixed-effect bias.

In functional linear quantile regression, kernel convolution is used to smooth the loss under an RKHS formulation. The smoothed empirical risk is
\[
\hat R_h(\theta;\tau)
=
\int \rho_\tau(u)\,d\hat F_h(u;\theta)
=
\frac1n\sum_{i=1}^n \ell_h(\epsilon_i(\theta);\tau),
\]
with representer-theorem reduction to finite dimensions and Fréchet derivatives
\[
S_{n,h\lambda}(\theta)\Delta\theta
=
\frac1n\sum_{i=1}^n
\left[
K\!\left(\frac{-\epsilon_i(\theta)}{h}\right)-\tau
\right]
\langle R_{X_i},\Delta\theta\rangle
+
\lambda J(\beta,\Delta\beta).
\]
The paper proves existence and uniqueness via the Banach fixed-point theorem, establishes a convergence rate, a functional Bahadur representation, pointwise asymptotic normality, and weak convergence of the estimated coefficient function [2202.11747].

A more elaborate functional extension is CLoSE, “Convolution-smoothing based Locally Sparse Estimation,” for scalar-on-function quantile regression with scalar and functional covariates. It uses
\[
\mathcal L^*(\alpha_\tau,\beta_\tau,\gamma_l,\lambda_l)
=
\frac1n\sum_{i=1}^n
(\rho_\tau*\mathcal K_h)\!\left(
Y_i-Z_i^\top\alpha_\tau-\int X_i^\top(t)\beta_\tau(t)\,dt
\right)
+
\text{roughness penalty}
+
\text{fSCAD},
\]
with gradient and Hessian driven by
\[
G_h(\cdot)-\tau
\quad\text{and}\quad
\mathcal K_h(\cdot),
\]
respectively. The method is designed to select significant functional covariates, identify nonzero regions of functional coefficients, and estimate those coefficients in one step; the paper establishes functional oracle properties, asymptotic normality, simultaneous confidence bands, and consistency of a split wild bootstrap [2512.01341].

In limited dependent variable models, smoothing is used to regularize threshold indicators arising in censored and binary QR. The Stata command `ldvqreg` replaces
\[
I\{x_i^\top b\ge 0\}
\quad\text{by}\quad
K(x_i^\top b/h_n),
\]
uses the cumulative distribution of a Gaussian kernel as the smoothing function, and adopts the rule
\[
h_n=0.9\hat\sigma_u n^{-1/5}
\]
for bandwidth selection [2112.06822]. The same paper reports that in binary examples the smoothed estimator “runs smoothly” where standard QR alternatives exhibit convergence failures.

## 5. Related smoothing paradigms and non-crossing quantile processes

The convolution-smoothed literature is explicit that different smoothing strategies target different objects and are not equivalent. One major contrast is between “smooth then estimate” and “estimate then smooth.” The former smooths the criterion before optimization, as in
\[
\widehat\beta_h(\tau)=\arg\min_b \widehat R_h(b;\tau),
\]
while the latter first computes ordinary QR at a grid of \(\tau\)-values and only then smooths the estimated quantile curve. “Convolution Mode Regression” states this contrast directly and presents the “smooth then estimate” route as technically advantageous for obtaining uniformity in bandwidth [2412.05736].

A second contrast is between smoothing the entire objective and smoothing only the indicator in the check function. “Smoothing quantile regressions” writes Horowitz’s smoothed objective as
\[
\widehat{\mathfrak R}_h(b;\tau)
=
\frac1n\sum_{i=1}^n e_i(b)\Big[\tau-K\!\left(-\frac{e_i(b)}{h}\right)\Big],
\]
whose derivative differs from the full-objective smoothing score by the extra term
\[
-\frac1n\sum_{i=1}^n X_i k\!\left(-\frac{e_i(b)}{h}\right)\frac{e_i(b)}{h}.
\]
The paper states that this difference leads asymptotically to larger bias and larger variance for Horowitz’s estimator, and more generally that not all smoothing schemes are equivalent [1905.08535].

Borrowing strength across nearby quantiles can also be imposed without kernel convolution. “Spline Quantile Regression” formulates
\[
\hat\beta(\cdot)
=
\arg\min_{\beta(\cdot)\in F^p}
\left\{
n^{-1}\sum_{\ell=1}^L\sum_{t=1}^n
\rho_{\tau_\ell}(y_t-x_t^\top\beta(\tau_\ell))
+
c\sum_{\ell=1}^L w_\ell \|\ddot\beta(\tau_\ell)\|_1
\right\},
\]
so smoothness is imposed directly on the coefficient process over \(\tau\) rather than by convolution of the quantile objective [2501.03883]. That paper explicitly presents spline QR as an alternative to ad hoc post-processing by kernel smoothing across quantiles. It also notes that smoothing across quantiles does not by itself guarantee noncrossing.

A different non-crossing mechanism appears in smoothed SGD for unconditional quantiles. There the indicator in the SGD score is replaced by the piecewise linear function \(g_{a\gamma_k}\), yielding the recursion
\[
Y_{i,k+1}(\tau)
=
Y_{i,k}(\tau)
+
\gamma_k\bigl(\tau-g_{a\gamma_k}(Y_{i,k}(\tau)-X_{i,k+1})\bigr),
\]
with \(a>1/2\). The paper proves that for \(\tau\ge \tau'\),
\[
Y_{i,k}(\tau)\ge Y_{i,k}(\tau')
\]
for all \(k\), so the estimated quantile curves do not cross [2505.13299]. The same work explicitly links its recursive scheme to convolution-type smoothed quantile regression by writing
\[
\ell_{\tau,h}(u)=(\rho_\tau*K_h)(u)
\]
with a uniform kernel and iteration-dependent bandwidth.

## 6. Modern uses, bandwidth practice, and open limitations

The contemporary literature applies convolution-smoothed quantile regression to settings where standard QR is either computationally prohibitive or inferentially unstable.

In transfer learning for high-dimensional QR, the \(\ell_1\)-penalized smoothed estimator is
\[
\hat w
\in
\arg\min_{w\in\mathbb R^p}
\left\{
\frac{1}{nh}\sum_{i=1}^n\int \rho_\tau(u)\,
K\!\left(\frac{u+x_i^\top w-y_i}{h}\right)\,du
+
\lambda\|w\|_1
\right\},
\]
and the paper proposes a two-step smoothed transfer algorithm with adaptive source selection. It establishes \(\ell_1/\ell_2\) error bounds and selection consistency under regular conditions [2212.00428].

In prediction-powered inference, convolution smoothing is used to regularize both a score-debiasing estimator and a predict-then-debias estimator. The smoothed score takes the form
\[
\psi_{h_n}(\beta;y,x)
=
\bigl(\mathcal K_{h_n}(x^\top\beta-y)-\tau\bigr)x,
\]
and the smoothed loss is
\[
\ell_{h_n}(\beta;y,x)
=
\int \rho_\tau(v)K_{h_n}(v-y-x^\top\beta)\,dv.
\]
The paper states that the proposed estimators are computationally tractable and that numerical studies show they mitigate the overcoverage observed with unsmoothed prediction-powered quantile regression [2606.04128].

In deep nonparametric QR, ConquerNet trains ReLU networks with the convolution-smoothed loss
\[
\ell_h(u)=\int_{-\infty}^{\infty}\rho_\tau(v)K_h(v-u)\,dv,
\]
and proves that, under Besov smoothness assumptions and suitable network scaling, the estimator achieves
\[
\|\hat f_h-f_\tau^*\|_{\ell_2}^2
=
O_p\!\left(n^{-\frac{2s}{2s+d}}\log^3 n\right),
\]
with bandwidth restriction
\[
h^2\lesssim n^{-s/(2s+d)}.
\]
The numerical studies reported there show improved estimation accuracy and training efficiency across multiple quantile levels, with especially pronounced gains at high and low quantiles [2605.06265].

Bandwidth choice is a recurrent practical issue. The literature repeatedly balances a smoother objective against smoothing bias. In standard smoothed QR, a Silverman-style rule of thumb
\[
h_{\mathrm{ROT}}=1.06\,\widehat s\,n^{-1/5}
\]
is recommended in simulations, with \(\widehat s\) computed from standard QR residuals [1905.08535]. In mode regression, a \(\tau\)-dependent rule
\[
h_{ROT}(\tau)
=
\frac{1.06}{\sqrt[5]{n}}\widehat S(\tau)
\]
is used, with \(\widehat S(\tau)=\min\{0.7199528\times iq(\tau),\sigma(\tau)\}\) [2412.05736]. In high-dimensional penalized SQR, the practical default is
\[
h=
\max\left\{
0.05,\,
\sqrt{\tau(1-\tau)}
\left(\frac{\log p}{n}\right)^{1/4}
\right\}
\]
[2205.02432]. In prediction-powered quantile regression, simulations use
\[
h_n=\left(\frac{p+\log n}{n}\right)^{2/5}
\]
[2606.04128].

The limitations emphasized in the literature are also consistent. Several foundational results depend on correctly specified linear quantile regression, bounded or compactly supported covariates, smooth conditional densities, and nontrivial kernel moment conditions [1905.08535] [2412.05736]. Some extensions remain specialized: the transfer-learning theory is high-dimensional and linear [2212.00428]; the deep-network minimax theory is developed for single-quantile estimation rather than full quantile processes [2605.06265]; and the fixed-effects panel paper explicitly leaves optimal data-driven bandwidth choice as future work [1911.04729]. These restrictions do not diminish the central role of convolution smoothing in the current literature, but they delimit where its strongest theory presently applies.

Source: https://www.emergentmind.com/topics/convolution-smoothed-quantile-regression