---
title: Highly Adaptive Lasso (HAL) Overview
url: https://www.emergentmind.com/topics/highly-adaptive-lasso-hal
type: topic
---

# Highly Adaptive Lasso (HAL) Overview

Searching arXiv for recent and foundational HAL papers to ground the article.
The **Highly Adaptive Lasso (HAL)** is a nonparametric minimum-loss estimation framework for infinite-dimensional target functions defined over classes of càdlàg functions with bounded sectional variation norm. Its defining construction replaces low-dimensional parametric structure or local smoothness assumptions with a global variation constraint that induces an \(L_1\)-type regularization on a very large dictionary of lower-orthant indicator bases, and, in higher-order variants, tensor-product spline bases. Across the HAL literature, the method is formulated as an empirical risk minimizer under a sectional variation norm bound, with cross-validation commonly used to select the effective complexity level. This yields a sparse, data-adaptive estimator that has been used for regression, density estimation, hazard estimation, causal nuisance estimation, plug-in estimation of non-pathwise differentiable functionals, and targeted learning, while supporting rates and inferential results that differ substantially from classical kernel- or sieve-based nonparametrics [1709.06256; 2301.13354; 2404.11083; 2602.10613].

## 1. Function class, representation, and optimization principle

HAL targets function classes consisting of multivariate càdlàg functions with bounded sectional variation norm. In one standard formulation, the parameter of interest is a population risk minimizer
\[
P_0L(\psi_0)=\min_{\psi\in{\bf \Psi}}P_0L(\psi),
\]
and the estimator is the empirical risk minimizer over a bounded-variation class,
\[
\hat{\Psi}(P_n)=\arg\min_{\psi\in{\bf \Psi}}P_nL(\psi).
\]
For a \(d\)-variate càdlàg function on a compact rectangle \([0,\tau]\), the uniform sectional variation norm is written as
\[
\psi_v = \psi(0)+ \sum_{s\subset\{1,\ldots,d\}} \int_{0_s}^{\tau_s}\left|\psi_s(du_s)\right|,
\]
where \(\psi_s(u_s)=\psi(u_s,0_{-s})\) denotes the section indexed by \(s\) [1709.06256].

A central structural fact is that bounded sectional variation induces an integral representation in lower-orthant indicators. For the zero-order setting, one representation used in the literature is
\[
\psi(x)=\psi(0)+\sum_{s\subset\{1,\ldots,d\}} \int_{0_s}^{x_s} d\psi_s(u_s),
\]
which shows that \(\psi\) can be written as an infinite linear combination of indicator basis functions \(I(u_s\le x_s)\). In regression-style notation, this becomes a discrete basis expansion over observed knot points,
\[
f_{n,\beta}(x) = \beta_0 + \sum_{\emptyset \neq s \subseteq \{1,\ldots,d\}} \sum_{i=1}^n \beta_{s,i}\,\phi_{s,i}(x), 
\qquad \phi_{s,i}(x)=\mathbf 1\{x_s\ge \tilde x_{i,s}\},
\]
with variation norm
\[
\|f_{n,\beta}\|_v = |\beta_0|+\sum_{\emptyset \neq s}\sum_i |\beta_{s,i}|.
\]
Thus the sectional variation norm is identified with the coefficient \(L_1\)-norm, which is why HAL is “lasso-like” despite being defined over an infinite-dimensional function class [2602.10613; 1709.06256].

The induced optimization problem is therefore an empirical risk minimization over a saturated, data-adaptive basis under an \(L_1\)-type complexity bound. In practical formulations, this appears either as a hard constraint \(\|\beta\|_1\le C\) or as an equivalent penalized objective. The class is fully nonparametric in the sense used in the literature: the basis is sufficiently rich to represent nonlinearities, interactions, and discontinuities, while the complexity control comes from sectional variation rather than Euclidean dimension [2301.12029; 2404.11083].

## 2. Zero-order HAL and higher-order spline HAL

The original HAL construction corresponds to the zero-order case, in which basis functions are step functions. Later work generalizes the framework to higher-order smoothness classes and tensor-product spline bases. In the higher-order formulation, a \(k\)-th order smoothness class \(D^{(k)}([0,1]^d)\) is defined by requiring recursively defined \(k\)-th order Radon–Nikodym derivatives to be càdlàg with bounded variation. The corresponding \(k\)-th order sectional variation norm is the sum of the variation norms of all \(k\)-th order sectional derivatives [2301.13354].

For zero-order HAL, one representation is
\[
Q(x)=\int_{[0, x]} d Q(u)=\int_{u \in[0,1]^d} \phi_u^0(x)\, d Q(u), 
\qquad \phi_u^0(x):=I(x\ge u)=\prod_{j=1}^d I(x_j\ge u_j).
\]
The discrete approximation becomes
\[
Q_{\beta}(x)=\beta_0+\sum_{s \subset \{1,\cdots,d\}} \sum_{i=1}^n \beta_{s,i} I(x(s)\ge x_i(s)).
\]
For first-order HAL,
\[
Q^{(1)}_{\beta}(x) = \beta_0+\sum_{s \subset \{1,\cdots,d\}} \sum_{i=1}^n \beta_{s,i}\phi_{x_i(s)}^1(x(s)), 
\qquad \phi_u^1(x)=(x-u)I(x\ge u),
\]
and for second order,
\[
\phi_u^2(x)=\frac12(x-u)^2 I(x\ge u).
\]
More generally, the higher-order theory represents the target as an infinite linear combination of tensor products of spline basis functions of order at most \(k\), including lower-order factors on boundary faces [2507.10511; 2301.13354].

This extension is not merely cosmetic. The higher-order papers state that for first and higher order smoothness classes, pointwise asymptotic normality and uniform convergence can be established at dimension-free rate
\[
n^{-k^*/(2k^*+1)}
\]
up to logarithmic factors, where \(k^*=k+1\). In a related inferential treatment, the \(m\)-th order spline HAL-MLE is described as smoothness-adaptive when \(m\) is selected by cross-validation, while still guaranteeing a rate faster than \(n^{-1/4}\) as long as the true function is càdlàg and has finite sectional variation norm [2301.13354; 1908.05607].

A useful special case arises in one dimension. Recent work on univariate density estimation states that bounded sectional variation coincides with classical bounded total variation, so in \(d=1\),
\[
\|f\|_v^*=|f(0)|+\mathrm{TV}(f),
\]
and for \(k\ge 1\),
\[
\|f\|_{v,k}^* = \sum_{i=0}^k |f^{(i)}(0)|+\mathrm{TV}\bigl(f^{(k)}\bigr).
\]
The exact spline representation
\[
f(x) = \sum_{i=0}^k \frac{1}{i!}f^{(i)}(0)x^i + \int_0^1 \frac{1}{k!}(x-u)_+^k\,d f^{(k)}(u)
\]
connects univariate HAL directly to classical total-variation penalized splines, local adaptive splines, and log-spline density models [2602.16259].

## 3. Statistical guarantees and convergence theory

A defining feature of HAL is that its theory is stated under minimal smoothness assumptions relative to much of nonparametric statistics. Early theory established that, in loss-based dissimilarity,
\[
d_0(\psi_n,\psi_0)=P_0L(\psi_n)-P_0L(\psi_0)=O_P\bigl(n^{-1/2-\alpha(d)}\bigr),
\]
with \(\alpha(d)>0\), and that under weak continuity conditions the estimator is uniformly consistent:
\[
\sup_{x\in [0,\tau]}|\psi_n(x)-\psi_0(x)| \rightarrow 0 \quad\text{in probability}.
\]
The uniform consistency result requires conditions labeled \(A0\)–\(A3\), including continuity of \(\psi_0\), bounded loss, identifiability through zero loss-based dissimilarity, and a weak continuity condition linking pointwise convergence of functions to pointwise convergence of the loss [1709.06256].

Subsequent papers present rate claims in more specialized settings. For zero-order HAL, one paper explicitly recalls the rate
\[
O_P\!\left(n^{-1/3}(\log n)^{2(d-1)/3}\right),
\]
and characterizes this as “almost dimension-free” because the dependence on dimension is only logarithmic in the leading rate expression [2602.10613]. In the multi-task formulation, under square-error loss and assumptions \(A1\) and \(A2\),
\[
\|\psi_{n,M}-\psi_{0,M}\|_{P_0} = O_P\!\left(n^{-1/4-1/(8(d+1))}\right),
\]
and with cross-validated \(M_n\),
\[
\|\psi_{n,M_n}-\psi_{0,M}\|_{P_0} = O_P\!\left(n^{-1/4-1/(8(d+1))}\right) + O_P\!\left(n^{-1/2\log B}\right).
\]
This is summarized there as dimension-free \(o_p(n^{-1/4})\) behavior [2301.12029].

Higher-order spline HAL sharpens this picture. The higher-order spline paper gives loss-based convergence at
\[
d_0(Q_n,Q_0)=O\!\big(n^{-2k^*/(2k^*+1)}(\log n)^m\big),
\]
implying \(L^2\)-type convergence at rate \(n^{-k^*/(2k^*+1)}\) up to logarithmic factors, and states pointwise asymptotic normality after normalization by \((n/d_{0,n})^{1/2}\) for an oracle working model with \(d_{0,n}\asymp n^{1/(2k^*+1)}\) [2301.13354]. The univariate log-spline HAL density paper states three new results in \(d=1\): asymptotic linearity, pointwise asymptotic normality, and uniform convergence at rate
\[
n^{-(k+1)/(2k+3)}
\]
up to logarithmic factors for smoothness order \(k \geq 1\) [2602.16259].

These results are repeatedly interpreted through the lens of plug-in efficiency. Several papers state that HAL nuisance fits converge fast enough—typically faster than \(n^{-1/4}\) in the relevant norm or loss—to support asymptotic linearity and semiparametric efficiency for smooth target parameters once the associated score or efficient influence curve equations are handled appropriately [2101.06290; 1908.05607].

## 4. Computational structure, basis explosion, and principal-component reductions

The main computational limitation of HAL is the size of the basis expansion. For zero-order HAL with \(d\) covariates and \(n\) observations, one paper states that the full dictionary contains
\[
p = n(2^d-1)
\]
non-intercept basis functions, while another notes that the canonical working model size is \(N=n(2^d-1)\), with higher-order versions even larger [2602.10613; 2603.18204]. This makes basis construction, cross-validation, and repeated lasso fitting computationally prohibitive in moderate to high dimensions.

Several recent papers address this bottleneck via kernel and principal-component representations. The Highly Adaptive Ridge (HAR) replaces the \(L_1\) constraint with an \(L_2\) penalty on the same dictionary, allowing kernelization through the Gram matrix
\[
K := HH^\top \in \mathbb R^{n\times n}.
\]
Building on this, Principal Component based HAL (PCHAL) and related PC-HA estimators use the singular value decomposition of the HAL design matrix \(H\),
\[
H = U D^{1/2}V^\top,
\]
and define low-dimensional scores
\[
Z_k := H V_k = U_k D_k^{1/2},
\]
so that fitting is performed in the orthogonal score space rather than on the original exponentially large basis [2602.10613].

For orthogonal scores, the resulting ridge and lasso problems admit closed forms. One paper states
\[
\hat\beta^{\mathrm{PCHAR}_{k,\lambda}} = (D_k+n\lambda I_k)^{-1}D_k^{1/2}U_k^\top Y,
\]
and
\[
\hat\beta^{\mathrm{PCHAL}_{k,\lambda}} = D_k^{-1}\,\mathrm{sign}(Z_k^\top Y)\,\bigl(|Z_k^\top Y|-n\lambda\bigr)_+,
\]
for coordinates with \(d_j>0\) [2602.10613]. Another PC-HA paper constructs an orthogonal PC basis
\[
\tilde\phi_m=\sum_{j\in\mathcal R_N}E_N^n(j,m)\,\phi_j
\]
from the eigenvectors of \(H^\top H\), and defines three estimators: PC-HAGL, PC-HAL, and PC-HAR, which constrain \(\|\beta(\alpha)\|_1\), \(\|\alpha\|_1\), and \(\|\alpha\|_2\), respectively [2603.18204].

The principal-component line of work emphasizes that the reduction is outcome-blind and theoretically justified. The 2026 PC-HAL paper derives an excess-risk bound relative to the full \(n\)-component fit,
\[
\mathbb E\!\left[\widehat R(\hat f_k)-\widehat R(\hat f_n)\right] \le \frac{2V(f_0)^2}{n}\sum_{j=k+1}^{r} d_j + \frac{2}{n}\|r_n\|_2^2,
\]
showing that truncation error is governed by the spectral tail of the HAL Gram operator [2602.10613]. The PC-HA paper further states that, under complexity control, PC-HA can inherit HAL-like loss rates and score-equation properties, allowing transfer of plug-in efficiency and pointwise asymptotic normality results [2603.18204].

A particularly striking observation appears in the one-dimensional ordered case, where the HAL Gram matrix has entries \((HH^\top)_{ij}=\min(i,j)\). The corresponding eigenvectors are discrete sine functions, yielding an explicit Fourier-type structure behind the leading principal components [2602.10613]. This suggests that the saturated step-function dictionary of HAL possesses a much lower-rank geometry than the raw feature count \(n(2^d-1)\) might indicate.

## 5. HAL in semiparametric and causal inference

HAL has been incorporated deeply into semiparametric inference, especially targeted learning. In the HAL-TMLE line, HAL-MLE provides initial nuisance estimates for outcome regression, treatment mechanism, density, or hazard components, and the targeted update is then designed to solve or approximately solve an efficient influence curve equation. One formulation of HAL-TMLE writes the target as
\[
\psi_n^*=\Psi(Q_n^*),
\]
where \(Q_n^*\) is a targeted update of the initial HAL estimate \(Q_n\) such that
\[
P_n D^*(Q_n^*,G_n)=r_n,
\]
with \(r_n=o_P(n^{-1/2})\) and \(r_n=0\) in the universal least favorable submodel case [1708.09502].

The asymptotic logic is standard semiparametric theory adapted to HAL rates. If the second-order remainder is \(o_P(n^{-1/2})\), then
\[
\Psi(Q_n^*)-\Psi(Q_0) = (P_n-P_0)D^*(Q_0,G_0)+o_P(n^{-1/2}),
\]
giving asymptotic efficiency under only sectional variation assumptions on the nuisances [2101.06290; 1708.09502]. However, multiple papers stress that finite-sample inference based solely on first-order asymptotics can be anti-conservative when the second-order remainder is not negligible, motivating nonparametric bootstrap procedures for HAL-TMLE and even higher-order TMLE constructions [1708.09502; 1905.10299; 2101.06290].

HAL has also been used directly to build efficient inverse probability weighted estimators without a separate outcome model in the estimator itself. In the IPW paper, the treatment mechanism \(G_0(1\mid W)\) is estimated by HAL, and the IPW estimator
\[
\widehat{\Psi}(P_n,G_n) = n^{-1}\sum_{i=1}^n \frac{A_iY_i}{G_n(A_i\mid W_i)}
\]
is shown to be asymptotically linear and efficient when HAL is deliberately undersmoothed so that bias terms are \(o_p(n^{-1/2})\) [2005.11303].

A related causal-inference development is the outcome highly adaptive lasso (OHAL), which modifies the propensity-score fit by weighting the HAL penalty using outcome-regression coefficients so as to downweight instrumental basis functions. The resulting ATE estimator is a robust TMLE whose limit expansion involves
\[
D(O\mid \bar Q_0,\bar G(\bar Q_0),Q_0) - D_r(O\mid \bar Q_0,\bar G_{r,0}),
\]
and simulation evidence in that paper reports lower MSE than standard HAL-TMLE and near-nominal coverage when cross-validated standard errors are used [1806.06784].

More recent work addresses targeting after HAL selection within the finite-dimensional working model implied by the active HAL basis. “Regularized Targeted Maximum Likelihood Estimation in Highly Adaptive Lasso Implied Working Models” proposes delta-method regHAL-TMLE and projection-based regHAL-TMLE, motivated by collinearity and computational instability in relaxed HAL and full HAL-TMLE implementations. The projection-based method approximates the efficient influence curve by regularized projection onto the HAL score space and is reported to be especially stable under positivity problems and in survival-curve estimation [2506.17214].

## 6. Extensions to multi-task learning, densities, hazards, and infinite-dimensional inference

HAL has been extended beyond single-task regression to settings where the target itself is structured.

The multi-task extension, MT-HAL, keeps the HAL basis expansion but shares it across tasks and applies a mixed \(l_{2,1}\) norm penalty,
\[
\|\alpha\|_{2,1} = \sum_{p=1}^d \|\alpha_p\|_2,
\]
so that predictor-basis groups are encouraged to drop out jointly across tasks. The fitted function is constructed from indicator basis functions shared across all tasks, with optional task interactions. The empirical studies reported in that paper compare MT-HAL to MT-lasso and MT-L21 in nonlinear and linear data-generating processes. Across the main nonlinear simulations, MT-HAL is reported to achieve the lowest mean squared error in every setting, with examples such as MSEs \(0.68\), \(0.45\), and \(0.67\) versus approximately \(1.4\)–\(1.6\) for MT-lasso and \(0.7\)–\(0.8\) for MT-L21 in \(d=6\), \(K=5\), \(n=600\) nonlinear cases [2301.12029].

In survival and density estimation, HAL is used not merely as a generic regression tool but as a remedy for ill-posed empirical risk minimization over the full càdlàg class. The conditional hazard and density paper states that, for likelihood-type losses, the empirical risk minimizer over the full bounded sectional variation class may be not well-defined or inconsistent. A data-adaptive sieve version of HAL,
\[
\hat f_n \in \arg\min_{f\in d_{M,n}} \mathbb P_n[L(f,O)],
\]
is then used to restore convexity and consistency. Under smoothness conditions, the rate
\[
\|\hat f_n-f^*\| = O_P\!\left(n^{-1/3}\log(n)^{2(d-1)/3}\right)
\]
is established for general HAL sieve estimation, with corollaries for conditional hazard estimation under right censoring and for a new direct conditional density parametrization [2404.11083].

Plug-in estimation of non-pathwise differentiable functional targets has also become a significant application area. For the causal dose-response curve with continuous treatment, the target is
\[
\Psi_P(a)= \mathbb E\!\left[\mathbb E_P(Y\mid A=a,W)\right],
\]
and HAL is used to estimate the conditional outcome regression \(Q_0(w,a)\), yielding the plug-in estimator
\[
\hat\Psi(a)=\frac{1}{n}\sum_{i=1}^n \hat Q(W_i,a).
\]
That paper emphasizes undersmoothing and smoothness-adaptive HAL fitting, and reports that the HAL-based estimator consistently outperforms GAM, polynomial regression, and npcausal in simulations, particularly for oscillatory and discontinuous dose-response curves [2406.05607].

A 2025 inferential paper extends HAL-based confidence interval construction for conditional mean functions and related infinite-dimensional parameters. It distinguishes regular HAL, relaxed HAL, and targeted HAL; introduces local and global undersmoothing criteria based on estimated bias-to-standard-error ratios; and uses delta-method Wald intervals built from the HAL working model. It also applies the same framework to conditional average treatment effect estimation through a doubly robust pseudo-outcome
\[
\eta(Q_n, g_n)(O) = \frac{2A - 1}{g_n(A \mid W)} \big[Y - Q_n(A, W)\big] + Q_n(1, W) - Q_n(0, W),
\]
which is then regressed on \(W\) using HAL [2507.10511].

## 7. Interpretation, misconceptions, and methodological position

HAL is often described as a lasso, but its defining object is not a standard linear model. The lasso aspect arises because bounded sectional variation becomes an \(L_1\)-constraint on the coefficients of an adaptive basis representation; the estimator itself is an empirical risk minimizer over a nonparametric càdlàg class. This distinction matters because the resulting guarantees are expressed in loss-based dissimilarity, sectional variation norms, and score-equation properties rather than solely in sparsity recovery or parametric restricted-eigenvalue conditions [1709.06256; 1708.09502].

Another common misconception is that HAL is only a zero-order step-function estimator. The higher-order spline literature shows that zero-order HAL is the \(m=0\) or \(k=0\) case in a hierarchy of spline-based classes with recursively defined smoothness. In that line of work, higher-order HAL retains the same essential regularization logic but replaces step-function bases by tensor-product splines, improving approximation for smoother targets while preserving the sectional-variation interpretation of the coefficient norm [1908.05607; 2301.13354; 2507.10511].

A further misconception is that cross-validated HAL is automatically suitable for inference. Several papers argue the opposite. Cross-validation selects the complexity level for predictive risk, whereas efficient inference may require undersmoothing so that bias terms or empirical score imbalances are negligible at the \(n^{-1/2}\) scale. This issue appears in efficient IPW with HAL-estimated propensity scores, in higher-order spline HAL plug-in efficiency, and in conditional mean confidence intervals based on HAL working models [2005.11303; 1908.05607; 2507.10511].

Finally, the main criticism of HAL in practice is computational rather than statistical. The dictionary size \(n(2^d-1)\) makes naïve implementations infeasible in moderate dimension. The recent PCHAL/PCHAR and PC-HA papers respond directly to this concern by providing outcome-blind, spectrally motivated dimension reduction with formal score-equation and excess-risk theory. This suggests that the current development of HAL is as much about computational reformulation as about new asymptotics [2602.10613; 2603.18204].

Taken together, the literature presents HAL as a general minimum-loss estimation framework over bounded sectional-variation classes, with a distinctive combination of properties: a nonparametric basis representation tied exactly to an \(L_1\)-type complexity measure; rates that are often described as near dimension-free up to logarithmic factors; compatibility with plug-in semiparametric efficiency theory; and extensibility to multi-task learning, density and hazard estimation, infinite-dimensional functionals, and modern targeted learning pipelines [2301.12029; 2404.11083; 2101.06290]. A plausible implication is that HAL occupies a methodological position between classical variation-penalized function estimation and modern machine-learning-based semiparametric inference: it is flexible enough to behave like a large-scale adaptive learner, yet structured enough to support exact or approximate score equations and detailed asymptotic theory.

Source: https://www.emergentmind.com/topics/highly-adaptive-lasso-hal