---
title: Data-Adaptive Factor Model (DAFM)
url: https://www.emergentmind.com/topics/data-adaptive-factor-model-dafm
type: topic
---

# Data-Adaptive Factor Model (DAFM)

to=arxiv_search.search  天天赢彩票  天天中彩票追号json
{"query":"all:\"data-adaptive factor model\" OR all:\"locally adaptive factor\" OR all:\"adaptive multi-factor model\"","max_results":10,"sort_by":"submittedDate","sort_order":"descending"}Japgollyassistant to=arxiv_search.search კომენტary  天天彩票网json
{"query":"id:1210.2022 OR id:2306.02848 OR id:1804.08472 OR id:2506.19945 OR id:2510.00558 OR id:2506.09776 OR id:2012.14708","max_results":10,"sort_by":"relevance","sort_order":"descending"}чилириassistant to=arxiv_search.search კომენტary  高频彩大发快三 】【。】【”】【json
{"query":"ti:\"Locally adaptive factor processes for multivariate time series\"","max_results":5,"sort_by":"relevance","sort_order":"descending"}
I’m not able to display raw tool internals here, but I used arXiv search to verify the topic area and the specific papers, including [1210.2022], [2306.02848], [1804.08472], [2506.19945], [2510.00558], [2506.09776], and [2012.14708].
Data-Adaptive Factor Model (DAFM) denotes a class of factor-modeling approaches in which the factor structure is learned from data rather than fixed ex ante. In the literature, the term refers explicitly to a continuous-time Locally Adaptive Factor (LAF) process for multivariate mean-covariance modeling and to a composite-quantile factor model for high-dimensional panels, while closely related work uses the same adaptive principle in high-dimensional asset pricing, online regime-aware stock prediction, manifold-based dynamic factor modeling, robust covariance decomposition, and locally stationary factor models with evolving loading spaces [1210.2022; 2510.00558; 1804.08472; 2306.02848; 2506.19945; 2506.09776; 2012.14708]. This suggests that DAFM is best understood as a methodological family rather than a single canonical specification.

## 1. Conceptual scope and distinguishing features

Across these formulations, the common objective is to avoid rigid factor structures that are known in advance to be misspecified in settings with heterogeneity, regime shifts, nonlinear dependence, or distributional asymmetry. The adaptive component can enter through locally varying smoothness, stock-specific sparse basis selection, regime-dependent decoding, nonlinear embedding geometry, uncertainty-ball robustification, or quantile-specific loadings.

This family differs from classical fixed-factor models in a recurrent way. Standard factor models often assume a small, fixed set of common factors; globally smooth Gaussian-process or spline models impose one global smoothness scale; PCA-based factors are variance-driven and are often difficult to interpret; and clustering-based regime methods may use the entire dataset in ways that are unsuitable for point-in-time deployment. DAFM-type approaches replace one or more of those fixed elements with data-dependent structure. In different papers, the adaptive object is the loading matrix, the loading span, the factor dimension, the factor-return mapping, the basis-asset subset, the regime label, the latent geometry, or the contribution of different quantiles.

A central misconception is that “data-adaptive” implies a single modeling recipe. The published work does not support that reading. Instead, the term spans continuous-time Bayesian factor processes, high-dimensional sparse basis-asset regressions, hierarchical VAE architectures, nonlinear diffusion-map embeddings, composite-quantile estimators, and robust low-rank-plus-diagonal covariance decompositions. The shared theme is not one estimator, but the decision to let data determine which factor representation is appropriate for the task.

## 2. Continuous-time locally adaptive factor processes

In multivariate time series, the LAF process models a time-varying mean vector and covariance matrix,
$$
\Gamma=\{\mu(t),\Sigma(t),\ t\in\mathcal T\},
$$
while allowing the degree of smoothness to change locally over time. The starting point is the latent factor representation
$$
Y_i=\Lambda(t_i)\eta_i+\epsilon_i,\qquad \epsilon_i\sim N_p(0,\Sigma_0),
$$
with
$$
\eta_i=\psi(t_i)+\nu_i,\qquad \nu_i\sim N_K(0,I_K),
$$
and a structured loading matrix
$$
\Lambda(t_i)=\Theta\,\xi(t_i).
$$
Here, $\Theta$ is a fixed sparse $p\times L$ loading matrix, $\xi(t)$ is an $L\times K$ matrix of latent dictionary functions, and $\psi(t)$ is a $K\times 1$ vector of latent mean factors. This yields
$$
\mu(t_i)=\Theta\,\xi(t_i)\psi(t_i),
$$
and
$$
\Sigma(t_i)=\Theta\,\xi(t_i)\xi(t_i)^T\Theta^T+\Sigma_0.
$$
The observed process is therefore driven by a low-rank, time-varying factor structure with diagonal residual covariance [1210.2022].

The model is “data-adaptive” because each dictionary function is governed by a nested Gaussian process (nGP), with a higher-level process acting like a local instantaneous mean smoothness trajectory for the derivative of the dictionary function. A parallel construction is used for the latent mean factors. The substantive point is that smoothness itself evolves over time. The model is designed to avoid over-smoothing across erratic intervals and under-smoothing across slowly varying intervals, a failure mode that otherwise leads to misleading inferences, predictions, and mis-calibration of predictive intervals.

A major technical contribution is the linear state-space representation obtained from discretized stochastic differential equations. For $m=2$ and $n=1$, the paper gives the approximate update
$$
\begin{bmatrix}
\xi_{lk}(t_{i+1})\\
\xi'_{lk}(t_{i+1})\\
A_{lk}(t_{i+1})
\end{bmatrix}
=
\begin{bmatrix}
1 & \delta_i & 0\\
0 & 1 & \delta_i\\
0 & 0 & 1
\end{bmatrix}
\begin{bmatrix}
\xi_{lk}(t_i)\\
\xi'_{lk}(t_i)\\
A_{lk}(t_i)
\end{bmatrix}
+
\begin{bmatrix}
0&0\\
1&0\\
0&1
\end{bmatrix}
\begin{bmatrix}
\omega_{i,\xi_{lk}}\\
\omega_{i,A_{lk}}
\end{bmatrix},
$$
with Gaussian innovations whose variance is proportional to $\delta_i=t_{i+1}-t_i$. This state-space form makes Kalman-style simulation smoothing feasible for irregular observation times and missing data.

Scalability is achieved through the sparse map $\Lambda(t)=\Theta\xi(t)$, which reduces the number of latent nGPs from $p\times K$ to $L\times K$. Sparsity in $\Theta$ is induced through the hierarchical shrinkage prior
$$
\theta_{jl}\mid \phi_{jl},\tau_l \sim N(0,\phi_{jl}^{-1}\tau_l^{-1}),\qquad \tau_l=\prod_{h=1}^l \vartheta_h.
$$
Inference uses a Gibbs sampler with simulation smoother updates for $\xi(t)$ and $\psi(t)$, updates for the latent factors $\nu_i$, conjugate updates for $\Theta$, $\Sigma_0$, and variance hyperparameters, and shrinkage-hyperparameter updates for $\phi_{jl}$ and $\tau_l$. The Durbin–Koopman simulation smoother is used, and the nGP state-space representation is reported to reduce the usual GP matrix inversion burden from about $O(T^3)$ to $O(T)$ in time series length $T$.

The empirical findings are equally specific. When the truth has locally varying smoothness, LAF clearly outperforms BCR, EWMA, DCC-GARCH, PC-GARCH, and GO-GARCH, while BCR tends to oversmooth and to produce overly narrow or miscalibrated credible intervals. When the truth is globally smooth, LAF performs about as well as BCR. In an application to 33 national stock indices from 2004–2012, the model identifies sharp volatility increases during the 2008 global financial crisis, regime changes around the Greek debt crisis and later Eurozone events, and changing cross-country correlation structures consistent with geo-economic blocs. The model also supports one-step-ahead prediction and conditional forecasting based on the multivariate covariance structure.

## 3. Adaptive basis-asset factor models in finance

A distinct DAFM lineage appears in high-dimensional asset pricing. The Adaptive Multi-Factor (AMF) model is motivated by generalized Arbitrage Pricing Theory and replaces a small fixed factor set with a large universe of tradable basis assets, especially ETFs. Its realized-return representation is
$$
R_i(t)-r_0(t)=\sum_{j=1}^{p}\beta_{i,j}\big[r_j(t)-r_0(t)\big]
=\bm{\beta}_i'[\bm{r}(t)-r_0(t)1],
$$
and the empirically testable version adds an intercept and noise,
$$
R_i(t)-r_0(t)=\alpha_i+\sum_{j=1}^{p}\beta_{i,j}(t)\big[r_j(t)-r_0(t)\big]+\epsilon_i(t),
$$
with $\epsilon_i(t)\overset{iid}{\sim}N(0,\sigma_i^2)$. After selection, each stock has its own sparse subset $S_i$ of basis assets:
$$
\bm{R}_i-\bm{r}_0=\alpha_i1_n+(\bm{r}_S-(\bm{r}_0)_S1'_{p})(\bm{\beta}_i)_S+\bm{\epsilon}_i.
$$
The contrast with Fama–French 5 is structural: FF5 uses the same five factors for all stocks, whereas AMF uses a large candidate set of tradable basis assets with stock-specific sparse selection [1804.08472].

Estimation proceeds through the Groupwise Interpretable Basis Selection (GIBS) algorithm. The candidate basis assets are first orthogonalized with respect to the market factor; then they are divided into economically defined groups such as bond/fixed income, commodity, currency, diversified portfolio, equity, alternative ETFs, inverse, leveraged, real estate, and volatility. Within each group, prototype clustering is performed using the correlation-based distance
$$
d(r_1,r_2)=1-|corr(r_1,r_2)|,
$$
with minimax linkage used to select representative prototypes. A second clustering stage is then applied across the prototypes to obtain a final low-correlated set $U$. For each stock, a modified LASSO is run on $U$,
$$
\hat{\beta}_{\lambda}=\arg\min_{\beta\in\mathbb{R}^p}\left\{\frac{1}{2n}\|y-X\beta\|_2^2+\lambda\|\beta\|_1\right\},
$$
with the conservative tuning rule
$$
\lambda=\max\{\lambda_{1se},\min\{\lambda:\#supp(\widetilde{\beta}_i)\le 20\}\}.
$$
The selected factors are then refit by OLS for post-selection inference.

The empirical results emphasize prediction, parsimony, and interpretability. Relative to FF5, the 2018 paper reports average adjusted $R^2$ of 0.348 for AMF/GIBS versus 0.234 for FF5, and out-of-sample prediction of 0.331 versus 0.209. It also reports that FF5 has 6.33% of stocks with significant intercepts at 5%, versus 4.12% for AMF, and that almost all significant intercepts disappear after BH/BHY false-discovery-rate adjustment. Average selected factors are 3.4 for GIBS, compared with 13.4 for LASSO and 186 for Ridge, indicating that the adaptive procedure remains sparse while standard regularization methods overfit [1804.08472].

The 2021 dissertation extends the same framework and makes the interpretability claim concrete. Using weekly data over 2014–2016, all ETFs in the CRSP database plus FF5 as candidate basis assets, and a final sample of about 5132 stocks, GIBS selects 182 basis assets that matter for at least one stock, with an average of 2.98 selected basis assets per stock and 1.92 significant basis assets after OLS. It reports that 77.47% of the selected basis assets have non-zero risk premia, that mean adjusted $R^2$ rises from 0.229 to 0.319, and that out-of-sample $R^2$ rises from 0.030 to 0.038, a 24.07% improvement. The dissertation also argues that the low-volatility anomaly essentially disappears after AMF adjustment, with excess return difference $=0.59\ (p=1.17\times10^{-5})$, FF5 residual difference $=0.18\ (p=9.25\times10^{-95})$, and AMF residual difference $=0.12\ (p=1.00)$ [2107.14410].

## 4. Online regime-adaptive deep factor models

HireVAE frames stock prediction as a factor-modeling problem in which
$$
f(\cdot;\Theta): \mathcal{F}_{<t}\mapsto \mathbf{y}_{t+\Delta_t},
\qquad
\hat{\mathbf{y}}_{t+\Delta_t}
=
\mathbb{E}\!\left[\mathbf{y}_{t+\Delta_t}\mid \mathcal{F}_{<t}\right].
$$
The inputs are stock-specific sequential features $\mathbf{X}_t\in\mathbb{R}^{T\times N\times C}$ and market/global sequential features $\mathbf{M}_t\in\mathbb{R}^{T\times C_m}$, with future return vector $\mathbf{y}_{t+\Delta t}\in\mathbb{R}^{N}$ as target. The paper’s premise is that factor effectiveness is regime-dependent and that practical stock prediction requires point-in-time, online, adaptive learning using only information available up to the current timestamp. It therefore proposes what it calls the first deep learning based online and adaptive factor model, built around a hierarchical latent space and a regime-switching decoder [2306.02848].

The hierarchical latent structure uses a market latent variable $\mathbf{m}$ and a stock latent variable $\mathbf{z}$, with the stock latent inferred conditionally on the market latent. The return decoder is regime-specific:
$$
\phi_{dec}(\mathbf{z},\mathbf{m},\mathbf{e}_s;c),
\qquad
c\in\{1,\dots,N_k\},
$$
where the selected regime $c$ is inferred from the market latent. The market encoder learns
$$
[\boldsymbol{\mu}^m_{post}, \boldsymbol{\sigma}^m_{post}] = \phi_{enc}^{m}(\mathbf{v}, \bar{\mathbf{y}}),
\qquad
\mathbf{m}\sim \mathcal{N}\!\left(\boldsymbol{\mu}^m_{post},
\operatorname{diag}(\boldsymbol{\sigma}^m_{post})\right),
$$
and the stock encoder learns
$$
[\boldsymbol{\mu}^s_{post}, \boldsymbol{\sigma}^s_{post}] = \phi_{enc}^{s}(\mathbf{v}, \mathbf{e}_s, \mathbf{y}),
\qquad
\mathbf{z}\sim \mathcal{N}\!\left(\boldsymbol{\mu}^s_{post},
\operatorname{diag}(\boldsymbol{\sigma}^s_{post})\right).
$$
At inference, priors are used instead of posteriors, and the predicted return is
$$
\hat{\mathbf{y}}=\phi_{dec}(\mathbf{z}_0,\mathbf{m}_0,\mathbf{e}_s;c),
\qquad
c=f(\mathbf{m}_0).
$$

The regime-switching module projects $\mathbf{m}$ into a one-dimensional score,
$$
s=\phi_{proj}(\mathbf{m}),
$$
then scores this against Gaussian regime distributions using
$$
l(s;\mu_r^i,\sigma_r^i)
=
-\ln(\sigma_r^i)-\frac{1}{2}\ln(2\pi)
-\frac{1}{2}\left(\frac{(s-\mu_r^i)^2}{\sigma_r^i}\right),
$$
and assigns
$$
c=\arg\max_i\ l(s;\mu_r^i,\sigma_r^i).
$$
To stabilize labels over time, regime centers are reordered by descending mean and updated by moving average. Training uses the total objective
$$
L_{overall}=L_{rec}+L_{hier}+L_{reg},
$$
combining predictive fit, hierarchical variational consistency, and regime separation.

The online learning procedure is explicitly chronological: obtain market latent $\mathbf{m}$, project to score $s$, predict candidate cluster centers, sort clusters by mean, update regime centers with a moving average, compute regime probabilities, and assign the regime by maximum likelihood. Experiments use the China stock market from the WIND database with price-volume data, fundamental data, 58 stock features, sequential length $T=20$, prediction horizon $\Delta t=20$, and 14 groups of train/validation/test constructed on a point-in-time basis. Across CSI 300, CSI 500, CSI 1000, CNI 2000, and CSI All, HireVAE reports the best predictive metrics among compared methods: IC $=0.058$, Rank IC $=0.066$, and Rank ICIR $=0.734$. Its reported active returns are CSI 300: AR 30.89, CSI 500: AR 29.05, CSI 1000: AR 27.14, CNI 2000: AR 25.94, and CSI All: AR 28.02. The ablations attribute the improvement primarily to adaptive regime switching rather than market information alone.

## 5. Nonparametric dynamic factor modeling via manifold learning

A further DAFM development abandons a fixed parametric factor dynamics specification and instead learns a low-dimensional latent state from the joint geometry of covariates and responses. In this framework, a latent state $\theta_t\in\mathbb{R}^d$ evolves according to
$$
d\theta_t = -\nabla U(\theta_t)\,dt + \sqrt{2}\,dW_t,
$$
with invariant density
$$
\mu(dw)=\frac{1}{Z}e^{-U(w)}\,dw,
$$
and observables
$$
X_t = F(\theta_t)\in\mathbb{R}^m,\qquad Y_t = G(\theta_t)\in\mathbb{R}^n.
$$
The model is designed for high-dimensional covariates and responses and explicitly argues that factors learned only from $x(t)$ may lose predictive power for $y(t)$. The resulting procedure is therefore joint rather than purely unsupervised, and the paper names it the Joint Diffusion Kalman Filter (JDKF) [2506.19945].

The embedding is built on anisotropic diffusion maps using a modified Mahalanobis distance,
$$
d(z(t_i),z(t_j))
=
\frac12 (z(t_i)-z(t_j))^\top
\big(C^{-1}(t_i)+C^{-1}(t_j)\big)
(z(t_i)-z(t_j)),
$$
with RBF kernel
$$
W_{ij}=\exp\!\left(-\frac{d(z(t_i),z(t_j))}{2}\right).
$$
If $P\psi_k=\kappa_k\psi_k$, then the corresponding eigenvalue of $-L$ is estimated by
$$
\lambda_k \approx -\epsilon^{-1}\log \kappa_k,
$$
and the diffusion coordinates are
$$
\psi(t_i)=[\psi_1(t_i),\dots,\psi_\ell(t_i)].
$$
A linear lifting operator $\mathbf H$ is estimated empirically through
$$
z(t_i)\approx \mathbf H \psi(t_i),
\qquad
\mathbf H_{j,k}=\langle z_j,\psi_k\rangle_{\text{data}}.
$$

A central step is the approximation of the learned eigenfunction dynamics by linear diffusions. The paper starts from
$$
d\varphi_i(\theta_t)
=
-\lambda_i \varphi_i(\theta_t)\,dt
+\sqrt{2}\,\|\nabla \varphi_i(\theta_t)\|_2\,dB_t^i,
$$
and approximates this by
$$
d\xi_t^i = -\lambda_i \xi_t^i\,dt + \sqrt{2}\gamma_i\,dB_t^i.
$$
The paper then uses the linear state-space model
$$
\psi_{t+1}=A\psi_t+w_t,\qquad w_t\sim\mathcal N(0,Q),
$$
with measurements
$$
x_t = H^x\psi_t+v_{x,t},\qquad y_t = H^y\psi_t+v_{y,t},
$$
and standard Kalman filter and RTS smoother recursions.

This formulation is not merely heuristic. The paper provides robustness results for the linear diffusion approximation, generalizes diffusion map consistency from i.i.d. manifold samples to Langevin diffusion time series, and proves convergence of ergodic averages under standard spectral assumptions. In equity-portfolio stress testing with 132 macroeconomic factors plus return data over 1967–2016, benchmarked against standard scenario analysis, static PCA, and dynamic PCA, the method reports mean absolute error reductions of up to 55% and 39% for scenario-based portfolio return prediction, respectively, together with VaR exception tests that are statistically consistent with realized returns.

## 6. Distribution-aware and robust formulations

The 2025 composite-quantile DAFM defines a common factor vector shared across quantiles while allowing loadings to vary with the quantile. For panel data $\{X_{it}\}$ and quantiles $0<\tau_1,\dots,\tau_K<1$, it assumes
$$
Q_{X_{it}}(\tau_k \mid f_t)=\lambda_{k,i}^\prime f_t,\qquad k=1,\dots,K,
$$
with common factors $f_t\in\mathbb{R}^r$ and quantile-specific loadings $\lambda_{k,i}\in\mathbb{R}^r$. Estimation minimizes the weighted composite check-loss
$$
M_{NT}(\theta)
=
\frac{1}{NT}\sum_{k=1}^K\sum_{i=1}^N\sum_{t=1}^T
w_k\,\rho_{\tau_k}\!\left(X_{it}-\lambda_{k,i}^\prime f_t\right),
$$
where
$$
\rho_\tau(u)=u\big(\tau-\mathbf{1}\{u\le 0\}\big).
$$
Because the objective is jointly nonconvex, the paper uses alternating minimization between loadings and factors, with ADMM mentioned for the factor step. The weights $w_k$ are presented as data-adaptive: they can be uniform or can emphasize quantiles that are more informative for the data distribution. The paper proves consistency at rate $O_p(L_{NT}^{-1})$, derives asymptotic normality via a kernel-smoothed objective, and proposes two consistent estimators of the number of factors. In simulations, DAFM generally outperforms single-quantile QFM, robust PCA, and CQF-H, especially under heavy tails, skewness, bimodality, or scale effects hidden from mean or median factor models. In empirical applications, it estimates two factors for CRSP stock return panels and reports that DAFM factors explain common idiosyncratic volatility much better than mean factors, CQF-H, or single-quantile QFM, with $R^2$ values above 0.7; it also improves short-horizon U.S. unemployment forecasting relative to AR, AR+PCA, AR+CQF-H, and AR+QFM [2510.00558].

A separate robust DAFM-style formulation addresses the low-rank-plus-diagonal covariance decomposition directly. Starting from the static factor model $\xi=\Phi\alpha+\omega$, with covariance decomposition $\Sigma=L+D$, it replaces the unknown covariance by an ambiguity set around the empirical covariance,
$$
B_{\dist}=\{\Sigma \succeq 0:\dist(\Sigma,\widehat{\Sigma})\le \varepsilon\},
$$
and solves
$$
J^\star
=
\min_{L,D}\ \operatorname{Tr}(L)
\quad
\text{s.t.}\quad
L\in S,\ D\in\mathbb D,\ L+D\in B_{\dist}.
$$
Dualization yields a saddle-point problem whose core subroutine is the linear minimization oracle
$$
O(\Lambda)=\arg\min_{\Sigma}\{\langle\Lambda,\Sigma\rangle:\Sigma\in B_{\dist}\}.
$$
The paper derives semi-closed-form LMOs for Frobenius, Kullback–Leibler, and Gelbrich distances, together with explicit Lipschitz constants for the dual function, and proposes the first-order iteration
$$
\Sigma_t = O(\Lambda_t),\qquad
\Lambda_{t+1} = \Pi_{S_1\cap S_2}\big[\Lambda_t+\delta \Sigma_t\big].
$$
The reported convergence guarantee is
$$
0 \le \langle \Lambda^\star,\Sigma^\star\rangle
-
\langle \bar{\Lambda}_T,\bar{\Sigma}_T\rangle
\le
\frac{\|\Lambda_1-\Lambda^\star\|^2}{2\delta T}
+\frac{\delta}{2}L^2.
$$
Numerically, for $n=20$, $r=4$, $N=15n$, and $\varepsilon=1$, normalized objective error after about 100 iterations is approximately 0.001% for Frobenius, 0.01% for KL, and $5\times10^{-5}\%$ for Gelbrich. The paper also reports improvements over the empirical covariance in 59% of Frobenius experiments, 56% of Gelbrich experiments, and 43% of KL experiments, and states that the method becomes more efficient than MOSEK in high dimensions, especially for $n\ge 150$–$200$ [2506.09776].

## 7. Time-varying loading spaces, inference, and interpretive issues

Adaptive estimation for non-stationary factor models extends the DAFM idea to locally stationary time series with evolving loading spaces. The model is
$$
x_{i,n}=A(i/n)\,z_{i,n}+e_{i,n}, \qquad i=1,\dots,n,
$$
where $A(t)\in\mathbb R^{p\times d(t)}$ may have both time-varying entries and time-varying dimension, while the latent factor and idiosyncratic components are locally stationary. The estimation target is the span of the varying loading matrix, recovered through the time-varying lag-$k$ covariance matrix
$$
\Sigma_x(t,k)=E\{G(t,F_{i+k})G^\top(t,F_i)\},
$$
and the nonnegative definite matrix
$$
\Lambda(t)=\sum_{k=1}^{k_0}\Sigma_x(t,k)\Sigma_x^\top(t,k).
$$
The paper approximates $\Sigma_x(t,k)$ by a sieve basis expansion and constructs
$$
\hat\Lambda(t)=\sum_{k=1}^{k_0}\hat M(J_n,t,k)\hat M^\top(J_n,t,k),
$$
so that the estimated loading space at time $t$ is given by the leading eigenvectors of $\hat\Lambda(t)$ [2012.14708].

The number of factors is estimated by a penalized eigen-ratio criterion. For fixed $d$,
$$
\hat d_n
=
\arg\max_{1\le i\le p}
\frac{\lambda_{i+1}(\bar{\hat\Lambda})+q_n}{\lambda_i(\bar{\hat\Lambda})+q_n},
$$
and in the time-varying case,
$$
\hat d_n(t)
=
\arg\max_{1\le i\le p}
\frac{\lambda_{i+1}(\hat\Lambda(t))+q_n}{\lambda_i(\hat\Lambda(t))+q_n}.
$$
The paper proves uniform consistency of the effective factor number and develops a bootstrap-assisted test for the hypothesis of static factor loadings,
$$
H_0:\ \mathrm{span}(A(t))=\mathrm{span}(A)\quad\text{for all }t.
$$
Its blockwise maximum-deviation statistic is
$$
\hat T_n
=
\sqrt{m_n}\max_{1\le h\le N_n}\max_{1\le i\le p-\tilde d_n}
\left|\hat f_i^\top S_h^X\right|,
$$
with critical values approximated by a multiplier bootstrap with overlapping blocks.

The theoretical message is that adaptation can concern not only coefficients but the latent dimension itself. The paper states uniform convergence rates for the sieve estimator, consistency of the factor-number estimator, asymptotic correctness of the bootstrap test, and power against local alternatives. In simulations, the adaptive sieve estimator improves over static-loading PCA and local PCA; the factor-number estimator is reported to be stable and accurate; and the static-loadings test shows correct size and good power. In a UK temperature application, monthly highest temperature strongly rejects static loading, whereas monthly lowest temperature does not.

Taken together, these lines of work indicate that DAFM should not be equated with a single statistical doctrine. Some DAFM formulations are latent and continuous-time; some are built from interpretable traded basis assets; some are explicitly online and point-in-time; some are supervised through joint covariate-response embeddings; some are distribution-aware through composite quantiles; and some are robustified through covariance ambiguity sets. A plausible implication is that the most precise use of the term is contextual: in any given paper, “data-adaptive” identifies the component of the factor model that is allowed to evolve with the empirical structure of the data, rather than being fixed in advance.

Source: https://www.emergentmind.com/topics/data-adaptive-factor-model-dafm