---
title: Multi-Level Dynamic Factor Models
url: https://www.emergentmind.com/topics/multi-level-dynamic-factor-models-ml-dfms
type: topic
---

# Multi-Level Dynamic Factor Models

Multi-Level Dynamic Factor Models (ML-DFMs) are dynamic latent-variable models that extend the standard dynamic factor model by imposing a hierarchical loading structure on a large panel of observed series. In the canonical formulation, a small set of pervasive factors affects all variables, while additional group-specific factors affect only variables inside their own block; some formulations also add semipervasive factors shared by subsets of blocks. Across recent work, ML-DFMs appear in at least three closely related forms: block-structured approximate factor models estimated by sequential least squares, multidimensional state-space systems linking macro and micro data, and matrix-structured models with global and local factors. At the same time, several nearby models—deep nonlinear DFMs, structured matrix DFMs, and static multilevel covariance models—are methodologically relevant but are not, strictly speaking, ML-DFMs [2602.14813][2507.10679][2301.12499][2310.13911].

## 1. Conceptual definition and scope

The basic ML-DFM enriches the standard approximate factor model by separating variation into common components that are pervasive across the full cross section and local components that are restricted to known groups. With groups indexed by \(s=1,\dots,S\), variables indexed by \(i=1,\dots,N_s\), and time indexed by \(t\), one canonical observation equation is
\[
y_{s,it}=\boldsymbol{\lambda}_{s,i}^{(g)\prime} \mathbf{G}_t+\boldsymbol{\lambda}_{s,i}^{(l)\prime} \mathbf{L}_{s,t}+\boldsymbol{\varepsilon}_{s,it},
\]
where \(\mathbf{G}_t\) collects global factors, \(\mathbf{L}_{s,t}\) collects group-specific factors, and the local loading blocks are zero outside their own group. In stacked form, the loading matrix is block-structured, with global loadings stacked across groups and local loadings arranged block diagonally [2602.14813].

A related block formulation partitions the observed vector as \(X_t=(X_{1,t},\dots,X_{K,t})'\) and writes
\[
X_t=P^{*}F_t^{*}+\epsilon_t,
\]
where \(F_t^{*}\) can include pervasive factors \(G_t\), block-specific factors \(F_{k,t}\), and, in overlapping-block designs, pairwise semipervasive factors such as \(F_{12,t},F_{13,t},F_{23,t}\). This extension is useful when some latent sources are shared only across selected blocks rather than the entire panel [2507.10679].

This suggests that “multi-level” refers first to the loading architecture rather than to any single estimation routine. In the papers considered here, the hierarchy may be non-overlapping, overlapping, additive, or embedded in a state-space system, but the defining feature is the imposition of structured zero restrictions or equality constraints that distinguish common and group-level information from idiosyncratic noise [2602.14813][2507.10679][2301.12499].

## 2. Core mathematical structure and identification

The compact stacked ML-DFM representation is
\[
\mathbf{Y}_t=\mathbf{\Lambda}^{\dagger} \mathbf{F}_t^{\dagger}+\boldsymbol{\varepsilon}_t,
\qquad
\mathbf{F}_t^{\dagger}=\left(\mathbf{G}_t,\mathbf{L}_{1,t},...,\mathbf{L}_{S,t} \right)^{\prime},
\]
with \(\mathbf{\Lambda}^{(l)}=\operatorname{diag}(\mathbf{\Lambda}_1^{(l)},\dots,\mathbf{\Lambda}_S^{(l)})\). The decisive structural feature is the block zero restriction: local factors load only within their own group [2602.14813].

In non-overlapping block systems, the measurement equation can be written explicitly as
\[
\begin{bmatrix} X_{1,\cdot t} \\ \vdots \\ X_{K,\cdot t} \end{bmatrix}
=
\begin{bmatrix}
\boldsymbol\mu_{1} & \boldsymbol\lambda_1 & 0 & \dots & 0 \\
\boldsymbol\mu_{2} & 0 & \boldsymbol\lambda_2 & \dots & 0 \\
\vdots &  & 0 & \ddots & 0 \\
\boldsymbol\mu_{K} & 0 & 0 & \dots & \boldsymbol\lambda_{K}
\end{bmatrix}
\begin{bmatrix}
G_{t}\\
F_{1,t}\\
F_{2,t}\\
\vdots\\
F_{K,t}
\end{bmatrix}
+
\begin{bmatrix}
\epsilon_{1,\cdot t}\\
\vdots\\
\epsilon_{K,\cdot t}
\end{bmatrix},
\]
while overlapping-block models augment the latent vector with pairwise semipervasive factors. The hierarchy is therefore not merely “more factors”; it is a factorization with economically meaningful loading restrictions [2507.10679].

Identification follows the usual factor-model logic, but admissible rotations must preserve the zero restrictions. One set of restrictions normalizes local factors and global factors separately:
\[
\frac{\mathbf{L}_s^{\prime} \mathbf{L}_s}{T}=\mathbf{I}_{r_s},
\qquad
\mathbf{\Lambda}_s^{(l)\prime} \mathbf{\Lambda}_s^{(l)} \text{ diagonal},
\]
\[
\frac{\mathbf{G}^{\prime}\mathbf{G}}{T}=\mathbf{I}_{r_g},
\qquad
\mathbf{\Lambda}^{(g)\prime} \mathbf{\Lambda}^{(g)} \text{ diagonal},
\]
and imposes orthogonality between global and local factors,
\[
\frac{\mathbf{L}_{s}^{\prime} \mathbf{G}}{T}=\mathbf{0}_{r_s\times r_g}.
\]
A notable technical point is that group-specific factors from different groups need not be orthogonal for identification, although enough normalization and orthogonality is imposed to place estimated and true factors in the same identified space [2602.14813].

In sequential least-squares implementations without an explicit normalization, only the common component \(P^{*}F_t^{*}\) is initially just-identified; the separate factors and loadings are identified only up to rotation. The normalization is then recovered through orthogonalization across levels and principal-components normalization of estimated common components, with the orthogonalization implemented recursively via regressions equivalent to Gram-Schmidt [2507.10679].

## 3. Estimation, factor extraction, and inference

A central estimation route is the sequential least squares (SLS) procedure associated with Breitung and Eickmeier. In the formulation studied for ML-DFM inference, estimation proceeds by: extracting \(r_g+r_s\) factors by PC within each group, using Canonical Correlation Analysis (CCA) across groups to estimate common global combinations, regressing the data on initial global factors to obtain residuals, extracting local factors within each group, then alternating least-squares updates for loadings and factors until the residual sum of squares converges. After convergence, the estimated factors are renormalized so that the factor covariance matrices equal identity and the loading cross-products are diagonal [2602.14813].

The FARS implementation generalizes the same sequential logic to non-overlapping and overlapping blocks. Initial global factors can be obtained by CCA or PCA; global components are filtered out block by block; lower-level factors are then extracted sequentially; block-specific factors are finally obtained by PC on residualized blocks; and the algorithm iterates between factor and loading updates until the residual sum of squares converges. The package exposes this structure through `mldfm()`, with `method = 0` for CCA initialization and `method = 1` for PCA initialization [2507.10679].

For inference, the key recent question is whether Bai’s asymptotic distribution for PC factors in standard DFMs can approximate the empirical distribution of SLS-estimated global and group-specific ML-DFM factors. The benchmark law is
\[
\sqrt{N}\left( \widetilde{\mathbf{F}_{t}-\mathbf{F}_{t}\right)\stackrel{d}{\rightarrow} N \left( 0,\mathbf{\Sigma} _{\Lambda}^{-1}\mathbf{\Gamma} _{t}\mathbf{\Sigma} _{\Lambda}^{-1}\right),
\]
with
\[
\mathbf{\Gamma}_t = \lim_{N\to\infty} \frac{1}{N} \sum_{i=1}^N\sum_{j=1}^N \boldsymbol{\lambda}_i'\boldsymbol{\lambda}_j E(\varepsilon_{it}\varepsilon_{jt}).
\]
Monte Carlo evidence shows that this approximation works well for moderate to large samples for both global and group-specific SLS factors, especially for the global factor and for local factors with stronger loadings. When local loadings are relatively weak, the asymptotic covariance can approximate empirical covariance better than full empirical mean squared error because finite-sample bias remains non-negligible [2602.14813].

The practical covariance estimator is
\[
\widehat{Avar}(\widetilde{\mathbf{F}_t})=\frac{1}{N}
\left(
\frac{\widetilde{\mathbf{\Lambda}^{\prime} \widetilde{\mathbf{\Lambda}}}{N}
\right)^{-1}
\widetilde{\mathbf{\Gamma}_t}
\left(
\frac{\widetilde{\mathbf{\Lambda}^{\prime} \widetilde{\mathbf{\Lambda}}}{N}
\right)^{-1}.
\]
Two alternatives for \(\widetilde{\mathbf{\Gamma}_t}\) are emphasized. The heteroscedasticity-robust Bai–Ng style estimator assumes no cross-sectional idiosyncratic correlation, whereas the FPR estimator uses a thresholded estimate of the idiosyncratic covariance matrix and therefore allows cross-sectional dependence. Across the simulations, the preferred uncertainty estimator is FPR with subsampling correction for loading-estimation uncertainty; when cross-sectional dependence is absent, HR and FPR perform similarly, but when dependence is present, FPR clearly dominates HR [2602.14813].

The same inferential logic underlies factor confidence regions in FARS. The corrected MSE combines the plug-in asymptotic term with a cross-sectional subsampling term,
\[
MSE^{\ast}_t =\frac{1}{N}\left(\frac{\hat{P}^{\prime}\hat{P}}{N} \right)^{-1}
\hat{\Gamma}_t
\left(\frac{\hat{P}^{\prime}\hat{P}}{N} \right)^{-1}
+
\frac{N^*}{NS}\sum_{s=1}^{S}
\left(
\hat{F}_{t}^{\ast (s)}-\hat{F}_{t}
\right)
\left(
\hat{F}_{t}^{\ast (s)}-\hat{F}_{t}
\right)^{\prime},
\]
and yields confidence ellipsoids
\[
g(F_t,\alpha)=\{F_t\in{\rm I\!R}^r \mid(F_t-\hat{F_t}) MSE_t^{*-1}(F_t-\hat{F_t})\leq \chi^2_{r(\alpha)}\}.
\]
These ellipsoids are then used for stress design [2507.10679].

## 4. State-space, multidimensional, and matrix-structured generalizations

A major extension replaces a fixed \(n\times T\) panel with multidimensional dependent data whose cross section changes over time. In the multidimensional dynamic factor model, the observed object at time \(t\) is a matrix \(H_t\in\mathbb{R}^{N_t\times K}\), which is vectorized and embedded into a fixed-dimension series \(Y_t=W_t h_t\), \(Y_t\in\mathbb{R}^{NK\times 1}\), with missingness treated as intrinsic rather than as zero padding. The generic model is
\[
Y_t = B(L)\Phi_t + e_t, \qquad e_t \overset{w.n.}{\sim} N(0_{NK\times 1},R),
\]
\[
\Phi_t = C(L)\Phi_{t-1} + D u_t, \qquad u_t \overset{w.n.}{\sim} N(0_{r\times 1},\Sigma),
\]
which yields a state-space ML-DFM capable of processing repeated cross-sections, missing values, and asynchronous releases [2301.12499].

In the household–macro specialization, the observed vector is
\[
Y_t = \big(X_t' \;\, Z_{1,t}' \;\, \ldots \;\, Z_{4,t}'\big)',
\]
with \(X_t\) the vector of 8 macroeconomic aggregates and \(Z_{i,t}\) the vector of real income per head for households in demographic group \(i\). The decomposition combines smooth trends \(\tau_{j,t}\), a common business-cycle factor \(\psi_t\), group- and series-specific idiosyncratic latent cycles \(\xi_{j,t}\), and measurement noise \(e_t\). For micro incomes, all households in the same demographic group share the same trend and business-cycle loading through replication by \(\iota_{\omega_g}\), yielding a hierarchical additive dynamic factor structure rather than a tensor factor model [2301.12499].

Estimation in that framework uses penalized quasi maximum likelihood with an ECM algorithm built around a Kalman smoother. The E-step accommodates arbitrary missingness through time-varying observed sets \(Y_t^{obs}\) and observation matrices \(B_t^{obs}\); the CM-step updates initial conditions, transition coefficients, innovation variances, and loadings under structural equality constraints. The model can therefore process rotating survey participation, ragged edges, and publication lags while remaining operational for real-time updating, localized predictions, counterfactuals, and impulse response functions [2301.12499].

Matrix-valued observations lead to a different but related generalization. The multilevel matrix factor model observes \(\boldsymbol{X}_{mt}\in\mathbb{R}^{N_m\times p}\) and writes
\[
\boldsymbol{X}_{mt} = \boldsymbol{R}_m \boldsymbol{G}_{t} \boldsymbol{C}_{m}^{'} +\boldsymbol{\Gamma}_{m} \boldsymbol{F}_{mt} \boldsymbol{\Lambda}_{m}^{'} +\boldsymbol{E}_{mt},
\]
where \(\boldsymbol{G}_t\) is global and \(\boldsymbol{F}_{mt}\) is local to group \(m\). In vectorized form,
\[
\mathrm{vec}(\boldsymbol{X}_{mt}) = (\boldsymbol{C}_m\otimes \boldsymbol{R}_m)\mathrm{vec}(\boldsymbol{G}_t) + (\boldsymbol{\Lambda}_m\otimes \boldsymbol{\Gamma}_m)\mathrm{vec}(\boldsymbol{F}_{mt}) + \mathrm{vec}(\boldsymbol{E}_{mt}),
\]
so the model is a multilevel factor model with Kronecker-structured loadings. Estimation is sequential: identify global row and column loading spaces from cross-group covariance, project them out, estimate local loading spaces from residual matrix dynamics, then recover global and local signal parts. The identified objects are loading spaces rather than unique loading matrices, reflecting the usual rotational indeterminacy [2310.13911].

## 5. Methodological boundaries and neighboring models

Several recent models are closely related to ML-DFMs but should not be conflated with them. The most explicit boundary case is the deep dynamic factor model \(D^2FM\), which replaces the linear factor-observation map with a deep autoencoder while preserving dynamic latent states and idiosyncratic AR components. Its core equations are
\[
\boldsymbol{f}_{t} = G(\boldsymbol{y}_{t}),
\qquad
\boldsymbol{y}_{t} = F(\boldsymbol{f}_{t}) + \boldsymbol{\varepsilon}_{t},
\]
with AR laws for factors and idiosyncratic terms. However, it has one shared latent bottleneck \(\boldsymbol{f}_t\), no grouped or nested decomposition, no global versus local blocks, and no identification of global versus local shocks. Its relevance to ML-DFMs is therefore methodological rather than structural [2007.11887].

Dynamic matrix factor models are another neighboring class. They model matrix-valued time series through
\[
X_t=U_1 F_t U_2^{\top} + E_t,
\qquad
F_t=A_1 F_{t-1} A_2^{\top}+\varepsilon_t,
\]
or, after vectorization, through Kronecker-structured loadings and transitions. This yields a structured dynamic factor model with row and column loading spaces, but not an explicit hierarchy of global and group-specific factors. The structure is two-way rather than multilevel in the standard nested sense [2407.05624].

Parshakova, Hastie, and Boyd study a static Gaussian multilevel factor model with covariance \(\Sigma=FF^T+D\) constrained to have multilevel low-rank structure. The model is
\[
y = Fz + e,\qquad z \sim \mathcal N(0, I_s),\qquad e \sim \mathcal N(0, D),
\]
with no time index and no state equation. Its main contribution is computational: a fast EM algorithm, recursive Sherman–Morrison–Woodbury inversion, and Cholesky/determinant machinery with linear time and storage complexities per iteration under fixed hierarchy and rank allocation. This is directly relevant to ML-DFM computation on the observation side, but it is not itself dynamic [2409.12067].

A further source of confusion is the phrase “level” in models for means and volatilities. In the joint level–volatility dynamic factor model,
\[
X_{it} = B_i f_t + v_{it},
\qquad
r_{it} = \frac{e^{\mathcal{B}_i \mathcal{F}_t}}{\lambda_{it}},
\qquad
F_t=\begin{pmatrix}f_t\\ \mathcal{F}_t\end{pmatrix},
\]
the baseline contribution is a joint dynamic factor model for first moments and second moments, not a hierarchical ML-DFM over blocks or regions. The international inflation application does introduce a global plus regional factor decomposition for both means and volatilities, making the connection to ML-DFMs direct there, but “level” in the title refers to the conditional mean rather than to hierarchical level [2604.03681].

These distinctions matter because the recent literature uses related language—multidimensional, multilevel, deep, matrix, level, volatility—for models that solve different problems. A common misconception is therefore to treat any structured latent dynamic model as an ML-DFM. The papers surveyed here instead support a narrower usage: ML-DFMs require an explicit hierarchical or grouped factor architecture, not merely nonlinear encoding, matrix geometry, or multilevel covariance structure [2007.11887][2407.05624][2409.12067][2604.03681].

## 6. Empirical roles, scenario design, and substantive applications

One applied role of ML-DFMs is macro-financial scenario design. FARS uses ML-DFM factors as inputs to factor-augmented quantile regressions,
\[
q_{\tau^*}(y_{t+h} \mid y_t, F_t) = \mu(\tau^*, h) + \phi(\tau^*, h) y_t + \sum_{k=1}^{r} \beta_k(\tau^*, h) F_{kt},
\]
then recovers predictive densities by fitting a skew-\(t\) distribution to the implied quantiles. Stress scenarios are defined by optimizing a target quantile over the boundary of the factor confidence ellipsoid; the stressed factors are then propagated through the quantile regression and density-recovery steps. In the U.S. GDP illustration with three blocks—63 global macro variables, 248 domestic macro variables, and 208 global financial variables—the specification uses one global factor, one pairwise factor for blocks 1 and 3, and one local factor for each block. The application reports that GiS is more negative than GaR and that this gap persists across horizons \(h=1,\dots,4\) and becomes more adverse at higher stress levels \(\alpha\) [2507.10679].

A second role is micro–macro integration. The multidimensional state-space formulation jointly models 8 macro aggregates from ALFRED and quarterly household income from the Consumer Expenditure Public Use Microdata over 1990–2020, with about 87,000 households restricted to prime-age urban households. The empirical hierarchy has four levels in practice: time, macro series, four household groups defined by education and ethnicity, and households within groups. The model finds persistent income differentials by education and ethnicity, common cyclical co-movement across demographics driven by the business cycle, additional group-specific idiosyncratic movements in income, and the ability to nowcast delayed household survey outcomes using timely macro releases [2301.12499].

A third role is the decomposition of commonality across cross sections and moments. In the international inflation application of the level–volatility model, inflation is decomposed into a world factor, an AE regional factor, and an EMDE regional factor for both levels and volatilities:
\[
\pi_{i,t} = b_i^{w} f_{w,t} + s_i\, b_i^{a} f_{a,t} + (1-s_i)\, b_i^{e} f_{e,t} + v_{i,t},
\]
\[
\log r_{i,t} = \beta_i^{w}\mathcal F_{w,t} + s_i\beta_i^{a}\mathcal F_{a,t} + (1-s_i)\beta_i^{e}\mathcal F_{e,t} - \log\lambda_{i,t}.
\]
The reported substantive pattern is a dominant global level component in advanced economies and stronger regional and volatility contributions in emerging and developing economies, indicating that a standard mean-only hierarchy can miss important heterogeneity in common uncertainty [2604.03681].

Across these applications, a consistent theme is that ML-DFMs are valuable when the panel has known structure that a single pervasive factor block would treat too coarsely. This suggests a unifying interpretation: the central object is not merely a low-dimensional latent state, but a low-dimensional latent state with imposed structure that determines which parts of the panel are allowed to move together. In current work, that structure may describe blocks of variables, subsets of blocks, households within demographic groups, or, in matrix formulations, groups of matrix time series with shared global and local dynamics [2507.10679][2301.12499][2310.13911].

Source: https://www.emergentmind.com/topics/multi-level-dynamic-factor-models-ml-dfms