---
title: Vector Multiplicative Error Models (vMEM)
url: https://www.emergentmind.com/topics/vector-multiplicative-error-models-vmem
type: topic
---

# Vector Multiplicative Error Models (vMEM)

Searching arXiv for the specified papers on Vector Multiplicative Error Models and related developments.
Vector Multiplicative Error Models (vMEMs) are multivariate extensions of the Multiplicative Error Model (MEM) for jointly modeling multiple nonnegative time series. They represent an observed positive-valued vector as the element-by-element product of a conditional mean vector and a nonnegative innovation vector with unit mean, thereby combining autoregressive mean dynamics, dynamic cross-lag effects, and contemporaneous dependence across series. In the vMEM literature, this framework is used for market activity variables such as realized volatility, volume, number of trades, durations, ranges, and related positive indicators, with recent developments addressing copula-based dependence, semiparametric innovation laws, latent common components, and formal specification tests for the innovation distribution [1604.01338] [2107.05923].

## 1. Formal definition and core representation

A vMEM specifies a nonnegative vector process \(x_t \in \mathbb{R}_+^K\) or \(\mathbb{R}_+^d\) as
\[
x_t = \mu_t \odot \varepsilon_t = \operatorname{diag}(\mu_t)\varepsilon_t,
\]
where \(\mu_t\) is the conditional mean vector and \(\varepsilon_t\) is a nonnegative innovation vector satisfying \(E[\varepsilon_t \mid \mathcal{F}_{t-1}] = \mathbf{1}\). Under this normalization,
\[
E(x_t \mid \mathcal{F}_{t-1}) = \mu_t,
\qquad
\operatorname{Var}(x_t \mid \mathcal{F}_{t-1}) = \operatorname{diag}(\mu_t)\Sigma\operatorname{diag}(\mu_t),
\]
or, equivalently in an alternative notation,
\[
\operatorname{Var}[\mathbf{x}_t \mid \mathscr{F}_{t-1}] = \bm{\mu}_t \bm{\mu}_t' \odot \bm{\Sigma},
\]
with constant covariance matrix \(\Sigma\) for the innovations [1604.01338] [2107.04354].

The canonical conditional mean recursion introduces both own dynamics and cross-variable spillovers. A general specification is
\[
\mu_t = \omega + \sum_{l=1}^L [\alpha_l x_{t-l} + \gamma_l x_{t-l}^{(-)} + \beta_l \mu_{t-l}],
\]
where \(\omega\) is \(K \times 1\), while \(\alpha_l\), \(\beta_l\), and \(\gamma_l\) are \(K \times K\). The term \(x_t^{(-)}\) introduces asymmetry through signed indicators, such as the sign of returns or a buy/sell dummy. A simpler baseline form widely used in the literature is
\[
\bm{\mu}_{t} = \bm{\omega} + \mathbf{B}\bm{\mu}_{t-1} + \mathbf{A}\mathbf{x}_{t-1},
\]
with off-diagonal entries in \(A\) and \(B\) allowing cross-series spillovers and interdependence [1604.01338] [2107.04354].

The framework is motivated by the same persistence and feedback considerations that underlie GARCH-type volatility models, but it is tailored to positive-valued observables. The literature explicitly states that MEM is borne of an extension of the popular GARCH approach for modeling and forecasting conditional volatility of asset returns, while vMEM generalizes this logic to multiple positive indicators. This makes vMEM particularly suitable when the target variables are positive by construction and when direct modeling on the original scale is desirable [2107.05923].

## 2. Conditional mean dynamics, positivity, and stationarity

Positivity of the conditional mean is central. In the linear vMEM, sufficient conditions are elementwise positivity of \(\omega\) and nonnegativity of the entries of \(\alpha_l\), \(\beta_l\), and \(\gamma_l\), which ensure \(\mu_t \ge 0\) for nonnegative \(x_t\). In alternative log-link formulations,
\[
\log \mu_t = \omega + A \log x_{t-1} + B \log \mu_{t-1} + D z_{t-1} + \cdots,
\]
positivity of \(\mu_t\) is automatic through exponentiation [1604.01338] [2107.05923].

For mean stationarity, the literature gives spectral-radius conditions analogous to multivariate GARCH restrictions. In the \(vMEM(1,1)\) with asymmetry, define \(A=\alpha+\beta+\gamma/2\); all eigenvalues of \(A\) must have modulus less than \(1\). Under \(L\) lags, define \(A_l=\alpha_l+\beta_l+\gamma_l/2\) and the companion matrix \(A^*\); stationarity holds if all eigenvalues of \(A^*\) lie inside the unit circle. In the canonical linear form without the asymmetry term, a sufficient condition is
\[
\rho(A+B)<1,
\]
where \(\rho(\cdot)\) denotes the spectral radius. Under stationarity, the unconditional mean exists and is
\[
E[x_t] = (I-A-B)^{-1}\omega,
\]
or, in the asymmetric expectation-targeted form,
\[
\mu = \left[I_K-\sum_{l=1}^L(\alpha_l+\beta_l+\gamma_l/2)\right]^{-1}\omega.
\]
These results make the unconditional mean an explicit function of the mean-dynamics parameters [1604.01338] [2107.05923].

The literature also develops component decompositions. In the components vMEM, or SpvMEM, the conditional mean is written as
\[
\mu_t = \tau_t \mu \odot \xi_t,
\]
where \(\tau_t\) is a scalar slow-moving component with \(E[\tau_t]=1\), \(\mu\) is the unconditional mean vector estimated by sample averages, and \(\xi_t\) is a short-run component with \(E[\xi_t]=1\). This decomposition is designed to separate slow/low-frequency variation from fast/high-frequency dynamics, and the applications in the survey paper state that the presence of a slow moving low-frequency component can improve the properties of the estimated models [2107.05923].

A further structural development is the 2026 log-vMEM with Spillover Effects and Co-movement, denoted vMEM-SeC. In that formulation,
\[
\ln \mu_t = \varsigma_t + \vartheta \xi_t,
\]
where \(\varsigma_t\) is an idiosyncratic component and \(\xi_t\) is a scalar common latent component driven by the first principal component of demeaned log-volatilities. The paper states that this rank-1 additive structure parsimoniously captures common co-movement with heterogeneous asset-specific loadings and separates pure spillovers from common market dynamics [2601.16837].

## 3. Innovation laws and contemporaneous dependence

Classical vMEM requires a nonnegative innovation law with unit means, but the multivariate case poses an identification and flexibility problem because sufficiently flexible multivariate nonnegative densities are scarce. The copula-based specification addresses this by combining tractable marginals with a copula density. If \(u_{i,t}=F_i(\varepsilon_{i,t};\phi_i)\), the joint density factorizes as
\[
f(\varepsilon_t) = c(u_{1t},\ldots,u_{Kt};\xi)\prod_{i=1}^K f_i(\varepsilon_{i,t};\phi_i),
\]
where \(c(\cdot;\xi)\) is the copula density and \(f_i\) are the marginal densities [1604.01338].

Gamma marginals with unit mean are standard. One specification is
\[
\varepsilon_{i,t}\sim \operatorname{Gamma}(\phi_i,\phi_i),
\]
with density
\[
f_i(\varepsilon)=\phi_i^{\phi_i}\Gamma(\phi_i)^{-1}\varepsilon^{\phi_i-1}e^{-\phi_i\varepsilon},\qquad \varepsilon>0,
\]
which has mean \(1\) and variance \(1/\phi_i\). The survey also lists Log-normal, Beta-prime, and Log-logistic marginals parameterized to ensure \(E[\varepsilon]=1\) and \(\operatorname{Var}[\varepsilon]=\sigma^2\). Other possibilities mentioned in the literature include Inverse-Gamma, Weibull, and zero-augmented mixtures [1604.01338] [2107.05923].

The copula literature for vMEM emphasizes elliptical copulas. The Gaussian copula with correlation matrix \(R\) is tractable, allows negative dependence, and has asymptotically independent tails. The Student-\(t\) copula uses correlation matrix \(R\) and degrees of freedom \(\nu\); it exhibits symmetric tail dependence, and with \(R=I\) it yields uncorrelated but dependent margins. This distinction is central because the copula is used to model contemporaneous dependence in innovations separately from lagged dependence in \(\mu_t\) [1604.01338].

More recent work treats the innovation law itself as an object of semiparametric inference. The Bayesian semiparametric vMEM models \(\varepsilon_t\) through a Dirichlet process mixture of multivariate log-normal kernels supported on the positive orthant. In stick-breaking form,
\[
f_{\varepsilon}(e)=\sum_{j=1}^{\infty} w_j\, \text{logN}_d(e\mid m_j,\Sigma_j),
\qquad
\mathbf{w}\sim \mathrm{GEM}(\alpha).
\]
To maintain identifiability of mean dynamics, the model recenters the mixture so that \(E_g[\varepsilon]=\iota_d\). The paper presents this as a way to preserve unit means while allowing skewness, multimodality, heavy tails, and cross-sectional dependence beyond classical parametric forms [2107.04354].

The 2026 log-vMEM-SeC adopts a different innovation strategy. It assumes
\[
\varepsilon_t \sim \ln N(m,V),
\]
with \(\ln \varepsilon_t \sim N(m,V)\) and \(m_i=-v_{ii}/2\) to guarantee \(E(\varepsilon_{i,t})=1\). Dependence across innovations is then encoded directly through the covariance matrix \(V\), without an independence restriction in the likelihood-based approach [2601.16837].

These approaches address the same structural issue from different angles. Copula-based models retain marginal tractability and flexible dependence; the semiparametric Bayesian model relaxes parametric shape restrictions on the innovation law; the log-normal log-vMEM obtains a tractable high-dimensional likelihood by working in logarithms. This suggests that the main trade-off in vMEM specification is between distributional flexibility, computational tractability, and dimensional scalability [1604.01338] [2107.04354] [2601.16837].

## 4. Estimation, inference, and diagnostic methodology

Likelihood-based inference starts from the transformation
\[
\varepsilon_t = x_t \odot \mu_t^{-1},
\]
with the Jacobian term induced by \(x_t=\operatorname{diag}(\mu_t)\varepsilon_t\). For the copula-based vMEM, the log-likelihood is
\[
\ell = \sum_{t=1}^T \left[\ln f(\varepsilon_t\mid \mathcal{F}_{t-1}) - \sum_{i=1}^K \ln \mu_{t,i}\right],
\]
and
\[
\ln f(\varepsilon_t\mid \mathcal{F}_{t-1})
=
\ln c(u_t;\xi)+\sum_{i=1}^K \ln f_i(\varepsilon_{i,t};\phi_i).
\]
The paper derives score expressions for the mean parameters, copula parameters, and marginal parameters, as well as expected information and Hessian formulas. Under correct specification, \(H^{(\varepsilon)}=-\mathcal{I}^{(\varepsilon)}\), and robust QML-type inference is available through the score/Hessian representation [1604.01338].

Expectation targeting is a major device for dimensionality reduction and numerical stability. Rather than estimating \(\omega\) directly, one reparameterizes it through the unconditional mean \(\mu\):
\[
\mu = \left[I_K-\sum_{l=1}^L(\alpha_l+\beta_l+\gamma_l/2)\right]^{-1}\omega.
\]
Then \(\mu\) is estimated first by the sample mean \(\bar{x}_T\), and the likelihood is maximized conditionally in a two-step procedure. The asymptotic covariance of the second-step estimator accounts for first-step uncertainty using Newey–McFadden two-step theory. The same principle appears in the log-vMEM literature, where expectation targeting provides the intercept reparameterization
\[
\omega=(I_n-A-B)\bar{x} + (I_n-B)\operatorname{diag}(V)/2
\]
in the log-normal specification [1604.01338] [2601.16837].

The survey paper emphasizes that GMM is recommended for multivariate components vMEM. With
\[
A_t = \frac{\partial \xi_t'}{\partial \theta}\,\operatorname{diag}(\xi_t)^{-1}\Sigma^{-1},
\]
efficient GMM solves
\[
\sum_{t=1}^T A_t(\varepsilon_t-\mathbf{1})=0,
\]
and its asymptotic covariance is the inverse of the limit of \(T^{-1}\sum_t E[A_t\Sigma A_t']\). The same source notes that in univariate MEM with Gamma errors, the first-order conditions for mean parameters do not depend on the Gamma scale, which motivates QML in the vector case as well [2107.05923].

Bayesian semiparametric inference introduces a different computational architecture. The parameter-expanded DPMLN2-vMEM uses a slice sampler with latent allocations and stick-breaking weights. To avoid the positivity and unit-mean constraints directly in sampling, the paper works with an unconstrained parameter-expanded model and post-processes posterior draws back to the constrained, identifiable formulation. Dynamic parameters are updated with an adaptive random-walk Metropolis–Hastings proposal, while mixture components are updated using Normal–Wishart conjugacy. The reported formulation is explicitly designed to make posterior inference feasible despite the positive-orthant constraint [2107.04354].

Diagnostics operate at both the mean-dynamics level and the innovation-law level. Standardized residuals
\[
u_t = x_t \oslash \mu_t
\]
should have unit mean and weak serial and cross correlation if the mean is correctly specified. The literature recommends Ljung–Box tests on \(u_{t,i}\), filtered series such as \(w_t\), and standardized residuals \(\varepsilon_{i,t}\). For copula models, probability integral transforms
\[
U_{i,t}=F_i(\varepsilon_{i,t};\phi_i)
\]
should be i.i.d. Uniform\((0,1)\), and dependence can be checked against Kendall’s \(\tau\) or tail-dependence indices. The survey also mentions Anderson–Darling and Cramér–von Mises tests for marginal distributional fit [1604.01338] [2107.05923].

A recent line of work treats the innovation law itself as the null hypothesis of a formal test. The 2025 specification-testing paper proposes weighted integrated distances between the empirical Laplace transform of residuals and the parametric Laplace transform under the null:
\[
S_{T,W}
=
T\int_{\mathbb{R}_+^d}
\big[\widehat{L}_T(u)-L_{\widehat{\varphi}_T}(u)\big]^2
W(u)\,du.
\]
When the null Laplace transform is unavailable in closed form, the paper proposes an artificial-sample statistic
\[
\Psi_{T,W}
=
\frac{TS}{T+S}
\int_{\mathbb{R}_+^d}
\big[\widehat{L}_T(u)-\widehat{L}_S^*(u)\big]^2
W(u)\,du,
\]
with bootstrap critical values. The Monte Carlo study reports that \(S_{T,W}\) is well-sized and generally more powerful than the artificial-sample alternative, while averaging over more artificial samples increases power [2509.06732].

## 5. Forecasting, model comparison, and empirical evidence

Under the unit-mean innovation restriction, point forecasting is straightforward:
\[
E(x_{t+1}\mid \mathcal{F}_t)=\mu_{t+1},
\qquad
E(x_{t+h}\mid \mathcal{F}_t)=\mu_{t+h}.
\]
Multi-step forecasts iterate the mean recursion. In detrended applications, forecasts of the original series are obtained by multiplying the detrended forecast by the last in-sample trend estimate, under the “frozen trend” assumption. Density forecasts are obtained by simulating future innovations from the estimated copula or semiparametric mixture and then forming \(x_{t+h}=\mu_{t+h}\odot \varepsilon_{t+h}\) [1604.01338] [2107.04354].

The copula-based paper evaluates forecast accuracy using two loss functions around \(\mu_t\):
\[
e_{N,t}=\tfrac12(x_t-\mu_t)^2,
\qquad
e_{G,t}=\ln(x_t/\mu_t)-x_t/\mu_t-1,
\]
interpreted respectively as Normal and Gamma loss functions, together with Diebold–Mariano tests. The survey additionally lists QLIKE,
\[
\operatorname{QLIKE}(x_{t+1},\hat{\mu}_{t+1|t})
=
x_{t+1}/\hat{\mu}_{t+1|t}
-
\log(x_{t+1}/\hat{\mu}_{t+1|t})
-
1,
\]
as a standard out-of-sample criterion for positive data [1604.01338] [2107.05923].

A widely cited empirical application studies daily realized kernel volatility, trading volume, and number of trades for Johnson & Johnson from January 3, 2007 to July 31, 2013, with \(T=1{,}656\). On detrended data, the conditional mean is specified as
\[
\mu_t=\omega+\alpha_1 x_{t-1}+\alpha_2 x_{t-2}+\gamma_1 x_{t-1}^{(-)}+\beta_1\mu_{t-1}.
\]
Competing mean structures include a diagonal benchmark, a model with full \(\alpha_1\) and diagonal \(\beta_1\), and a model with both \(\alpha_1\) and \(\beta_1\) full. With Gamma marginals and independent, Normal, or Student-\(t\) copulas, the paper reports that the Student-\(t\) copula is preferred, with \(\nu \approx 9\) in the A-T specification and \(\nu \approx 8.7\) in the AB-T specification, both highly significant. Estimated contemporaneous correlations are strong and stable: \(\operatorname{corr}(\text{rkv},\text{vol}) \approx 0.48\)–\(0.50\), \(\operatorname{corr}(\text{rkv},\text{nt}) \approx 0.61\)–\(0.63\), and \(\operatorname{corr}(\text{vol},\text{nt}) \approx 0.90\)–\(0.92\). Among the full models, AB-T attains the highest log-likelihood, approximately \(2{,}168\), and out-of-sample Diebold–Mariano statistics versus the diagonal independent benchmark reach about \(2.09\) for \(e_N\) and about \(1.81\) for \(e_G\) on the original series. The paper concludes that copula-based contemporaneous correlation and the inclusion of volume and number of trades yield significantly superior realized volatility forecasts [1604.01338].

The Bayesian semiparametric vMEM presents both simulation and empirical evidence. In a trivariate simulation with \(T=3000\), posterior means recover the true parameters well, credible intervals cover the truth, effective sample sizes exceed \(327\), and the mixture typically uses about \(6\) active components, with only two carrying most weight. The average log pseudo marginal likelihood is reported as \(-9.0578\) versus \(-9.1023\) for the parametric log-normal vMEM, favoring the semiparametric model. In empirical analysis on bivariate \((|r_t|,rv_t)\) series for the S&P 500, DJIA, and FTSE 100, in-sample predictive performance measured by LPS and LPML favors DPMLN2-vMEM over LN1-vMEM across all three indices [2107.04354].

The 2026 vMEM-SeC application uses daily high–low range volatility proxies for \(29\) DJIA assets from March 19, 2008 to April 22, 2024, with \(T=4{,}051\). The first principal component of demeaned log-volatilities explains \(57\%\) of total variance. Among the SeC models, d-vMEM-SeC achieves the highest log-likelihood, \(417441.2\), and the lowest AIC, \(-206.05\), whereas c-vMEM-SeC achieves the lowest BIC, \(-206.01\). In-sample, d-vMEM-SeC attains the lowest MSE, \(3.0211\), and QLIKE, \(-3.2348\); out-of-sample over 327 observations from January 2023 to April 2024, s-vMEM-SeC attains the best MSE, \(2.3757\), and QLIKE, \(-3.3790\). The paper states that SeC models dominate vMEM models in sample and out-of-sample, while the clustered parameterization benefits from parsimony [2601.16837].

## 6. Extensions, limitations, and research directions

The vMEM literature has developed along three main axes: richer innovation laws, richer dependence structures, and scalable high-dimensional parameterizations. Copula-based models circumvent the limitations of multivariate Gamma formulations and allow negative dependence and tail dependence; components vMEM adds low-frequency multiplicative structure; Bayesian semiparametric vMEM replaces the parametric innovation law with a Dirichlet process mixture; and vMEM-SeC introduces a common latent component to distinguish spillovers from co-movements in large volatility panels [1604.01338] [2107.05923] [2107.04354] [2601.16837].

Several limitations recur across the literature. The copula-based paper notes computational complexity for high \(K\), symmetry restrictions in elliptical copulas, the possibility that a static copula misses time-varying dependence, and collinearity among activity indicators. The semiparametric Bayesian paper notes computational burden, including about \(13\) hours for one \(200\)k-run in \(d=3\), and sensitivity to hyperparameters such as \(\alpha\) and the Normal–Wishart priors. The high-dimensional log-vMEM-SeC paper points to the single common component, static loadings, and diagonal \(A,B\) as restrictions that may be relaxed by richer factor structures, time-varying loadings, block-diagonal or sparse off-diagonal dynamics, regime dependence, and nonlinearities [1604.01338] [2107.04354] [2601.16837].

The literature also warns against a common misconception that multivariate MEM is simply a collection of equation-by-equation univariate MEMs. The defining point of vMEM is that it leverages cross-series lagged effects and contemporaneous dependence among innovations. The survey explicitly contrasts vMEM with univariate MEM, ad hoc multivariate densities, and GARCH/HEAVY-type systems, emphasizing that vMEM directly models nonnegative indicators without logging, preserves interpretable conditional means, and supports multi-step forecasts. The newer log-vMEM papers show that logging can nonetheless be advantageous when the objective is tractable likelihood-based estimation or large-dimensional scaling, so the distinction is methodological rather than doctrinal [2107.05923] [2601.16837].

Open directions are explicit in the recent papers. Proposed extensions include time-varying copulas, asymmetric or vine copulas, zero-augmented marginals, exogenous regressors, higher-order lags, semiparametric GMM estimation, more efficient sampling schemes exploiting sparsity, alternative positive-support kernels, more general stick-breaking processes, and ICA-based common components. The 2025 error-law testing paper identifies further open questions concerning optimal weight selection, local asymptotic power, scalability in high dimensions, bootstrap validity under general vMEM settings with estimated copula parameters, and serial dependence in innovations. Taken together, these developments suggest that vMEM has evolved from a parsimonious multivariate positive-data model into a family of models in which conditional mean structure, innovation law, and cross-sectional dependence can be specified and tested with increasing granularity [1604.01338] [2107.04354] [2509.06732].

Source: https://www.emergentmind.com/topics/vector-multiplicative-error-models-vmem