---
title: Intensive Mixed Skewness Models
url: https://www.emergentmind.com/topics/intensive-mixed-skewness
type: topic
---

# Intensive Mixed Skewness Models

Intensive mixed skewness denotes a class of statistical phenomena and modelling strategies in which asymmetry is generated, amplified, or decomposed through latent mixing, unequal left/right scaling, contaminated components, or skewed random effects. In the cited literature, such settings are studied jointly with heavy tails, bimodality, cluster heterogeneity, semicontinuity, and longitudinal dependence. The resulting models range from one-piece and two-piece density constructions to mean-mixtures of multivariate normals, mixtures of multivariate skew Laplace components, canonical fundamental skew-\(t\) linear mixed models, and contaminated generalized asymmetric Laplace mixed-effects quantile regressions [1512.03341], [2109.12152], [1702.00628], [2504.14515], [2401.04603], [2006.10018].

## 1. Formal characterization of skewness

For a univariate random variable \(X\) with mean \(\mu\) and variance \(\sigma^2\), the Fisher–Pearson skewness is
\[
\gamma_1 = \frac{E[(X-\mu)^3]}{\sigma^3}.
\]
Positive \(\gamma_1\) indicates right skew and negative \(\gamma_1\) indicates left skew [2011.09152].

For multivariate data, the literature emphasizes Mardia’s skewness and kurtosis. If \(X\) is \(p\)-variate with mean vector \(\mu\) and covariance \(\Sigma\), Mardia’s multivariate skewness is
\[
\beta_{1,p} = \sum_{r,s,t} \sum_{r',s',t'} \sigma^{rr'} \sigma^{ss'} \sigma^{tt'} \mu_{111}^{(rst)} \mu_{111}^{(r's't')},
\]
where
\[
\mu_{111}^{(rst)} = E[(X_r-\mu_r)(X_s-\mu_s)(X_t-\mu_t)],
\]
and Mardia’s multivariate kurtosis is
\[
\beta_{2,p} = E\left\{\left[(X-\mu)' \Sigma^{-1} (X-\mu)\right]^2\right\}.
\]
Under multivariate normality, \(\beta_{1,p}=0\) and \(\beta_{2,p}=p(p+2)\) [2011.09152].

The mean-mixture literature broadens this measurement framework. For the MMN family, scalar and vector-valued indices are both considered, including Mardia’s skewness, Malkovich–Afifi, Srivastava, Móri–Roy–Goswami, Kollo, Balakrishnan–Brito–Quiroz, and Isogai measures [2006.10018]. A central structural result is that, after a suitable canonical transformation, the MMN family can be written with one skewed coordinate and \(p-1\) independent standard normals. This implies that several multivariate skewness measures reduce to functions of the univariate skewness of the canonical non-Gaussian coordinate [2006.10018].

This body of work suggests that “intensive” skewness is not merely a large third standardized moment. It is often a consequence of model architecture: latent truncation, scale mixtures, mean mixtures, unequal-side rescaling, or component-specific asymmetry can each induce qualitatively different skew structures.

## 2. Construction by unequal scales and bimodal perturbation

A direct route to skewness starts from a continuous density \(f(x)\) symmetric about \(0\) and unimodal at the origin. The Fernandez–Steel transformation introduces a positive skew parameter \(\gamma>0\) by stretching one side of the density and compressing the other:
\[
g(x\mid \gamma)
=\frac{2}{\gamma+\gamma^{-1}}
\left\{
f(x/\gamma)\mathbf{1}_{\{x\ge 0\}}
+
f(\gamma x)\mathbf{1}_{\{x<0\}}
\right\}.
\]
When \(\gamma=1\), one recovers \(g(x)=f(x)\); when \(\gamma>1\), mass is pushed to the right; and when \(0<\gamma<1\), mass is pushed to the left [1512.03341].

Ehlers then perturbs this skewed unimodal density to induce bimodality:
\[
h(x\mid \alpha,\gamma)
=
\frac{1+\alpha x^2}{1+\alpha b_\gamma}\,g(x\mid\gamma),
\qquad
b_\gamma=\int_{-\infty}^{\infty} x^2 g(x\mid\gamma)\,dx.
\]
The density remains proper because \((1+\alpha b_\gamma)^{-1}\) normalizes the factor \(1+\alpha x^2\). The threshold behavior is explicit: if \(0\le \alpha<\tfrac12\), \(h\) remains unimodal, whereas for \(\alpha\ge \tfrac12\), it develops two modes [1512.03341].

The associated raw moments retain closed form. If
\[
M_r(\gamma)=E_g[X^r]
=
\frac{\gamma^{r+1}+(-1)^r\gamma^{-(r+1)}}{\gamma+\gamma^{-1}}\,m_r,
\]
with \(m_r=2\int_0^\infty x^r f(x)\,dx\), then
\[
E_h[X^r]
=
\frac{M_r(\gamma)+\alpha M_{r+2}(\gamma)}{1+\alpha b_\gamma}.
\]
Mean, second moment, variance, skewness, and kurtosis follow from this expansion [1512.03341].

A more general unequal-scale construction is the skewed pivot–blend density. Let \(g(\cdot;\theta)\) be any continuous base density with distribution function \(G(\cdot;\theta)\), let \(\tau\in(0,1)\) be a pivot quantile, and define \(m=G^{-1}(\tau)\). With left and right scales \(\sigma_l,\sigma_r>0\),
\[
f_{SP}(x;\tau,\sigma_l,\sigma_r,\theta)
=
\frac{1}{\tau \sigma_l + (1-\tau)\sigma_r}
\left\{
g\!\left(\frac{x-m}{\sigma_l}+m;\theta\right) I\{x\le m\}
+
g\!\left(\frac{x-m}{\sigma_r}+m;\theta\right) I\{x>m\}
\right\}.
\]
If \(\sigma_l<\sigma_r\), the right tail is heavier and the model is right-skewed; if \(\sigma_l=\sigma_r\), the model collapses to an ordinary location–scale transform of \(g\); and varying \(\tau\) moves the point of asymmetry away from the median or mode [2401.04603].

These constructions isolate two distinct mechanisms. The Ehlers family starts from a symmetric unimodal base, introduces one-piece skewing, and then disturbs unimodality. The pivot–blend framework allows any continuous base density, including asymmetric and nonunimodal \(g\), and treats the pivot quantile itself as a parameter of the skew structure.

## 3. Latent mixtures, normal mean–variance mixtures, and componentwise asymmetry

A second major route to intensive mixed skewness uses latent variables. In the MMN family,
\[
X\mid\{T=t\}\sim N_p(\mu+\alpha t,\Sigma),
\qquad
T\sim g(t),\quad t\ge 0,
\]
so the marginal density is
\[
f_X(x)=\int_0^\infty \phi_p(x;\mu+\alpha t,\Sigma)\,g(t)\,dt.
\]
This yields
\[
E[X]=\mu+m_1\alpha,
\qquad
\mathrm{Var}(X)=\Sigma+(m_2-m_1^2)\alpha\alpha^\top,
\]
where \(m_k=E[T^k]\) [2006.10018]. Skewness is therefore tied directly to the mixing law \(g\) and the direction vector \(\alpha\).

Cluster analysis literature distinguishes two broader strategies for skewed heterogeneous data. The first uses mixtures of flexible skewed distributions,
\[
f(x\mid \Psi)=\sum_{g=1}^G \pi_g f_g(x\mid \theta_g),
\]
with component densities drawn from families such as variance-Gamma, generalized hyperbolic, skew-\(t\), skew-normal, NIG, or SAL. A common normal mean–variance mixture representation is
\[
X=\mu+W\alpha+\sqrt{W}\,V,
\qquad
V\sim N_p(0,\Sigma),\quad W>0,
\]
with the law of \(W\) determining the family, tail behavior, and concentration parameters [2011.09152].

The second strategy is transformation-based. Here one seeks a coordinatewise transform \(T:\mathbb R^p\to\mathbb R^p\) such that \(T(x\mid \Lambda)\) is approximately \(N_p(\mu,\Sigma)\). The transformed density is
\[
f_T(x\mid \mu,\Sigma,\Lambda)=\phi_p(T(x\mid\Lambda);\mu,\Sigma)\,J_T(x\mid\Lambda),
\]
with Yeo–Johnson and Manly transformations as prominent examples [2011.09152].

Finite mixtures of multivariate skew Laplace components provide a related but distinct formulation. If \(Y\) follows an MSL distribution with parameters \((\mu,\Sigma,\gamma)\), its density depends on
\[
Q(y)=(y-\mu)^\top \Sigma^{-1}(y-\mu),
\qquad
a=1+\gamma^\top \Sigma^{-1}\gamma,
\]
and satisfies
\[
E[Y]=\mu+(p+1)\gamma,
\qquad
\mathrm{Var}[Y]=(p+1)\bigl(2\Sigma+2\gamma\gamma^\top\bigr).
\]
The hierarchical form is
\[
Y\mid V=v\sim N_p(\mu+v^{-1}\gamma,\,v^{-1}\Sigma),
\qquad
V\sim \mathrm{IG}\!\left(\frac{p+1}{2},\frac12\right).
\]
A \(K\)-component MSL mixture then takes the form
\[
p(y_i)=\sum_{k=1}^K \pi_k f_{\rm MSL}(y_i;\mu_k,\Sigma_k,\gamma_k)
\]
[1702.00628].

Across these models, skewness is not an external correction applied after Gaussian modelling. It is encoded structurally through latent scales, latent means, or component-specific asymmetry parameters. A plausible implication is that the phrase “mixed skewness” refers as much to the generative mechanism as to the marginal shape.

## 4. Mixed-effects and longitudinal formulations

For clustered and longitudinal data, intensive mixed skewness arises when both subject-specific effects and within-subject errors depart from Gaussian symmetry. The canonical fundamental skew-\(t\) linear mixed model (ST-LMM) is defined by
\[
y_i=X_i\beta+Z_i b_i+\varepsilon_i,
\]
where \(y_i\) is the \(n_i\times 1\) response vector, \(X_i\) and \(Z_i\) are design matrices, \(b_i\) are random effects, and \(\varepsilon_i\) are errors [2109.12152].

Instead of Gaussian assumptions, the joint \((q+n_i)\)-vector \(\begin{pmatrix} b_i \\ \varepsilon_i \end{pmatrix}\) is assumed to follow a canonical fundamental skew-\(t\) distribution with location, scale, shape, and degrees-of-freedom parameters. The model uses \(D\) for the random-effects covariance, \(\Sigma_i\) for the error covariance, \(\Delta\) for the random-effects skewness matrix, \(r\le q\) for the latent half-normal dimension, and \(\nu>1\) for tail-heaviness. The constant
\[
b=-\sqrt{\nu/\pi}\cdot \Gamma((\nu-1)/2)/\Gamma(\nu/2)
\]
ensures \(E[b_i]=0\) [2109.12152].

Its hierarchical representation introduces two latent variables:
\[
U_i\sim \mathrm{Gamma}(\nu/2,\nu/2),
\qquad
S_i\mid U_i=u_i\sim HN_r(0,u_i^{-1}I_r),
\]
with Gaussian conditional layers for \(Y_i\) and \(b_i\). The \(U_i\) layer yields \(t\)-tails, and the \(S_i\) layer yields skewness [2109.12152].

The mixed-effects quantile literature uses a different route. In the contaminated generalized asymmetric Laplace framework,
\[
y_{ij}=X_{ij}^\top\beta+Z_{ij}^\top b_i+\epsilon_{ij},
\qquad
b_i\sim \mathcal N(0,\Sigma),
\qquad
\epsilon_{ij}\sim \mathrm{cGAL}(0,\sigma,\kappa,\nu,\pi).
\]
The cGAL density is
\[
f_{\rm cGAL}(y\mid \mu,\sigma,\kappa,\nu,\pi)
=
\pi f_{\rm GAL}(y\mid \mu,\sigma,\kappa,p_0)
+
(1-\pi)f_{\rm GAL}(y\mid \mu,\nu\sigma,\kappa,p_0),
\]
where \(\nu>1\) inflates the scale of the contamination component and \(\pi\in(0,1)\) is the good-data weight [2504.14515].

The GAL component itself augments the asymmetric Laplace by a shape parameter \(\kappa\neq 0\), while preserving \(\mu\) as the \(p_0\)th quantile. A Gaussian–exponential–truncated-normal mixture representation supports MCMC data augmentation. In the mixed-effects model, \(\delta_{ij}\sim \mathrm{Bernoulli}(\pi)\) indicates whether an observation belongs to the main or contamination component, and latent variables \((z_{ij},s_{ij})\) are introduced for the GAL structure [2504.14515].

These two mixed-effects families target different inferential goals. ST-LMM directly models asymmetric and heavy-tailed random effects and errors in a likelihood framework, whereas cGAL targets conditional quantiles and robustness to outliers without explicit outlier deletion. Both treat skewness as a hierarchical latent effect rather than a residual nuisance.

## 5. Estimation, identifiability, and model assessment

Likelihood-based inference is central across the literature. For the skewed bimodal family,
\[
\ell(\alpha,\gamma)=\sum_{i=1}^n \log h(x_i\mid \alpha,\gamma)
\]
is maximized numerically over \(\alpha\ge 0\) and \(\gamma>0\). Two practical issues are emphasized: identifiability between \(\alpha\) and \(\gamma\) when sample size is small, and numerical stability when \(\alpha\) is large, since the factor \(1+\alpha x^2\) can overflow in the tails. A common reparameterization is \(\phi=\gamma^2\), interpreted as the mass-ratio above and below zero, and tail computations are carried out on the log-scale [1512.03341].

For ST-LMM, estimation proceeds by ECME with \(\{b_i,S_i,U_i\}\) treated as missing. The E-step computes conditional expectations such as \(E[U_i\mid Y_i]\), \(E[U_i b_i\mid y_i]\), \(E[U_i S_i\mid y_i]\), and second-order cross-moments. These are available in closed form using ratios of multivariate \(t\)-pdfs and cdfs. The M-step updates \(\beta\) by generalized least squares,
\[
\beta^{(k+1)}
=
\left(\sum_i X_i^\top [u_i\Sigma_i^{-1}]X_i\right)^{-1}
\sum_i X_i^\top \Sigma_i^{-1}\bigl[u_i y_i-Z_iE(U_i b_i)\bigr],
\]
while \(D\), \(\Delta\), and \(\Sigma_i\) admit closed or partial-closed form updates, and \(\nu\) is updated by direct maximization of the observed-data log-likelihood over \(\nu>1\) [2109.12152].

Posterior means of random effects in ST-LMM are given by
\[
\widehat b_i
=
E[b_i\mid y_i,\widehat\theta]
=
b\,1_r+M_i Z_i^\top \Sigma_i^{-1}(y_i-\mu_i)+M_i D^{-1}E[S_i\mid y_i,\widehat\theta],
\]
with
\[
M_i=(D^{-1}+Z_i^\top \Sigma_i^{-1}Z_i)^{-1}.
\]
Standard errors are obtained from Louis’s formula for the observed information [2109.12152].

For finite mixtures of MSL distributions, EM uses responsibilities
\[
\tau_{ik}
=
\frac{\pi_k^{(t)}f_{\rm MSL}(y_i;\mu_k^{(t)},\Sigma_k^{(t)},\gamma_k^{(t)})}
{\sum_{j=1}^K \pi_j^{(t)}f_{\rm MSL}(y_i;\mu_j^{(t)},\Sigma_j^{(t)},\gamma_j^{(t)})},
\]
and conditional expectations \(d_{ik}=E[V_i^{-1}\mid y_i,Z_{ik}=1]\), \(b_{ik}=E[V_i\mid y_i,Z_{ik}=1]\), followed by closed-form updates for \(\pi_k\), \(\mu_k\), \(\gamma_k\), and \(\Sigma_k\). Model selection is performed with AIC or BIC, and convergence may be monitored by log-likelihood, parameter change, Aitken acceleration, or stability of cluster assignments [1702.00628].

For the pivot–blend model, the MLE minimizes a piecewise differentiable objective in \((\beta,m,\sigma_l,\sigma_r,\theta)\); off-the-shelf gradient or quasi-Newton optimizers are used, with possible alternating updates between location/pivot parameters and side-specific scales [2401.04603].

For cGAL mixed-effects quantile regression, inference is Bayesian. Gibbs updates are available for \(\beta\), \(b_i\), \(\Sigma\), and \(\pi\mid\{\delta_{ij}\}\sim \mathrm{Beta}(1+\sum\delta_{ij},\,9+\sum(1-\delta_{ij}))\), while \(\sigma\) and \(\kappa\) typically require slice or Metropolis–Hastings steps [2504.14515]. Model checking uses leave-one-out cross-validation, LOOIC, WAIC, Kullback–Leibler divergence per observation, and DHARMa residual tests [2504.14515].

Taken together, these methods show that skewness modelling is inseparable from computation. The key technical issues are not only flexibility of the family, but also tractable latent representations, stable optimization, and separability of parameters that may simultaneously affect skewness, tails, and modality.

## 6. Empirical behavior, applications, and interpretive issues

Illustrative densities in the skewed bimodal family show how intensity of skewness and bimodality interact. With \(f(x)=\phi(x)\), \(\gamma=1.5\), and \(\alpha=1.0\), the density is weakly bimodal, with modes at approximately \(x\approx -1.1\) and \(x\approx 2.3\), and the right peak is higher. With \(\gamma=2.0\) and \(\alpha=5.0\), the modes are clearly separated at approximately \(x\approx -2.5\) and \(x\approx 3.5\), and the right-hand peak dominates heavily. Replacing the normal base by a standardized Student-\(t(\nu=4)\) with \(\gamma=1.5\) and \(\alpha=3.0\) yields thicker shoulders and fatter tails while preserving roughly comparable modal locations [1512.03341].

In the schizophrenia application for ST-LMM, Brief Psychiatric Rating Scale scores were observed on \(N=118\) subjects at up to \(6\) visits. The fitted model was
\[
Y_{ij}/10 = (\beta_0+b_{0i}) + (\beta_1+b_{1i})x_{ij} + \beta_2 x_{ij}^2 + \beta_3 NT_i + \beta_4 NT_i x_{ij} + \epsilon_{ij},
\]
with \(x_{ij}=(\text{visit}-3)/10\) and \(NT_i=1\) for new treatment. The ST-LMM with \(r=2\) attained the lowest AIC, \(1499.02\), versus SN and SDB variants. The reported estimates included \(\hat\beta_0\approx 2.66\) (SE \(0.14\)), \(\hat\beta_1\approx -1.32\) (SE \(0.35\)), and \(\hat\beta_2\approx 6.71\) (SE \(0.51\)); \(\hat\beta_3\) and \(\hat\beta_4\) were non-significant, confirming equivalence of the new drug. The fitted asymmetric, heavy-tailed contours for the random effects matched the empirical BLUPs well [2109.12152].

In cluster analysis, benchmark comparisons showed that no single approach uniformly dominated. On Iris, all methods selected \(G=2\) and achieved identical ARI \(=0.568\). On Wine and Diabetes, all methods under-fitted the true number of groups or chose \(G=2\), with similar ARI near \(0.45\). On Crabs, skewed-distribution mixtures perfectly separated species, whereas transformation methods separated sexes. Transformation methods were also more parsimonious because they required fewer tail or concentration parameters [2011.09152].

The MSL mixture study reported that, for Swiss bank-note data, a two-component FM-MSL fit achieved higher log-likelihood and lower AIC/BIC than the corresponding FM-MSN fit, indicating improved handling of pronounced skewness and heavy tails [1702.00628]. In the MMN study, AIS and olive oil data similarly favored MMN over skew-normal and skew-\(t\) by log-likelihood, AIC, and BIC, while the multivariate skewness measures supported right-skewed structure in the observed coordinates [2006.10018].

For HIV viral-load decay, the cGAL mixed-effects quantile model was preferred to AL and GAL in predictive LOOIC and showed reduced influence of outliers via lower Kullback–Leibler divergence. In simulation, data were generated from cGAL with contamination levels \(\alpha\in\{0.001,0.05\}\) at quantiles \(p_0\in\{0.50,0.85\}\), with \(300\) replicates per setting; cGAL showed substantially lower bias and RMSE under contamination, tighter HPD intervals, and nominal or conservative coverage relative to GAL [2504.14515].

Several recurring interpretive issues follow from these results. First, skewness and heavy tails are often statistically entangled, so parameter identifiability may be weak in small samples. Second, different model classes may recover different latent structure from the same data, as illustrated by the Crabs example. Third, scalar and vector-valued skewness measures may emphasize different aspects of asymmetry. A plausible implication is that intensive mixed skewness should be treated as a model-selection problem over mechanisms of asymmetry, not as a single numerical descriptor.

Source: https://www.emergentmind.com/topics/intensive-mixed-skewness