---
title: 'Coskewness: Third-Order Dependence Measure'
url: https://www.emergentmind.com/topics/coskewness
type: topic
---

# Coskewness: Third-Order Dependence Measure

Coskewness is a third-order dependence measure, but the term is used in several non-equivalent ways across the recent literature. In a standardized centered mixed-moment framework, it is the first nontrivial higher-order centered mixed moment beyond covariance, given in the bivariate case by
\[
\mu_2(X_1,X_2)=\mathbb{E}\left[\left(\frac{X_1-\mu_1}{\sigma_1}\right)\left(\frac{X_2-\mu_2}{\sigma_2}\right)^2\right].
\]
In a trivariate standardized-moment framework it appears as
\[
S(X_i,X_j,X_k)=\frac{\mathbb{E}\big[(X_i-\mu_i)(X_j-\mu_j)(X_k-\mu_k)\big]}{\sigma_i\sigma_j\sigma_k},
\]
while alternative formulations include \(L\)-coskewness, standardized rank coskewness, and normalized three-mode correlators in quantum optics [2508.16600] [2412.13362] [2503.11289] [2204.03520]. Across these settings, coskewness quantifies asymmetric co-movement beyond covariance or correlation, and its attainable values depend on the marginals and the dependence structure rather than obeying a universal distribution-free range [2412.13362].

## 1. Formal definitions and competing frameworks

The recent literature uses several formalizations of coskewness, each tailored to a different statistical object.

| Framework | Formula | Scope |
|---|---|---|
| Centered mixed moment | \(\mu_2(X_1,X_2)=\mathbb{E}\!\left[\left(\frac{X_1-\mu_1}{\sigma_1}\right)\left(\frac{X_2-\mu_2}{\sigma_2}\right)^2\right]\) | Bivariate asymmetric dependence |
| Trivariate standardized moment | \(S(X_i,X_j,X_k)=\frac{\mathbb{E}[(X_i-\mu_i)(X_j-\mu_j)(X_k-\mu_k)]}{\sigma_i\sigma_j\sigma_k}\) | Three-variable dependence |
| \(L\)-coskewness | \(L_{3(i,j)}=\operatorname{Cov}\!\left(X_i,\ 6F_j(X_j)^2-6F_j(X_j)+1\right)\) | Quantile-based bivariate comoment |
| Standardized rank coskewness | \(RS(X_1,X_2,X_3)=32\,\mathbb{E}\!\left[\left(F_1(X_1)-\frac12\right)\left(F_2(X_2)-\frac12\right)\left(F_3(X_3)-\frac12\right)\right]\) | Monotone-invariant trivariate dependence |
| Quantum three-mode coskewness | \(\mathcal C_{abc}=\frac{\langle \hat x_a\hat x_b\hat x_c\rangle}{\sqrt{\langle \hat x_a^2\rangle\langle \hat x_b^2\rangle\langle \hat x_c^2\rangle}}\) | Non-Gaussian three-mode correlations |

In the centered mixed-moment hierarchy, the standardized centered mixed moment of order \(d+1\) is
\[
\mu_d(X_1,X_2)=\mathbb{E}\left[\left(\frac{X_1-\mu_1}{\sigma_1}\right)\left(\frac{X_2-\mu_2}{\sigma_2}\right)^d\right],
\]
with \(d=1\) giving a Pearson correlation-type comoment, \(d=2\) coskewness, and \(d=3\) cokurtosis [2508.16600]. In this formulation, the means and standard deviations themselves do not matter for the standardized comoments; what matters is the dependence structure between the variables [2508.16600].

The quantile-based literature replaces classical centered moments by \(L\)-comoments. In the bivariate quantile-function model of “A quantile-based bivariate distribution,” coskewness is specifically \(L\)-coskewness in the sense of Serfling and Xiao (2007), not the classical multivariate coskewness tensor [2503.11289]. The dependence is embedded directly in the conditional quantile-density factor \((1+\theta u_1)\), so the coskewness is a derived functional of the distributional parameters rather than a primitive parameter [2503.11289].

In quantum optics, coskewness is used in a different but structurally analogous sense: a normalized third-order correlator of quadrature fluctuations. The paper on three-body ultrastrong coupling treats \(\mathcal C_{abc}\) as a witness of non-Gaussianity and of correlations that cannot be captured within semiclassical and Gaussian approximations [2204.03520]. This suggests that “coskewness” is best viewed as a family of third-order dependence descriptors whose exact normalization and interpretation depend on the ambient modeling framework.

## 2. Dependence uncertainty and sharp extremal bounds

A central recent theme is the behavior of coskewness under fixed marginals and unknown copula. In the actuarial setting, the sharp bounds for the raw mixed moment
\[
\mathbb{E}(X_1X_2^d)
\]
are defined by
\[
m=\inf_{\forall i\, X_i\sim F_i}\mathbb{E}(X_1X_2^d), \qquad M=\sup_{\forall i\, X_i\sim F_i}\mathbb{E}(X_1X_2^d),
\]
and Theorem 1 gives explicit sharp bounds together with the dependence structures that attain them [2508.16600]. For odd \(d\), the extremizers simplify to classical monotone couplings: the maximum is attained by the comonotonic coupling \(U_1=U_2=U\), and the minimum by the countermonotonic coupling \(U_1=1-U,\ U_2=U\) [2508.16600]. For even \(d\), the sign of \(X_2\) matters because \(X_2^d\ge 0\); the optimizer is then a piecewise copula built from two inverse branches rather than a single monotone copula [2508.16600].

The structural reason is that the distribution function of \(X_2^d\) is not one-to-one when \(d\) is even and \(X_2\) can take both positive and negative values. In that case, the optimal dependence structure becomes a mixture of two monotone branches. For symmetric, zero-mean \(X_2\), the inverse branches simplify to
\[
g^{-1}(U)=\frac{1+U}{2}, \qquad f^{-1}(U)=\frac{1-U}{2},
\]
and the maximizing and minimizing copulas become correspondingly explicit [2508.16600].

The passage from raw to centered moments is direct. Proposition 4.1 states that bounding
\[
\mu_d(X_1,X_2)
\]
is equivalent to bounding \(\mathbb{E}(X_1X_2^d)\) for standardized variables, so the same extremal dependence structures apply after standardization [2508.16600]. In this sense, coskewness inherits the full dependence-uncertainty geometry of the raw mixed moment.

For fully trivariate coskewness, the dependence-uncertainty problem is formulated through
\[
m=\inf_{\forall i\, X_i\sim F_i}\mathbb{E}(X_1X_2X_3),\qquad
M=\sup_{\forall i\, X_i\sim F_i}\mathbb{E}(X_1X_2X_3),
\]
and, for symmetric zero-mean marginals, the sharp bounds are
\[
M=\mathbb{E}\!\left(\prod_{i=1}^d G_i^{-1}(U)\right), \qquad
m=-\mathbb{E}\!\left(\prod_{i=1}^d G_i^{-1}(U)\right),
\]
with attainment by explicit copulas; in odd dimension the optimizer is a cross product copula [2303.17266]. Specializing to \(d=3\), the same machinery yields sharp coskewness bounds. For standardized symmetric marginals, explicit upper bounds reported in the literature include
\[
\overline{S}=\frac{2\sqrt{2\pi}}{\pi}\quad\text{for the normal case},\qquad
\overline{S}=\frac{3\sqrt2}{2}\quad\text{for Laplace},\qquad
\overline{S}=\frac{3\sqrt3}{4}\quad\text{for uniform},
\]
with the lower bounds given by their negatives [2303.17266].

These results make a precise distinction between marginal information and dependence information. In particular, they show that sharp coskewness statements require explicit control of the copula and, when only marginals are known, the relevant object is an extremal dependence problem rather than a single descriptive statistic.

## 3. Rank-based and quantile-based analogues

Several recent papers introduce monotone-invariant or quantile-based versions of coskewness. In the trivariate case, standardized rank coskewness is defined by
\[
RS(X_1,X_2,X_3)=\frac{4\sqrt3}{9}\,S\big(F_1(X_1),F_2(X_2),F_3(X_3)\big)
=32\,\mathbb{E}\left[\left(F_1(X_1)-\frac12\right)\left(F_2(X_2)-\frac12\right)\left(F_3(X_3)-\frac12\right)\right].
\]
It is invariant under strictly increasing transformations, always lies in \([-1,1]\), and equals \(0\) under independence [2303.17266]. The same paper positions \(RS\) as a three-dimensional analogue of Spearman’s rho [2303.17266].

A bivariate higher-order rank analogue appears in the actuarial literature as the standardized rank coefficient
\[
RS_d(X_1,X_2)=\frac{2^{d+1}(d+1)(d+2)}{d}\,
\mathbb{E}\left[\left(F_{1}(X_1)-\frac12\right)\left(F_{2}(X_2)-\frac12\right)^d\right]
\]
for even \(d\), and
\[
RS_d(X_1,X_2)=2^{d+1}(d+2)\,
\mathbb{E}\left[\left(F_{1}(X_1)-\frac12\right)\left(F_{2}(X_2)-\frac12\right)^d\right]
\]
for odd \(d\). These coefficients satisfy \(-1\le RS_d(X_1,X_2)\le 1\), the bounds are sharp, and they are invariant under strictly increasing transformations [2508.16600]. In that setting, they serve as a rank-based, scale-free analogue of coskewness, similar in spirit to Spearman correlation [2508.16600].

The quantile-based bivariate distribution literature uses \(L\)-coskewness rather than standardized central moments. For \(i\neq j\),
\[
L_{3(i,j)}=\operatorname{Cov}\!\left(X_i,\ 6F_j(X_j)^2-6F_j(X_j)+1\right),
\]
and, using the covariance representation of Cuadras (2002), the paper derives the quantile form
\[
L_{3(1,2)}=\int_0^1\!\!\int_0^1
\Big[\big(1-u_{21}\big)(1-u_1)-\big(1-u_1\big)(1-u_2)\Big](12u_2-6)\,q_1(u_1)\,du_1\,du_2,
\]
where \(u_{21}=F_{21}(Q_{21}(u_1,u_2))\) [2503.11289]. In this model, positive \(L\)-coskewness is interpreted as indicating that the variables undergo positive deviation at the same time, and the first cable-lifetime dataset yields the reported values
\[
L\text{-covariance}=0.24,\quad L\text{-correlation}=0.53,\quad L\text{-coskewness}=0.30,\quad L\text{-cokurtosis}=0.25
\]
[2503.11289].

These constructions separate two aims that are often conflated. Classical centered coskewness measures asymmetric dependence on the original scale, whereas rank-based and \(L\)-moment formulations seek monotone invariance or quantile robustness. This suggests a division of labor: centered coskewness is natural when original-scale tail asymmetry matters, while rank and \(L\)-comoment analogues are natural when dependence should be isolated from marginal shape.

## 4. Decoupling from correlation and the leakage problem

A recurrent misconception is that correlation and coskewness should be linked. The recent literature explicitly rejects that claim. “Modeling coskewness with zero correlation and correlation with zero coskewness” proves two separation results: one can have arbitrary admissible coskewness with zero pairwise correlations, and one can have arbitrary correlations with zero coskewness [2412.13362]. For symmetric marginals, the first construction uses mixtures of extremal coskewness copulas while maintaining
\[
\rho_{12}=\rho_{13}=\rho_{23}=0,
\]
and the second uses trivariate Gaussian vectors, for which
\[
S(X_1,X_2,X_3)=0
\]
for any admissible correlation structure [2412.13362].

The same paper extends the separation to rank dependence. For arbitrary continuous strictly increasing marginals, standardized rank coskewness can range arbitrarily in \([-1,1]\) while all pairwise rank correlations vanish, and Gaussian copulas permit arbitrary pairwise rank correlations with
\[
RS(X_1,X_2,X_3)=0
\]
[2412.13362]. It also gives a striking example: under comonotonicity,
\[
F_i(X_i)=U \implies RS=32\mathbb{E}(U-\tfrac12)^3=0,
\]
while all pairwise rank correlations are \(1\) [2412.13362]. The conclusion is exact: there is no universal monotone or deterministic relationship between correlation and coskewness, or between rank correlation and standardized rank coskewness [2412.13362].

A different but related issue arises in conditional maximum-entropy models of neural activity. For sparse correlated binary inputs, higher-order features can project onto first-order features through a coskewness leakage channel [2606.01661]. With
\[
x_i=\mu_i+\delta_i,\qquad \mu_i=E[x_i],\qquad \delta_i=x_i-\mu_i,
\]
the centered interaction \( \delta_i\delta_j \) projects onto a first-order coordinate through
\[
\operatorname{Cov}(\delta_i\delta_j,\delta_i)=E[\delta_i^2\delta_j].
\]
For binary inputs \(x_i\in\{0,1\}\), the identity
\[
\delta_i^2=(1-2p_i)\delta_i+p_i(1-p_i)
\]
implies
\[
E[\delta_i^2\delta_j]=(1-2p_i)\operatorname{Cov}(x_i,x_j),
\]
which the paper identifies as the explicit coskewness form of leakage [2606.01661]. In sparse data, \(p_i\ll \tfrac12\), so \(1-2p_i\approx 1\), making covariance almost directly convertible into a leakage channel [2606.01661].

The corresponding information-geometric statement is
\[
\hat\theta-\theta_0=I_{FF}^{-1}I_{FG}\lambda+O(\|\lambda\|^2),
\]
so omitted higher-order, temporal, or hidden-state terms can shift fitted first-order parameters whenever the omitted and included sufficient statistics are correlated under the sampled input distribution [2606.01661]. The paper’s interpretive conclusion is that entropy explained by a direct MaxEnt fit is a prediction measure under the observed state distribution, not a mechanism-identification statistic [2606.01661]. In this setting, coskewness is not only a dependence descriptor but also the algebraic mechanism by which higher-order structure becomes statistically confounded with lower-order models.

## 5. Tensor structure, computational burden, and algorithmic surrogates

In multivariate settings, coskewness naturally becomes a third-order tensor. In principal skewness analysis for hyperspectral imagery, after centering and whitening,
\[
\mathcal S=\frac1N\sum_{i=1}^{N}\mathbf r_i\circ\mathbf r_i\circ\mathbf r_i
\]
is the supersymmetric coskewness tensor, and directional skewness is
\[
\operatorname{skew}(\mathbf u)=\mathcal S\times_1\mathbf u\times_2\mathbf u\times_3\mathbf u.
\]
The associated optimization leads to the tensor eigenpair equation
\[
\mathcal S\times_1\mathbf u\times_3\mathbf u=\lambda\mathbf u
\]
[1907.09811]. The key technical point is that, unlike matrix eigenvectors, eigenvectors of a supersymmetric tensor are not inherently orthogonal in general [1907.09811]. Nonorthogonal Principal Skewness Analysis therefore replaces the usual orthogonal deflation strategy by a Kronecker-structured projection, enlarging the feasible search space and reducing the improved update complexity from storage \(L^3\times L^3\), complexity \(O(L^6)\), to storage \(L\times L\times L\), complexity \(O(L^3)\) [1907.09811].

The same tensor-growth problem appears in high-dimensional learning. In self-supervised action recognition, the moment-descriptor viewpoint treats mean, covariance, coskewness, and cokurtosis as first-, second-, third-, and fourth-order summaries, but the paper explicitly notes that the numbers of coefficients grow linearly, quadratically, cubically, and quartically with feature dimension [2001.04627]. Its practical response is to keep mean, a low-rank covariance subspace, skewness, and kurtosis rather than full coskewness and cokurtosis tensors [2001.04627].

Large-scale higher-moment portfolio optimization exhibits the same bottleneck. In the sample-moment MVSK model, the coskewness tensor is
\[
\mathcal S:=\frac{1}{T}\sum_{t=1}^T a_t^{\otimes 3},
\qquad
m_3(x)=\langle \mathcal S,x^{\otimes 3}\rangle,
\]
but explicit higher-order tensors require \(\Theta(n^3)\) storage for coskewness and \(O(Tn^3)\) construction cost [2604.25378]. The Yau affine-normal descent method avoids explicit construction of \(\mathcal S\) and \(\mathcal K\) by working directly with the centered return matrix \(A\); the objective, gradient, Hessian-vector product, and third-order directional action can then each be evaluated in \(O(Tn)\) arithmetic and \(O(T+n)\) working memory [2604.25378].

A further response is regularization rather than explicit tensor estimation. In Discrete World Models via Regularization, coskewness is penalized through
\[
M_{ijk}=\mathbb{E}_n[\tilde p_{ni}\tilde p_{nj}\tilde p_{nk}],
\qquad
\mathcal L_{\text{cos}}=
\frac{1}{K(K-1)(K-2)}
\sum_{i\neq j,\;j\neq k,\;i\neq k}|M_{ijk}|,
\]
restricted to distinct indices so that only genuine triplet interactions are penalized [2603.01748]. The paper’s rationale is explicit: pairwise decorrelation alone does not eliminate structured higher-order dependence, and minimizing \(\mathcal L_{\text{cos}}\) pushes latent bits toward joint factorization beyond second order [2603.01748].

The simulation literature offers yet another compressed representation. In Random Orthogonal Matrix simulation, the full co-skewness tensor of standardized components is
\[
c_{ijk}=\mathbb{E}[Y_i^*Y_j^*Y_k^*],
\]
assembled into \(\mathbf m_3\), while the Kollo skewness vector
\[
\tau(\mathbf Y_n)=
\begin{bmatrix}
\sum_{i,k}^{n}c_{i1k},\,
\sum_{i,k}^{n}c_{i2k},\,
\cdots,\,
\sum_{i,k}^{n}c_{ink}
\end{bmatrix}'
\]
is used as a lower-dimensional exact target for simulation [2004.06586]. This paper makes precise the tradeoff between preserving the full third-order object and targeting a structured summary that remains computationally manageable.

## 6. Applications in actuarial science, finance, quantum physics, and representation learning

In actuarial science, coskewness is treated as a material risk driver rather than a purely descriptive statistic. Under a copula-based mixture model with exponential marginals,
\[
\mu_d(X_1,X_2)=\underline{S_d}+\lambda\big(\overline{S_d}-\underline{S_d}\big),
\]
so dependence moves continuously from the lower-extremal to the upper-extremal structure [2508.16600]. Within that model, expected shortfall
\[
\mathrm{ES}_p(S)=\mathbb{E}[S\mid S>\mathrm{VaR}_p(S)]
\]
increases monotonically with even-order centered mixed moments such as coskewness, whereas marginal expected shortfall
\[
\mathrm{MES}_p(L_1,S)=\mathbb{E}[L_1\mid S>\mathrm{VaR}_p(S)]
\]
is more nuanced, sometimes nearly unchanged and in general possibly non-monotone or flattening as \(p\to1\) [2508.16600]. In life contingencies, last-survivor annuity values increase with dependence and coskewness, joint-life annuity values decrease, and the difference between minimal and maximal dependence grows with the annuity term \(n\) [2508.16600]. The practical implication is that mis-specifying higher-order dependence can induce significant annuity mispricing [2508.16600].

In portfolio selection, coskewness enters as the third central sample moment of returns. The unrestricted MVSK objective
\[
f(x)= -c_1 m_1(x) + c_2 m_2(x) - c_3 m_3(x) + c_4 m_4(x)
\]
rewards positive portfolio skewness through the \(-c_3 m_3(x)\) term, and the empirical study on a 5-minute A-share panel with \(5{,}440\) stocks finds that higher moments add the most value at moderate return targets [2604.25378]. This indicates that coskewness can be economically useful, but not uniformly so across target regimes [2604.25378]. A related option-implied object is the skew stickiness ratio
\[
R_t=\frac{1}{\sigma_t'}\frac{d\langle \sigma^S,\log S\rangle_t}{d\langle \log S\rangle_t},
\]
which the rough-volatility literature treats as a normalized return–volatility covariation and interprets as a coskewness-like statistic; in Bergomi-type models its short-maturity limit is
\[
\lim_{T\to t}R_t(T)=H+\frac32
\]
[2602.05241].

In quantum optics, coskewness becomes a diagnostic of genuinely non-Gaussian criticality. The three-body ultrastrong-coupling Hamiltonian exhibits a first-order superradiant transition with broken \(\mathbb Z_2\times\mathbb Z_2\) symmetry, and the ground-state coskewness \(\mathcal C_{abc}\) diverges near the critical point in a way not seen in ordinary two-body models [2204.03520]. The steady state under dissipation inherits the same qualitative signature [2204.03520]. Here coskewness functions as a witness that the critical state cannot be reduced to Gaussian fluctuations around a semiclassical displaced solution [2204.03520].

In machine learning and representation learning, coskewness is used both positively and negatively: positively as a tensor or regularizer that captures third-order structure, and negatively as a complexity source to be compressed or avoided. The action-recognition literature treats full coskewness as theoretically informative but practically too large [2001.04627]; the world-model literature penalizes it to suppress triplet entanglement [2603.01748]; and the neural identifiability literature shows that unmodeled coskewness can leak into first-order fits under sparse correlated sampling [2606.01661]. Across these domains, a common theme is that third-order dependence is neither negligible nor automatically interpretable from second-order summaries.

Taken together, these developments present coskewness as a structurally rich but context-dependent notion. It may denote a centered mixed moment, a rank-based dependence coefficient, an \(L\)-comoment, a tensor of third-order interactions, or a normalized three-mode correlator. What unifies these uses is the attempt to quantify asymmetric dependence beyond covariance. What distinguishes them is the object being normalized, the invariances being sought, and the role played by dependence uncertainty, computation, and application-specific semantics.

Source: https://www.emergentmind.com/topics/coskewness