---
title: Joint Projected Normal & Skew-Normal (JPSN)
url: https://www.emergentmind.com/topics/joint-projected-normal-and-skew-normal-jpsn
type: topic
---

# Joint Projected Normal & Skew-Normal (JPSN)

The Joint Projected Normal and Skew-Normal (JPSN) distribution is a class of flexible multivariate models for joint circular-linear or poly-cylindrical data. It combines the projected normal (PN) distribution for circular components and the skew-normal (SN) distribution for linear components, extending multivariate modeling in applications such as animal movement studies where angular (directional) and linear (e.g., speed or length) variables co-occur. The construction leverages latent variable augmentations to facilitate Bayesian inference and Gibbs sampling, addresses non-identifiability in the projected normal component via post-processing, and naturally supports quantification of arbitrary multivariate dependence structures [1512.00487], [1711.10463].

## 1. Constructive Definition

Let $p$ denote the number of circular variables, $\Theta = (\Theta_1, \ldots, \Theta_p)^\top \in [0,2\pi)^p$, and $q$ the number of linear variables, $Y = (Y_1, \ldots, Y_q)^\top \in \mathbb{R}^q$. The latent variable construction introduces $W_i = (W_{i1}, W_{i2})^\top \in \mathbb{R}^2$ for each angular component, with $R_i = \|W_i\| > 0$ and $\Theta_i = \text{atan2}(W_{i2}, W_{i1})$. The linear variables incorporate a skew-normal structure via a latent random effect $d \sim \mathcal{HN}_q(0, I_q)$, so that conditional on $d$, $Y \mid d \sim N_q(\mu_y + \Lambda d, \Sigma_y)$, where $\Lambda = \text{diag}(\lambda_1, \ldots, \lambda_q)$.

Stacking $W = (W_1, ..., W_p)^\top \in \mathbb{R}^{2p}$, the augmented vector $(W, Y)^\top$ is modeled jointly as:
\[
(W, Y)^\top \mid d \sim N_{2p+q}\left( \mu + (0_{2p}, \Lambda d)^\top,\; \Sigma \right),
\]
where
\[
\mu = \begin{pmatrix} \mu_w \\ \mu_y \end{pmatrix}, \quad 
\Sigma = \begin{pmatrix} \Sigma_w & \Sigma_{wy} \\ \Sigma_{wy}^\top & \Sigma_y \end{pmatrix}.
\]
The full (augmented) likelihood is:
\[
f(\theta, r, y, d) = \left[\prod_{i=1}^p r_i\right]\, 2^q\, \varphi_{2p+q}\left( (w, y)^\top \;\Big|\; \mu + (0_{2p}, \Lambda d)^\top, \Sigma \right)\, \varphi_q(d | 0, I_q),
\]
where $w$ is determined by $(\theta, r)$. The marginal joint density $f(\theta, y)$, which integrates out the latent $r, d$, is intractable; hence inference proceeds using the augmented representation [1512.00487], [1711.10463].

## 2. Relation to Projected Normal and Skew-Normal Distributions

The JPSN construction unifies two well-established frameworks:

- **Projected Normal (PN):** For $W \sim N_{2p}(\mu_w, \Sigma_w)$, the circular component $\Theta$ arises by projection: $\Theta_i = \text{atan2}(W_{i2}, W_{i1})$ for $i = 1, \ldots, p$. Marginalization over the radius recovers the PN density for angles [1711.10463].

- **Sahu-Sahu Skew-Normal (SSN):** If $H \sim N_q(0, \Sigma_y)$ and $D \sim \mathcal{HN}_q(0, I_q)$, the linear component $Y = \mu_y + \Lambda D + H$ has a multivariate skew-normal distribution, with augmented density $2^q \varphi_q(y | \mu_y + \Lambda d, \Sigma_y)\, \varphi_q(d|0,I)$.

The JPSN replaces the product of separate circular and linear densities for $(W,Y)$ with a single $(2p+q)$-variate normal, coupling circular and linear components via $\Sigma_{wy}$, thus encoding full multivariate dependence [1711.10463].

## 3. Parameters, Identifiability, and Posterior Representation

The parameters consist of:
- $\mu_w \in \mathbb{R}^{2p}$ and $\Sigma_w \in \mathbb{R}^{2p \times 2p}$ for circular latent normals;
- $\mu_y \in \mathbb{R}^{q}$ and $\Sigma_y \in \mathbb{R}^{q \times q}$ for the linear component;
- $\Sigma_{wy}$ capturing block covariance;
- $\lambda \in \mathbb{R}^q$, the skew-normal parameter vector for linear part.

A structural non-identifiability arises from the PN: scaling $W_i$ by any $c_i>0$ leaves $\Theta_i$ invariant. Identifiability is often enforced post hoc by setting $\text{Var}(W_{i2}) = 1$ for each $i$ (white noise constraint), so after sampling unconstrained $(\mu, \Sigma)$, draws are mapped to an identifiable space:
\[
c_i = \sqrt{(\Sigma)_{2i,2i}}, \quad C = \text{diag}(c_1,c_1,\ldots,c_p,c_p,1,\ldots,1),
\]
\[
\tilde\mu = C^{-1}\mu, \quad \tilde\Sigma = C^{-1}\Sigma C^{-1}.
\]
All posterior summaries are then reported for $(\tilde\mu,\tilde\Sigma, \lambda)$ [1711.10463], [1512.00487].

## 4. Bayesian Inference and Augmented Gibbs Sampling

The Bayesian model employs conjugate priors: $(\mu, \Sigma)$ follows a normal-inverse-Wishart prior, and $\lambda$ is Gaussian a priori. Inference leverages the tractability of the augmented framework, enabling Gibbs updates as if for standard multivariate normal models:

- Given the latent radii $R_t$ and skew-normals $D_t$, updating $(\mu, \Sigma)$ reduces to standard normal-inverse-Wishart conjugate forms.
- Full conditional for $\lambda$ is Gaussian, due to its linear regression idiom in $Y_t - \mu_{y|w_t}$ versus $d_t$. 
- Each $D_t$ is sampled from a truncated normal with mean and covariance determined by the linear block of $\Sigma$ and $\lambda$.
- $R_{ti}$ is sampled by slice sampling or Metropolis-Hastings, exploiting the $1$-dimensional structure.

Posterior draws are mapped to the identifiable space. Convergence and mixing are evaluated via trace plots, effective sample size (ESS), and Geweke’s diagnostic [1512.00487], [1711.10463].

## 5. Dependence, Closure, and Interpretability

The JPSN enjoys closure under marginalization: any subset of $(\Theta_1, ..., \Theta_p, Y_1, ..., Y_q)$ is again JPSN-distributed for corresponding parameter sub-blocks. This property is inherited from the stability of both the projected normal and skew-normal distributions under subsetting.

Multivariate dependence can be quantified as follows:
- **Circular–circular correlation (Jammalamadaka–Sarma):**
\[
\rho_{cc}(\Theta_i, \Theta_j) = \frac{\mathbb{E}\big[\sin(\Theta_i - \Theta_i^*) \sin(\Theta_j - \Theta_j^*)\big]}{\sqrt{\mathbb{E}[\sin^2(\Theta_i - \Theta_i^*)]\, \mathbb{E}[\sin^2(\Theta_j - \Theta_j^*)] }},
\]
where $\Theta_i^*$ is an i.i.d. copy of $\Theta_i$.

- **Circular–linear dependence (Mardia’s measure):**
\[
\rho_{cl}^2(\Theta_i, Y_j) = \frac{\text{Cor}(\cos\Theta_i, Y_j)^2 + \text{Cor}(\sin\Theta_i, Y_j)^2 - 2\,\text{Cor}(\cos\Theta_i, Y_j)\,\text{Cor}(\sin\Theta_i, Y_j)\,\text{Cor}(\cos\Theta_i, \sin\Theta_i)}{1-\text{Cor}(\cos\Theta_i, \sin\Theta_i)^2}.
\]
- **Linear–linear dependence** is characterized by Pearson correlation.

This allows comprehensive interpretability of joint circular-linear structures [1711.10463].

## 6. Extensions to Time Series: Hidden Markov Models

Time-dependent extensions model sequences $(\theta_t, x_t)$ as emissions of a $K$-state hidden Markov model (HMM), where each state $k$ admits a state-specific JPSN$_{p,q}(\mu_k, \Sigma_k, \lambda_k)$ emission law. The latent discrete state sequence $z_t$ evolves by a Markov chain with transition matrix $\Pi$ (finite $K$) or as a stick-breaking process (infinite-state sticky HDP-HMM; sHDP-HMM).

Inference employs beam sampling (slice sampling for the state sequence), combining the JPSN Gibbs steps for within-state parameters. This enables state-dependent modeling of heterogeneous and dependent circular and linear behavior in, e.g., animal movement data, supporting automatic (Bayesian nonparametric) inference for the effective number of states [1512.00487].

## 7. Applications and Empirical Performance

Applications include modeling movement of free-ranging Maremma Sheepdogs (six animals, $p=6$ angles and $q=6$ log-step lengths, $T\approx 3000$ observations) with the sHDP–JPSN–HMM. The posterior mode for the number of behavioral states was three, corresponding to interpretable behaviors: (1) low-speed, tortuous motion (rest or livestock attending), (2) intermediate speed with moderate alignment, and (3) high-speed, nearly straight boundary patrol. The JPSN–HMM reveals intra- and inter-individual dependence in both turning and step-length, whereas simpler HMM models with independent marginals tend to over-segment the behavior and obscure social structure.

Out-of-sample predictive scores (Continuous Ranked Probability Score, CRPS) demonstrated that the JPSN–HMM outperformed alternative models using independent von Mises or wrapped Cauchy (for angles) and Gamma or Weibull (for lengths): mean circular CRPS $\approx 0.489$ versus $0.49$–$0.50$, mean linear CRPS $\approx 2.34$ versus $2.47$–$3.23$ for the alternatives [1512.00487].

A comparable analysis on zebra movement data ($p=4$, $q=4$, $T=442$) confirmed improved recovery of empirical dependence and predictive accuracy, with the JPSN attaining the lowest CRPS across held-out circular and linear segments [1711.10463].

## 8. Computational and Diagnostic Considerations

The principal computational bottleneck is the per-iteration cost: $\mathcal{O}(T K_\text{active} (p+q)^3)$ for normal-inverse-Wishart updates in HMM extensions, driven by the covariance structure. Metropolis steps for $R$ require tuning. The model’s structure allows for block-wise sparse or diagonal approximations in high dimension. Mixing and convergence are assessed via trace plots and ESS; identifiability is reliably enforced via post-processing [1512.00487], [1711.10463].

---

**Primary sources:**  
- [1512.00487] The Joint Projected and Skew Normal  
- [1711.10463] The joint projected normal and skew-normal: a distribution for poly-cylindrical data

Source: https://www.emergentmind.com/topics/joint-projected-normal-and-skew-normal-jpsn