---
title: Folded Normal Distribution
url: https://www.emergentmind.com/topics/folded-normal
type: topic
---

# Folded Normal Distribution

Searching arXiv for recent and foundational papers on the folded normal distribution.
The folded normal distribution, denoted \(FN(\mu,\sigma^2)\), is the distribution of \(X=|Y|\) when \(Y\sim N(\mu,\sigma^2)\). It is supported on \([0,\infty)\), inherits its parameters from the underlying normal law, and reduces to the half-normal at \(\mu=0\). Because folding removes the sign of the latent Gaussian variable, the statistical model is identifiable only up to \(\pm \mu\). Classical analyses develop its density, distribution function, transforms, moments, entropy, and maximum-likelihood estimation, while recent work treats it as a prototypical nonregular model whose asymptotic behavior changes qualitatively at the kink \(\mu=0\) [1402.3559; 2509.12206].

## 1. Definition, parameterization, and elementary representations

Let
\[
Y\sim N(\mu,\sigma^2), \qquad \mu\in\mathbb R,\ \sigma>0,
\]
and define
\[
X=|Y|.
\]
Then \(X\sim FN(\mu,\sigma^2)\). For \(x\ge 0\), the density is
\[
f(x;\mu,\sigma)
=\frac{1}{\sigma}\left\{\phi\!\left(\frac{x-\mu}{\sigma}\right)+\phi\!\left(\frac{x+\mu}{\sigma}\right)\right\}
=\frac{1}{\sigma\sqrt{2\pi}}
\exp\!\left(-\frac{x^2+\mu^2}{2\sigma^2}\right)\,
2\cosh\!\left(\frac{x\mu}{\sigma^2}\right),
\]
and the distribution function is
\[
F(x)=\Phi\!\left(\frac{x-\mu}{\sigma}\right)-\Phi\!\left(\frac{-x-\mu}{\sigma}\right),
\qquad x\ge 0.
\]
Equivalently, the law is obtained by summing the left and right Gaussian densities after reflection onto the positive half-axis. The support is \(x\in[0,\infty)\), and \(f(x;\mu,\sigma)=f(x;-\mu,\sigma)\), so only \(|\mu|\) is identified [2509.12206].

The parameter \(\mu\) is the mean of the underlying normal variable, not the mean of the observed nonnegative variable \(X\). For modeling, the ratio
\[
\theta=\frac{\mu}{\sigma}
\]
is especially important, because many qualitative properties depend primarily on \(\mu/\sigma\). When \(\mu=0\), the folded normal reduces to the half-normal. If \(Z=X/\sigma\), then \(Z\) is a noncentral \(\chi\) distribution with \(1\) degree of freedom and noncentrality parameter \(|\mu|/\sigma\) [1402.3559].

## 2. Distributional properties, transforms, and information measures

Tsagris, Beneki, and Hassani derived the characteristic function and moment generating function explicitly. The characteristic function is
\[
\varphi_X(t)
=
e^{-\frac{\sigma^2 t^2}{2}+i\mu t}
\left[1-\Phi\!\left(-\frac{\mu}{\sigma}+i\sigma t\right)\right]
+
e^{-\frac{\sigma^2 t^2}{2}-i\mu t}
\left[1-\Phi\!\left(\frac{\mu}{\sigma}+i\sigma t\right)\right],
\]
and the moment generating function is
\[
M_X(t)
=
e^{\frac{\sigma^2 t^2}{2}+\mu t}
\left[1-\Phi\!\left(-\frac{\mu}{\sigma}-\sigma t\right)\right]
+
e^{\frac{\sigma^2 t^2}{2}-\mu t}
\left[1-\Phi\!\left(\frac{\mu}{\sigma}-\sigma t\right)\right].
\]
These formulas imply existence of moments of all orders and provide a direct route to cumulants and higher-order expansions [1402.3559].

The first two moments take the standard folded-normal form
\[
\mathbb E[X]
=
\sigma\sqrt{\frac{2}{\pi}}\,
\exp\!\left(-\frac{\mu^2}{2\sigma^2}\right)
+
\mu\left[1-2\Phi\!\left(-\frac{\mu}{\sigma}\right)\right],
\]
\[
\mathbb E[X^2]=\mu^2+\sigma^2,
\qquad
\operatorname{Var}(X)=\mu^2+\sigma^2-\big(\mathbb E[X]\big)^2.
\]
The mode does not admit a closed form in general; differentiating the density yields the nonlinear equation
\[
x=-\frac{\sigma^2}{2\mu}\log\!\left(\frac{\mu-x}{\mu+x}\right).
\]
Numerically, the maximum is at \(x=0\) when \(|\mu|<\sigma\), while for \(|\mu|\ge \sigma\) the mode occurs at some \(x>0\). For \(|\mu|\gtrsim 3\sigma\), the mode approaches \(|\mu|\), and the folded normal becomes very close to the underlying normal truncated to the positive axis. The distribution is not an exponential family and is not stable under addition [1402.3559].

Entropy and Kullback–Leibler divergences to the normal and half-normal do not simplify to elementary closed forms in the treatment of [1402.3559]. They are approximated by Taylor expansions involving the same basic integral
\[
\int_0^\infty f(x)\log\!\left(1+e^{-2\mu x/\sigma^2}\right)\,dx.
\]
The reported numerical behavior is that these Taylor approximations work well for moderate or large \(\mu/\sigma\), whereas for small \(\mu/\sigma\) direct numerical integration is preferable [1402.3559].

## 3. Likelihood geometry, identifiability, and maximum likelihood estimation

For observations \(y_1,\dots,y_n\ge 0\), the sample log-likelihood can be written as
\[
\ell_n(\mu,\sigma)
=
\sum_{i=1}^n
\Bigg\{
-\log\sigma
-\frac{y_i^2+\mu^2}{2\sigma^2}
+\log\!\Big(2\cosh\frac{y_i\mu}{\sigma^2}\Big)
-\frac12\log(2\pi)
\Bigg\}.
\]
Because \(f(y;\mu,\sigma)=f(y;-\mu,\sigma)\), the true identification set is
\[
\Theta_0=\{(+\mu_0,\sigma_0),(-\mu_0,\sigma_0)\}.
\]
Identifiability up to sign can be proved from two invariants of the distribution: the second moment,
\[
\mathbb E_\theta[Y^2]=\mu^2+\sigma^2,
\]
and the density at zero,
\[
f(0;\mu,\sigma)=\frac{2}{\sigma\sqrt{2\pi}}
\exp\!\left(-\frac{\mu^2}{2\sigma^2}\right).
\]
These determine \(\sigma\) and \(|\mu|\) uniquely [2509.12206].

A central structural fact is boundary coercivity of the profiled likelihood. Let
\[
\ell_{n,p}(\sigma)=\sup_{\mu\in\mathbb R}\ell_n(\mu,\sigma),
\qquad
s^2=\frac1n\sum_{i=1}^n (y_i-\bar y)^2.
\]
If \(s^2>0\), then for all \(\sigma\in(0,1]\),
\[
\sup_{\mu\in\mathbb R}\ell_n(\mu,\sigma)
\le
-n\log\sigma-\frac{n s^2}{2\sigma^2}+C_0,
\qquad
C_0=n\log 2-\frac{n}{2}\log(2\pi),
\]
and
\[
\lim_{\sigma\downarrow 0}\ell_{n,p}(\sigma)
=
\lim_{\sigma\to\infty}\ell_{n,p}(\sigma)
=
-\infty.
\]
This yields existence of a maximizer in \(\sigma\) for nonconstant data. For fully degenerate samples \(y_i\equiv 0\) or \(y_i\equiv c>0\), the supremum of the likelihood is \(+\infty\), so those cases are excluded from the main asymptotic theory [2509.12206].

For fixed \(\sigma>0\), the likelihood in \(\mu\) is unimodal on \(\mu\ge 0\). Writing
\[
S_y=\sum_{i=1}^n y_i^2,
\qquad
k(\sigma,\mu)=\sum_{i=1}^n y_i\tanh\!\left(\frac{y_i\mu}{\sigma^2}\right)-n\mu,
\]
the maximizer satisfies \(k(\sigma,\hat\mu(\sigma))=0\). If \(\sigma^2\ge S_y/n\), then the unique maximizer is \(\hat\mu(\sigma)=0\); if \(\sigma^2<S_y/n\), there is a unique positive maximizer. The implicit-function analysis in [2509.12206] shows that \(\hat\mu(\sigma)\) is \(C^1\) on the region where it is positive and is strictly decreasing in \(\sigma\). The profiled derivative has a one-crossing property, so there is a unique \(\hat\sigma>0\), and the full MLE set is
\[
\{(\hat\mu_n,\hat\sigma_n),(-\hat\mu_n,\hat\sigma_n)\}.
\]

## 4. Nonregularity at the kink and asymptotic regimes

The folded normal is nonregular because the usual regularity conditions fail at \(\mu=0\). The pointwise log-likelihood has \(\mu\)-score
\[
S_\mu(y;\mu,\sigma)
=
-\frac{\mu}{\sigma^2}
+
\frac{y}{\sigma^2}
\tanh\!\left(\frac{y\mu}{\sigma^2}\right),
\]
so at \(\mu=0\),
\[
S_\mu(y;0,\sigma)=0 \quad \text{for all } y,
\qquad
I_{\mu\mu}(0,\sigma)=\mathbb E[S_\mu^2(Y;0,\sigma)]=0.
\]
Thus the Fisher information matrix is singular in the \(\mu\)-direction at the kink, although it is continuous in \(\theta\) and positive definite when \(\mu_0\neq 0\) [2509.12206].

The sharp asymptotic distinction is summarized below.

| Regime | Main asymptotic statement | Consequence |
|---|---|---|
| \(\mu_0\neq 0\) | \(\sqrt n(\hat\theta_n-\theta_0)\overset{d}{\to}\mathcal N(0,I(\theta_0)^{-1})\) | Regular asymptotic normality |
| \(\mu_0=0\) | \(n^{1/4}\hat\mu_n\overset{d}{\to}T^\star\), while \(\hat\sigma_n-\sigma_0=O_p(n^{-1/2})\) | Mixed rates and non-Gaussian limit for location |

At \(\mu_0=0\), the relevant local expansion comes from a sixth-order expansion of \(\log(2\cosh t)\),
\[
\log(2\cosh t)=\log 2+\frac{t^2}{2}-\frac{t^4}{12}+r_6(t),
\qquad
|r_6(t)|\le C_6|t|^6.
\]
On the local scale \(\mu=n^{-1/4}t\), the likelihood contrast converges to a random quadratic-minus-quartic limit
\[
\Psi_Z(t)=aZt^2-bt^4,
\]
with \(a=1/(2\sigma_0^4)\) and \(b=\mathbb E[Y^4]/(12\sigma_0^8)>0\). The argmax of this limit yields the non-Gaussian distribution of \(n^{1/4}\hat\mu_n\). By contrast, the \(\sigma\)-direction remains regular and retains a \(\sqrt n\) rate. The Hausdorff distance from the estimated identification set to the true set is therefore governed by the slower rate:
\[
d_H\big(\{(\pm\hat\mu_n,\hat\sigma_n)\},\Theta_0\big)=O_p(n^{-1/4})
\quad \text{when } \mu_0=0.
\]
Consistency itself follows from coercivity, explicit uniform laws of large numbers, and strict Kullback–Leibler separation away from \(\Theta_0\) [2509.12206].

## 5. Multivariate folding and the extended skew-normal framework

The folded normal is the one-dimensional instance of a more general componentwise folding operation. If \(X\in\mathbb R^p\) has pdf \(f_X(x;\boldsymbol\theta)\) and
\[
Y=|X|=(|X_1|,\ldots,|X_p|)^\top,
\]
then for \(y\ge \mathbf 0\),
\[
f_Y(y)=\sum_{\mathbf s\in\{-1,1\}^p} f_X(\mathbf D_s y;\boldsymbol\theta),
\qquad
F_Y(y)=\sum_{\mathbf s\in\{-1,1\}^p}\pi_s F_X(\mathbf D_s y;\boldsymbol\theta),
\]
where \(\mathbf D_s=\operatorname{Diag}(\mathbf s)\) and \(\pi_s=\prod_i s_i\). This sign-sum representation gives the multivariate folded distribution in full generality [2009.13488].

A particularly useful extension is the folded extended skew-normal (ESN) family. If
\[
X\sim ESN_p(\boldsymbol\mu,\boldsymbol\Sigma,\boldsymbol\lambda,\tau),
\]
then \(Y=|X|\) has density
\[
f_Y(y)=\sum_{\mathbf s\in\{-1,1\}^p}
ESN_p(y;\mathbf D_s\boldsymbol\mu,\mathbf D_s\boldsymbol\Sigma\mathbf D_s,\mathbf D_s\boldsymbol\lambda,\tau),
\qquad y\ge \mathbf 0.
\]
The ordinary folded normal is the special case
\[
p=1,\qquad \lambda=0,\qquad \tau=0.
\]
Within this framework, arbitrary truncated and folded ESN moments can be computed by recurrence relations, and any truncated ESN moment can be expressed as a corresponding moment of a higher-dimensional truncated multivariate normal. This reduction is computationally important because it converts ESN moment evaluation into a truncated-normal problem for which optimized algorithms are available. The R package **MomTrunc** implements these methods for truncated normal, truncated ESN, folded ESN, and, as a special case, the folded normal [2009.13488].

## 6. Statistical practice, applications, and terminological scope

In likelihood computation, two approaches appear in the literature. One is direct numerical maximization of the log-likelihood, for example via `optim`; the other uses the score equations
\[
\sum_{i=1}^n \frac{x_i}{1+e^{2\mu x_i/\sigma^2}}
=
\frac12\sum_{i=1}^n (x_i-\mu),
\qquad
\hat\sigma^2=\frac1n\sum_{i=1}^n x_i^2-\hat\mu^2,
\]
as the basis of an iterative MLE scheme. An EM formulation based on latent signs was attempted in [1402.3559] but did not perform as well as direct maximization in the reported experiments. For interval estimation, the same paper found that asymptotic normal confidence intervals perform poorly when \(\theta=\mu/\sigma\) is small, whereas percentile bootstrap intervals are clearly preferable when \(\theta<1\). When \(\theta\ge 1\), both methods perform similarly, and for \(\theta\ge 1.5\) the asymptotic intervals for \(\mu\) achieve close to nominal \(95\%\) coverage across the reported sample sizes [1402.3559].

Applications mentioned for the folded normal include absolute deviations, process capability indices, and magnitudes of random effects or errors. Within the broader folded and truncated ESN setting, related applications include censored and truncated data in environmental science, finance, longitudinal censored models, and geostatistical models. A specific illustration in [1402.3559] fitted a folded normal model to body mass index data for \(700\) New Zealand adults, with reported estimates
\[
\hat\mu=26.685,\qquad \hat\sigma^2=21.324.
\]
That example reflects a regime in which the estimated probability of a negative latent value is essentially zero, so the folded normal behaves very similarly to a normal law on the positive axis [1402.3559].

The term also has an unrelated use outside statistics. In the slow–fast dynamical-systems literature, “folded normal” can refer informally to a folded equilibrium of node type, namely a folded node on a critical manifold. That usage concerns canards, loss of normal hyperbolicity, and piecewise-smooth analogues obtained by pinching, and it is unrelated to the probability distribution \(X=|Y|\) with \(Y\sim N(\mu,\sigma^2)\) [1506.00831].

Source: https://www.emergentmind.com/topics/folded-normal