---
title: Site Frequency Spectrum (SFS) Overview
url: https://www.emergentmind.com/topics/site-frequency-spectrum-sfs
type: topic
---

# Site Frequency Spectrum (SFS) Overview

The site frequency spectrum (SFS) is a summary of DNA sequence variation that records how many polymorphic sites have a derived allele at each possible frequency in a sample. For a sample of \(n\) sequences, the unfolded SFS is the vector \((\xi_1,\dots,\xi_{n-1})\), where \(\xi_i\) is the number of sites at which the mutant allele appears in exactly \(i\) copies. Under the standard neutral coalescent with infinite-sites mutation, the SFS links mutational counts to the branch-length decomposition of the genealogy and yields the classical expectation \(E[\xi_i]=\theta/i\), but a large literature has extended this object to linked sites, triallelic sites, multiple populations, multiple-merger coalescents, spatial genealogies, branching populations, and gene-presence/absence models [2103.00335; 1604.06713].

## 1. Definitions and genealogical representation

For a sample of size \(n\), the unfolded SFS consists of \(n-1\) frequency classes, with \(\xi_1\) counting singletons, \(\xi_2\) doubletons, and so on up to \(\xi_{n-1}\) [2103.00335]. In a large population, allele frequency may instead be treated as continuous: if \(f\in(0,1)\) is the derived-allele frequency, then \(\xi(f)\) denotes the density of mutations at frequency \(f\) along a sequence of length \(L\) [1604.06713]. When ancestral and derived states are not known, one works with the folded SFS, where counts \(i\) and \(n-i\) are merged as \(\eta_j=\xi_j+\xi_{n-j}\) for \(j=1,\dots,\lfloor (n-1)/2\rfloor\) [2308.07533].

Under the infinite-sites model, the SFS is a Poisson randomization of branch lengths. If \(L_i\) denotes the total length of all branches subtending exactly \(i\) leaves in the genealogy, then mutations occurring on such branches are observed in exactly \(i\) sampled sequences, and \(E[\xi_i]=u\,E[L_i]\) under mutation rate \(u\) per generation per lineage [2103.00335]. The same representation appears in more general genealogical settings: in coalescing Brownian motion, for example, \(L_{n,m}\) denotes the total length of all branches supporting exactly \(m\) leaves, and conditional on the tree the number of mutations inherited by exactly \(m\) individuals is Poisson with mean \(\nu L_{n,m}\) [2308.07533].

This branch-length view makes the SFS a genealogical rather than merely tabular object. It encodes how much tree length sits near the tips, at intermediate depths, or near the root, and consequently it is sensitive to the coalescent law, demographic history, linkage structure, mutation model, and even whether the genomic feature under study is always present in all sampled genomes [1407.2488].

## 2. Standard neutral theory and the invariance property

Under selective neutrality, random mating, constant population size, the Kingman coalescent, and the infinite-sites mutation model, the expected unfolded SFS has the classical form
\[
E[\xi_i]=\frac{\theta}{i},\qquad i=1,2,\dots,n-1,
\]
with \(\theta=4Nu\) in the diploid scaling used in the coalescent note of Rogers and Wooding [2103.00335]. Equivalently,
\[
E[(\xi_1,\dots,\xi_{n-1})]=\left(\theta,\frac{\theta}{2},\dots,\frac{\theta}{n-1}\right),
\]
so rare derived alleles are expected to be more numerous than common derived alleles [2103.00335]. Summing over classes gives the expected total number of segregating sites,
\[
E[S]=\theta H_{n-1},\qquad H_{n-1}=\sum_{k=1}^{n-1}\frac1k.
\]

A striking feature of this result is its invariance under sample-size increase. If one moves from sample size \(n\) to \(n+1\), then the expectations of the existing classes \(i=1,\dots,n-1\) do not change; one only appends the new class \(E[\xi_n]=\theta/n\) [2103.00335]. Rogers and Wooding explain this through coalescent recursions for the expected number of sites at frequency \(i\) at the end of the interval with \(k\) lineages, \(s_{i,k}\), and show that for fixed \(i\), \(s_{i,k}\) does not depend on \(k\) once \(k>i\) [2103.00335]. In genealogical terms, the expected total branch length subtending exactly \(i\) leaves is independent of the overall sample size beyond \(i\), so adding more samples creates new higher-frequency classes without changing the expected amount of branch length supporting the old ones [2103.00335].

A common misconception is that the \(1/i\) law is a universal neutral baseline. The literature represented here is more restrictive: the formula is specific to the infinite-sites, neutral, constant-size, panmictic model, and the later sections show multiple distinct ways in which the spectrum departs from \(1/i\) once those assumptions are relaxed [2103.00335; 1407.2488].

## 3. Extensions beyond the one-site spectrum

The one-site SFS admits several higher-dimensional generalizations. For two linked neutral sites without recombination, the expected neutral two-site spectrum (2-SFS) records either the population density \(\xi(f_1,f_2)\) or the sample counts \(\xi_{k,l}\) of pairs of sites with derived frequencies \(f_1,f_2\) or counts \(k,l\) [1604.06713]. A central decomposition separates **nested** mutation pairs, in which one mutation lies inside the background of the other, from **disjoint** pairs, in which the derived alleles never co-occur on the same chromosome:
\[
\xi(f_1,f_2)=\xi^N(f_1,f_2)+\xi^D(f_1,f_2),\qquad
\xi_{k,l}=\xi^N_{k,l}+\xi^D_{k,l}.
\]
This decomposition captures haplotype structure that is not identifiable from \(\xi_{k,l}\) alone [1604.06713].

Conditioning on a focal mutation of known frequency yields a conditional SFS of linked sites. In the no-recombination setting, linked mutations relative to the focal can be strictly nested, co-occurring, enclosing, complementary, or strictly disjoint, and exact formulas are available both in finite samples and in the population limit [1604.06713]. These expressions provide a neutral reference for linked variation around structural variants, introgressed haplotypes, or chromosomal inversions. In the inversion application, the inversion frequency itself acts as the focal-site frequency, and the expected spectrum of neutral mutations inside the inverted region is exactly the conditional linked-site SFS [1604.06713].

A different extension allows up to two mutation events at a site and thereby produces a triallelic SFS. In that setting, the sample configuration \((n_a,n_b,n_c)\) of ancestral allele \(a\) and two derived alleles \(b,c\) depends on whether the two mutation events are nested or nonnested on the underlying genealogy [1310.3444]. Under general variable population size, the triallelic SFS can be written in closed form in terms of the second moments \(E[T_jT_k]\) of inter-coalescent times, in contrast to the usual biallelic SFS, which depends on first moments \(E[T_k]\) [1310.3444]. This suggests a sharper sensitivity to demographic history, and the numerical comparisons in that work indeed show larger Kullback–Leibler divergences between demographic models for the triallelic spectrum than for the diallelic one [1310.3444].

## 4. Alternative genealogies and nonclassical spectra

When the genealogy is not Kingman, the SFS changes qualitatively. In \(\Lambda\)-coalescents, multiple ancestral lineages can merge at once, with merger rates determined by a finite measure \(\Lambda\) on \([0,1]\) [1305.6043]. The expected SFS no longer has the simple \(\theta/i\) form; instead it is computed from recursions involving the block-counting process, the expected sojourn times \(g(n,k)\), and the probability \(p^{(n)}[k,b]\) that a branch at level \(k\) subtends \(b\) leaves [1305.6043]. For regularly varying \(\Lambda\)-coalescents with exponent \(1\), including the Bolthausen–Sznitman coalescent, branch lengths of fixed order and hence fixed-frequency SFS entries have asymptotics very different from Kingman:
\[
\widehat\ell_{n,1}\sim \frac{n}{\log n},\qquad
\widehat\ell_{n,a}\sim \frac{1}{(a-1)a}\frac{n}{\log^2 n},\ a\ge2,
\]
so
\[
E[\xi_1^{(n)}]\sim \theta\frac{n}{\log n},\qquad
E[\xi_a^{(n)}]\sim \frac{\theta}{(a-1)a}\frac{n}{\log^2 n},\ a\ge2
\]
for the Bolthausen–Sznitman case [1804.00961]. These asymptotics imply an extreme excess of low-frequency variants relative to Kingman.

Spatial expansion can also alter the spectrum. For genealogies modeled by coalescing Brownian motion on the line or circle, the total branch length \(L_{n,m}\) supporting exactly \(m\) leaves satisfies
\[
nL_{n,m}\xrightarrow[]{\mathbb P}1
\]
for each fixed \(m\) [2308.07533]. Thus, for fixed \(m\), the \(m\)-class branch length is typically of order \(1/n\), in strong contrast to the order-one Kingman behavior for fixed frequency classes [2308.07533].

Branching-population models replace coalescent sampling formulas by forward-time growth formulas. In a supercritical birth–death process with mutations only at birth, the large-sample SFS for a uniform sample of size \(n\) taken at time \(T_n\) has
\[
\frac{M^k_{n,T_n}}{n}\xrightarrow{P}\frac{\lambda\nu}{r\,k(k-1)},\qquad k\ge2,
\]
under \(n e^{-rT_n}\to0\), so the large-\(n\) shared-mutation spectrum has the \(1/[k(k-1)]\) form familiar from exponential-growth branching models despite the underlying time-inhomogeneous ancestral reproduction rate [2402.16306]. For whole-population SFS in exponentially growing branching processes, almost sure limits are available both at fixed times and at hitting times of large population size, and in the pure-birth case the limiting per-individual SFS is proportional to \(1/[j(j+1)]\) [2307.03346].

Tumor models sharpen this picture. In a birth–death branching process with cell division rate \(r_0\), death rate \(d_0\), net growth \(\lambda_0=r_0-d_0\), and extinction probability \(p_0=d_0/r_0\), the expected SFS of the total population at the fixed-time scale \(t_N\) defined by \(E[Z_0(t_N)\mid Z_0(t_N)>0]=N\) is
\[
E[S_j(t_N)\mid Z_0(t_N)>0]=wN\int_0^{1-1/N}\frac{1-y}{1-p_0y}y^{j-1}\,dy
\]
[2102.11959]. For large \(j\) this behaves like \(1/j^2\), whereas as \(p_0\to1\) with fixed \(j\) it approaches \(wN/j\), producing a transition between a \(1/j^2\) spectrum and a \(1/j\) spectrum that reflects cell viability [2102.11959]. In a different rescued-population model with rare resistance mutations, fixed-\(i\) classes among resistant cells satisfy
\[
\mathbb E[S_i^N(t_N)]\sim I(i)\,\frac{2b_0\gamma\omega}{\lambda_1+\lambda_0}N^{\lambda_1 t+1-\alpha},
\]
showing how rescue dynamics alter the normalization while preserving the branching-process kernel \(I(i)\) for small families [2303.04069].

## 5. Coupled genomic features, linked markers, and dispensable genes

The classical SFS presumes that all sampled sequences carry homologous material at the locus under study. Dispensable genes violate that assumption. In the infinitely many genes plus infinitely many sites framework, each gene can be gained, lost, and, while present, accumulate site mutations [1407.2488]. The joint gene-and-site spectrum \(G_{k,s}\) counts mutated sites whose gene is present in exactly \(k\) sampled genomes and whose mutation is present in exactly \(s\) of those \(k\) carriers [1407.2488]. From this, the conditional within-gene SFS for a gene present in exactly \(k\) out of \(n\) individuals is derived explicitly; for \(s<k\),
\[
\mathbb E[S_s^u\mid F(u)=k]
=
\frac{\theta_2}{s}\; \frac{k}{n}\; \binom{n-1}{s}^{-1}
\sum_{j=0}^{n-s-1}\frac{j+1}{j+1+\rho}\binom{n-j-2}{s-1},
\]
and even when \(k=n\), the expectation differs from the classical \(\theta_2/s\) formula if the gene-loss rate \(\rho>0\) [1407.2488].

This has immediate inferential consequences. Watterson’s estimator and Tajima’s \(\pi\), when naively applied to dispensable genes, are both biased downward; the paper gives the explicit expectations and shows
\[
\mathbb E[\widehat\theta_{W,k}] \le \frac{k}{n}\theta_2,\qquad
\mathbb E[\widehat\pi_k] \le \frac{k}{n}\theta_2
\]
[1407.2488]. The same work reports that the normalized SFS of dispensable genes shows an excess of low-frequency variants relative to the classical neutral SFS and that Tajima’s \(D\) is shifted toward negative values even under strict neutrality, especially for rare genes and large \(\rho\) [1407.2488]. A common interpretation of negative Tajima’s \(D\) as evidence for purifying selection or expansion is therefore not portable to dispensable genes without adjusting the neutral baseline.

Linked markers generate a different kind of coupling. In the no-recombination 2-SFS framework, the SFS conditional on a focal neutral mutation gives a natural null model for linked variation around structural variants, introgressed segments, and inversions [1604.06713]. The same paper notes an immediate reinterpretation in ancestral-recombination-graph language: because recombination events follow a Poisson process on branches like mutations, the conditional SFS of linked sites can describe the distribution of recombination events along a sequence [1604.06713].

## 6. Statistical uses, computation, and fundamental limits

Linear functionals of the SFS underpin mutation-rate estimation and neutrality testing. Under the standard coalescent, any statistic of the form
\[
T=c^\top \boldsymbol{\xi}=\sum_{i=1}^{n-1} c_i\xi_i
\]
with \(c^\top \boldsymbol v=1\) and \(v_i=1/i\) is an unbiased estimator of \(\theta\) [2101.04941]. Watterson’s estimator, the average pairwise-difference estimator, Fay and Wu’s \(\hat\theta_H\), Zeng et al.’s \(\hat\theta_L\), and unnormalized versions of Tajima’s \(D\) all fit this framework [2101.04941]. Multivariate phase-type theory provides an explicit probability generating function for the full SFS under the standard coalescent, and many 0–1 or non-negative integer weighted linear statistics are discrete phase-type distributed [2101.04941]. Statistics with mixed-sign weights, such as Tajima’s \(D\), are generally not discrete phase-type distributed, but their characteristic functions remain explicit and can be numerically inverted [2101.04941]. The same framework is implemented in the R package **phasty** [2101.04941].

For structured demographies, the expected joint SFS across multiple populations can be computed efficiently from truncated single-population spectra and dynamic programming on a demographic DAG. New recursions for the truncated SFS reduce the core computation to \(O(n^2)\) per subpopulation, and these ideas are implemented in **momi** for joint SFS inference under arbitrary population-size histories, including piecewise exponential growth [1503.01133]. In tree-shaped demographies, forward-time Moran-model representations and FFT-based convolutions further reduce the cost of computing the expected joint spectrum [1503.01133].

For \(\Lambda\)-coalescents, exact finite-\(n\) recursions for expectations, variances, and covariances of SFS components extend the classical formulas of Fu from Kingman to multiple-merger genealogies [1305.6043]. The normalized expected SFS
\[
\varphi_n(i)=\frac{E[\xi_i^{(n)}]}{\sum_{\ell=2}^n \ell E[T_\ell]}
\]
is interpreted there as the probability that a randomly chosen mutation is seen in \(i\) copies, and it becomes the basis of pseudo-likelihood inference for parameters such as the Beta-coalescent \(\alpha\) or the Eldon–Wakeley point-mass parameter \(\psi\) [1305.6043].

The SFS is also the basis of a large class of neutrality tests, but its informativeness has hard limits. A decomposition of neutrality tests into waiting-time and topology components shows that statistics such as Tajima’s \(D\) and Fay and Wu’s \(H\) depend directly on a measure of tree balance that is largely determined by the root balance of the genealogy [1510.06748]. At the same time, there are information-theoretic lower bounds on what SFS-based demographic inference can achieve. For population histories with bottlenecks, the minimax error for estimating the size history from \(s\) independent segregating sites is at least of order \(1/\log s\), and the bound does not depend on the SFS dimension \(n-1\) [1505.04228]. This implies that, for fixed \(s\), increasing the number of sampled individuals does not improve the minimax lower bound. A plausible implication is that linkage-aware or genealogy-aware summaries are necessary when one wishes to resolve demographic features that the SFS compresses too aggressively.

The contemporary view of the SFS is therefore twofold. It remains a central and highly tractable summary of mutational variation, with exact formulas, recursion schemes, and matrix methods available for many models [2103.00335; 2101.04941; 1503.01133]. But its classical \(1/i\) form, its standard inferential interpretations, and even its attainable accuracy are all model-dependent, and the modern literature treats the SFS less as a universal law than as a family of genealogy-dependent spectra whose shape changes in systematic and diagnostically useful ways across coalescent, branching, spatial, and pangenomic settings [1604.06713; 1407.2488; 1804.00961; 1505.04228].

Source: https://www.emergentmind.com/topics/site-frequency-spectrum-sfs