---
title: Pair Correlation Functions Overview
url: https://www.emergentmind.com/topics/pair-correlation-functions-pcfs
type: topic
---

# Pair Correlation Functions Overview

Pair correlation functions (PCFs) are two-point statistics that quantify how pairs are distributed relative to an uncorrelated reference. In liquids, plasmas, glasses, and electron systems they are commonly written as radial distribution functions \(g(r)\) or \(g_{\alpha\beta}(r)\); in spatial point-process theory they appear as \(g_0(h)\); in heavy-ion femtoscopy they are written as two-particle correlation functions \(C(q)\); in one-dimensional model sets they are averaged pair correlations \(\nu(y)\); and in the analytic theory of the Riemann zeta function they appear as pair-correlation sums such as \(F(\alpha,T)\) [2308.11151] [2004.00527] [2510.23030] [2502.20487] [2502.20569]. Across these settings, the normalization is typically chosen so that the statistic tends to \(1\) in an uncorrelated or asymptotic regime, while deviations from \(1\) encode ordering, exclusion, clustering, screening, or oscillatory structure.

## 1. Definitions, normalizations, and basic objects

In atomistic statistical mechanics, the PCF measures the probability of finding a second particle at a separation \(r\) relative to an ideal gas at the same density. For glasses, the pair correlation function or radial distribution function \(g_{\alpha\beta}(r)\) is defined for atom types \(\alpha\) and \(\beta\) through interatomic separations \(r_{ij}\) and an ensemble average over molecular-dynamics trajectories [2308.11151]. For the unpolarized uniform electron gas, the spin-averaged pair correlation \(g(r)\) is defined so that \(n\,g(r)\,4\pi r^2\,dr\) is the probability to find a second electron between \(r\) and \(r+dr\), with spin-resolved functions \(g_{uu}(r)\), \(g_{dd}(r)\), and \(g_{ud}(r)\) entering the unpolarized average [2203.12272]. In crystals and in classical density-functional theory, one starts from the one-particle density \(\rho(\mathbf r)\), the two-particle density \(\rho^{(2)}(\mathbf r_1,\mathbf r_2)\), the pair distribution function \(g(\mathbf r_1,\mathbf r_2)\), and the total pair correlation \(h(\mathbf r_1,\mathbf r_2)=g(\mathbf r_1,\mathbf r_2)-1\) [1405.0561] [2605.18326].

For spatial point processes under second-order intensity-reweighted stationarity, the central relation is
\[
\rho^{(2)}(x,y)=\rho(x)\rho(y)g_0(y-x),
\]
where translation invariance is imposed only after reweighting by the intensity field \(\rho\) [2004.00527]. The associated inhomogeneous \(K\)-function is
\[
K(t)=\int_{\|h\|\le t} g_0(h)\,dh,
\]
and in isotropic settings one writes \(g_0(h)=g_1(\|h\|)\) [2004.00527].

In heavy-ion femtoscopy, the object of interest is the correlation function of emitted fragments as a function of relative momentum,
\[
q \equiv \mu\left|\frac{p_1}{m_1}-\frac{p_2}{m_2}\right|,\qquad \mu=\frac{m_1m_2}{m_1+m_2},
\]
with
\[
C(q)=C_{12}\,\frac{Y_{\mathrm{same}}(q)}{Y_{\mathrm{mix}}(q)},
\]
where \(Y_{\mathrm{same}}\) is the same-event histogram, \(Y_{\mathrm{mix}}\) is the mixed-event reference, and \(C_{12}\) fixes \(C(q\to\infty)=1\) [2510.23030]. This formulation makes explicit that femtoscopic PCFs are experimentally constructed from coincidence yields rather than directly from spatial coordinates.

In one-dimensional regular model sets, the averaged pair correlation is
\[
\nu(y)=\lim_{R\to\infty}\frac{1}{2R}\#\{x\in \Lambda\cap[-R,R]:x+y\in\Lambda\},
\]
and it is equivalently the mass of the autocorrelation measure at \(y\) [2502.20487]. For \(n\)-type decompositions \(\Lambda=\bigsqcup_{i=1}^n\Lambda_i\), one refines this to cross-correlations \(\nu_{ij}(y)\) [2502.20487].

In Montgomery-type pair correlation for zeta zeros, assuming RH so that nontrivial zeros are \(1/2+i\gamma\), one defines
\[
F(\alpha,T)= \frac{2\pi}{T\log T}\sum_{\gamma,\gamma'\in(0,T]} T^{i\alpha(\gamma-\gamma')}w(\gamma-\gamma'),
\]
with \(w(u)=4/(u^2+4)\) in the classical case [2502.20569]. The terminology is the same, but the “pairs” are pairs of zero ordinates rather than particles or points in Euclidean space.

## 2. Physical interpretation and representative regimes

The most standard interpretation of a PCF is that \(g(r)=1\) signals the ideal-gas limit or decay of correlations, while peaks signal preferred length scales. This is explicit in the classical radial PCF, where \(\rho\,g(r)\,a(r)\,dr\) gives the expected number of particles in a shell at distance \(r\), and where peaks in \(g(r)\) encode preferred separations [2210.09731]. In cosmology, the galaxy pair-correlation function \(\xi(r)\) is the excess probability, relative to a Poisson distribution, of finding two galaxies separated by \(r\); oscillations in the range \(20\)–\(200\ \mathrm{Mpc}/h\) are identified with baryon acoustic oscillations [1401.7218].

In warm dense matter, the onset of short-range order is visible as a first peak in \(g(r)\). In the unpolarized uniform electron gas, SMPIMC calculations show that when the classical coupling \(\Gamma\) exceeds roughly \(5\)–\(10\), depending on degeneracy, a pronounced first peak develops, and this appearance of short-range order is accompanied by a high-momentum “quantum tail” in the momentum distribution [2203.12272]. The paper interprets these as two aspects of the same many-body effect: transient effective wells formed by neighboring electrons and tunneling through those wells [2203.12272].

In two-temperature strongly coupled plasmas, PCFs resolve interspecies screening and the validity of closure theories. Molecular-dynamics comparisons show that the Seuferling–Vogel–Teopffer model reproduces electron-electron, ion-ion, and electron-ion radial distribution functions well over a wide range of mass ratios, temperature ratios, and couplings, while the Yukawa OCP description of ion-ion correlations is accurate only up to \(\Gamma_e\approx1\), beyond which it rapidly breaks down [1707.01509].

In fractional quantum Hall states, the pair-correlation function \(g(r)\) is a reduced radial object constructed from the trial wavefunction after fixing one particle at the origin and another at separation \(r\). Its thermodynamic-limit form feeds directly into the projected static structure factor and the single-mode-approximation estimate of the magneto-roton gap [2211.00371]. In the two-dimensional Hubbard model, the phrase “pair correlation” refers instead to superconducting pairing correlators such as \(P_d(\ell)=\langle \Delta_d^\dagger(i+\ell)\Delta_d(i)\rangle_H\), where enhancement away from half-filling and growth with system size are used as a criterion for superconductivity [1304.1273]. The shared terminology therefore covers both geometric pair distributions and operator-valued pairing correlators.

## 3. Integral equations, closures, and broken symmetry

A large part of PCF theory is organized by the Ornstein–Zernike equation. In a general inhomogeneous system,
\[
h(\mathbf r_1,\mathbf r_2)=c(\mathbf r_1,\mathbf r_2)+\int d\mathbf r_3\; c(\mathbf r_1,\mathbf r_3)\rho(\mathbf r_3)h(\mathbf r_3,\mathbf r_2),
\]
with \(c\) the direct pair correlation function [1405.0561] [2605.18326]. In homogeneous fluids this reduces to the usual convolution form, and a closure relation such as Percus–Yevick, hypernetted-chain, Rogers–Young, or Zerah–Hansen is then required [2605.18326].

For crystals and other ordered phases, the essential step is to decompose one- and two-particle quantities into symmetry-conserving and symmetry-broken parts. The density is written as \(\rho(\mathbf r)=\rho+\rho^b(\mathbf r)\), and the pair correlations as
\[
h=h^{(0)}+h^{(b)},\qquad c=c^{(0)}+c^{(b)},
\]
where \(h^{(0)}\) and \(c^{(0)}\) are homogeneous-fluid correlations at the average density, and \(h^{(b)}\), \(c^{(b)}\) are periodic in the center-of-mass coordinate and encode the broken translational symmetry [1405.0561]. The broken direct correlation function is then expanded in powers of the order parameters by using higher-body direct correlation functions of the uniform fluid [1405.0561] [2605.18326]. In the two-dimensional hexagonal soft-disk example, truncation at the three-body term already captures more than \(80\%\) of the full \(c^{(b)}\) near freezing [1405.0561].

Within classical density-functional theory, these PCFs are not merely diagnostic. The excess free-energy functional generates the direct correlations through functional differentiation, and the grand-potential difference between competing phases contains both a symmetry-conserving contribution and a symmetry-breaking contribution involving \(\bar c^{(b)}\) [2605.18326]. This is the basis for the “exact” DFT described in the review, where directly computed PCFs of broken-symmetry phases are used to predict freezing, nematic–isotropic transitions, and crystal–crystal transitions [2605.18326].

Heavy-ion femtoscopy uses a different dynamical framework but addresses the same two-point object. The classical trajectory approximation CTA-I combines a thermal equilibrium emission source, sequential emission with an exponential time delay, a self-consistent mean field that is harmonic at small \(r\) and Coulombic at large \(r\), and the Coulomb interaction between the emitted fragments [2510.23030]. The resulting three-body Monte Carlo propagates the source nucleus and two emitted particles to asymptotic momenta before constructing \(C(q)\) from accepted events and mixed events [2510.23030].

## 4. Estimation, normalization, and corrections in empirical settings

The numerical value of a PCF depends critically on the reference measure used in its denominator. In inhomogeneous spatial point processes, the “global” estimator replaces pointwise intensity products by the overlap integral
\[
\gamma(h)=\int_{W\cap(W-h)} \rho(u)\rho(u+h)\,du,
\]
leading to the unbiased estimators
\[
\hat K_{\rm global}(t)=\sum_{x\neq y\in X\cap W}\frac{1\{\|y-x\|\le t\}}{\gamma(y-x)},
\qquad
\hat g_{\rm global}(h)=\frac{1}{\gamma(h)}\sum_{x\neq y\in X\cap W}\kappa_b(h-(y-x)).
\]
Under SOIRS, these remain unbiased, and the simulations in the paper show reduced bias and variance relative to local normalization, especially when the intensity must itself be estimated [2004.00527].

On discrete lattices, exact normalization requires counting all site pairs at a given lattice distance. For occupied sites \(\mathbb L\), metric \(d\), and separation \(m\), the discrete PCF is
\[
f_d(m)=\frac{c_d(m)}{\mathbb E[\bar c_d(m)]},\qquad
\mathbb E[\bar c_d(m)]=\frac{N}{Z}\frac{N-1}{Z-1}s_d(m),
\]
where \(c_d(m)\) counts occupied-occupied pairs and \(s_d(m)\) counts site-site pairs at distance \(m\) [1804.03452]. On square lattices, taxicab and uniform metrics yield isotropic, correctly normalized PCFs, unlike the annular PCF, which is poorly calibrated on a lattice, or the rectilinear PCF, which is anisotropic and misses diagonal patterns such as chessboards and diagonal stripes [1804.03452].

When obstacles are present, naive normalization is no longer valid because straight-line lattice distances do not respect connectivity. The corrected pair-correlation function replaces the standard pair count \(D(m)\) by
\[
D_{\rm corr}(m)=D_{\rm NO}(m)-A(m)+I(m)-L(m)+G(m),
\]
thereby removing accessible-obstacle and obstacle-obstacle artifacts and reassigning pairs whose shortest paths are altered by obstacle avoidance [1811.07518]. The paper shows that standard PCFs can produce spurious oscillations and false long-range aggregation in obstacle-strewn domains, while the corrected PCF recovers the genuine short-range peak at \(m=1\) associated with birth clustering [1811.07518].

A related critique applies to the classical radial PCF more generally. Because it averages all pairs in metric shells, it can wash out subtle local rearrangements, rare defects, or local symmetries. The Voronoi PCF addresses this by replacing Euclidean shells with topological shells in the Delaunay graph: \(d_V(i,j)\) is the minimum number of Delaunay edges in a path from \(i\) to \(j\), \(u(k)\) is the mean number of \(k\)-neighbors, and
\[
v_S(k)=\frac{u_S(k)}{u_{\rm id}(k)}.
\]
This produces a discrete, topology-based PCF that is scale-invariant, robust under infinitesimal perturbations, and sensitive to local connectivity [2210.09731]. The same paper also notes a limitation: averaging over all particles still loses higher-order distributional information, so one may need the variance or full distribution of shell populations for finer discrimination [2210.09731].

## 5. Aperiodic order and analytic number theory

In regular model sets, PCFs are linked directly to covariograms of internal-space windows. For a cut-and-project set with windows \(W_i\), one has
\[
\nu_{ij}(y)=\frac{1}{\mathrm{Vol}(W)}\,(1_{-W_i}*1_{W_j})(y^\star)
=\frac{1}{\mathrm{Vol}(W)}\,C_{W_i,W_j}(y^\star),
\]
with \(\nu_{ij}(y)=0\) outside the corresponding difference set [2502.20487]. This identification converts pair-correlation questions into convolution questions in internal space.

When the model set comes from a primitive inflation, the cross-correlations satisfy exact renormalisation equations. In the paper’s formulation,
\[
\nu_{ij}(y)
=\frac{1}{\lambda}\sum_{k,\ell=1}^n\sum_{r\in T_{ik}}\sum_{s\in T_{j\ell}}
\nu_{k\ell}\!\left(\frac{y+r-s}{\lambda}\right),
\]
and the window covariograms satisfy the corresponding internal-space contraction equation [2502.20487]. In the sister silver mean example, recursion on a self-consistent finite set determines all pair-correlation values, yielding a continuous but extremely spiky graph. In the “splitting” Cantorval example, the numerical sample reveals two disjoint branches associated with an even or odd number of \(b\)-tiles in the digital decomposition, and convergence to the continuous curve is very slow [2502.20487]. The paper characterizes the resulting covariograms as unexpectedly complex and wild.

In analytic number theory, the PCF language refers to distributions of differences between zero ordinates. Banks introduces a two-parameter family of weights
\[
w_{\nu,k}(u)
=(2\nu)^{2k+1}\,\Re\!\left\{\frac{(2\nu-iu)^{2k+1}}{(u^2+4\nu^2)^{2k+1}}\right\},
\]
leading to weighted pair-correlation functions \(F_{\nu,k}(\alpha,T)\) [2502.20569]. Although these weights alter the asymptotic behavior of the weighted PCF, the paper states that they do not yield any new information about the simplicity of the zeros [2502.20569]. The same work extends Montgomery’s approach from differences of ordinates to sums \(\Sigma\gamma=\gamma_1+\cdots+\gamma_\mu\), defining \(G_\mu(\alpha,T)\) and formulating a generalized strong pair-correlation conjecture for any integer \(\mu\ge2\) [2502.20569]. This suggests that “pair correlation” in the number-theoretic literature is best understood as a spectral two-point statistic rather than a geometric radial function.

## 6. Computational representations, surrogate models, and domain-specific sensitivities

PCFs are increasingly treated as high-dimensional objects that can be learned, compressed, or stably parametrized. In glass modeling, the machine-learning pipeline of a CNN autoencoder plus random-forest regression compresses each \(64\times64\) grayscale PCF image into an \(8\)-dimensional latent vector and predicts that vector from a \(58\)-dimensional composition-and-pair fingerprint [2308.11151]. Direct regression on the full \(100\)-point \(g(r)\) curves yields \(R^2\approx0.4\), whereas encoding plus regression plus decoding attains \(R^2\approx0.92\) in latent-space prediction and reconstructed-curve coefficients of determination exceeding \(0.9\) on all nine unseen glass compositions [2308.11151]. A plausible implication is that low-dimensional latent structure captures the dominant peak-position and peak-height variability more efficiently than naive coordinatewise regression.

For fractional quantum Hall wavefunctions, stable representation rather than generic regression is the central issue. Girvin’s original basis is physically appealing but numerically unstable because its overlap matrix is ill-conditioned; the orthogonalized spherical basis built from modified Jacobi polynomials yields coefficients that can be obtained stably by projection, and these coefficients extrapolate linearly in \(1/N\) to the thermodynamic limit [2211.00371]. The paper further reports that, for all states considered, the large-\(n\) expansion coefficients are fit remarkably well by a cosine oscillation with exponentially decaying amplitude [2211.00371].

In femtoscopic heavy-ion applications, the most important practical sensitivity in CTA-I is not the thermal spectrum but the source geometry. The model shows that \(C(q)\) is highly sensitive to the Gaussian source size \(\sigma_R\) and only weakly dependent on the temperature parameter \(k_BT\). For Ar+Au\(\to\)B–B at \(35\ \mathrm{MeV/u}\), changing \(\sigma_R\) from \(6\) to \(8\ \mathrm{fm}\) visibly modifies the Coulomb anti-correlation dip around \(q\approx30\)–\(80\ \mathrm{MeV}/c\), whereas varying \(k_BT\) from \(10\) to \(20\ \mathrm{MeV}\) leaves the correlation nearly unchanged; similar behavior appears in \(^{86}\mathrm{Kr}+\mathrm{Pb}\to t\)–\(t\) and \(^3\mathrm{He}\)–\(^3\mathrm{He}\) at \(25\ \mathrm{MeV/u}\) [2510.23030]. The proposed workflow is therefore to determine \(k_BT\) from single-particle spectra and use the PCF to scan \(\sigma_R\) and extract the spatio-temporal size of the emission region [2510.23030].

A recurrent methodological point is that a PCF is powerful precisely because it is low-order and broadly measurable, but that same averaging can be a limitation. This is explicit in the contrast between shell-averaged radial PCFs and Voronoi-topological PCFs [2210.09731], in the need for obstacle corrections on lattices [1811.07518], and in the need for orthogonal bases or renormalisation equations when direct numerical representations become unstable or fractal-like [2211.00371] [2502.20487]. The cumulative picture is that PCFs remain foundational descriptors of two-point structure, yet their most informative use depends on choosing the normalization, geometry, closure, and representation that match the underlying domain.

Source: https://www.emergentmind.com/topics/pair-correlation-functions-pcfs