---
title: Average Minimal Hamming Distance in Codes & Quantum Data
url: https://www.emergentmind.com/topics/average-minimal-hamming-distance
type: topic
---

# Average Minimal Hamming Distance in Codes & Quantum Data

Searching arXiv for the cited works to ground the article in the relevant literature.
Average minimal Hamming distance denotes a family of closely related functionals on Hamming spaces. In binary coding theory, the principal quantity is the **minimum average Hamming distance** of an \((n,M)\) code \(C\subset F_2^n\), defined by minimizing the average pairwise distance
\[
a(C)=\frac{1}{M^2}\sum_{c,c'\in C} d(c,c')
\]
over all codes of size \(M\); the resulting extremal value is denoted \(B(n,M)\) [0706.3295]. In quantum-sampling analysis, the term refers instead to the average, over a set of distinct measured bit-strings, of each sample’s nearest-neighbor Hamming distance,
\[
\langle d_{\min}\rangle(N_b)=\frac{1}{N_b}\sum_{i=1}^{N_b} d_{\min}(x^{(i)}),
\qquad
d_{\min}(x^{(i)})=\min_{j\neq i} d(x^{(i)},x^{(j)}),
\]
which is used as a sample-level descriptor of bit-string complexity [2606.04558]. Related literature also studies the expectation or ensemble average of the **minimum distance** itself, rather than an average of pairwise or nearest-neighbor distances [1912.12833], [0905.4545].

## 1. Definitions and distinctions

For binary codes, \(F_2^n\) is the set of binary words of length \(n\), and an \((n,M)\) code is a subset \(C\subset F_2^n\) with \(|C|=M\). The average Hamming distance of \(C\) is
\[
a(C)=\frac{1}{M^2}\sum_{c,c'\in C} d(c,c'),
\]
and the minimum average distance over all \((n,M)\) codes is
\[
B(n,M)=\min_{C\subset F_2^n,\ |C|=M} a(C).
\]
This is an isodiametric extremal problem in Hamming space: the optimization variable is the entire code, and the objective averages over all ordered pairs in the code [0706.3295].

A distinct quantity appears in the analysis of quantum measurement data. There one begins with a collection \(\{x^{(i)}\}_{i=1}^{N_b}\) of \(N_b\) distinct bit-strings of length \(N\), defines the Hamming distance
\[
d(x^{(i)},x^{(j)})=\sum_{k=1}^N |x_k^{(i)}-x_k^{(j)}|,
\]
then assigns to each sample its nearest-neighbor distance \(d_{\min}(x^{(i)})\), and finally averages these nearest-neighbor distances to obtain \(\langle d_{\min}\rangle(N_b)\) [2606.04558]. Here the functional is not an extremum over all codes of fixed size; it is a statistic of an observed sample set.

A common source of confusion is the distinction between **minimum average distance** and **minimum distance**. In the coding-theoretic notation of [0706.3295] and [1910.09416], the optimized object is an average over all pairs. By contrast, the random linear-code literature studies
\[
d_{\min}(C)=\min\{ \mathrm{wt}(c)\;|\;c\in C\setminus\{0\}\},
\]
and the Hamming-accumulate-accumulate literature studies probabilistic or ensemble-average behavior of that minimum distance [1912.12833], [0905.4545]. These quantities are related only indirectly.

## 2. Delsarte linear programming for minimum average distance

The basic linear-programming framework for \(B(n,M)\) is Delsarte’s approach via the dual distance distribution. If \((A_0,\ldots,A_n)\) is the usual distance distribution of a binary code \(C\), and \(P_k(x)\) is the binary Krawtchouk polynomial of degree \(k\), then the dual distance distribution is
\[
B_k=\frac{1}{M}\sum_{i=0}^n P_k(i)\,A_i,\qquad k=0,\ldots,n.
\]
The key facts are
\[
B_0=1,\qquad B_k\ge 0\ \text{for }k=1,\ldots,n,\qquad \sum_{k=0}^n B_k=\frac{2^n}{M},
\]
together with the identity
\[
a(C)=n-B_1.
\]
Consequently, a lower bound on \(B(n,M)\) follows from an upper bound on \(B_1\) [0706.3295].

This yields the linear program
\[
\text{maximize } B_1
\]
subject to
\[
\sum_{k=1}^n B_k=\frac{2^n}{M}-1,
\]
\[
\sum_{k=1}^n P_j(k)\,B_k\ge -P_j(0),\qquad j=1,2,\ldots,n,
\]
and
\[
B_k\ge 0,\qquad k=1,2,\ldots,n.
\]
The proof method uses dual feasible polynomials \(X(x)=\sum_{j=0}^n \alpha_j P_j(x)\) satisfying sign constraints such as \(X(1)=-1\), \(X(i)\le 0\) on forbidden \(i\), and \(\alpha_j\ge 0\) on allowed \(j\). By standard LP duality this yields an upper bound on \(B_1\), hence a lower bound on \(B(n,M)\) [0706.3295].

A later reformulation works in the binary Hamming space \(\{-1,+1\}^n\). If \(A\subset\{-1,1\}^n\) has size \(|A|=M\) and distance distribution
\[
P_A(i)=\Pr_{x,x'\in A}[d_H(x,x')=i],
\]
then its average Hamming distance is
\[
D(A)=\sum_{i=0}^n i\,P_A(i).
\]
With \(a=M/2^n\le \tfrac12\), one defines
\[
Q_A(k)=\sum_{i=0}^n P_A(i)\,K_k^{(n)}(i),
\]
where \(K_k^{(n)}\) is the \(k\)th Krawtchouk polynomial, and obtains
\[
\sum_{k=0}^n Q_A(k)=\frac{2^n}{M}=\frac1a,
\qquad
D(A)-\frac n2=\frac12\Bigl[1-\frac1a+\sum_{i=2}^n Q_A(i)\Bigr].
\]
This gives a primal–dual LP pair whose feasible dual vectors directly generate lower bounds on \(D(A)\) [1910.09416].

## 3. Explicit lower bounds and sparse-code regimes

The linear-programming method yields several closed-form lower bounds on \(B(n,M)\) that are valid for every integer pair \((n,M)\) with \(2\le M\le 2^n-1\). For very small codes,
\[
B(n,M)\ge 1-\frac1M.
\]
For small codes,
\[
B(n,M)\ge \frac{n-2}{M}.
\]
More significantly, for codes of moderate size with \(M\approx n\), the paper gives formulas that are nontrivial as soon as \(M\) is of order \(n\) [0706.3295].

If \(n\) is even, then
\[
B(n,M)\ge \frac{3n}{n+2}-\frac{3n+2}{(n+2)M}.
\]
If \(n\) is odd, then
\[
B(n,M)\ge \frac{3(n+1)}{n+3}-\frac{3n+3}{(n+3)M}.
\]
Further refinements depend on the residue class of \(n \bmod 4\):
\[
B(n,M)\ge
\begin{cases}
\displaystyle\frac{7n+2}{2(n+2)}-\frac{2n}{(n+2)M}, & n\equiv 0\pmod 4,\\[1ex]
\displaystyle\frac{7n-5}{2(n+1)}-\frac{2n}{(n+1)M}, & n\equiv 1\pmod 4,\\[1ex]
\displaystyle\frac{2(n+2)}{n+3}-\frac{2n}{(n+3)M}, & n\equiv 2\pmod 4,\\[1ex]
\displaystyle\frac{7n+9}{2(n+3)}-\frac{2(n+1)}{(n+3)M}, & n\equiv 3\pmod 4.
\end{cases}
\]

These formulas are important because the earlier “classical” bound of Althöfer–Sillke,
\[
B(n,M)\ge \frac{n+1}{2}-\frac{2^n}{M},
\]
is only useful when \(M\gg 2^n/n\). For codes of size \(M=O(n)\) that bound is vacuous, whereas the new bounds give a constant-order lower bound [0706.3295].

The numerical example \(n=10\), \(M=20\) illustrates the difference. The classical bound gives
\[
B(10,20)\ge 5.5-\frac{1024}{20}=-45.7,
\]
whereas the even-\(n\) bound gives
\[
B(10,20)\ge \frac{3\cdot 10}{12}-\frac{32}{12\cdot 20}=2.3667.
\]
The paper further proves the asymptotic statement
\[
\lim_{n\to\infty} B(n,2n)=\frac52.
\]
This establishes the exact constant \(2.5\) in the limit for the sparse code size \(M=2n\) [0706.3295].

## 4. Improved LP bounds, dense-code asymptotics, and Fourier weights

For code size \(M=\lceil a2^n\rceil\) with fixed \(0<a\le \tfrac12\), the improved LP analysis of Yu and Tan constructs a two-point dual feasible solution supported at \(k\) and \(k+1\), with \(k=2\lfloor \beta n/2\rfloor\) and \(\beta\in(\tfrac12,1)\), rather than the single-point solution used earlier [1910.09416]. Evaluating the dual objective and optimizing over \(\beta\) yields the asymptotic bound
\[
\bar d_{\min}(n,M)=\min_{|A|=M} D(A)\ge \frac n2-\varphi(a),
\]
where
\[
\varphi(a)=
\begin{cases}
\displaystyle \frac1{\sqrt a}-1, & 0<a<\tfrac14,\\[4pt]
\displaystyle \frac1{4a}, & \tfrac14\le a\le \tfrac12.
\end{cases}
\]

The same paper shows that this is asymptotically optimal for the Fu–Wei–Yeung LP: for large \(n\), no nonnegative dual feasible vector can asymptotically do better than the function \(\theta(a)\) obtained by the two-point construction. In that sense, all possible asymptotic bounds that can be derived by that linear program have been characterized [1910.09416].

A further aspect of the theory is its Fourier-analytic reformulation. If \(f:\{-1,1\}^n\to\{-1,1\}\) and \(A=f^{-1}(+1)\) has size \(M=a2^n\), then the degree-\(m\) Fourier weight is
\[
W_m(f)=\sum_{|S|=m}\hat f_S^2.
\]
For \(m=1\),
\[
W_1(f)=4a^2\bigl(n-2D(A)\bigr),
\]
so the average-distance bound is equivalent to
\[
W_1(f)\le 8a^2\varphi(a).
\]
The paper also gives higher-degree bounds:
\[
W_m(f)\le 4a(1-a)\quad \text{for even }m\ge 2,
\qquad
W_m(f)\le 2a\quad \text{for odd }m\ge 3.
\]
This places the coding-theoretic LP inside a broader MacWilliams–Delsarte and Fourier-analytic framework [1910.09416].

## 5. Expectation and ensemble averages of minimum distance

Closely related literature studies the expectation or ensemble average of the **minimum** Hamming distance, rather than the minimum over codes of an average pairwise distance. For a random linear code of dimension \(k\) in \(\mathbb{F}_q^n\), with
\[
M=\frac{q^k-1}{q-1},
\qquad
\rho_d=P\{\mathrm{Binomial}(n,1/q)\le d\},
\]
the paper on random linear codes shows that
\[
E[d_{\min}]
=
\sum_{d=0}^{n-1} \bigl(1-\rho_d\bigr)^M + O(e^{-c\sqrt n}),
\]
and that the c.d.f. of \(d_{\min}\) is \(O(e^{-c\sqrt n})\)-close to the c.d.f. of the minimum of \(M\) independent Binomial\((n,1/q)\) variables [1912.12833]. When \(k/n\to R\in(0,1)\), the minimum exhibits a Gumbel-limit description on an arithmetic lattice, and the expectation has the asymptotic form
\[
E[d_{\min}]
=
n\delta_0-\frac{\log n+O(1)}{\log\bigl((1/\delta_0-1)(q-1)\bigr)}+o(1),
\qquad
R=1-H_q(\delta_0).
\]
The leading-order term is therefore \(n\delta_0\) [1912.12833].

For Hamming-accumulate-accumulate ensembles, the relevant object is the ensemble-average weight enumerator and the induced probabilistic lower bound on \(d_{\min}\). If \(C=C_2\circ \pi_2\circ C_1\circ \pi_1\circ C_0\) is formed by serial concatenation of an \((n,k)\) Hamming or extended-Hamming code with two rate-1 accumulate codes, then the uniform-interleaver ensemble-average weight enumerator is
\[
\bar A_h^C
=
\sum_{h_0=0}^N \sum_{h_1=0}^N
\frac{A_{h_0}^{C_0}\,A_{h_0,h_1}^{C_1}\,A_{h_1,h}^{C_2}}
{\binom{N}{h_0}\binom{N}{h_1}}.
\]
A union bound gives
\[
\Pr\{d_{\min}<d\}\le \sum_{h=1}^{d-1} \bar A_h^C.
\]
Asymptotically, if the spectral-shape function
\[
r(\delta)=\limsup_{N\to\infty}\frac1N\ln \bar A_{\lfloor \delta N\rfloor}^C
\]
satisfies \(r(\delta)<0\) for all \(\delta<\delta_{\min}\), then with high probability \(d_{\min}(N)\gtrsim \delta_{\min}N\). Numerically, the paper reports \(\delta_{\min}=0.0197\) for \((32,26)\)AA, \(0.0140\) for \((31,26)\)AA, \(0.0091\) for \((64,57)\)AA, \(0.0067\) for \((63,57)\)AA, \(0.0042\) for \((128,120)\)AA, and \(0.0032\) for \((127,120)\)AA [0905.4545].

These results do not define “average minimal Hamming distance” in the same way as \(B(n,M)\) or \(\langle d_{\min}\rangle(N_b)\). They nevertheless show that neighboring notions of averaging—expectation over a random-code ensemble or over a concatenated-code ensemble—play an important role in minimum-distance analysis.

## 6. Sample-based average minimal Hamming distance in quantum data

In quantum sampling data, the average minimal Hamming distance is computed from a dataset of \(M\) unique bit-strings in the measurement basis, with \(M=N_b^{\max}\). For a target sample size \(N_b\le M\), one chooses at random, or by stratified sampling, a subset \(S\) of size \(N_b\), computes the Hamming distances between every pair of elements of \(S\), records each sample’s nearest-neighbor distance, and averages. Repeating this procedure typically \(5\)–\(10\) times reduces sampling noise [2606.04558].

The naive implementation requires \(O(N_b^2\times N)\) bit-operations. The paper notes two concrete optimizations: use bitwise XOR and population-count hardware instructions to compute Hamming distances extremely quickly, and for large \(N_b\) use approximate nearest-neighbor methods in Hamming space, such as locality-sensitive hashing or trie-based indexes, to reduce the \(O(N_b^2)\) scaling [2606.04558].

Empirically, over several decades in \(N_b\), a wide variety of quantum states obey
\[
\langle d_{\min}\rangle(N_b)\approx a\,N_b^{-\alpha}+c.
\]
The prefactor satisfies \(a\simeq \langle d_{\min}\rangle(1)\), the scaling exponent lies in \([0.1,0.35]\), and the offset obeys \(c\ge 0\). In practice, \(c\to 0\) for large Hilbert spaces, although occasionally \(c\approx 2\) when the minimal distance saturates at the smallest possible value [2606.04558].

The reported examples are specific. For a 24-qubit Dicke state with \(D=6\) excitations, fitting over \(N_b\in[10,500]\) gives \(\langle d_{\min}\rangle\approx (8.59)N_b^{-0.15}\), with \(c\approx 0\), and the fit remains good until saturation at \(\langle d_{\min}\rangle=2\) for \(N_b\gtrsim 5000\). For the exact ground state of the \(J_1\)–\(J_2\) model on a \(6\times 6\) lattice at \(J_2=0\), the reported values are \(\alpha_{\mathrm{ED}}(J_2=0)\approx 0.14\) and \(A_{\mathrm{ED}}(J_2=0)\approx 7.8\), with \(c\approx 0\). For Haar-random 53-qubit Sycamore data, a typical exponent is \(\alpha\approx 0.30\) with prefactor \(a\approx 11\). For a D-Wave \(18\times 18\) spin glass at quench time \(t_a=7\,\mathrm{ns}\), fitting over \(N_b\in[700,900]\) yields \(\alpha\approx 0.14\), \(a\approx 9.5\), and negligible \(c\); at \(t_a=20\,\mathrm{ns}\), one finds \(\alpha\approx 0.19\) and \(a\approx 9.5\) [2606.04558].

The interpretation proposed in that work is operational. The prefactor \(a\) reflects the “typical” nearest-neighbor distance when only one sample is drawn: larger \(a\) means the wave function is more delocalized in bit-string space. The exponent \(\alpha\) measures how “connected” the manifold of high-probability bit-strings is: smaller \(\alpha\) implies that even large numbers of samples rarely find very close neighbors. In the frustrated \(J_1\)–\(J_2\) model, \(\alpha(J_2)\) exhibits a pronounced maximum in the intermediate regime \(J_2\approx 0.5\)–\(0.6\), coinciding with the quantum phase transition between Néel and stripe orders. This suggests that changes in \(\alpha\) and \(a\) can signal phase boundaries without requiring computation of order parameters [2606.04558].

The same paper also states the main practical limitations. To extract \(\alpha\) reliably, one needs at least \(N_b\approx 10^2\)–\(10^3\) unique bit-strings. At small \(N_b\lesssim 10\), \(\langle d_{\min}\rangle\) has large sample-to-sample variance, so averaging over multiple random subsets is necessary. When the minimal Hamming distance saturates at the theoretical lower bound, such as \(2\) for balanced Dicke states, one must include \(c\) in the fit or restrict to \(N_b\) below saturation. The parameters \(\alpha\) and \(a\) are empirical descriptors and may depend weakly on measurement basis, so comparisons require identical bases and fitting windows [2606.04558].

## 7. Conceptual placement

Taken together, these works place average minimal Hamming distance at the intersection of extremal coding theory, probabilistic minimum-distance analysis, and data-driven characterization of high-dimensional discrete distributions. In coding theory, the central problem is to determine or bound \(B(n,M)\), and the Delsarte LP with dual-feasible Krawtchouk polynomials yields explicit nontrivial lower bounds in the regime \(M=\Theta(n)\), including the asymptotic identity \(B(n,2n)\to 5/2\) [0706.3295]. In the dense regime \(M=\lceil a2^n\rceil\), the improved LP of Yu and Tan characterizes the asymptotically optimal output of that linear program and connects the problem to Fourier weights of Boolean functions [1910.09416].

In adjacent settings, “average” may refer instead to averaging over a random-code ensemble or over repeated subset selections from measured bit-strings. The random linear-code results show that linear dependencies have only an \(O(e^{-c\sqrt n})\) effect on the c.d.f. of the minimum distance beyond the identification of proportional codewords [1912.12833]. The quantum-sampling results show that averaging nearest-neighbor distances across unique bit-strings can produce a robust power-law descriptor \((a,\alpha,c)\) across experimentally and numerically generated states [2606.04558].

A plausible implication is that the phrase “average minimal Hamming distance” should always be read with its precise ambient definition: an extremal average over all code pairs, an expectation of a minimum over a random ensemble, or a nearest-neighbor average over an observed sample set. The cited literature treats all three with technically different objectives, constraints, and asymptotic regimes.

Source: https://www.emergentmind.com/topics/average-minimal-hamming-distance