---
title: Semi-Classical Encoded Subset States
url: https://www.emergentmind.com/topics/semi-classical-encoded-subset-states
type: topic
---

# Semi-Classical Encoded Subset States

Searching arXiv for recent and foundational papers on subset states, pseudorandomness, and related semi-classical encodings.
Semi-classical encoded subset states are quantum states whose full amplitude vector is induced by classical combinatorial support data. In the standard form, one fixes a subset \(S\subseteq\{0,1\}^n\) and defines the corresponding \(n\)-qubit state by
\[
|S\rangle=\frac{1}{\sqrt{|S|}}\sum_{x\in S}|x\rangle.
\]
All nonzero amplitudes have the same magnitude, all phases are identical, and the only free datum is the support set \(S\). This support-only structure makes subset states a particularly sharp model of semi-classical encoding: they are genuinely coherent quantum superpositions, yet their specification is far more rigid than that of generic pure states. Recent work shows that such phase-free states can nevertheless exhibit Haar-like pseudorandomness and pseudoentanglement in bounded-copy settings [2312.09206], [2312.15285], while earlier work established that subset states already suffice as witnesses for the full class \(QMA\) [1410.2882]. A separate simulator-oriented line treats sparse support states as classical data structures storing only occupied basis states and amplitudes [2204.11042].

## 1. Definition and semi-classical structure

For \(d=2^n\), a subset state is a uniform superposition over a subset \(S\subseteq[d]\), equivalently \(S\subseteq\{0,1\}^n\):
\[
|S\rangle=\frac{1}{\sqrt{|S|}}\sum_{i\in S}|i\rangle.
\]
This definition appears in the cryptographic, complexity-theoretic, and simulation-oriented treatments of subset states [2312.09206], [2312.15285], [1410.2882].

The semi-classical character is precise. The amplitudes are supported only on \(S\); all nonzero amplitudes have the same magnitude \(1/\sqrt{|S|}\); and all phases are equal. There is therefore no hidden phase function, no sign pattern, and no arbitrary complex amplitude profile. The classical datum \(S\) determines the support of the wavefunction and, in the uniform case, the complete state. This sharply distinguishes subset states from generic Haar-random states, whose amplitudes are unconstrained complex numbers.

This does not make subset states classical in the QCMA sense. They remain quantum states, not classical samples from \(S\), and the support \(S\) may be exponentially large or lack an efficient classical description. The relevant distinction is that the quantum coherence is concentrated in a highly structured corner of Hilbert space. In that sense, semi-classical encoded subset states are support-defined coherent states rather than arbitrary amplitude-defined states.

A broader sparse-state variant appears in hybrid simulation work, where one stores
\[
\ket{\psi}=\sum_{x\in S}\alpha_x\ket{x},
\qquad
S=\{x:\alpha_x\neq 0\},
\]
as a classical list or database of occupied basis strings and amplitudes [2204.11042]. Uniform subset states are the special case in which the nonzero amplitudes are all equal.

## 2. Haar-like pseudorandomness from random support

Two 2023 papers established that random subset states can be information-theoretically close to Haar-random states when only polynomially many copies are available, resolving the question of whether random or pseudorandom phases are necessary [2312.09206], [2312.15285].

In one formulation, for \(N=2^n\), \(t=O(\mathrm{poly}(n))\), and subset size \(m\) satisfying
\[
\omega(\mathrm{poly}(n)) < m < o(2^n),
\]
the \(t\)-copy moment operator of the random-subset ensemble obeys
\[
TD\!\left(
\mathbb{E}_{S \in \binom{[N]}{m}} |S\rangle\langle S|^{\otimes t},
\;
\mathbb{E}_{|\phi\rangle\sim \mathrm{Haar}([N])} |\phi\rangle\langle \phi|^{\otimes t}
\right)
\le O\!\left(\frac{tm}{N}\right)+O\!\left(\frac{t^2}{m}\right),
\]
where \(TD(A,B)=\frac12\|A-B\|_1\) [2312.09206]. This is an approximate \(t\)-design statement in \(t\)-copy trace distance. Negligible error requires
\[
\frac{tm}{N}\to 0
\qquad\text{and}\qquad
\frac{t^2}{m}\to 0,
\]
so \(m\) must be neither too small nor too large.

A concurrent independent result studies
\[
\Psi_k := \int \psi^{\otimes k}\, d\mu(\psi),
\qquad
\Phi_{s,k} := \mathbb{E}_{S\subseteq[d],\, |S|=s}\,\phi_S^{\otimes k},
\]
and proves
\[
\left\|
\int \psi^{\otimes k}\, d\mu(\psi)
-
\mathbb{E}_{S\subseteq[d], |S|=s}\phi_S^{\otimes k}
\right\|_1
\le
O\!\left(\frac{k^2}{d}+\frac{k}{\sqrt{s}}+\frac{sk}{d}\right)
\]
[2312.15285]. For \(d=2^n\), \(k=\mathrm{poly}(n)\), and
\[
s=\omega(\mathrm{poly}(n))
\qquad\text{and}\qquad
s=\frac{2^n}{\omega(\mathrm{poly}(n))},
\]
all three terms are negligible, yielding information-theoretic indistinguishability even given polynomially many copies.

Both analyses identify the same operational obstruction regime. If the subset is too small, repeated computational-basis measurements expose collisions. In the \(m\)-parameterization, this gives the \(O(t^2/m)\) term; in the \(s\)-parameterization, if \(s=p(n)\), measuring \(p(n)+1\) copies yields repeated outcomes with probability \(1\) by the pigeonhole principle [2312.09206], [2312.15285]. If the subset is too large, the state has noticeable overlap with the uniform superposition. For example,
\[
|\langle S|+^n\rangle|^2=\frac{m}{N},
\]
and when \(s=2^n/p(n)\),
\[
\langle u|S\rangle = \sqrt{\frac{s}{2^n}}=\frac{1}{\sqrt{p(n)}},
\]
which yields an efficient swap-test distinguisher [2312.09206], [2312.15285].

The conceptual consequence is that support randomness alone can mimic Haar low-order statistics. Random phases are not necessary in the broad intermediate regime.

## 3. Representation-theoretic mechanism

The most detailed proof proceeds by analyzing the random-subset moment operator on the symmetric subspace via the representation theory of the symmetric group [2312.09206]. Let
\[
\rho := \mathbb{E}_{S\sim \binom{[N]}{m}} |S\rangle\langle S|^{\otimes t},
\qquad
\rho_0 := \mathbb{E}_{|\phi\rangle\sim \mathrm{Haar}([N])} |\phi\rangle\langle\phi|^{\otimes t}.
\]
The central basis is the type basis of \(\mathrm{Sym}^t([N])\), indexed by multisets \(\theta\in \mathrm{MSet}([N],t)\). Among these, the unique types correspond to ordinary \(t\)-subsets \(\alpha\in\binom{[N]}{t}\), spanning
\[
\mathrm{unique}=\mathrm{span}\{|\alpha\rangle:\alpha\in\tbinom{[N]}{t}\}.
\]
Its dimension is
\[
\dim(\mathrm{unique})=\binom{N}{t},
\]
while
\[
\dim(\mathrm{Sym}^t([N]))=\binom{N+t-1}{t},
\]
and
\[
\binom{N}{t}
=
\binom{N+t-1}{t}\left(1+O\!\left(\frac{t^2}{N}\right)\right).
\]
Thus for \(t\ll N\), almost all of the symmetric subspace is already captured by unique types.

Inside this basis, the matrix elements of \(\rho\) depend only on the overlap structure of \(\alpha\) and \(\beta\). For \(\alpha,\beta\in\binom{[N]}{t}\),
\[
\langle \alpha|\rho|\beta\rangle
=
\frac{t!}{m^t}
\frac{\binom{N-|\alpha\cup\beta|}{\,m-|\alpha\cup\beta|\,}}{\binom{N}{m}},
\]
and after approximation this becomes a function of \(|\alpha\cup\beta|\), equivalently of Johnson-graph distance.

The group action is induced by basis permutations. Writing
\[
X = S_N / (S_t\times S_{N-t}),
\qquad
(G,K)=\bigl(S_N,\;S_t\times S_{N-t}\bigr),
\]
the pair \((G,K)\) is a Gelfand pair. The double-cosets \(K\backslash G/K\) are indexed by subset distance \(p\in\{0,\dots,t\}\), which is precisely the Johnson scheme. Because \(\rho\) is invariant under basis permutations, its restriction to the unique subspace is group-circulant, with circulant function
\[
\nu(p)=\left(\frac{m}{N}\right)^p\left(1+O\!\left(\frac{t^2}{m}\right)\right).
\]

The permutation representation decomposes multiplicity-freely as
\[
L(X)\simeq \bigoplus_{q=0}^t V_{[N-q,q]},
\]
with dimensions
\[
d_{[N-q,q]}=
\begin{cases}
\binom{N}{q}-\binom{N}{q-1}, & 1\le q\le t,\\[4pt]
1, & q=0.
\end{cases}
\]
The dominant block is \(V_{[N-t,t]}\), whose dimension satisfies
\[
\dim V_{[N-t,t]}
=
\binom{N}{t}-\binom{N}{t-1}
=
\binom{N+t-1}{t}\left(1+O\!\left(\frac{t^2}{N}\right)\right).
\]
Its eigenvalue is
\[
\mu_{[N-t,t]}
=
\left(1-\frac{m}{N}\right)^t
\left(1+O\!\left(\frac{t^2}{m}\right)\right)
=
1+O\!\left(\frac{tm}{N}\right)+O\!\left(\frac{t^2}{m}\right).
\]

This shows that on the overwhelmingly dominant irreducible block, the rescaled subset-state moment operator is asymptotically the identity, hence the original moment operator is asymptotically the Haar moment. A plausible implication is that the relevant pseudorandomness is driven not by amplitude complexity but by the combinatorics of random support intersections, organized by the Johnson scheme.

## 4. Computational pseudorandom states and pseudoentanglement

The information-theoretic results become pseudorandom-state constructions by replacing uniformly random subsets with pseudorandomly generated subsets [2312.15285]. Using a quantum-secure pseudorandom permutation family \(\{\mathrm{PRP}_k\}\) on \([2^n]\), one defines
\[
|\phi_k\rangle
=
\frac{1}{\sqrt{s}}\sum_{x\in[s]} |\mathrm{PRP}_k(x)\rangle,
\]
equivalently the subset state corresponding to
\[
S_k = \{\mathrm{PRP}_k(1),\dots,\mathrm{PRP}_k(s)\}.
\]
If the PRP is replaced by a truly random permutation, the image of \([s]\) is a uniformly random \(s\)-subset of \([2^n]\), so the resulting state is exactly a random subset state [2312.15285].

Efficient preparation is obtained by preparing \(\frac{1}{\sqrt{s}}\sum_{x\in[s]}|x\rangle\), coherently computing \(\mathrm{PRP}_k(x)\), and then uncomputing the first register. The construction depends on the existence of an efficiently computable quantum-secure PRP; the paper notes that quantum-secure PRPs exist assuming quantum-secure one-way functions [2312.15285]. In a closely related formulation, the same minimalist preparation strategy uses only a PRP and no PRF for sign generation, unlike earlier subset-phase constructions [2312.09206].

The pseudoentanglement consequence is immediate. Since \(|S\rangle\) is a sum of only \(|S|\) computational-basis product states, its Schmidt rank across any cut is at most \(|S|\), and therefore the entanglement entropy across any cut is at most
\[
O(\log |S|).
\]
Yet in the admissible regime the ensemble is computationally or information-theoretically indistinguishable from Haar-random states given polynomially many copies [2312.15285]. For
\[
h(n)=\omega(\log n)
\qquad\text{and}\qquad
h(n)=n-\omega(\log n),
\]
choosing \(s=2^{h(n)}\) yields PRS families with entanglement entropy \(O(h(n))\) across all cuts [2312.15285].

One paper also gives a biased-phase interpolation. For subset-phase states with i.i.d. signs of bias parameter \(b\in[-1,1]\),
\[
TD\!\left(
\mathbb{E}_{\substack{S\sim \binom{[N]}{m}\\ f\sim \mathrm{Ber}\left[\frac{1+b}{2}\right]^{[N]}}}
|S,f\rangle\langle S,f|^{\otimes t},
\;
\mathbb{E}_{|\phi\rangle\sim \mathrm{Haar}} |\phi\rangle\langle \phi|^{\otimes t}
\right)
\le
O\!\left(\frac{tmb^2}{N}\right)+O\!\left(\frac{t^2}{m}\right).
\]
The subset-state case is \(b=1\), while unbiased random phases correspond to \(b=0\). This suggests that phases mainly help in the very-large-subset regime by suppressing the overlap with the global uniform superposition [2312.09206].

## 5. Sparse classical encodings in hybrid simulation

A simulator-oriented treatment uses sparse computational-basis support as an explicit classical encoding of quantum states [2204.11042]. The proposed representations are: Array, Database, Qiskit, and Mixed. In the sparse representations, the simulator “encode[s] the entire state as a standard Numpy array containing all classical states occurring within the superposition and their respective amplitudes,” or saves “all classical states with non-zero amplitude within the quantum state” in an `sqlite3` database [2204.11042].

The core encoded form is
\[
\ket{\psi}=\sum_{x\in S}\alpha_x\ket{x},
\]
where only the occupied basis strings and their amplitudes are stored. The implied memory model is full state-vector simulation with \(O(2^n)\) amplitudes versus sparse support encoding with \(O(|S|)\) records. This is useful when \(|S|\ll 2^n\).

The canonical uniform-subset benchmark is
\[
\ket{\psi}
=
\left(H^{\otimes r}\otimes I^{\otimes (n-r)}\right)\ket{0^n}
=
\frac{1}{\sqrt{2^r}}\sum_{x\in S_r}\ket{x},
\]
with \(|S_r|=2^r\) [2204.11042]. The reported empirical pattern is that database runtime depends mainly on \(r\), the number of “Nondet qubits,” whereas dense state-vector simulation depends on the total number of qubits \(n\). This is the most direct simulator manifestation of semi-classical encoded subset states: complexity scales with support size rather than ambient Hilbert-space dimension.

The same framework is applied to an addition benchmark with total qubit count
\[
3k+5,
\]
and to Grover’s algorithm, where the diffusion operator is implemented “within a single SQL query to the database” rather than by a standard gate sequence [2204.11042]. A mixed heuristic chooses the database encoding if fewer than \(\frac{2}{3}\) of the qubits have a Hadamard somewhere in the circuit, and otherwise uses Qiskit.

The limitations are explicit. The method is useful only when circuits remain sparse or near-classical in the computational basis. Support growth under many Hadamards destroys the advantage, and the approximate “state drop” heuristic—after each quantum operation, keep only the 1000 largest-amplitude support entries—causes errors that “deteriorate quickly” [2204.11042]. This suggests that simulator-side semi-classical encodings are operationally valuable, but only in regimes where support sparsity is preserved.

## 6. Subset-state witnesses and complexity-theoretic significance

Subset states also play a structural role in quantum complexity theory. The class \(SQMA\) restricts the completeness witness of a \(QMA\) protocol to a subset state, while soundness remains against arbitrary quantum witnesses [1410.2882]. The main theorem is
\[
SQMA = QMA,
\]
with the two-prover analogue
\[
SQMA(2)=QMA(2)
\]
[1410.2882].

The technical core is the Subset State Approximation Lemma: for any \(n\)-qubit state \(|\psi\rangle\), there exists a subset \(S\subseteq [2^n]\) such that
\[
|\langle S|\psi\rangle| \ge \frac{1}{8\sqrt{n+3}}.
\]
A geometric lemma establishes the corresponding statement for arbitrary \(v\in\mathbb{C}^d\). The proof partitions amplitudes into dyadic level sets and shows that one level set carries enough \(\ell_2\)-mass to yield non-negligible overlap with a uniform vector on its support [1410.2882].

This approximation is sufficient because amplified \(QMA\) verifiers have exponentially small error. Replacing an optimal witness by the approximating subset state yields inverse-polynomial acceptance probability, and standard amplification restores the usual constant gap. The result is structural rather than algorithmic: it does not imply efficient preparation, succinct classical descriptions, or a collapse to \(QCMA\) [1410.2882].

The paper also defines \(oSQMA\), where a yes-instance admits a subset state that is optimal among all witnesses, and proves
\[
SQMA_1 = oSQMA_1 = oSQMA,
\qquad
oSQMA \subseteq QMA_1 \subseteq QMA.
\]
It introduces a new \(QMA\)-complete problem, \(BSCSS(\alpha)\), whose promise is phrased directly in terms of subset states [1410.2882].

A common misconception is that subset-state witnesses are merely classical witnesses in disguise. The formal results rule that out. What is reduced is amplitude complexity and phase freedom, not the presence of coherent superposition itself. The broader lesson across complexity theory, cryptography, and simulation is that support information alone can already sustain substantial quantum power. In bounded-copy pseudorandomness it can mimic Haar low-order statistics [2312.09206], [2312.15285]; in \(QMA\) it already captures the full witness power of the class [1410.2882]; and in hybrid simulation it provides a practical sparse-state representation whenever computational-basis support remains small [2204.11042].

Source: https://www.emergentmind.com/topics/semi-classical-encoded-subset-states