Semi-Classical Encoded Subset States
- Semi-classical encoded subset states are uniform quantum superpositions defined solely by their classical support sets, displaying structured amplitude and phase behavior.
- Recent work demonstrates that such states mimic Haar-random low-order moments, achieving pseudorandomness and pseudoentanglement under bounded-copy measurements.
- They also enable efficient simulation via sparse classical encoding and serve as witnesses in QMA protocols, highlighting their complexity-theoretic significance.
Searching arXiv for recent and foundational papers on subset states, pseudorandomness, and related semi-classical encodings. Semi-classical encoded subset states are quantum states whose full amplitude vector is induced by classical combinatorial support data. In the standard form, one fixes a subset and defines the corresponding -qubit state by
All nonzero amplitudes have the same magnitude, all phases are identical, and the only free datum is the support set . This support-only structure makes subset states a particularly sharp model of semi-classical encoding: they are genuinely coherent quantum superpositions, yet their specification is far more rigid than that of generic pure states. Recent work shows that such phase-free states can nevertheless exhibit Haar-like pseudorandomness and pseudoentanglement in bounded-copy settings (Giurgica-Tiron et al., 2023, Jeronimo et al., 2023), while earlier work established that subset states already suffice as witnesses for the full class (Grilo et al., 2014). A separate simulator-oriented line treats sparse support states as classical data structures storing only occupied basis states and amplitudes (Gabor et al., 2022).
1. Definition and semi-classical structure
For , a subset state is a uniform superposition over a subset , equivalently : This definition appears in the cryptographic, complexity-theoretic, and simulation-oriented treatments of subset states (Giurgica-Tiron et al., 2023, Jeronimo et al., 2023, Grilo et al., 2014).
The semi-classical character is precise. The amplitudes are supported only on ; all nonzero amplitudes have the same magnitude 0; and all phases are equal. There is therefore no hidden phase function, no sign pattern, and no arbitrary complex amplitude profile. The classical datum 1 determines the support of the wavefunction and, in the uniform case, the complete state. This sharply distinguishes subset states from generic Haar-random states, whose amplitudes are unconstrained complex numbers.
This does not make subset states classical in the QCMA sense. They remain quantum states, not classical samples from 2, and the support 3 may be exponentially large or lack an efficient classical description. The relevant distinction is that the quantum coherence is concentrated in a highly structured corner of Hilbert space. In that sense, semi-classical encoded subset states are support-defined coherent states rather than arbitrary amplitude-defined states.
A broader sparse-state variant appears in hybrid simulation work, where one stores
4
as a classical list or database of occupied basis strings and amplitudes (Gabor et al., 2022). Uniform subset states are the special case in which the nonzero amplitudes are all equal.
2. Haar-like pseudorandomness from random support
Two 2023 papers established that random subset states can be information-theoretically close to Haar-random states when only polynomially many copies are available, resolving the question of whether random or pseudorandom phases are necessary (Giurgica-Tiron et al., 2023, Jeronimo et al., 2023).
In one formulation, for 5, 6, and subset size 7 satisfying
8
the 9-copy moment operator of the random-subset ensemble obeys
0
where 1 (Giurgica-Tiron et al., 2023). This is an approximate 2-design statement in 3-copy trace distance. Negligible error requires
4
so 5 must be neither too small nor too large.
A concurrent independent result studies
6
and proves
7
(Jeronimo et al., 2023). For 8, 9, and
0
all three terms are negligible, yielding information-theoretic indistinguishability even given polynomially many copies.
Both analyses identify the same operational obstruction regime. If the subset is too small, repeated computational-basis measurements expose collisions. In the 1-parameterization, this gives the 2 term; in the 3-parameterization, if 4, measuring 5 copies yields repeated outcomes with probability 6 by the pigeonhole principle (Giurgica-Tiron et al., 2023, Jeronimo et al., 2023). If the subset is too large, the state has noticeable overlap with the uniform superposition. For example,
7
and when 8,
9
which yields an efficient swap-test distinguisher (Giurgica-Tiron et al., 2023, Jeronimo et al., 2023).
The conceptual consequence is that support randomness alone can mimic Haar low-order statistics. Random phases are not necessary in the broad intermediate regime.
3. Representation-theoretic mechanism
The most detailed proof proceeds by analyzing the random-subset moment operator on the symmetric subspace via the representation theory of the symmetric group (Giurgica-Tiron et al., 2023). Let
0
The central basis is the type basis of 1, indexed by multisets 2. Among these, the unique types correspond to ordinary 3-subsets 4, spanning
5
Its dimension is
6
while
7
and
8
Thus for 9, almost all of the symmetric subspace is already captured by unique types.
Inside this basis, the matrix elements of 0 depend only on the overlap structure of 1 and 2. For 3,
4
and after approximation this becomes a function of 5, equivalently of Johnson-graph distance.
The group action is induced by basis permutations. Writing
6
the pair 7 is a Gelfand pair. The double-cosets 8 are indexed by subset distance 9, which is precisely the Johnson scheme. Because 0 is invariant under basis permutations, its restriction to the unique subspace is group-circulant, with circulant function
1
The permutation representation decomposes multiplicity-freely as
2
with dimensions
3
The dominant block is 4, whose dimension satisfies
5
Its eigenvalue is
6
This shows that on the overwhelmingly dominant irreducible block, the rescaled subset-state moment operator is asymptotically the identity, hence the original moment operator is asymptotically the Haar moment. A plausible implication is that the relevant pseudorandomness is driven not by amplitude complexity but by the combinatorics of random support intersections, organized by the Johnson scheme.
4. Computational pseudorandom states and pseudoentanglement
The information-theoretic results become pseudorandom-state constructions by replacing uniformly random subsets with pseudorandomly generated subsets (Jeronimo et al., 2023). Using a quantum-secure pseudorandom permutation family 7 on 8, one defines
9
equivalently the subset state corresponding to
0
If the PRP is replaced by a truly random permutation, the image of 1 is a uniformly random 2-subset of 3, so the resulting state is exactly a random subset state (Jeronimo et al., 2023).
Efficient preparation is obtained by preparing 4, coherently computing 5, and then uncomputing the first register. The construction depends on the existence of an efficiently computable quantum-secure PRP; the paper notes that quantum-secure PRPs exist assuming quantum-secure one-way functions (Jeronimo et al., 2023). In a closely related formulation, the same minimalist preparation strategy uses only a PRP and no PRF for sign generation, unlike earlier subset-phase constructions (Giurgica-Tiron et al., 2023).
The pseudoentanglement consequence is immediate. Since 6 is a sum of only 7 computational-basis product states, its Schmidt rank across any cut is at most 8, and therefore the entanglement entropy across any cut is at most
9
Yet in the admissible regime the ensemble is computationally or information-theoretically indistinguishable from Haar-random states given polynomially many copies (Jeronimo et al., 2023). For
0
choosing 1 yields PRS families with entanglement entropy 2 across all cuts (Jeronimo et al., 2023).
One paper also gives a biased-phase interpolation. For subset-phase states with i.i.d. signs of bias parameter 3,
4
The subset-state case is 5, while unbiased random phases correspond to 6. This suggests that phases mainly help in the very-large-subset regime by suppressing the overlap with the global uniform superposition (Giurgica-Tiron et al., 2023).
5. Sparse classical encodings in hybrid simulation
A simulator-oriented treatment uses sparse computational-basis support as an explicit classical encoding of quantum states (Gabor et al., 2022). The proposed representations are: Array, Database, Qiskit, and Mixed. In the sparse representations, the simulator “encode[s] the entire state as a standard Numpy array containing all classical states occurring within the superposition and their respective amplitudes,” or saves “all classical states with non-zero amplitude within the quantum state” in an sqlite3 database (Gabor et al., 2022).
The core encoded form is
7
where only the occupied basis strings and their amplitudes are stored. The implied memory model is full state-vector simulation with 8 amplitudes versus sparse support encoding with 9 records. This is useful when 0.
The canonical uniform-subset benchmark is
1
with 2 (Gabor et al., 2022). The reported empirical pattern is that database runtime depends mainly on 3, the number of “Nondet qubits,” whereas dense state-vector simulation depends on the total number of qubits 4. This is the most direct simulator manifestation of semi-classical encoded subset states: complexity scales with support size rather than ambient Hilbert-space dimension.
The same framework is applied to an addition benchmark with total qubit count
5
and to Grover’s algorithm, where the diffusion operator is implemented “within a single SQL query to the database” rather than by a standard gate sequence (Gabor et al., 2022). A mixed heuristic chooses the database encoding if fewer than 6 of the qubits have a Hadamard somewhere in the circuit, and otherwise uses Qiskit.
The limitations are explicit. The method is useful only when circuits remain sparse or near-classical in the computational basis. Support growth under many Hadamards destroys the advantage, and the approximate “state drop” heuristic—after each quantum operation, keep only the 1000 largest-amplitude support entries—causes errors that “deteriorate quickly” (Gabor et al., 2022). This suggests that simulator-side semi-classical encodings are operationally valuable, but only in regimes where support sparsity is preserved.
6. Subset-state witnesses and complexity-theoretic significance
Subset states also play a structural role in quantum complexity theory. The class 7 restricts the completeness witness of a 8 protocol to a subset state, while soundness remains against arbitrary quantum witnesses (Grilo et al., 2014). The main theorem is
9
with the two-prover analogue
00
The technical core is the Subset State Approximation Lemma: for any 01-qubit state 02, there exists a subset 03 such that
04
A geometric lemma establishes the corresponding statement for arbitrary 05. The proof partitions amplitudes into dyadic level sets and shows that one level set carries enough 06-mass to yield non-negligible overlap with a uniform vector on its support (Grilo et al., 2014).
This approximation is sufficient because amplified 07 verifiers have exponentially small error. Replacing an optimal witness by the approximating subset state yields inverse-polynomial acceptance probability, and standard amplification restores the usual constant gap. The result is structural rather than algorithmic: it does not imply efficient preparation, succinct classical descriptions, or a collapse to 08 (Grilo et al., 2014).
The paper also defines 09, where a yes-instance admits a subset state that is optimal among all witnesses, and proves
10
It introduces a new 11-complete problem, 12, whose promise is phrased directly in terms of subset states (Grilo et al., 2014).
A common misconception is that subset-state witnesses are merely classical witnesses in disguise. The formal results rule that out. What is reduced is amplitude complexity and phase freedom, not the presence of coherent superposition itself. The broader lesson across complexity theory, cryptography, and simulation is that support information alone can already sustain substantial quantum power. In bounded-copy pseudorandomness it can mimic Haar low-order statistics (Giurgica-Tiron et al., 2023, Jeronimo et al., 2023); in 13 it already captures the full witness power of the class (Grilo et al., 2014); and in hybrid simulation it provides a practical sparse-state representation whenever computational-basis support remains small (Gabor et al., 2022).