- The paper introduces a sequential qubit-reuse protocol that estimates all partial-transpose moments through order K using at most 2n+1 active qubits, independent of K.
- The protocol achieves uniform additive error ε with copy complexity O(K log K/ε²), using cumulative ancilla parities so one experimental dataset estimates the entire moment hierarchy.
- Two converse bounds of Ω(K/ε²) establish near-optimality, while a PT-specific hard family shows that partial-transpose moments can remain difficult even when ordinary state moments reveal no information.
Problem and contribution
This paper studies the copy complexity of simultaneously estimating the partial-transpose (PT) moment hierarchy pj(ρAB)=Tr[(ρABTB)j] for j=2,…,K, given independent copies of an unknown bipartite n-qubit state, under an explicit active-memory constraint. The active-memory model counts all simultaneously occupied qubits (including ancillas) but not classical memory or feed-forward, and permits adaptive dynamic circuits with mid-circuit measurement and reset. The main result is a sequential qubit-reuse protocol that uses at most $2n+1$ active qubits — independent of K — and estimates the entire hierarchy to uniform additive error ε with total copy complexity O(KlogK/ε2) (2606.14204). Two converse bounds of order Ω(K/ε2) show this is optimal up to a logarithmic factor.
The problem is distinct from prior work in two respects. First, most PT-moment literature addresses what the moments reveal once available — entanglement certification from few moments, symmetry-resolved diagnostics, or phase diagrams — rather than the cost of acquiring them. Second, prior qubit-reuse protocols target ordinary state moments, whose observables are single forward cycles; the PT observable is instead the counter-propagating permutation Πj=Sj⊗(Sj)−1, which requires a different routing structure.
Sequential realization and the parity estimator
The protocol rests on the identity Tr[(ρABTB)j]=Tr[ΠjρAB⊗j], proved by term-by-term index comparison: subsystem j=2,…,K0 indices contract along a forward cycle while j=2,…,K1 indices contract along the inverse cycle. Although j=2,…,K2 need not be Hermitian for j=2,…,K3, the target moment is real because j=2,…,K4 is Hermitian.
The circuit maintains one ancilla, one storage register holding a single bipartite copy, and one transient register reloaded with a fresh copy at each layer; each depth-j=2,…,K5 execution consumes j=2,…,K6 copies while never holding more than two simultaneously. Each layer applies the ancilla-controlled unitary j=2,…,K7, where j=2,…,K8 and j=2,…,K9 are the subsystem swaps between storage and transient registers on n0 and n1 respectively. Measuring the ancilla in the n2 basis yields Kraus operators n3.
The central lemma establishes that cumulative ancilla parities n4 are unbiased: n5. The proof proceeds by a superoperator-level induction. The parity-weighted storage recursion n6 closes exactly as a symmetrized cross-term expression, and the target family n7 with n8 obeys the same recursion with the same initial condition — both cross-term maps n9 and $2n+1$0 act identically on $2n+1$1. A parity bookkeeping identity over measurement branches then converts the operator equality into the expectation formula. This exactness is what allows a single bitstring per shot to serve all moment orders at once.
Achievability
Because each depth-$2n+1$2 execution produces the full vector $2n+1$3, one dataset serves all targets. Each $2n+1$4 lies in $2n+1$5, so Hoeffding's inequality plus a union bound over the $2n+1$6 outputs gives a sup-norm guarantee $2n+1$7 with probability at least $2n+1$8 using $2n+1$9 executions, hence total copy complexity K0. The paper is explicit that the K1 factor originates entirely from uniform control over the full hierarchy, not from the sequential architecture or memory accounting.
Converse bounds
Two lower bounds establish near-optimality. The universal minimax converse reduces PT-moment estimation to ordinary moment estimation on a commuting separable diagonal family K2, for which partial transpose is trivial. Choosing parameters separated by K3 produces a PT-moment gap of order K4 via the mean-value theorem, while Pinsker's inequality applied to the resulting Bernoulli product distributions gives KL divergence K5, forcing K6 copies by the Holevo–Helstrom theorem.
The PT-specific converse is the sharper structural statement. On the two-qubit pure-state family K7, every ordinary moment satisfies K8 for all K9, yet the PT spectrum ε0 varies with ε1. On the window ε2, the derivative bound ε3 holds — obtained by isolating the dominant contribution from ε4 and showing the even-ε5 cross term is uniformly smaller there. A parameter separation ε6 then yields a moment gap of order ε7, while pure-state discrimination of tensor powers costs ε8 copies. Since the hard instances are NPT states invisible to ordinary spectral-moment estimation, the hardness is genuinely PT-specific. Notably, ε9 is excluded: for pure states O(KlogK/ε2)0 carries no parameter dependence.
Together these results characterize the copy complexity of the hierarchy up to the logarithmic factor.
Separation from downstream functional reconstruction
The paper deliberately separates acquisition of PT moments from reconstruction of nonsmooth PT functionals such as negativity. For the monomial erf–Taylor route, approximating O(KlogK/ε2)1 by O(KlogK/ε2)2 truncated at degree O(KlogK/ε2)3, the sampling overhead is governed by the rescaled coefficient weight O(KlogK/ε2)4, entering Hoeffding bounds as O(KlogK/ε2)5. For O(KlogK/ε2)6 this weight can be small, but in general it can be large, so efficient moment acquisition does not automatically yield efficient negativity estimation. A coherent PT-adapted linear-combination-of-state-powers (PT-LCSP) route via an LCU/Hadamard-test selector encounters the same coefficient-weight bottleneck; the realization model changes but the representation-instability constraint persists. These statements are coefficient-stability analyses, not information-theoretic lower bounds over all representations. The authors also note that if a certification task needs only O(KlogK/ε2)7 moments, running the depth-O(KlogK/ε2)8 circuit with O(KlogK/ε2)9 and retaining those coordinates gives Ω(K/ε2)0 copies, connecting acquisition cost to few-moment certification frameworks.
Experimental compatibility
A small-scale demonstration on IBM's ibm_pittsburgh backend estimated PT moments up to Ω(K/ε2)1 for a three-qubit ansatz state using mid-circuit measurement and reset, with readout-matrix calibration, Pauli twirling, and Clifford data regression mitigation. Mitigated estimates track exact values more closely than raw data, though residual bias remains at higher orders due to depth-dependent noise accumulation. The authors correctly frame this as a cloud compatibility check rather than an asymptotic hardware claim, noting the mitigation overhead exceeds the target shot budget at this scale.
Limitations and open questions
The principal open question is the Ω(K/ε2)2 gap between the upper bound Ω(K/ε2)3 and the lower bounds Ω(K/ε2)4. The upper-bound logarithm arises from Hoeffding plus union bound over correlated cumulative-parity estimators; the converses control only the single moment Ω(K/ε2)5, and near-optimality of the simultaneous estimator follows only because simultaneous estimation contains single-moment estimation as a special case. Extending the converse to several orders is structurally nontrivial since the hierarchy consists of coupled power sums of one partially transposed spectrum — the coordinates are not independent. Whether correlations among the Ω(K/ε2)6 collapse the uniform complexity to the single-output scale, or whether some logarithmic overhead is intrinsic, remains unresolved. Additional open directions include hard-instance families beyond the two-qubit NPT construction, genuinely simultaneous hierarchy lower bounds, and restricted settings (local access, shadow-like models, bounded memory beyond the present architecture) where additional lower-bound phenomena may appear even though unrestricted copy complexity is nearly characterized.
Conclusion
The paper formulates simultaneous PT-moment estimation as a copy-complexity problem with explicit active-memory accounting and resolves it up to a logarithmic factor. Its technical core is a sequential realization of the counter-propagating PT permutation under qubit reuse, with a superoperator induction proving cumulative ancilla parities unbiased for the full hierarchy, complemented by a PT-specific converse on an isospectral NPT family whose ordinary moments carry no information. By separating moment acquisition from downstream functional reconstruction, the work clarifies which resource questions — acquisition versus certification versus stable reconstruction — must be analyzed independently.