---
title: Self-Testing Datasets
url: https://www.emergentmind.com/topics/self-testing-datasets
type: topic
---

# Self-Testing Datasets

“Self-testing datasets” names a field-dependent class of data resources in which the recorded observations are themselves sufficient for certification, correctness checking, or discrimination among systems. In quantum information, the data may be Bell-correlation tables, assemblages, or related operator-system states that certify target states and measurements up to local isometries; in autonomous-systems and code-generation research, the data may carry embedded oracles, deterministic test suites, or temporal splits that enable evaluation without re-running expensive simulations or manually reconstructing ground truth. Experimental Bell correlations $P(a,b|x,y)$ were explicitly described as “self-testing datasets” because they are sufficient to certify bipartite pure entangled states up to local isometries [1803.10961], while SensoDat is presented as an “information-rich, self-testing dataset” for self-driving cars [2401.09808], and LeetCodeDataset as a “contamination-aware, self-testing benchmark and training testbed” for code LLMs [2504.14655]. This suggests a family of dataset designs in which verification is pushed into the data representation itself rather than delegated entirely to trusted internals or repeated execution [2106.00840].

## 1. Conceptual scope and recurring design pattern

In current usage, the term does not denote a single formalism. Instead, it covers several research traditions that share a common structural idea: a dataset is paired with enough metadata, constraints, or observables to support certification from the dataset alone. In quantum self-testing, the central object is a correlation or assemblage together with an equivalence notion under local isometries. In software and ML evaluation, the central object is a benchmark with built-in correctness checks, such as canonical outputs, hidden tests, or oracle-derived PASS/FAIL labels. In benchmark diagnostics, the central object is a test set whose ability to distinguish current systems is itself quantified.

| Research area | Data object | Self-testing mechanism |
|---|---|---|
| Quantum information | Bell correlations, assemblages, operator-system states | Certification up to local isometries |
| Autonomous systems | Sensor time series, trajectories, PASS/FAIL metadata | Oracle-based evaluation via OOB safety metric |
| Code LLM evaluation | Problems, canonical outputs, 100+ tests per problem | Sandboxed execution and correctness checking |
| NLP benchmark diagnostics | Per-example correctness across model pools | IRT-based discrimination and headroom analysis |

A plausible implication is that “self-testing dataset” is best understood functionally rather than taxonomically: the defining feature is not the domain, but the presence of a data-internal certification route.

## 2. Bipartite and one-sided quantum self-testing datasets

In the one-sided device-independent setting based on EPR-steering, the client’s subsystem $C$ is trusted with known finite-dimensional Hilbert space $\mathcal{H}_C$, while the provider’s subsystem $P$ is untrusted with potentially unrestricted dimension $\mathcal{H}_P$. The joint physical pure state is $|\psi\rangle \in \mathcal{H}_C \otimes \mathcal{H}_P$, and the central steering object is the assemblage
$$
\sigma_{a|x} = \mathrm{Tr}_P[(I_C \otimes E_{a|x}) |\psi\rangle\langle\psi|].
$$
Equivalence to a reference experiment is defined by a local isometry on the provider and a fixed ancilla “junk” state. Robustness is measured in trace distance and the Schatten $1$-norm, with
$$
D(\rho,\sigma) = \frac{1}{2}\|\rho-\sigma\|_1.
$$
For the EPR experiment targeting the ebit
$$
|\Phi^+\rangle = \frac{|00\rangle + |11\rangle}{\sqrt{2}},
$$
the main assemblage-based one-sided self-testing theorem gives
$$
f(\epsilon) = 24\sqrt{\epsilon} + \epsilon,
$$
while a steering-inequality route gives $f(\eta)=13\sqrt{\eta}$, and an SDP-based CHSH-as-steering analysis gives the numerical fit $D \le 1.19\sqrt{\eta}$; the comparable device-independent bound quoted in the paper is $D \le 1.59\sqrt{\eta}$. The improvement is by constant factors only, because $O(\epsilon)$ robustness is impossible in general and the fundamental scaling remains $O(\sqrt{\epsilon})$ [1601.01552].

This one-sided formulation is dataset-like in a literal sense. The client can either reconstruct $\sigma_{a|x}$ and $\rho_C$ tomographically, or measure a fixed set of trusted POVM elements to estimate $p(a,b|x,y)$. The practical protocol assumes a trusted qubit client, an untrusted provider with $x \in \{0,1\}$ and $a \in \{0,1\}$, many i.i.d. rounds, and classical communication. Acceptance can be based directly on assemblage closeness,
$$
\|\sigma_{a|x}-\tilde{\sigma}_{a|x}\|_1 \le \epsilon,\qquad D(\rho_C,\tilde{\rho}_C)\le \epsilon,
$$
or on steering-inequality violation and the resulting fidelity lower bound. The paper also formalizes a SWAP isometry and extends the same logic to GHZ scenarios and tensor-product certification on the untrusted side [1601.01552].

The experimental literature made the dataset interpretation explicit. In a bipartite Bell experiment, the recorded probabilities $P(a,b|x,y)$ themselves were presented as “self-testing datasets”: in the two-qubit case, the experiment used a $[\{2,2\},\{2,2\}]$ Bell scenario; in the $d=4$ case, Alice implemented $12$ projective measurements and Bob $16$, yielding $192$ recorded probabilities for full no-signalling checks, while the four $2\times 2$ blocks used for the self-test required $64$ probabilities in total [1803.10961].

## 3. Multipartite, parallel, and experimental quantum datasets

A major line of work generalizes bipartite self-testing datasets to multipartite settings by reducing them to certified two-qubit subtests. One construction combines projections onto two-qubit subspaces with maximal violation of tilted CHSH inequalities. This yields self-testing of partially entangled GHZ states, Dicke states, and graph states with $2$ measurements per party, and multipartite qudit Schmidt states with $3$ measurements per party for parties $1,\ldots,N-1$ and $4$ for party $N$. The target qubit GHZ family is
$$
|\mathrm{GHZ}^{(n)}_\theta\rangle = \cos\theta\,|0\rangle^{\otimes n} + \sin\theta\,|1\rangle^{\otimes n},
$$
and the bipartite building block is the tilted CHSH expression
$$
\beta_\alpha = \alpha\langle A_0\rangle + \langle A_0 B_0\rangle + \langle A_0 B_1\rangle + \langle A_1 B_0\rangle - \langle A_1 B_1\rangle,
$$
with quantum maximum $\sqrt{8+2\alpha^2}$. The same paper gives the first self-test of a class of multipartite qudit states, namely all multipartite states that admit a Schmidt decomposition [1707.06534].

Parallel self-testing addresses a different dataset regime: many EPR pairs are tested simultaneously rather than sequentially. The target is $n$ maximally entangled qubit pairs shared between two players, and the certification goal is the existence of local isometries such that
$$
(V_A \otimes V_B)|\psi'\rangle \approx |junk\rangle \otimes |\Phi^+\rangle^{\otimes n}.
$$
The paper provides a general reduction from approximate commutation and anti-commutation conditions to an isometry theorem, together with two concrete constructions: a parallelized Mayers–Yao test and a strictly parallel CHSH-like test. The former uses only $O(\log n)$ measurements and has robustness scaling polynomially in $n$; the latter has an exponential number of measurement settings and exponential robustness scaling factors [1511.04194].

Experimental multi-photon self-testing then turned these constructions into measurement datasets for genuinely multipartite entanglement. For four-photon graph states, the GHZ target and the four-qubit linear cluster target were certified from observed input-output statistics. The experiment reported a four-photon GHZ state with a fidelity of $0.957(2)$ and a four-photon linear cluster state with a fidelity of $0.945(2)$, and then derived device-independent fidelity lower bounds from Bell violations. For the GHZ inequalities, the observed values were $\langle\mathcal{B}_1\rangle=4.74(2)$, $\langle\mathcal{B}_2\rangle=6.50(4)$, and $\langle\mathcal{B}_3\rangle=8.27(5)$, yielding $F_{\mathrm{DI}}\ge 0.91(2)$, $0.89(3)$, and $0.89(3)$. For the cluster inequalities, the observed values were $\langle\mathcal{B}_4\rangle=4.66(4)$, $\langle\mathcal{B}_5\rangle=6.43(7)$, and $\langle\mathcal{B}_6\rangle=5.43(6)$, yielding $F_{\mathrm{DI}}\ge 0.84(4)$, $0.84(6)$, and $0.86(4)$ [2105.10298].

Experimentally robust self-testing for bipartite and tripartite states focused on analytic lower bounds linking Bell violation to extractability. For CHSH,
$$
F_{\mathrm{singlet}}(\beta) \ge \alpha_{\mathrm{CHSH}} \beta - \gamma_{\mathrm{CHSH}},
$$
with $\alpha_{\mathrm{CHSH}} = (4 + 5\sqrt{2})/16$ and $\gamma_{\mathrm{CHSH}} = (1 + 2\sqrt{2})/4$, while for the tripartite GHZ state under the Mermin inequality,
$$
F_{\mathrm{GHZ}}(\beta) = \frac{1}{2} + \frac{1}{2}\cdot \frac{\beta - 2\sqrt{2}}{4 - 2\sqrt{2}}.
$$
The CHSH threshold for a nontrivial fidelity lower bound is reported as $\beta \approx 2.11$, improving on a previous $\approx 2.37$, and the Mermin bound is described as tight [1804.01375].

## 4. Generalized certification frameworks and relaxed assumptions

Several works broaden the formal notion of what a quantum self-testing dataset may contain. One operator-system formulation describes bipartite correlations as states on the commuting tensor product of operator systems, $S_A \otimes_c S_B$. In that framework, a quantum commuting model is
$$
S = ({}_A H_B,\varphi_A,\varphi_B,\xi),
$$
with correlation
$$
f_S(u)=\langle (\varphi_A\cdot\varphi_B)(u)\xi,\xi\rangle,\qquad u\in S_A\otimes_c S_B.
$$
The paper defines local isometries in the commuting operator model, distinguishes self-testing from abstract self-testing, proves that self-tests are always abstract self-tests, and shows converse results in some cases. The framework covers correlations with quantum inputs and outputs, quantum commuting correlations for CHSH, synchronous correlations, contextuality scenarios, quantum colourings, and Schur quantum channels, but it does not provide explicit quantitative robustness bounds [2506.17980].

A graph-theoretic framework replaces direct analysis of the quantum set by analysis of the theta body of a weighted exclusivity graph. For a Bell witness written as $S=\sum_i w_i p_i$, the weighted Lovász theta number is given by the SDP
$$
\theta(G,w)=\max \sum_{i=1}^N w_i X_{ii}
$$
subject to
$$
X_{ii}=X_{0i},\quad X_{ij}=0\ \text{for all edges } i\sim j,\quad X_{00}=1,\quad X\in S^{1+N}_+.
$$
If the quantum maximum $S_Q$ equals $\theta(G,w)$ and the primal SDP has a unique optimizer, self-testing follows from uniqueness of the associated Gram decomposition. This framework recovers CHSH and three-party Mermin, proves self-testability of chained Bell inequalities for rank-one projective measurements, and adds the Abner Shimony inequality for rank-one projective measurements. It also yields the closed-form expression
$$
\theta(Ci_{4N}(1,2N)) = N\left[1 + \cos\left(\frac{\pi}{2N}\right)\right]
$$
for the Möbius ladders arising from chained inequalities [2104.13035].

A different relaxation concerns input generation. Standard self-testing assumes measurement independence, but self-testing with untrusted random number generators replaces this by the residual randomness condition
$$
p(st\mid abxy)>0\qquad \text{for all } s,t,a,b,x,y.
$$
Under this condition, all pure bipartite partially entangled states
$$
|\psi\rangle=\sum_{i<d} c_i |ii\rangle,\qquad c_i>0,
$$
can be self-tested up to local isometries using one measurement with $d$ outcomes and $(d-1)$ dichotomic measurements per party. The same work shows that Hardy-type, possibilistic self-tests survive under arbitrarily weak independence, whereas Bell-value-only self-testing can fail: in the untrusted-source model, even a maximal CHSH value $2\sqrt{2}$ does not self-test a unique state when $l<1/4<u$ [2603.10663]. This directly qualifies a common misconception that large Bell values alone are always sufficient.

## 5. Embedded-oracle datasets in autonomous systems and code generation

Outside quantum information, self-testing datasets are built by embedding an oracle or executable correctness criterion into the dataset. SensoDat is a large-scale, simulation-based dataset of self-driving-car test executions produced in BeamNG.tech. It contains $32{,}580$ executed test cases across $14$ simulation campaigns, with $19{,}926$ PASS and $12{,}654$ FAIL outcomes, and records time series from $81$ distinct simulated sensors together with trajectory logs. The primary oracle is the Out-of-Bound safety metric with threshold $0.5$:
$$
OOB_{\mathrm{rate}} = \frac{1}{T}\sum_{t=1}^{T} oob(t),\qquad
outcome = PASS\ \text{if}\ OOB_{\mathrm{rate}} < 0.5,\ FAIL\ \text{otherwise}.
$$
The dataset is stored as JSON mapped $1{:}1$ to MongoDB documents, totals approximately $3.34$ GB, and is organized around OpenDRIVE metadata and execution data. Its stated uses include regression testing, flakiness detection, scenario coverage analysis, test selection, anomaly detection on time series, and oracle construction [2401.09808].

LeetCodeDataset implements the same idea for code LLMs. It contains $2{,}869$ curated Python problems with verified test outputs, covering over $90\%$ of LeetCode Python problems, and provides over $100$ test cases per problem. It uses a strict temporal split: post-cutoff evaluation problems are those released after $2024$-$07$-$01$, and the post-cutoff test set contains $256$ newly released problems. The harness uses generated JSON-like structured inputs, canonical-solution execution in a sandbox, deterministic recorded outputs, and custom validators for linked lists and binary trees. Evaluation uses the standard estimator
$$
pass@k = 1 - \frac{\binom{n-c}{k}}{\binom{n}{k}}.
$$
The same paper reports that self-testing SFT with only $2.6$K model-generated, automatically verified samples is comparable to or better than several $75$K–$110$K baselines on HumanEval and MBPP, while reasoning-oriented models substantially outperform non-reasoning ones on the post-cutoff set [2504.14655].

In both cases, the dataset is not only a passive archive. It also specifies how correctness is derived: OOB-based PASS/FAIL in the autonomous-systems case, and deterministic unit-test execution against canonical outputs in the code case. This suggests that embedded oracles are the non-quantum analogue of the certification map supplied by Bell inequalities, steering inequalities, or swap isometries.

## 6. Dataset quality, saturation, and limitations

Whether a dataset remains useful as a self-testing resource depends on how much discrimination it retains. A direct analysis of this question was carried out for NLP test sets using Item Response Theory. The study fit a $3$PL model by variational inference in Pyro to predictions from $18$ pretrained Transformer models over $29$ datasets, treating models as examinees and test examples as items. The item-response function is
$$
P_i(\theta)=c_i + (1-c_i)\cdot\frac{1}{1+\exp(-a_i(\theta-b_i))},
$$
and the paper’s main discriminative statistic is locally estimated headroom,
$$
LEH_i = \frac{d}{d\theta}P_i(\theta)\Big|_{\theta=\theta^\star},
$$
evaluated at the ability of the best-performing model. By this criterion, Quoref, HellaSwag, and MC-TACO are best suited for distinguishing among strong models, whereas SNLI, MNLI, and CommitmentBank appear saturated. The paper also reports that the correlation between LEH computed from all models and from a weaker-only subset is $95.5\%$ at the $75$th percentile, and that approximately $56\%$ of BoolQ items have inferred $c_i>0.5$ [2106.00840].

Across domains, several limitations recur. In multipartite quantum self-testing through projected two-qubit subtests, explicit robustness bounds are not provided [1707.06534]. The operator-system and graph-theoretic formalisms likewise do not supply quantitative robustness bounds [2506.17980; 2104.13035]. For self-testing with untrusted random number generators, robust self-testing under residual randomness and finite statistics is identified as an open problem [2603.10663]. These absences matter because exact self-testing statements at maximal violation do not automatically translate into experimentally stable dataset pipelines.

Non-quantum datasets exhibit analogous constraints. SensoDat does not include explicit collision or disengagement annotations, centers scenario diversity on road geometry and vehicle configurations, and retains the usual simulation-to-reality domain gap [2401.09808]. LeetCodeDataset is restricted to Python, excludes multi-function and design problems because of missing judgment code, does not validate time or space complexity beyond functional correctness, and does not specify timeouts or Python version in the paper [2504.14655]. Benchmark-diagnostic work also depends on the chosen model pool, metric, and task format, so estimated discrimination can shift with new model families [2106.00840].

A common misconception is that larger datasets or stronger aggregate scores necessarily imply stronger self-testing. The literature is more restrictive. Steering improves constants but not the $O(\sqrt{\epsilon})$ scaling of robustness [1601.01552]; maximal CHSH violation can fail to self-test in the presence of measurement dependence [2603.10663]; and benchmark datasets can become saturated even while remaining large and popular [2106.00840]. The most stable interpretation, therefore, is that self-testing datasets are not defined by scale alone, but by the existence, sharpness, and persistence of a verification mechanism that remains informative under realistic noise, distribution shift, or model improvement.

Source: https://www.emergentmind.com/topics/self-testing-datasets