Self-Testing Datasets
- Self-testing datasets are data resources that include built-in certification mechanisms enabling verification using internal metadata, constraints, or observables.
- They span domains like quantum information, where Bell correlations and steering inequalities certify states, and autonomous systems, where embedded oracles drive deterministic evaluations.
- These datasets focus on reproducibility and robust performance assessment, integrating internal checks to overcome challenges like noise, simulation-to-reality gaps, and model improvements.
“Self-testing datasets” names a field-dependent class of data resources in which the recorded observations are themselves sufficient for certification, correctness checking, or discrimination among systems. In quantum information, the data may be Bell-correlation tables, assemblages, or related operator-system states that certify target states and measurements up to local isometries; in autonomous-systems and code-generation research, the data may carry embedded oracles, deterministic test suites, or temporal splits that enable evaluation without re-running expensive simulations or manually reconstructing ground truth. Experimental Bell correlations were explicitly described as “self-testing datasets” because they are sufficient to certify bipartite pure entangled states up to local isometries (Zhang et al., 2018), while SensoDat is presented as an “information-rich, self-testing dataset” for self-driving cars (Birchler et al., 2024), and LeetCodeDataset as a “contamination-aware, self-testing benchmark and training testbed” for code LLMs (Xia et al., 20 Apr 2025). This suggests a family of dataset designs in which verification is pushed into the data representation itself rather than delegated entirely to trusted internals or repeated execution (Vania et al., 2021).
1. Conceptual scope and recurring design pattern
In current usage, the term does not denote a single formalism. Instead, it covers several research traditions that share a common structural idea: a dataset is paired with enough metadata, constraints, or observables to support certification from the dataset alone. In quantum self-testing, the central object is a correlation or assemblage together with an equivalence notion under local isometries. In software and ML evaluation, the central object is a benchmark with built-in correctness checks, such as canonical outputs, hidden tests, or oracle-derived PASS/FAIL labels. In benchmark diagnostics, the central object is a test set whose ability to distinguish current systems is itself quantified.
| Research area | Data object | Self-testing mechanism |
|---|---|---|
| Quantum information | Bell correlations, assemblages, operator-system states | Certification up to local isometries |
| Autonomous systems | Sensor time series, trajectories, PASS/FAIL metadata | Oracle-based evaluation via OOB safety metric |
| Code LLM evaluation | Problems, canonical outputs, 100+ tests per problem | Sandboxed execution and correctness checking |
| NLP benchmark diagnostics | Per-example correctness across model pools | IRT-based discrimination and headroom analysis |
A plausible implication is that “self-testing dataset” is best understood functionally rather than taxonomically: the defining feature is not the domain, but the presence of a data-internal certification route.
2. Bipartite and one-sided quantum self-testing datasets
In the one-sided device-independent setting based on EPR-steering, the client’s subsystem is trusted with known finite-dimensional Hilbert space , while the provider’s subsystem is untrusted with potentially unrestricted dimension . The joint physical pure state is , and the central steering object is the assemblage
Equivalence to a reference experiment is defined by a local isometry on the provider and a fixed ancilla “junk” state. Robustness is measured in trace distance and the Schatten $1$-norm, with
For the EPR experiment targeting the ebit
the main assemblage-based one-sided self-testing theorem gives
0
while a steering-inequality route gives 1, and an SDP-based CHSH-as-steering analysis gives the numerical fit 2; the comparable device-independent bound quoted in the paper is 3. The improvement is by constant factors only, because 4 robustness is impossible in general and the fundamental scaling remains 5 (Šupić et al., 2016).
This one-sided formulation is dataset-like in a literal sense. The client can either reconstruct 6 and 7 tomographically, or measure a fixed set of trusted POVM elements to estimate 8. The practical protocol assumes a trusted qubit client, an untrusted provider with 9 and 0, many i.i.d. rounds, and classical communication. Acceptance can be based directly on assemblage closeness,
1
or on steering-inequality violation and the resulting fidelity lower bound. The paper also formalizes a SWAP isometry and extends the same logic to GHZ scenarios and tensor-product certification on the untrusted side (Šupić et al., 2016).
The experimental literature made the dataset interpretation explicit. In a bipartite Bell experiment, the recorded probabilities 2 themselves were presented as “self-testing datasets”: in the two-qubit case, the experiment used a 3 Bell scenario; in the 4 case, Alice implemented 5 projective measurements and Bob 6, yielding 7 recorded probabilities for full no-signalling checks, while the four 8 blocks used for the self-test required 9 probabilities in total (Zhang et al., 2018).
3. Multipartite, parallel, and experimental quantum datasets
A major line of work generalizes bipartite self-testing datasets to multipartite settings by reducing them to certified two-qubit subtests. One construction combines projections onto two-qubit subspaces with maximal violation of tilted CHSH inequalities. This yields self-testing of partially entangled GHZ states, Dicke states, and graph states with 0 measurements per party, and multipartite qudit Schmidt states with 1 measurements per party for parties 2 and 3 for party 4. The target qubit GHZ family is
5
and the bipartite building block is the tilted CHSH expression
6
with quantum maximum 7. The same paper gives the first self-test of a class of multipartite qudit states, namely all multipartite states that admit a Schmidt decomposition (Šupić et al., 2017).
Parallel self-testing addresses a different dataset regime: many EPR pairs are tested simultaneously rather than sequentially. The target is 8 maximally entangled qubit pairs shared between two players, and the certification goal is the existence of local isometries such that
9
The paper provides a general reduction from approximate commutation and anti-commutation conditions to an isometry theorem, together with two concrete constructions: a parallelized Mayers–Yao test and a strictly parallel CHSH-like test. The former uses only 0 measurements and has robustness scaling polynomially in 1; the latter has an exponential number of measurement settings and exponential robustness scaling factors (McKague, 2015).
Experimental multi-photon self-testing then turned these constructions into measurement datasets for genuinely multipartite entanglement. For four-photon graph states, the GHZ target and the four-qubit linear cluster target were certified from observed input-output statistics. The experiment reported a four-photon GHZ state with a fidelity of 2 and a four-photon linear cluster state with a fidelity of 3, and then derived device-independent fidelity lower bounds from Bell violations. For the GHZ inequalities, the observed values were 4, 5, and 6, yielding 7, 8, and 9. For the cluster inequalities, the observed values were 0, 1, and 2, yielding 3, 4, and 5 (Wu et al., 2021).
Experimentally robust self-testing for bipartite and tripartite states focused on analytic lower bounds linking Bell violation to extractability. For CHSH,
6
with 7 and 8, while for the tripartite GHZ state under the Mermin inequality,
9
The CHSH threshold for a nontrivial fidelity lower bound is reported as 0, improving on a previous 1, and the Mermin bound is described as tight (Zhang et al., 2018).
4. Generalized certification frameworks and relaxed assumptions
Several works broaden the formal notion of what a quantum self-testing dataset may contain. One operator-system formulation describes bipartite correlations as states on the commuting tensor product of operator systems, 2. In that framework, a quantum commuting model is
3
with correlation
4
The paper defines local isometries in the commuting operator model, distinguishes self-testing from abstract self-testing, proves that self-tests are always abstract self-tests, and shows converse results in some cases. The framework covers correlations with quantum inputs and outputs, quantum commuting correlations for CHSH, synchronous correlations, contextuality scenarios, quantum colourings, and Schur quantum channels, but it does not provide explicit quantitative robustness bounds (Crann et al., 22 Jun 2025).
A graph-theoretic framework replaces direct analysis of the quantum set by analysis of the theta body of a weighted exclusivity graph. For a Bell witness written as 5, the weighted Lovász theta number is given by the SDP
6
subject to
7
If the quantum maximum 8 equals 9 and the primal SDP has a unique optimizer, self-testing follows from uniqueness of the associated Gram decomposition. This framework recovers CHSH and three-party Mermin, proves self-testability of chained Bell inequalities for rank-one projective measurements, and adds the Abner Shimony inequality for rank-one projective measurements. It also yields the closed-form expression
$1$0
for the Möbius ladders arising from chained inequalities (Bharti et al., 2021).
A different relaxation concerns input generation. Standard self-testing assumes measurement independence, but self-testing with untrusted random number generators replaces this by the residual randomness condition
$1$1
Under this condition, all pure bipartite partially entangled states
$1$2
can be self-tested up to local isometries using one measurement with $1$3 outcomes and $1$4 dichotomic measurements per party. The same work shows that Hardy-type, possibilistic self-tests survive under arbitrarily weak independence, whereas Bell-value-only self-testing can fail: in the untrusted-source model, even a maximal CHSH value $1$5 does not self-test a unique state when $1$6 (Morán et al., 11 Mar 2026). This directly qualifies a common misconception that large Bell values alone are always sufficient.
5. Embedded-oracle datasets in autonomous systems and code generation
Outside quantum information, self-testing datasets are built by embedding an oracle or executable correctness criterion into the dataset. SensoDat is a large-scale, simulation-based dataset of self-driving-car test executions produced in BeamNG.tech. It contains $1$7 executed test cases across $1$8 simulation campaigns, with $1$9 PASS and 0 FAIL outcomes, and records time series from 1 distinct simulated sensors together with trajectory logs. The primary oracle is the Out-of-Bound safety metric with threshold 2:
3
The dataset is stored as JSON mapped 4 to MongoDB documents, totals approximately 5 GB, and is organized around OpenDRIVE metadata and execution data. Its stated uses include regression testing, flakiness detection, scenario coverage analysis, test selection, anomaly detection on time series, and oracle construction (Birchler et al., 2024).
LeetCodeDataset implements the same idea for code LLMs. It contains 6 curated Python problems with verified test outputs, covering over 7 of LeetCode Python problems, and provides over 8 test cases per problem. It uses a strict temporal split: post-cutoff evaluation problems are those released after 9-0-1, and the post-cutoff test set contains 2 newly released problems. The harness uses generated JSON-like structured inputs, canonical-solution execution in a sandbox, deterministic recorded outputs, and custom validators for linked lists and binary trees. Evaluation uses the standard estimator
3
The same paper reports that self-testing SFT with only 4K model-generated, automatically verified samples is comparable to or better than several 5K–6K baselines on HumanEval and MBPP, while reasoning-oriented models substantially outperform non-reasoning ones on the post-cutoff set (Xia et al., 20 Apr 2025).
In both cases, the dataset is not only a passive archive. It also specifies how correctness is derived: OOB-based PASS/FAIL in the autonomous-systems case, and deterministic unit-test execution against canonical outputs in the code case. This suggests that embedded oracles are the non-quantum analogue of the certification map supplied by Bell inequalities, steering inequalities, or swap isometries.
6. Dataset quality, saturation, and limitations
Whether a dataset remains useful as a self-testing resource depends on how much discrimination it retains. A direct analysis of this question was carried out for NLP test sets using Item Response Theory. The study fit a 7PL model by variational inference in Pyro to predictions from 8 pretrained Transformer models over 9 datasets, treating models as examinees and test examples as items. The item-response function is
00
and the paper’s main discriminative statistic is locally estimated headroom,
01
evaluated at the ability of the best-performing model. By this criterion, Quoref, HellaSwag, and MC-TACO are best suited for distinguishing among strong models, whereas SNLI, MNLI, and CommitmentBank appear saturated. The paper also reports that the correlation between LEH computed from all models and from a weaker-only subset is 02 at the 03th percentile, and that approximately 04 of BoolQ items have inferred 05 (Vania et al., 2021).
Across domains, several limitations recur. In multipartite quantum self-testing through projected two-qubit subtests, explicit robustness bounds are not provided (Šupić et al., 2017). The operator-system and graph-theoretic formalisms likewise do not supply quantitative robustness bounds (Crann et al., 22 Jun 2025, Bharti et al., 2021). For self-testing with untrusted random number generators, robust self-testing under residual randomness and finite statistics is identified as an open problem (Morán et al., 11 Mar 2026). These absences matter because exact self-testing statements at maximal violation do not automatically translate into experimentally stable dataset pipelines.
Non-quantum datasets exhibit analogous constraints. SensoDat does not include explicit collision or disengagement annotations, centers scenario diversity on road geometry and vehicle configurations, and retains the usual simulation-to-reality domain gap (Birchler et al., 2024). LeetCodeDataset is restricted to Python, excludes multi-function and design problems because of missing judgment code, does not validate time or space complexity beyond functional correctness, and does not specify timeouts or Python version in the paper (Xia et al., 20 Apr 2025). Benchmark-diagnostic work also depends on the chosen model pool, metric, and task format, so estimated discrimination can shift with new model families (Vania et al., 2021).
A common misconception is that larger datasets or stronger aggregate scores necessarily imply stronger self-testing. The literature is more restrictive. Steering improves constants but not the 06 scaling of robustness (Šupić et al., 2016); maximal CHSH violation can fail to self-test in the presence of measurement dependence (Morán et al., 11 Mar 2026); and benchmark datasets can become saturated even while remaining large and popular (Vania et al., 2021). The most stable interpretation, therefore, is that self-testing datasets are not defined by scale alone, but by the existence, sharpness, and persistence of a verification mechanism that remains informative under realistic noise, distribution shift, or model improvement.