Papers
Topics
Authors
Recent
Search
2000 character limit reached

Self-Testing Datasets

Updated 6 July 2026
  • Self-testing datasets are data resources that include built-in certification mechanisms enabling verification using internal metadata, constraints, or observables.
  • They span domains like quantum information, where Bell correlations and steering inequalities certify states, and autonomous systems, where embedded oracles drive deterministic evaluations.
  • These datasets focus on reproducibility and robust performance assessment, integrating internal checks to overcome challenges like noise, simulation-to-reality gaps, and model improvements.

“Self-testing datasets” names a field-dependent class of data resources in which the recorded observations are themselves sufficient for certification, correctness checking, or discrimination among systems. In quantum information, the data may be Bell-correlation tables, assemblages, or related operator-system states that certify target states and measurements up to local isometries; in autonomous-systems and code-generation research, the data may carry embedded oracles, deterministic test suites, or temporal splits that enable evaluation without re-running expensive simulations or manually reconstructing ground truth. Experimental Bell correlations P(a,bx,y)P(a,b|x,y) were explicitly described as “self-testing datasets” because they are sufficient to certify bipartite pure entangled states up to local isometries (Zhang et al., 2018), while SensoDat is presented as an “information-rich, self-testing dataset” for self-driving cars (Birchler et al., 2024), and LeetCodeDataset as a “contamination-aware, self-testing benchmark and training testbed” for code LLMs (Xia et al., 20 Apr 2025). This suggests a family of dataset designs in which verification is pushed into the data representation itself rather than delegated entirely to trusted internals or repeated execution (Vania et al., 2021).

1. Conceptual scope and recurring design pattern

In current usage, the term does not denote a single formalism. Instead, it covers several research traditions that share a common structural idea: a dataset is paired with enough metadata, constraints, or observables to support certification from the dataset alone. In quantum self-testing, the central object is a correlation or assemblage together with an equivalence notion under local isometries. In software and ML evaluation, the central object is a benchmark with built-in correctness checks, such as canonical outputs, hidden tests, or oracle-derived PASS/FAIL labels. In benchmark diagnostics, the central object is a test set whose ability to distinguish current systems is itself quantified.

Research area Data object Self-testing mechanism
Quantum information Bell correlations, assemblages, operator-system states Certification up to local isometries
Autonomous systems Sensor time series, trajectories, PASS/FAIL metadata Oracle-based evaluation via OOB safety metric
Code LLM evaluation Problems, canonical outputs, 100+ tests per problem Sandboxed execution and correctness checking
NLP benchmark diagnostics Per-example correctness across model pools IRT-based discrimination and headroom analysis

A plausible implication is that “self-testing dataset” is best understood functionally rather than taxonomically: the defining feature is not the domain, but the presence of a data-internal certification route.

2. Bipartite and one-sided quantum self-testing datasets

In the one-sided device-independent setting based on EPR-steering, the client’s subsystem CC is trusted with known finite-dimensional Hilbert space HC\mathcal{H}_C, while the provider’s subsystem PP is untrusted with potentially unrestricted dimension HP\mathcal{H}_P. The joint physical pure state is ψHCHP|\psi\rangle \in \mathcal{H}_C \otimes \mathcal{H}_P, and the central steering object is the assemblage

σax=TrP[(ICEax)ψψ].\sigma_{a|x} = \mathrm{Tr}_P[(I_C \otimes E_{a|x}) |\psi\rangle\langle\psi|].

Equivalence to a reference experiment is defined by a local isometry on the provider and a fixed ancilla “junk” state. Robustness is measured in trace distance and the Schatten $1$-norm, with

D(ρ,σ)=12ρσ1.D(\rho,\sigma) = \frac{1}{2}\|\rho-\sigma\|_1.

For the EPR experiment targeting the ebit

Φ+=00+112,|\Phi^+\rangle = \frac{|00\rangle + |11\rangle}{\sqrt{2}},

the main assemblage-based one-sided self-testing theorem gives

CC0

while a steering-inequality route gives CC1, and an SDP-based CHSH-as-steering analysis gives the numerical fit CC2; the comparable device-independent bound quoted in the paper is CC3. The improvement is by constant factors only, because CC4 robustness is impossible in general and the fundamental scaling remains CC5 (Šupić et al., 2016).

This one-sided formulation is dataset-like in a literal sense. The client can either reconstruct CC6 and CC7 tomographically, or measure a fixed set of trusted POVM elements to estimate CC8. The practical protocol assumes a trusted qubit client, an untrusted provider with CC9 and HC\mathcal{H}_C0, many i.i.d. rounds, and classical communication. Acceptance can be based directly on assemblage closeness,

HC\mathcal{H}_C1

or on steering-inequality violation and the resulting fidelity lower bound. The paper also formalizes a SWAP isometry and extends the same logic to GHZ scenarios and tensor-product certification on the untrusted side (Šupić et al., 2016).

The experimental literature made the dataset interpretation explicit. In a bipartite Bell experiment, the recorded probabilities HC\mathcal{H}_C2 themselves were presented as “self-testing datasets”: in the two-qubit case, the experiment used a HC\mathcal{H}_C3 Bell scenario; in the HC\mathcal{H}_C4 case, Alice implemented HC\mathcal{H}_C5 projective measurements and Bob HC\mathcal{H}_C6, yielding HC\mathcal{H}_C7 recorded probabilities for full no-signalling checks, while the four HC\mathcal{H}_C8 blocks used for the self-test required HC\mathcal{H}_C9 probabilities in total (Zhang et al., 2018).

3. Multipartite, parallel, and experimental quantum datasets

A major line of work generalizes bipartite self-testing datasets to multipartite settings by reducing them to certified two-qubit subtests. One construction combines projections onto two-qubit subspaces with maximal violation of tilted CHSH inequalities. This yields self-testing of partially entangled GHZ states, Dicke states, and graph states with PP0 measurements per party, and multipartite qudit Schmidt states with PP1 measurements per party for parties PP2 and PP3 for party PP4. The target qubit GHZ family is

PP5

and the bipartite building block is the tilted CHSH expression

PP6

with quantum maximum PP7. The same paper gives the first self-test of a class of multipartite qudit states, namely all multipartite states that admit a Schmidt decomposition (Šupić et al., 2017).

Parallel self-testing addresses a different dataset regime: many EPR pairs are tested simultaneously rather than sequentially. The target is PP8 maximally entangled qubit pairs shared between two players, and the certification goal is the existence of local isometries such that

PP9

The paper provides a general reduction from approximate commutation and anti-commutation conditions to an isometry theorem, together with two concrete constructions: a parallelized Mayers–Yao test and a strictly parallel CHSH-like test. The former uses only HP\mathcal{H}_P0 measurements and has robustness scaling polynomially in HP\mathcal{H}_P1; the latter has an exponential number of measurement settings and exponential robustness scaling factors (McKague, 2015).

Experimental multi-photon self-testing then turned these constructions into measurement datasets for genuinely multipartite entanglement. For four-photon graph states, the GHZ target and the four-qubit linear cluster target were certified from observed input-output statistics. The experiment reported a four-photon GHZ state with a fidelity of HP\mathcal{H}_P2 and a four-photon linear cluster state with a fidelity of HP\mathcal{H}_P3, and then derived device-independent fidelity lower bounds from Bell violations. For the GHZ inequalities, the observed values were HP\mathcal{H}_P4, HP\mathcal{H}_P5, and HP\mathcal{H}_P6, yielding HP\mathcal{H}_P7, HP\mathcal{H}_P8, and HP\mathcal{H}_P9. For the cluster inequalities, the observed values were ψHCHP|\psi\rangle \in \mathcal{H}_C \otimes \mathcal{H}_P0, ψHCHP|\psi\rangle \in \mathcal{H}_C \otimes \mathcal{H}_P1, and ψHCHP|\psi\rangle \in \mathcal{H}_C \otimes \mathcal{H}_P2, yielding ψHCHP|\psi\rangle \in \mathcal{H}_C \otimes \mathcal{H}_P3, ψHCHP|\psi\rangle \in \mathcal{H}_C \otimes \mathcal{H}_P4, and ψHCHP|\psi\rangle \in \mathcal{H}_C \otimes \mathcal{H}_P5 (Wu et al., 2021).

Experimentally robust self-testing for bipartite and tripartite states focused on analytic lower bounds linking Bell violation to extractability. For CHSH,

ψHCHP|\psi\rangle \in \mathcal{H}_C \otimes \mathcal{H}_P6

with ψHCHP|\psi\rangle \in \mathcal{H}_C \otimes \mathcal{H}_P7 and ψHCHP|\psi\rangle \in \mathcal{H}_C \otimes \mathcal{H}_P8, while for the tripartite GHZ state under the Mermin inequality,

ψHCHP|\psi\rangle \in \mathcal{H}_C \otimes \mathcal{H}_P9

The CHSH threshold for a nontrivial fidelity lower bound is reported as σax=TrP[(ICEax)ψψ].\sigma_{a|x} = \mathrm{Tr}_P[(I_C \otimes E_{a|x}) |\psi\rangle\langle\psi|].0, improving on a previous σax=TrP[(ICEax)ψψ].\sigma_{a|x} = \mathrm{Tr}_P[(I_C \otimes E_{a|x}) |\psi\rangle\langle\psi|].1, and the Mermin bound is described as tight (Zhang et al., 2018).

4. Generalized certification frameworks and relaxed assumptions

Several works broaden the formal notion of what a quantum self-testing dataset may contain. One operator-system formulation describes bipartite correlations as states on the commuting tensor product of operator systems, σax=TrP[(ICEax)ψψ].\sigma_{a|x} = \mathrm{Tr}_P[(I_C \otimes E_{a|x}) |\psi\rangle\langle\psi|].2. In that framework, a quantum commuting model is

σax=TrP[(ICEax)ψψ].\sigma_{a|x} = \mathrm{Tr}_P[(I_C \otimes E_{a|x}) |\psi\rangle\langle\psi|].3

with correlation

σax=TrP[(ICEax)ψψ].\sigma_{a|x} = \mathrm{Tr}_P[(I_C \otimes E_{a|x}) |\psi\rangle\langle\psi|].4

The paper defines local isometries in the commuting operator model, distinguishes self-testing from abstract self-testing, proves that self-tests are always abstract self-tests, and shows converse results in some cases. The framework covers correlations with quantum inputs and outputs, quantum commuting correlations for CHSH, synchronous correlations, contextuality scenarios, quantum colourings, and Schur quantum channels, but it does not provide explicit quantitative robustness bounds (Crann et al., 22 Jun 2025).

A graph-theoretic framework replaces direct analysis of the quantum set by analysis of the theta body of a weighted exclusivity graph. For a Bell witness written as σax=TrP[(ICEax)ψψ].\sigma_{a|x} = \mathrm{Tr}_P[(I_C \otimes E_{a|x}) |\psi\rangle\langle\psi|].5, the weighted Lovász theta number is given by the SDP

σax=TrP[(ICEax)ψψ].\sigma_{a|x} = \mathrm{Tr}_P[(I_C \otimes E_{a|x}) |\psi\rangle\langle\psi|].6

subject to

σax=TrP[(ICEax)ψψ].\sigma_{a|x} = \mathrm{Tr}_P[(I_C \otimes E_{a|x}) |\psi\rangle\langle\psi|].7

If the quantum maximum σax=TrP[(ICEax)ψψ].\sigma_{a|x} = \mathrm{Tr}_P[(I_C \otimes E_{a|x}) |\psi\rangle\langle\psi|].8 equals σax=TrP[(ICEax)ψψ].\sigma_{a|x} = \mathrm{Tr}_P[(I_C \otimes E_{a|x}) |\psi\rangle\langle\psi|].9 and the primal SDP has a unique optimizer, self-testing follows from uniqueness of the associated Gram decomposition. This framework recovers CHSH and three-party Mermin, proves self-testability of chained Bell inequalities for rank-one projective measurements, and adds the Abner Shimony inequality for rank-one projective measurements. It also yields the closed-form expression

$1$0

for the Möbius ladders arising from chained inequalities (Bharti et al., 2021).

A different relaxation concerns input generation. Standard self-testing assumes measurement independence, but self-testing with untrusted random number generators replaces this by the residual randomness condition

$1$1

Under this condition, all pure bipartite partially entangled states

$1$2

can be self-tested up to local isometries using one measurement with $1$3 outcomes and $1$4 dichotomic measurements per party. The same work shows that Hardy-type, possibilistic self-tests survive under arbitrarily weak independence, whereas Bell-value-only self-testing can fail: in the untrusted-source model, even a maximal CHSH value $1$5 does not self-test a unique state when $1$6 (Morán et al., 11 Mar 2026). This directly qualifies a common misconception that large Bell values alone are always sufficient.

5. Embedded-oracle datasets in autonomous systems and code generation

Outside quantum information, self-testing datasets are built by embedding an oracle or executable correctness criterion into the dataset. SensoDat is a large-scale, simulation-based dataset of self-driving-car test executions produced in BeamNG.tech. It contains $1$7 executed test cases across $1$8 simulation campaigns, with $1$9 PASS and D(ρ,σ)=12ρσ1.D(\rho,\sigma) = \frac{1}{2}\|\rho-\sigma\|_1.0 FAIL outcomes, and records time series from D(ρ,σ)=12ρσ1.D(\rho,\sigma) = \frac{1}{2}\|\rho-\sigma\|_1.1 distinct simulated sensors together with trajectory logs. The primary oracle is the Out-of-Bound safety metric with threshold D(ρ,σ)=12ρσ1.D(\rho,\sigma) = \frac{1}{2}\|\rho-\sigma\|_1.2:

D(ρ,σ)=12ρσ1.D(\rho,\sigma) = \frac{1}{2}\|\rho-\sigma\|_1.3

The dataset is stored as JSON mapped D(ρ,σ)=12ρσ1.D(\rho,\sigma) = \frac{1}{2}\|\rho-\sigma\|_1.4 to MongoDB documents, totals approximately D(ρ,σ)=12ρσ1.D(\rho,\sigma) = \frac{1}{2}\|\rho-\sigma\|_1.5 GB, and is organized around OpenDRIVE metadata and execution data. Its stated uses include regression testing, flakiness detection, scenario coverage analysis, test selection, anomaly detection on time series, and oracle construction (Birchler et al., 2024).

LeetCodeDataset implements the same idea for code LLMs. It contains D(ρ,σ)=12ρσ1.D(\rho,\sigma) = \frac{1}{2}\|\rho-\sigma\|_1.6 curated Python problems with verified test outputs, covering over D(ρ,σ)=12ρσ1.D(\rho,\sigma) = \frac{1}{2}\|\rho-\sigma\|_1.7 of LeetCode Python problems, and provides over D(ρ,σ)=12ρσ1.D(\rho,\sigma) = \frac{1}{2}\|\rho-\sigma\|_1.8 test cases per problem. It uses a strict temporal split: post-cutoff evaluation problems are those released after D(ρ,σ)=12ρσ1.D(\rho,\sigma) = \frac{1}{2}\|\rho-\sigma\|_1.9-Φ+=00+112,|\Phi^+\rangle = \frac{|00\rangle + |11\rangle}{\sqrt{2}},0-Φ+=00+112,|\Phi^+\rangle = \frac{|00\rangle + |11\rangle}{\sqrt{2}},1, and the post-cutoff test set contains Φ+=00+112,|\Phi^+\rangle = \frac{|00\rangle + |11\rangle}{\sqrt{2}},2 newly released problems. The harness uses generated JSON-like structured inputs, canonical-solution execution in a sandbox, deterministic recorded outputs, and custom validators for linked lists and binary trees. Evaluation uses the standard estimator

Φ+=00+112,|\Phi^+\rangle = \frac{|00\rangle + |11\rangle}{\sqrt{2}},3

The same paper reports that self-testing SFT with only Φ+=00+112,|\Phi^+\rangle = \frac{|00\rangle + |11\rangle}{\sqrt{2}},4K model-generated, automatically verified samples is comparable to or better than several Φ+=00+112,|\Phi^+\rangle = \frac{|00\rangle + |11\rangle}{\sqrt{2}},5K–Φ+=00+112,|\Phi^+\rangle = \frac{|00\rangle + |11\rangle}{\sqrt{2}},6K baselines on HumanEval and MBPP, while reasoning-oriented models substantially outperform non-reasoning ones on the post-cutoff set (Xia et al., 20 Apr 2025).

In both cases, the dataset is not only a passive archive. It also specifies how correctness is derived: OOB-based PASS/FAIL in the autonomous-systems case, and deterministic unit-test execution against canonical outputs in the code case. This suggests that embedded oracles are the non-quantum analogue of the certification map supplied by Bell inequalities, steering inequalities, or swap isometries.

6. Dataset quality, saturation, and limitations

Whether a dataset remains useful as a self-testing resource depends on how much discrimination it retains. A direct analysis of this question was carried out for NLP test sets using Item Response Theory. The study fit a Φ+=00+112,|\Phi^+\rangle = \frac{|00\rangle + |11\rangle}{\sqrt{2}},7PL model by variational inference in Pyro to predictions from Φ+=00+112,|\Phi^+\rangle = \frac{|00\rangle + |11\rangle}{\sqrt{2}},8 pretrained Transformer models over Φ+=00+112,|\Phi^+\rangle = \frac{|00\rangle + |11\rangle}{\sqrt{2}},9 datasets, treating models as examinees and test examples as items. The item-response function is

CC00

and the paper’s main discriminative statistic is locally estimated headroom,

CC01

evaluated at the ability of the best-performing model. By this criterion, Quoref, HellaSwag, and MC-TACO are best suited for distinguishing among strong models, whereas SNLI, MNLI, and CommitmentBank appear saturated. The paper also reports that the correlation between LEH computed from all models and from a weaker-only subset is CC02 at the CC03th percentile, and that approximately CC04 of BoolQ items have inferred CC05 (Vania et al., 2021).

Across domains, several limitations recur. In multipartite quantum self-testing through projected two-qubit subtests, explicit robustness bounds are not provided (Šupić et al., 2017). The operator-system and graph-theoretic formalisms likewise do not supply quantitative robustness bounds (Crann et al., 22 Jun 2025, Bharti et al., 2021). For self-testing with untrusted random number generators, robust self-testing under residual randomness and finite statistics is identified as an open problem (Morán et al., 11 Mar 2026). These absences matter because exact self-testing statements at maximal violation do not automatically translate into experimentally stable dataset pipelines.

Non-quantum datasets exhibit analogous constraints. SensoDat does not include explicit collision or disengagement annotations, centers scenario diversity on road geometry and vehicle configurations, and retains the usual simulation-to-reality domain gap (Birchler et al., 2024). LeetCodeDataset is restricted to Python, excludes multi-function and design problems because of missing judgment code, does not validate time or space complexity beyond functional correctness, and does not specify timeouts or Python version in the paper (Xia et al., 20 Apr 2025). Benchmark-diagnostic work also depends on the chosen model pool, metric, and task format, so estimated discrimination can shift with new model families (Vania et al., 2021).

A common misconception is that larger datasets or stronger aggregate scores necessarily imply stronger self-testing. The literature is more restrictive. Steering improves constants but not the CC06 scaling of robustness (Šupić et al., 2016); maximal CHSH violation can fail to self-test in the presence of measurement dependence (Morán et al., 11 Mar 2026); and benchmark datasets can become saturated even while remaining large and popular (Vania et al., 2021). The most stable interpretation, therefore, is that self-testing datasets are not defined by scale alone, but by the existence, sharpness, and persistence of a verification mechanism that remains informative under realistic noise, distribution shift, or model improvement.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Self-Testing Datasets.