Unifying test-inadequacy diagnosis with targeted test synthesis

Develop a unified framework that integrates test-inadequacy diagnosis with targeted test synthesis in a single end-to-end system to detect under-constrained behaviors in regression test suites and generate focused regression tests that close those gaps.

Background

Benchmark evaluations commonly accept patches that pass existing regression tests, but these suites often under-constrain intended behaviors, allowing semantically incorrect patches to be accepted. Existing countermeasures either rely on manual analysis, focus on distinguishing AI- from human-written code rather than strengthening tests, or treat gap detection and test generation as separate steps.

The authors highlight that without a diagnostic phase, generated tests can redundantly cover already-tested behavior, leaving critical gaps unaddressed. They explicitly identify the unification of test-inadequacy diagnosis and targeted test generation within a single framework as an open challenge motivating their work.

References

Unifying test-inadequacy diagnosis with targeted test synthesis in a single framework remains an open challenge.

We leave the question of how to write EBTs to communicate the user's intended condition for future work.

— Retrofitting Code Using LLMs to Support Exceptional Behavior  (2609.10397 - Zhong et al., 9 Sep 2026) in Section 6.2, subsection “Equivalence to Ground Truth”

Track B's 8 testbenches are not yet certified under this protocol at all, and extending certification to Track B, along with strengthening the 14 below-floor Track A testbenches, are both future work --- the area-reduction Track B objective's lack of sec coverage (Section~\ref{sec:trackb}) is a related, disclosed gap.

— GateTruth: Auditing the Rigor of RTL Design Benchmarks via Mutation Testing  (2608.12635 - Bhadra, 12 Aug 2026) in Section 10, Limitations and Threats to Validity, Applying the same denominator scrutiny to our own suite