How test effectiveness should be measured for LLM-generated code

Establish how test effectiveness should be measured in fully LLM-driven software development, and determine whether new adequacy criteria are required beyond statement coverage, branch coverage, and mutation testing.

Background

The study finds that statement coverage, branch coverage, and mutation testing provide weak guidance about the fault-detection capability of LLM-generated tests applied to LLM-generated code. In particular, generated tests may trigger faulty behavior without detecting it because their oracles are incomplete or incorrect.

The authors therefore do not regard the study as providing a definitive account of how test effectiveness should be assessed in this setting. They leave unresolved both the broader measurement problem and the question of whether entirely new adequacy criteria are necessary for LLM-driven software development.

References

While our results do not provide definitive answers regarding how test effectiveness should be measured in this new paradigm, they offer strong empirical evidence that coverage- and mutation-based adequacy criteria are poor indicators of fault-detection capability for LLM-generated tests. These findings highlight the need for further research to better understand how test effectiveness should be assessed in LLM-driven development and whether new adequacy criteria are required to support this emerging software engineering paradigm.

— How effective are traditional test criteria at detecting bugs in large language models generated code?  (2609.09315 - Hamidi et al., 8 Sep 2026) in Section 8, Conclusion