How test effectiveness should be measured for LLM-generated code
Establish how test effectiveness should be measured in fully LLM-driven software development, and determine whether new adequacy criteria are required beyond statement coverage, branch coverage, and mutation testing.
References
While our results do not provide definitive answers regarding how test effectiveness should be measured in this new paradigm, they offer strong empirical evidence that coverage- and mutation-based adequacy criteria are poor indicators of fault-detection capability for LLM-generated tests. These findings highlight the need for further research to better understand how test effectiveness should be assessed in LLM-driven development and whether new adequacy criteria are required to support this emerging software engineering paradigm.