Establish whether one hundred generated attempts adequately characterize task convergence
Determine whether generating one hundred large-language-model solution attempts is sufficient to characterize solution convergence for every programming task.
References
Our results do not establish that one hundred attempts are sufficient for every task, but the observed cost suggests that preliminary testing can be affordable.
— Correctness, Convergence, and AI-Generated Code Detection: A Longitudinal Study of Student and Large Language Model Code in Introductory Programming
(2610.00863 - Ye et al., 1 Oct 2026) in Section 5.3, “What Instructors Can Do Before Releasing an Assignment”