Establish whether one hundred generated attempts adequately characterize task convergence

Determine whether generating one hundred large-language-model solution attempts is sufficient to characterize solution convergence for every programming task.

Background

The paper recommends generating and inspecting reference solutions before releasing an assignment, but does not establish a universally adequate sample size for the reference bank. Although the observed generation cost makes a preliminary bank of one hundred attempts affordable, task-specific convergence may require more attempts, and the appropriate number remains unresolved.

References

Our results do not establish that one hundred attempts are sufficient for every task, but the observed cost suggests that preliminary testing can be affordable.

— Correctness, Convergence, and AI-Generated Code Detection: A Longitudinal Study of Student and Large Language Model Code in Introductory Programming  (2610.00863 - Ye et al., 1 Oct 2026) in Section 5.3, “What Instructors Can Do Before Releasing an Assignment”