Assess the generalizability of generative-AI learning effects

Establish whether the observed effects of ChatGPT assistance on programming performance, retention, cognitive effort, ownership, and authorship generalize across universities, geographic and socioeconomic regions, genders, and minority groups.

Background

Participants self-selected into the experiment and came from a single university’s first-year computer science cohort. This restricted sampling frame may limit the extent to which the reported performance, retention, cognitive-load, and ownership effects apply to other student populations.

Replication across more diverse institutional, geographic, socioeconomic, gender, and minority-group contexts is needed to determine whether the findings reflect a broader phenomenon or are specific to this sample.

References

This could introduce self-selection bias and these results may not be generalisable to all students. Future studies should look to replicate this experimental design across different geographic and socioeconomic regions. Specifically gender and other traditionally minority group specific focused replications would be beneficial to a more holistic understanding of the impacts of GenAI across gender and minority groups.

— Your Programming Students' Cognition with ChatGPT: Higher Performance, Lower Retention, and Reduced Ownership  (2609.21194 - Bergh et al., 18 Sep 2026) in Section 5, Limitations and Future Work