- The paper’s main finding is that ChatGPT availability has minimal measurable effects on high school test performance, with effect sizes near zero.
- Researchers employed cross-sectional fixed-effects regression and placebo tests, ensuring robust control over demographic and temporal variations.
- The results challenge both highly optimistic and pessimistic views by showing that sporadic LLM use does not significantly alter standardized exam outcomes.
Introduction
The paper "Little Impact of ChatGPT Availability on High School Student Test Score Performance" (2605.08812) addresses the empirical effects of LLM availability on educational outcomes, specifically high school standardized test performance. While concerns about LLMs—particularly ChatGPT—center on both their potential as educational aids and as vehicles for academic dishonesty, the aggregate impact in real-world, unsupervised student use remains largely unexplored, especially at the secondary education level. This study leverages regional and temporal variation in ChatGPT traffic to assess its association with high school test performance from 2023–2024, providing a critical non-experimental, population-level assessment.
Literature Context
Prior LLM education research is predominantly experimental or descriptive, focused on university settings with AI tool interventions rather than voluntary student-driven usage. Meta-analyses—such as Deng et al.—find generally positive effects for learning with AI instructional tools among college students, with stronger results frequently concentrated at the lower end of performance distributions. Contradictory findings exist: several studies demonstrate negative learning effects upon withdrawal of AI access or note declines in information retention and self-efficacy (e.g., Bastani et al., Barcaui, Shi et al.). Observational causal work remains sparse, with Hausman et al. [2025] and Yu et al. [2024] offering the most rigorous evidence of positive, but possibly superficial, gains in measured student performance, often without evidence of durable learning increases. Importantly, little prior work examines K-12 education or the effects of unsupervised, non-interventionist LLM use.
Data and Methodology
This analysis is founded on a quasi-experimental approach utilizing multiple independent sources:
- Test Score Data: Drawn from the Stanford Education Data Archive (SEDA) for grades 3-8, SAT/ACT state-level reports for 2023–2025, and SchoolDigger.com for high school subject exams (PSAT, Smarter Balanced, STAAR EOC, Regents, etc.), with the primary outcome variables being standardized scores or proficiency rates.
- ChatGPT Usage Data: Inferred from Google Trends (relative search activity for “chatgpt” at state and metropolitan levels) and Similarweb (web visits to chatgpt.com/openai.com), leveraging the distinct seasonality in usage—high during the school year, reduced during summer—to identify regions of educational (as opposed to recreational) LLM engagement.
- Controls and Covariates: Demographic and educational attainment controls are introduced at the metro level (from ACS and IPEDS), with state and year fixed effects employed to isolate local variation independent from national trends and post-pandemic educational recovery.
The principal empirical strategy is a cross-sectional, fixed effects regression where the main regressor is the “school-year ChatGPT usage bump” (school-year usage vs. summer). The analytical design is supported by a placebo test on younger students (unlikely to use LLMs), addressing concerns of COVID-recovery or socioeconomic endogeneity.
Main Findings
Placebo and State-Level Results
- Placebo Test: Among grades 3–8, no significant effect of ChatGPT traffic is observed, supporting the assertion that the model is not confounded by extraneous trends in educational recovery or shifting demographics.
- State-Level (SAT/ACT): Across all specifications, both Google Trends and Similarweb indicators yield estimated ChatGPT effects on test scores that are statistically indistinguishable from zero, with effect sizes often less than 0.05 SD, even for large usage differentials. Estimates are consistently non-significant, with wide confidence intervals, reflecting low power at this level.
- School-Level Subject Exams: The overwhelming majority of effect sizes across grades and subjects are small and not statistically significant. Where significance emerges (e.g., 10th-grade English/Reading, certain science grades), both magnitude and consistency are low—typically involving proficiency shifts on the order of 0.1–0.2 percentage points per 10% increase in school-year ChatGPT usage. These effects are substantively negligible vis-à-vis normal inter-annual (or inter-state) outcome volatility.
- Distributional Analysis: Examining impacts on students above the lowest proficiency threshold or at the highest performance level reveals no systematic evidence of LLM-induced score compaction or polarization. Positive and negative coefficients appear scattered with low magnitude and unreliable statistical support.
- Robustness: Introducing comprehensive metro-level demographic and education attainment controls does not materially alter the null results.
Interpretation and Implications
The null aggregate effect stands in contrast to some experimental evidence (mainly in higher education settings) and challenges both utopian (massive LLM-facilitated learning gains) and dystopian (widespread erosion of authentic learning via “cheating”) narratives. The results suggest several possibilities:
- Aggregate Cancelation or Heterogeneity: Potential positive and negative mechanisms (e.g., authentic help vs. avoidance of active learning) may cancel out, or LLM usage could be highly heterogeneous—some subgroups benefit while others are unaffected or harmed, yielding a net zero at population level.
- Insufficient Usage Intensity: Despite survey claims of ubiquity, real-world ChatGPT use may be too sporadic, superficial, or last-minute to measurably influence outcomes in high-stakes, standardized assessment contexts.
- Test Robustness: Standardized tests may be sufficiently insulated from LLM-driven changes in regular homework or class engagement, either because they assess retained skill/knowledge that cannot be replaced by LLM “shortcuts,” or because students exert compensatory effort in advance of high-stakes testing.
- Transferability and Generalization: The results are robust evidence that neither dire nor transformational effects of LLMs on learning are manifest—at least for the high school aggregate U.S. population in the years immediately following ChatGPT’s public availability.
Future Trajectories
The research restricts inference to 2023–2024. A steep rise in non-educational usage (Chatterji et al., 2025) subsequently obfuscates the traffic signal, necessitating methodological advances for future observational causal inference. As LLMs become more integrated into classroom practice—not merely extracurricular or clandestine tools—longitudinal and experimental designs targeting specific interventions, usage patterns, or subpopulations will be crucial.
From a policy and practice standpoint, the null effect provides latitude for considered experimentation with AI augmentation in curriculum without immediate fear of score deflation. However, vigilance is warranted: as both LLMs and educational strategies adapt, and as norms around their deployment and assessment integrity evolve, effects in later cohorts or under different regulatory and instructional regimes may diverge.
Conclusion
The analysis decisively rejects both large positive and large negative impacts of ChatGPT availability on high school test scores at the U.S. population level in 2023–2024, providing a crucial corrective to speculative claims about LLMs in secondary education. While the results do not preclude differential subgroup effects, evolving patterns of use, or shifts in educational practice, they place empirically-grounded bounds on the initial real-world educational impact of open-domain LLMs like ChatGPT on standardized test performance. Continued research is required to track dynamic effects as LLM technology and its integration into educational practice advance.