- The paper finds that study time on AI-susceptible math problems fell by 26.9% for college students and 31.3% for high school students over eleven post-ChatGPT quarters, while graph-based problem times remained stable.
- The paper shows that proctoring eliminated the decline in response times and that students’ odds of answering proctored retention items correctly fell by 25%, suggesting reduced cognitive engagement during earlier AI-assisted learning.
- The paper finds an 85% increase in correct-answer odds on non-proctored retention items, a reversal consistent with AI-generated answers masking weaker durable knowledge and potentially corrupting adaptive learning systems.
Motivation and research questions
Self-report surveys conducted after ChatGPT's release find little change in students' reported academic dishonesty relative to pre-ChatGPT baselines, while small behavioral studies document widespread AI reliance but lack the scale or duration to measure learning consequences. This paper closes that evidentiary gap using ten years of behavioral trace data from ALEKS, an adaptive mathematics learning and assessment platform serving more than four million students annually. The authors pose four sequential questions: (RQ1) whether study time on AI-susceptible problems shifted after ChatGPT's release; (RQ2) whether that shift disappears under proctoring; (RQ3) whether the shift is associated with weaker durable knowledge on proctored retention items; and (RQ4) whether the same estimator applied to non-proctored assessments produces the opposite-signed result that only an AI-use mechanism would predict.
Identification strategy
The design exploits within-curriculum variation in AI susceptibility. Text-based word problems — proportional reasoning, algebraic word problems, rate and mixture problems — can be transcribed into a LLM prompt in seconds and constitute the treated group. Graph-based problems requiring visual interpretation and interactive manipulation of platform widgets are substantially harder to outsource and serve as the comparison group. The unit of analysis is the topic × calendar quarter cell, with topic fixed effects absorbing time-invariant difficulty differences and quarter fixed effects absorbing seasonality and platform-wide trends. The treatment variable is a linear post-ChatGPT ramp interacted with AI susceptibility, capturing gradual technology diffusion rather than an immediate discontinuity.
The learning-time analysis draws on 3,197,803 interactions from Grade 5 through College Algebra spanning 2015 Q3 through 2025 Q3. The proctoring contrast uses 12.2 million response-time observations from ALEKS PPL placement assessments administered at more than 400 U.S. institutions, split between roughly 4.5 million proctored and 7.7 million non-proctored responses. The retention analysis uses 6,698 item-level observations on randomly assigned "extra" retention questions — items appearing in 1 of every 30 session items and assigned independently of the adaptive algorithm's mastery estimates — which removes adaptive item-selection bias at the cost of statistical power. Item classification was author-validated on 100 randomly selected items, with an LLM classifier replicating labels at 86% agreement (Cohen's κ=0.79); under non-differential misclassification, classical measurement-error theory implies attenuation toward zero, making reported estimates conservative.
Learning time on AI-susceptible topics declines by 2.80% per quarter among college students (β=−0.0284, p<0.001), cumulating to 26.9% over eleven post-ChatGPT quarters (bootstrap CI: −34.2%, −19.0%). High school students show a larger effect: 3.35% per quarter, cumulating to 31.3% (CI: −37.3%, −25.4%). Middle school shows a smaller but significant 9.0% cumulative decline, and Grade 5 shows no detectable change (p=0.79). Randomization inference with 1,000 placebo permutations yields p<0.001 for every non-null subset, and estimates are stable when the treatment break is moved across four quarters spanning 2022-Q4 through 2023-Q3.
Two features of this pattern deserve emphasis. First, learning times for graph-based problems remain stable throughout the observation period in all age groups, isolating the effect from platform-wide changes in cohort ability, pacing, or curriculum. Second, the age gradient — large for college and high school, small for middle school, null for Grade 5 — aligns with differential student autonomy and unsupervised access to AI tools, and mirrors pre-AI age gradients in the academic dishonesty literature. One caveat applies here: the College and High School subsets exhibit a small positive pre-trend (+0.0071 and +0.0046 log-time per quarter respectively) as word problems were already becoming relatively slower before ChatGPT. Because this drift is opposite in sign to the post-period effect, the unadjusted estimates are conservative; trend-adjusted specifications yield cumulative effects of −35.8% (College) and −37.0% (High School). The paper treats the proctored-PPL and retention results, which pass formal parallel-trends tests (joint Wald p>0.3), as the primary causal claims and reports the learning-time ramps as supportive evidence with the pre-trend disclosed.
RQ2: Proctoring eliminates the divergence
In non-proctored PPL assessments, response times on AI-susceptible topics decline by 1.11% per quarter relative to graph-based topics (β=−0.0112, p<0.001), cumulating to 11.6%. In proctored assessments, the estimated slope is effectively zero under every inference method applied, including randomization inference (p=0.998). This supervision-sensitivity functions as a falsification test: general efficiency gains, cohort changes, or motivational shifts would affect both problem types equally regardless of supervision, whereas only AI-assisted problem-solving predicts a topic-selective effect that vanishes when AI access is restricted and detection risk is high.
RQ3: Retention declines under proctored conditions
On randomly assigned proctored retention items, a logistic item-and-quarter fixed-effects model estimates a 2.61% per-quarter decline in log-odds of correct response on AI-susceptible items (β=−0.0260, β=−0.02840), accumulating to a 25.0% cumulative reduction in odds of correct response (odds-ratio multiplier 0.75; bootstrap CI: 0.56–0.98). Stratified randomization inference within baseline-difficulty quartiles yields β=−0.02841. Because these items were administered under proctoring where AI use was prohibited at the point of testing, the decline cannot reflect AI assistance during assessment itself; it reflects the downstream consequence of reduced cognitive engagement during the earlier learning phase. The temporal ordering — AI use during unproctored learning, retention failure in subsequent proctored assessment — supports a causal interpretation.
The paper frames this result through the distinction between cognitive offloading (strategic outsourcing of a narrow subtask while retaining executive control, as with calculators) and cognitive surrender (adopting an AI-generated output with minimal scrutiny). Prior tools produced efficiency gains without retention losses because they left reasoning in the student's hands; generative AI can surrender the reasoning itself. The finding also carries a structural implication for intelligent tutoring systems: platforms calibrate mastery estimates assuming responses reflect genuine effort, so AI-generated correct answers corrupt the diagnostic signal on which adaptivity depends.
RQ4: Non-proctored reversal falsifies alternatives
Applying the identical logistic estimator to non-proctored retention items yields β=−0.02842 (β=−0.02843), a cumulative 85% increase in odds of correct response — opposite in sign and comparable in magnitude to the proctored estimate. No mechanism based on curriculum change, platform evolution, cohort composition, or general ability trends can produce retention decline under proctoring alongside retention improvement without it. Only AI-assisted answer production at the time of assessment predicts this reversal: apparent retention rises while durable knowledge falls.
Robustness
The supplementary analyses include a functional-form horse race (step, linear, quadratic, logistic), break-date sensitivity, pre-period placebo breaks, parallel-trends tests, randomization inference, delta-method and cluster-bootstrap confidence intervals, window cuts (COVID removal, endpoint removal, attempt floors), two-way clustering, trend-adjusted ramps, and attempt-level re-fits. All non-null results survive; the wild-cluster bootstrap is disclosed as degenerate and uninformative for three subsets due to within-cluster residual-variance absorption, with cluster-robust, two-way, and randomization inference providing concordant checks instead. Notably, the AIC-minimizing quadratic form would imply larger cumulative effects (−42%), but its non-monotonic trajectory is inconsistent with any diffusion mechanism, and the authors explicitly decline to exploit it.
Limitations
The paper concedes three constraints plainly. First, the learning-time and retention analyses use different populations under different conditions, so the individual-level causal chain — a specific student used AI during learning and later failed a retention test — cannot be established; the retention result is population-level. Second, PPL placement performance reflects prior learning, test familiarity, and platform experience, not purely retention of practiced concepts, leaving a gap between PPL performance and laboratory-grade retention measurement. Third, AI use is never observed directly; all inference rests on behavioral signatures validated by the two falsification tests, and alternative explanations cannot be fully excluded — though none identified predicts the joint pattern of selectivity, proctoring sensitivity, age gradient, temporal trajectory, and non-proctored reversal. The paper also notes that multimodal and agentic AI systems capable of controlling browsers make the boundary between AI-susceptible and AI-resistant formats a moving target, requiring re-evaluation as capabilities advance.
Conclusion
Using 3.2 million learning interactions and 12.2 million assessment observations spanning a decade, this paper documents a substantial, monotonically growing, topic-selective decline in mathematics study time following ChatGPT's release — 26.9% cumulative for college and 31.3% for high school students — that disappears entirely under proctoring and carries a 25% cumulative decline in retention odds on randomly assigned proctored items, with an opposite-signed 85% increase in non-proctored conditions ruling out non-AI mechanisms. The findings extend randomized experimental evidence on AI-assisted practice to population scale, multi-year horizons, and authentic commercial settings, and they identify a concrete tension for policy: proctoring preserves authentic behavioral signals but imposes costs likely to fall disproportionately on under-resourced institutions, while AI-resistant task design offers a scalable but capability-dependent alternative.