- The paper demonstrates that proactive-critical learning behavior mediates the advantage of background factors in AI-assisted education using a rigorous RCT design.
- It identifies distinct AI-use archetypes such as active trial, verification, and error-correction, each showing varied impacts on exam performance.
- Findings imply that scaffolding critical AI engagement may help reduce educational disparities and enhance learning outcomes.
Experimental Paradigm and Conceptual Framework
The study conducts a randomized controlled trial (RCT) with 318 undergraduate students stratified by university tier and prior academic preparation, spanning Python programming and game theory instructional modules. Participants are randomized within each course to either GPT-4o access (experimental condition) or a control group without any LLM accessibility across the learning, assignment, and review phases; no group receives LLM access during the final exam, ensuring the evaluation metric directly captures near-transfer knowledge acquisition rather than rote completion or momentary scaffolding.
This experimental framework enables isolation of both the direct and indirect effects of AI assistance on learning outcomes, dissecting these through the mediational role of student learning behaviors. Heterogeneity is systematically integrated into the model, acknowledging that both exogenous (university ranking, prior knowledge) and endogenous (interaction modality with LLMs) characteristics modulate the efficacy of AI augmentation.

Figure 1: Overview of the experimental design, taxonomy of AI-use behaviors, and the hypothesized mediation of background-advantaged learning gains by proactive-critical engagement with AI.
Taxonomy and Distribution of Learning Behaviors in AI-Assisted Contexts
Empirical annotation of chat, platform, and answer modification logs establishes a taxonomy of five learning behavior archetypes: abstention, rote-adoption, active-trial, verification, and error-correction. Limited engagement encompasses abstention and rote-adoption—either not using AI at all or copy-pasting LLM output with minimal cognitive filtering. In contrast, proactive and critical engagement subsumes active-trial (attempt before consult), verification (explicit checking/understanding of AI output), and error-correction (identification and correction of LLM mistakes).

Figure 2: Schematic and quantitative decomposition of AI-use learning patterns and their association with downstream exam performance; violin plots emphasize robust performance differentials based on engagement style.
The distribution analysis reveals that proactive-critical strategies are empirically sparse among students with weaker preparation or from less selective institutions, but highly enriched among academically advantaged students. Notably, the error-correction archetype is nearly absent in Python due to near-perfect LLM output, implying task-reliability interacts with the observable behavioral repertoire.
Exam performance, serving as a proxy for actual mastery (not just task completion), is tightly coupled to the behavioral engagement mode. Students with GPT access but exhibiting abstention do not outperform controls, and those using simple rote-adoption also show negligible or negative effects on subsequent knowledge assessments. In contrast, students employing proactive and critical interaction strategies achieve clear and significant gains, with one-sided Brunner–Munzel test results and adjusted P-values confirming the statistical robustness of this effect.
Background variables (university ranking, prior knowledge as measured by a pre-study quiz) strongly predict both the likelihood of adopting proactive-critical behaviors and raw exam gains in the unadjusted models. OLS mediation analyses show that, when learning behavior is explicitly modeled, coefficients linking university ranking and prior knowledge with exam performance are attenuated by 13–32% (depending on course and background metric), with the indirect effect of behavior accounting for a substantial and statistically significant fraction of the total advantage, especially in Python.

Figure 3: Stratified analysis of behavior profiles and exam outcomes; mediation illustrated by the convergence of profile-stratified performance once behavior is held fixed.
Moreover, interaction models suggest a positive directional effect wherein higher-background students derive greater marginal benefit from GPT-assisted learning, although limited power renders the statistical inference here more exploratory (i.e., wide CIs encompassing zero in several cells).
Survey Alignment and Self-Reported Integration Strategies
Survey data confirm that observed behavioral patterns correspond to self-reported intentions and attitudes: proactive-critical users disproportionately report using AI as a just-in-time support tool after independent effort, whereas rote-adoption students report higher willingness to use AI but also greater concerns about overreliance and plagiarism. Risk perceptions regarding AI accuracy diminish following direct exposure, consistent with prior observations that trusted AI mitigates epistemic uncertainty for the user even if it reinforces dependency.

Figure 4: Connections between attitudinal change, perceived risk, and finalized integration practices, partitioned by annotated behavior pattern.
Implications and Future Research Trajectories
The findings assert that undifferentiated LLM access does not yield uniform benefits in educational settings. Instead, the principal pathway through which background-related academic advantage manifests in AI-augmented learning is via the adoption of proactive and critical behavior patterns—behaviors more common among the academically advantaged due to either greater metacognitive skill or higher AI literacy. These findings contradict simplistic equalizing narratives regarding AI as a universal support vector in education. The mediation results are strong: adjusting for learning behavior markedly reduces but does not fully eliminate the effect of background, especially in domains with highly reliable LLM output (e.g., Python code correctness).
Theoretically, this work extends the understanding of human–AI collaboration by specifying the engagement interface as a rate-limiting factor modulating transfer and learning, rather than access alone. Practically, it suggests that future AI tutors must scaffold not just answer provision but also the critical engagement process, and that interventions raising AI literacy among less-prepared students may be as important as algorithmic advancement for equitable gains.
Conclusion
This RCT rigorously demonstrates the mediating role of learning behavior in translating background-related academic advantage into realized gains from AI-assisted education. While GPT-4o and similar systems amplify positive effects for proactive and critically engaged students, the same systems offer little to no benefit—and occasionally negative impact—to those engaging via rote-adoption or abstention. The implication is that closing (rather than widening) educational gaps via AI requires scaffolding behavioral engagement, not merely broadening access. Future research should address long-term effects, extension across domains and populations, and explore algorithmic interventions for differential scaffolding based on dynamic engagement profile estimation.
(2607.10101)