Papers
Topics
Authors
Recent
Search
2000 character limit reached

Generative AI Availability, Grades, and Student Satisfaction at a Large University

Published 23 Jul 2026 in cs.CY and econ.GN | (2607.21534v1)

Abstract: The spread of generative AI (GenAI) in higher education has raised concerns that students offload cognitive effort to AI, earning high grades without learning. If this "GenAI substitution hypothesis" is true, grades should rise disproportionately in GenAI-susceptible courses--those relying more on assessments like take-home problem sets and essays rather than in-class exams. Substitution could also affect student satisfaction, measured here as self-reported understanding and interest in the subject, which prior research links to assessments. We test the substitution hypothesis using syllabus and administrative data from a large U.S. university (2015-2025; 156,135 students; 87,936 course offerings). We measure courses' GenAI susceptibility using a human-validated LLM pipeline to extract assessment types from syllabi, and use a differences-in-differences design comparing outcomes across courses before and after ChatGPT's release, while modeling COVID-19 pandemic effects as either persistent or transient. We find no significant differential effect of GenAI availability on grades overall or among previously lower-performing students. Effects on self-reported understanding are likewise insignificant; effects on interest are significant only assuming transient pandemic effects. Our findings temper concerns that GenAI inflates grades and reduces students' satisfaction.

Summary

  • The paper rigorously tests the GenAI substitution hypothesis using a difference-in-differences design on a 10-year institutional dataset.
  • It employs an LLM-based pipeline to accurately measure course susceptibility to GenAI and controls for COVID-related disruptions.
  • Results indicate no significant effect of GenAI availability on academic performance or course evaluations, challenging prior positive findings.

GenAI Availability, Academic Outcomes, and Student Satisfaction in Higher Education

Introduction and Framing

This study conducts a rigorous, large-scale empirical test of the “GenAI substitution hypothesis” in higher education: the argument that generative AI tools, by enabling students to automate cognitive effort on assessments, threaten the validity of grades and may erode student satisfaction and learning. Leveraging an extensive, ten-year institutional data set from the University of Michigan (2015–2025), the analysis exploits exogenous variation in course assessment structures and the discrete introduction of ChatGPT (November 2022) to estimate differential shifts in grades and satisfaction-related outcomes as a function of assessment susceptibility to GenAI-enabled substitution. Notably, the study counters prior research reporting positive grade effects of GenAI and presents robust null findings across academic performance and course evaluation measures once COVID-era disruptions are analytically separated.

Data and Assessment Susceptibility Measurement

The analytic corpus consists of 156,135 students across 87,936 course offerings and 6,836 unique courses, enriched with syllabus metadata, detailed grade records, and course evaluations. Critical to causal identification is the measurement of “GenAI susceptibility”—the final grade fraction dependent on assessments that can, in principle, be completed via GenAI tools (e.g., take-home essays, problem sets, open-book online exams), as opposed to in-person, proctored, or performative assignments. An LLM-based pipeline (based on GPT-5.1) reconstructs assignment types and weightings from syllabus PDFs, validated against 525 human-coded syllabi, achieving a mean absolute error of 0.063 on the susceptibility measure. Figure 1

Figure 1: Distributions of contemporaneous GenAI susceptibility across courses for 2019 (pre-COVID/pre-GenAI), 2021 (COVID), and 2023 (post-ChatGPT), with substantial persistence and elevated post-pandemic susceptibility.

The operationalization of susceptibility is anchored to each course’s 2019 offering (pre-GenAI, pre-COVID) to prevent post-treatment contamination by either pandemic-induced or GenAI-prompted assessment changes.

Empirical Strategy

A difference-in-differences design estimates semester-by-semester, within-course changes in outcomes for high-susceptibility vs. low-susceptibility courses before and after ChatGPT’s release. The COVID-19 pandemic period (2020–2022) is modeled as a discrete disruption, with GenAI effects assessed relative to both the pre-COVID (2015–2019) baseline and as net of persistent pandemic effects. The key identification assumption is that absent GenAI, outcome trends across courses of differing susceptibility would have remained parallel, conditional on course and semester fixed effects. Figure 2

Figure 2: Time evolution of course susceptibility terciles, showing strong year-over-year persistence in assessment susceptibility category despite pandemic perturbations.

Strong robustness checks ensure that spurious positive post-GenAI effects arising from COVID-driven grading or assessment changes are not misattributed to GenAI.

Results: Grades, Withdrawal, and Failure

Average Grade Effects

Across all model specifications and susceptibility measures, there is no statistically significant effect of GenAI availability on course average grades once pandemic shocks are appropriately modeled and differenced. Whether assessing the average effect, or heterogeneity by students’ prior academic performance (using cohort- and course-residualized first-term GPA ranks), the results are consistently null. Figure 3

Figure 3: Event-study: per-semester difference-in-differences estimates of susceptibility on final grade, relative to pre-COVID baseline, show no systematic post-ChatGPT uplift.

Figure 4

Figure 4: Event-study: no evidence of grade effects stratified by prior ability (AbilityRank terciles), refuting hypothesized equalization or exacerbation of grade disparities.

Notably, even at the lowest ability tercile, there is no evidence of GenAI-driven grade improvement (see Figure 4) and null findings hold across alternative ability proxies, including high school GPA and standardized test ranks.

Grade Distribution, Withdrawal, and Failures

No impact is measurable at the grade distribution tails: the GenAI effect on the probability of achieving an A or avoiding failure/withdrawal is statistically indistinguishable from zero under conservative model assumptions. Marginal descriptive increases in the A grade share are not robust to parallel trend diagnostics and vanish after COVID adjustment. Figure 5

Figure 5: Event-study: threshold effects for “at least A” and “at least D” grades reveal only transient, non-causal shifts during COVID and no persistent GenAI impact.

Figure 6

Figure 6: Event-study: withdrawal and failure rates are invariant to GenAI susceptibility after controlling for pandemic dynamics.

Overall, the null findings are robust to alternative susceptibility measures, model specifications, and are not driven by instructor adaptation of grading standards.

Student Satisfaction: Course Evaluations

The analysis extends to self-reported understanding, interest, and perceived workload (reverse-coded) from standardized course evaluations (2016–2025). Again, there is no evidence that GenAI availability is associated with statistically significant erosion or improvement in student satisfaction, interest, or perceptions of workload after accounting for persistent pandemic effects. Figure 7

Figure 7: Event-study: course evaluation outcomes (median self-reported understanding, interest, and workload)—no post-ChatGPT differentials by susceptibility outside of preexisting pandemic trends.

Null effects extend to all sub-measures. Some positive shifts in the less conservative, pre-COVID comparison specification might suggest modestly increased interest or reduced workload, but these do not survive the more credible difference-within-pandemic-level modeling.

Theoretical Implications and Relation to Prior Work

This null result stands in marked contrast to previous institutional-scale studies that attribute positive grade shifts to GenAI availability in high-susceptibility courses [hausmanGenerativeAIsImpact, chirikovAIGradeInflation2026]. The divergence is traced to improved susceptibility measurement, cleaner treatment anchoring, and explicit modeling of pandemic disruptions, as well as superior validation of LLM-based extraction. When treatment anchoring is relaxed or modeled contemporaneously—as in prior studies—positive, significant (but likely artifactual) effects reappear, highlighting the risk of conflation between COVID-induced and AI-induced shocks.

The absence of a GenAI-driven collapse in grade signaling or satisfaction suggests that, at least within the three years post-ChatGPT, grade inflation, signal degradation, or academic disengagement in U.S. higher education are not yet systemically detectable at scale. These results complicate broad claims of rapid GenAI disruption in education, instead pointing to substantial institutional, pedagogical, or behavioral factors buffering against immediate substitution effects.

Limitations and Future Directions

Key limitations include potential selection bias in course sampling (as syllabus uploads are voluntary), instructor behavioral adaptation (potentially tightening grading standards post-GenAI), and noise from imperfect syllabus parsing. The strong year-over-year persistence in susceptibility quantification, as indicated by both continuous and tercile-level correlation analyses, mitigates concerns regarding treatment misclassification. Figure 8

Figure 8

Figure 8

Figure 8: Raw dynamics of average grades by susceptibility group, further demonstrating the absence of post-GenAI divergence.

Developing more granular, within-assignment measures of substitution or triangulating with direct measures of AI usage would further refine estimates. Longitudinal tracking will be necessary to detect possible lagged or cumulative effects as GenAI adoption becomes even more entrenched and instructor or administrative policies stabilize.

Conclusion

Contrary to the GenAI substitution hypothesis, this study finds no robust evidence that the introduction and widespread adoption of generative AI in higher education has degraded the informational value of grades, increased grade inflation, or reduced student satisfaction even in the most susceptible assessment environments. These results remain stable across compositional, temporal, and ability-based heterogeneity and are resilient to a variety of treatment measurement strategies. The findings challenge prevailing narratives of rapid, universal GenAI-induced academic disruption, and underscore the necessity of careful institutional, methodological, and pandemic-aware modeling for future research on AI’s longitudinal effects in educational settings. Continued surveillance is warranted as technology and policy baselines further shift.


Reference:

"Generative AI Availability, Grades, and Student Satisfaction at a Large University" (2607.21534)

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

Explain it Like I'm 14

What this paper is about

This study asks a simple, timely question: Since tools like ChatGPT became widely available, did students start getting higher grades (without learning more) in classes where AI can easily help with homework and essays? And did students feel less satisfied with their learning?

The main questions, in plain language

The researchers focus on three big questions:

  • Do grades rise more in classes that are easier to “do with AI” (like take‑home problem sets and essays) compared to classes that rely on in‑person exams or live presentations?
  • Are any grade changes bigger for students who previously struggled?
  • Do students report lower understanding, less interest, or different workload in those more “AI‑friendly” classes?

How they studied it

The basic idea

Think of two kinds of classes:

  • “GenAI‑susceptible” classes: grades mostly come from work you can do at home or on a computer (homework, take‑home exams, papers). AI could help a lot here.
  • “Less susceptible” classes: grades mostly come from things you must do live, in person (closed‑book exams, presentations, performances). AI can’t take the test for you.

If AI really lets students “substitute” their own effort with the tool (the “substitution hypothesis”), then grades should jump more in the susceptible classes after ChatGPT became available.

What data they used

They combined three big data sources at a large U.S. university (2015–2025):

  • Nearly 1.5 million student–course records (156,000+ students in 87,900+ course offerings)
  • Course syllabi (to see how grades are determined)
  • Course evaluations (students’ reports of their understanding, interest, and workload)

How they measured “AI‑susceptible” classes

Syllabi are like a course’s contract. The team used an AI program to read 36,000+ syllabi and pick out:

  • Which types of assignments were used (e.g., take‑home essays vs. in‑class exams)
  • How much each assignment type counted toward the final grade

They then built a “susceptibility score” for each course: the higher the share of take‑home work, the more susceptible. To make this fair, they “froze” each course’s score at its 2019 version (before COVID and before ChatGPT), so they wouldn’t mistake changed policies for the effect they were trying to measure. Humans checked a sample of the AI’s labels to ensure accuracy.

How they compared before vs. after ChatGPT

They used a common, simple comparison logic:

  • Look at grade trends in more‑ and less‑susceptible classes before ChatGPT (pre‑2023).
  • Look again after ChatGPT (2023–2025).
  • Ask: Did the susceptible classes improve more than the less‑susceptible ones?

You can imagine two groups of runners on the same track. If both were improving at similar rates before, and suddenly one group gets new shoes (ChatGPT), you’d expect that group to speed up more after. That extra boost is the “difference‑in‑differences.”

Handling the COVID complication

COVID (2020–2022) scrambled teaching and grading. To avoid blaming or crediting AI for COVID‑era changes, they tried two sensible assumptions:

  1. COVID effects were temporary (they faded after 2022).
  2. COVID effects were persistent (some changes lingered into 2023+).

Reporting both gives a “range” of possible AI effects.

What they found (the short version)

Here are the main results:

  • Grades: No clear rise in grades in the more AI‑susceptible classes after ChatGPT, compared to less‑susceptible classes. This holds even when they look:
    • At withdrawals and failures (no clear differences)
    • At the bottom vs. top of the grade distribution (no convincing AI‑related boost)
    • At students who previously performed worse (no special gains)
  • Student satisfaction:
    • Self‑reported understanding: No clear change linked to AI.
    • Interest in the subject: No change if COVID effects are assumed to persist; a small increase if COVID effects are assumed to fade.
    • Workload: No change if COVID effects persist; a small drop (felt easier) if COVID effects fade.

Bottom line: They do not find strong evidence that AI made grades inflate more in AI‑friendly classes or that it hurt students’ satisfaction. If anything, under one COVID assumption, students reported slightly higher interest and slightly lower workload.

Why this matters

  • Calming a major worry: Many feared that AI would let students get higher grades without learning, especially in take‑home‑heavy classes. This large, careful study does not find strong support for that worry at this university.
  • Keeping grades meaningful: If grades had jumped mainly in AI‑friendly classes, grades would become less useful as signals of skill. The results suggest grades remained fairly stable across class types.
  • Policy and teaching: Instructors and schools may not need drastic grading overhauls just to “fight” AI. Instead, they can focus on thoughtful assignment design, clear AI policies, and support for learning.
  • Future research: Other schools or subjects might differ, and AI tools keep changing. Still, this study’s methods (reading syllabi with AI, comparing class types over time, and carefully handling COVID) offer a strong model for tracking real‑world impact.

Takeaway in one sentence

At this large university, after ChatGPT’s release, classes that should be easiest to “do with AI” did not show extra grade inflation or worse student satisfaction compared to other classes—suggesting AI hasn’t broadly undermined grades or students’ learning experience.

Knowledge Gaps

Unresolved gaps, limitations, and open questions

Below is a focused list of concrete gaps and open questions that remain after this study and that future research can act on:

  • External validity: Replicate the design across diverse institutions (community colleges, less-selective publics, liberal arts colleges, HBCUs, international universities) to test whether null effects generalize beyond one large U.S. flagship.
  • Discipline and course-level heterogeneity: Estimate effects within finer strata (e.g., writing-intensive humanities vs lab STEM; coding-heavy vs problem-set courses; lower- vs upper-division; small seminars vs large lectures) to detect domain-specific substitution or mitigation.
  • Adoption and policy measurement: Directly measure student and instructor GenAI use and policies (e.g., survey-linked usage diaries, LMS plug-in logs, AI policy text on syllabi, proctoring modality, device rules, detection tools, grading policy changes) and test them as moderators/mediators.
  • Actual GenAI usage data: Link outcomes to verified AI-use indicators (browser/app telemetry with consent, assignment-level AI-detection flags, AI-enabled IDE use, prompt logs) to replace “availability” with “use intensity” and identify dose–response.
  • Instructor response as an outcome: Model time-varying changes in assessment weights and formats post-ChatGPT as outcomes (not just fix 2019 anchors) to quantify how much mitigation (e.g., shifts to proctored in-person exams) offset potential substitution.
  • Learning vs performance: Incorporate direct learning measures (common concept inventories, proctored standardized tests, rubric-level mastery checks, performance on subsequent unassisted assessments) to test if understanding diverges from grades.
  • Downstream effects: Examine persistence, major choice, course sequencing, time-to-degree, gateway-course pass rates, and performance in follow-on prerequisites to detect delayed or spillover learning impacts.
  • Equity and heterogeneity: Test heterogeneous effects by demographics (first-gen status, Pell eligibility, race/ethnicity, gender, English-learner status, disability accommodations), metacognition/self-regulation, and prior preparation beyond GPA rank (using validated scales).
  • Pretrend validity: Address failed parallel-trend tests for grades by applying alternative identification (interactive fixed effects, matrix-completion DiD, synthetic controls at department level, event studies with flexible group-specific trends, randomization inference).
  • COVID overlap modeling: Move beyond transient/persistent bracketing by estimating overlapping-shock models (e.g., dynamic factor models, interactive fixed effects) to separate lingering pandemic effects from AI-induced changes more credibly.
  • Power and MDEs: Report minimum detectable effects under course-level clustering and measurement error to clarify whether nulls rule out policy-relevant effect sizes.
  • Measurement error in susceptibility: Quantify attenuation from LLM reconstruction error; implement SIMEX or reliability-adjusted IV using human-audited subsamples; test robustness to alternative susceptibility taxonomies and treatment of “unknown” exams.
  • Susceptibility construct validity: Enrich the construct to account for partially susceptible formats (timed online quizzes, laptop-allowed exams, group work), real-time AI access constraints, and discipline-specific automation potential.
  • Syllabi vs practice gap: Validate that syllabi reflect implemented assessment weights/formats (random audits of gradebooks, assignment prompts, proctoring records) to bound misclassification.
  • Evaluation data limitations: Adjust for changing response rates/composition over time; analyze full response distributions instead of medians; apply item-response models to address scale drift and comparability across versions.
  • Grade-scale coarseness: Where possible, use numeric/percentage grades or assignment-level scores to detect subtle shifts masked by letter-grade bins.
  • Integrity and enforcement outcomes: Incorporate academic-integrity reports, AI-detection referrals, sanctions, and proctoring-flag rates to map deterrence/mitigation channels.
  • Social-learning channel: Measure shifts in help-seeking (office hours, tutoring/writing center usage, LMS forum activity, peer-network surveys) to test the hypothesized social displacement by GenAI.
  • Instructor AI use: Track instructors’ own GenAI use (assignment design, grading, feedback generation) as a confounder that could alter both grades and satisfaction.
  • Temporal dynamics: Estimate term-by-term post effects (Winter 2023 vs Fall 2023 vs 2024–2025), acknowledging rapid capability and adoption changes (GPT-4/4o, Claude, Copilot) and evolving institutional policies.
  • Spillovers across courses: Test whether GenAI-induced time savings in susceptible courses reallocate effort to other courses (general-equilibrium spillovers on grades and evaluations elsewhere).
  • Robustness to course-sample selection: Assess how restricting to courses with 2019 syllabi biases the analytic sample toward larger/more stable offerings; reweight or impute missing susceptibility where feasible.
  • Unknown-exam handling: Evaluate sensitivity to classifying “unknown exam” as non-susceptible; sample and hand-code unknowns to estimate misclassification direction and magnitude.
  • Causal mechanisms for nulls: Decompose null average effects into offsetting forces (student substitution vs instructor mitigation vs grading norms) via structural or mediation analysis using time-varying policy and format data.
  • Cross-institution policy experiments: Leverage natural experiments (department-level AI policy rollouts, proctoring tech adoption, device bans, standardized AI-permitted assignments) for cleaner causal identification.
  • Reproducibility and LLM dependence: Document prompt templates, model versions, and guardrail behaviors; assess reconstruction stability across LLMs and over time; release redacted syllabi excerpts and code where permissible.
  • Long-run cohort effects: Follow cohorts who experienced AI throughout K–12 into university to detect whether baseline preparation, study habits, and susceptibility responses differ from earlier cohorts.
  • Non-cognitive outcomes: Measure motivation, self-efficacy, academic integrity norms, and AI literacy to understand behavioral adaptations that grades/evaluations may not capture.
  • International contexts and languages: Examine settings with different assessment cultures and non-English instruction, where GenAI capabilities and susceptibility may differ materially.

Practical Applications

Immediate Applications

The following applications can be deployed with today’s tools and data practices, leveraging the paper’s findings (no systematic GenAI-driven grade inflation in susceptible courses; no robust decline in student satisfaction) and its validated LLM-based syllabus parsing pipeline.

  • Campus-wide “GenAI Susceptibility” audit and dashboard
    • Sector(s): Higher education administration; Institutional research; EdTech (LMS vendors)
    • What: Use the validated LLM pipeline to parse syllabi into structured assessment-weight data, compute course-level susceptibility, and monitor outcomes (grades, withdrawals, evaluations) over time.
    • Tools/workflows: Syllabus-to-JSON extractor (PDF → Markdown → structured labels); continuous susceptibility scoring; DiD/event-study monitoring by term; course/department roll-ups; risk flags.
    • Assumptions/dependencies: Centralized syllabus archive access; data-sharing agreements; LLM privacy/compliance; stable versioning of prompts/models; faculty participation in timely syllabus uploads.
  • Evidence-based assessment policy and resource allocation
    • Sector(s): Academia (departments, teaching & learning centers)
    • What: Temper blanket restrictions on take-home/problem-set/essay assessments; prioritize in-person demonstrations or proctoring only where susceptibility and risk indicators are high; align integrity interventions with measured risk rather than fear of inflation.
    • Tools/workflows: Departmental susceptibility reports; rubric templates for AI-aware assignments; triage playbook for proctoring resources; faculty workshops.
    • Assumptions/dependencies: Faculty governance buy-in; local integrity norms; accessibility/accommodation constraints; continued monitoring of outcomes post-policy change.
  • Admissions, advising, and employer signaling calibration
    • Sector(s): Admissions; Career services; Employers (finance, tech, consulting, public sector)
    • What: Use null grade-inflation results to avoid overcorrecting GPA-based decisions; supplement GPA with course rigor/susceptibility context where available; communicate that post-2022 GPAs remain broadly comparable in aggregate.
    • Tools/workflows: GPA context tags (e.g., % credits in low/high-susceptibility courses); employer briefings; updated recommendation letter guidance.
    • Assumptions/dependencies: Access to course metadata; employer willingness to integrate contextual fields; careful messaging to avoid stigmatizing certain fields.
  • Course evaluation interpretation and student-experience QA
    • Sector(s): Academia; Quality assurance; Accreditation
    • What: Re-interpret post-2022 evaluation trends using the paper’s COVID-persistence lens; track understanding/interest/workload medians by susceptibility without presuming GenAI harm.
    • Tools/workflows: Evaluation analytics with “COVID window” vs “persistent effect” toggles; department scorecards; item-level trend alerts.
    • Assumptions/dependencies: Stable evaluation instruments; sufficient response rates; disaggregations that protect student privacy.
  • LLM-based syllabus management products
    • Sector(s): EdTech; Publishing; LMS
    • What: Productize the annotated syllabus pipeline as: (a) “Syllabus QA” (detect missing policies, weight mis-sums), (b) “Assessment Mix Optimizer” (benchmark against peers), and (c) “AI-readiness Badging” (transparent use-of-AI policies).
    • Tools/workflows: API for syllabus ingestion; admin console; faculty-facing suggestions; integrations with Canvas/Moodle/Blackboard.
    • Assumptions/dependencies: Model reliability across formats/disciplines; institutional security reviews; clear UX for faculty edits and human-in-the-loop correction.
  • Replicable monitoring studies across institutions
    • Sector(s): Research; Institutional consortia; Policy labs
    • What: Re-run the paper’s DiD/event-study design at peer institutions; share benchmarks; publish annual “GenAI & Outcomes” reports.
    • Tools/workflows: Shared codebooks; pre-registered analysis plans; crosswalks for grading scales; consortium data enclaves.
    • Assumptions/dependencies: Inter-institutional data MOUs; IRB approvals; consistent anchoring year (pre-COVID) for susceptibility.
  • Student-facing AI usage guidance that emphasizes learning
    • Sector(s): Student success; Daily life
    • What: Guidance that positions GenAI as a complement (planning, scaffolding, feedback) rather than a substitute; self-checks before closed-book assessments.
    • Tools/workflows: Study-planning prompts; retrieval practice checklists; “AI reflection” logs tied to assignments.
    • Assumptions/dependencies: Instructor alignment; availability of low-friction tools; attention to equity in access.

Long-Term Applications

The following require further research, scaling, or development (e.g., multi-institution data standards, sustained policy shifts, or new product categories).

  • National or system-level “AI Susceptibility and Outcomes” observatory
    • Sector(s): State/federal education agencies; Accreditation; Research consortia
    • What: Standardize syllabus schemas; mandate periodic susceptibility audits; longitudinally track grades/evaluations/withdrawals with COVID-adjustment options.
    • Tools/workflows: Open syllabus metadata standards; secure data pipelines; public dashboards with stratifications (field, level, assessment mix).
    • Assumptions/dependencies: Policy mandates; funding; harmonized privacy regimes; robust governance for model drift and re-anchoring beyond 2019.
  • Dynamic causal monitoring platforms in LMS/analytics suites
    • Sector(s): EdTech; Analytics
    • What: Built-in event-study/DiD modules to quantify impacts of shocks (AI, policy changes) with bookended disruptions (e.g., persistent vs transient COVID effects).
    • Tools/workflows: Causal inference widgets; auto-generated parallel-trend diagnostics; semester-over-semester alerts.
    • Assumptions/dependencies: High-quality time-stamped data; trained users; guardrails against misinterpretation.
  • Assessment redesign toward demonstrable mastery and authentic tasks
    • Sector(s): Academia; EdTech; Professional certification
    • What: Scale in-person demonstrations, oral exams, code reviews, viva-style defenses, and authentic projects that resist pure automation but allow AI-augmented preparation.
    • Tools/workflows: Rubrics for authenticity and collaboration; scalable oral assessment scheduling tools; AI-supported formative feedback that fades assistance over time.
    • Assumptions/dependencies: Faculty workload support; class-size constraints; accessibility considerations; cultural shift in assessment norms.
  • AI-augmented formative learning tools with “assistance tapering”
    • Sector(s): EdTech; K–12 and higher education
    • What: Tutors that scaffold problem solving and progressively reduce help, measuring unassisted performance to protect learning while allowing efficiency.
    • Tools/workflows: Adaptive prompting; metacognitive checkpoints; unassisted challenge phases; learning analytics tied to later closed-book results.
    • Assumptions/dependencies: Rigorous efficacy trials; guardrails against over-reliance; interoperability with LMS gradebooks.
  • Employer-side recalibration of academic signals
    • Sector(s): Labor markets (tech, finance, healthcare, public sector)
    • What: ATS and hiring rubrics that incorporate course-context metadata (assessment mix, level, field) and emphasize samples from non-susceptible assessments, portfolios, and proctored certifications.
    • Tools/workflows: GPA-plus context feeds from institutions; secure transcript APIs; skills assessments aligned to closed-book competencies.
    • Assumptions/dependencies: Data-sharing standards; legal/privacy compliance; employer adoption incentives.
  • Policy frameworks that require AI-impact audits but avoid punitive overreach
    • Sector(s): University governance; Accreditation; Legislatures
    • What: Policies that (a) require periodic AI-impact analyses using validated methods, (b) publish transparency reports, and (c) support faculty development rather than blanket bans.
    • Tools/workflows: Audit templates; public reporting norms; continuous improvement cycles.
    • Assumptions/dependencies: Stable governance; funding for capacity building; stakeholder trust.
  • Generalization to other educational contexts (community colleges, MOOCs, K–12)
    • Sector(s): Education systems; Online learning platforms
    • What: Adapt susceptibility definitions to different modalities (e.g., auto-graded quizzes, discussion forums), re-validate LLM parsers, and re-estimate effects.
    • Tools/workflows: Context-specific label ontologies; multilingual parsing; platform-native telemetry.
    • Assumptions/dependencies: Heterogeneous student populations; platform data access; revised anchoring years.
  • Institutional data engineering and governance upgrades
    • Sector(s): University IT; Data governance
    • What: Make syllabi and assessment schemas first-class, machine-readable data assets; enable reproducible analytics with versioning and audit trails.
    • Tools/workflows: Syllabus APIs; document version control; prompt/version registries; human-in-the-loop correction pipelines.
    • Assumptions/dependencies: Investment in data infrastructure; clear ownership; ongoing model evaluation as LLMs evolve.
  • Research on heterogeneous effects and social dynamics
    • Sector(s): Academia; Social science research
    • What: Extend prior-performance heterogeneity analyses to metacognition, help-seeking networks, field-specific norms; measure long-run learning via unassisted capstones/licensure.
    • Tools/workflows: Linked survey/behavioral telemetry; network analyses; quasi-experiments with policy variation.
    • Assumptions/dependencies: Longitudinal data; IRB approvals; careful causal identification strategies.
  • Alternative credentialing that certifies unassisted competence
    • Sector(s): Professional bodies; Testing/certification; Employers
    • What: Micro-credentials or badges based on supervised, AI-restricted tasks validating mastery independent of AI assistance.
    • Tools/workflows: Secure proctoring; oral/practical exams; blockchain-backed credential registries.
    • Assumptions/dependencies: Stakeholder acceptance; psychometric validation; equitable access.

Notes on overarching assumptions

  • Generalizability: Findings are from a single large U.S. university; replication elsewhere is advised before high-stakes decisions.
  • Modeling choice: COVID effects may be persistent; tools should allow both “transient” and “persistent” scenarios.
  • Data sensitivity: Student records and syllabi require strict privacy, governance, and IRB oversight.
  • Model drift: LLM extraction accuracy may vary with future model changes; maintain validation sets and human review loops.

Glossary

  • Active learning: An instructional approach emphasizing student engagement and effort to enhance learning. "This GenAI substitution hypothesis supposes exerting effort improves learning (i.e., ``active learning,'' see \citet{Prince2004-cw})."
  • Clustered standard errors: A method of computing standard errors that accounts for correlation within clusters (e.g., courses) to avoid overstating precision. "Standard errors are clustered at the course level."
  • Cognitive offloading: Shifting mental effort to external tools or aids, potentially reducing deep learning. "a key concern is that students are offloading the cognitive effort necessary for learning (see \citet{baldeo_generative_2026} and \citet{gerlichAIToolsSociety2025} on ``cognitive offloading'')."
  • Confusion matrix: A table summarizing correct and incorrect classifications across categories in a prediction/labeling task. "Appendix~\ref{app:LLM-pipeline} reports the full validation summary, confusion matrices, and category-by-category evaluation details."
  • Difference-in-differences (DiD): A causal inference design comparing changes over time between treated and control groups to estimate an intervention’s effect. "and use a differences-in-differences design comparing outcomes across courses before and after ChatGPT's release, while modeling COVID-19 pandemic effects as either persistent or transient."
  • Event study: An analysis tracing dynamic effects over time around a specific event or shock. "Observe the average grades and withdrawal rate event studies presented in Figures \ref{fig:es-grade} and \ref{fig:es-withdraw-fail}."
  • Exogenous: Arising from external factors not determined by the model or system, often implying as-good-as-random variation. "courses varied in their susceptibility to GenAI use for direct substitution of student ability in a plausibly exogenous way before ChatGPT."
  • Fixed effects: Controls in panel models that absorb time-invariant characteristics of entities (e.g., courses, students) or periods to reduce omitted-variable bias. "we use course, semester, and, where possible, student fixed effects."
  • GenAI susceptibility: The extent to which a course’s assessments can be completed with generative AI in ways that may substitute for student effort. "We define a course's GenAI susceptibility as the proportion of the final grade allocated to susceptible assessments."
  • Grade inflation: A rise in average grades over time without corresponding increases in learning or ability. "To account for common trends, like overall grade inflation, or changes to student composition, we use course, semester, and, where possible, student fixed effects."
  • Heterogeneous effects: Differences in treatment effects across subgroups or parts of the distribution. "Null average effects could mask underlying heterogeneous effects, so we also estimate grade effects on grade distributions (i.e., share of A's versus share of D's) and between students based on prior academic performance measured six ways."
  • Human capital: The stock of skills and knowledge that increases individuals’ productivity. "Education aims to build human capital, not just to produce outputs, which often requires productive struggle."
  • Likert scale: A psychometric response scale (typically 5 points) used in questionnaires to measure attitudes or perceptions. "The standard course evaluation survey includes required five-point Likert-scale questions and additional instructor-selected questions, each with open-ended text boxes."
  • Marginal effects: The change in an outcome associated with a one-unit change in a predictor, holding other factors constant. "We use linear models throughout for computational efficiency and interpretability of coefficients as marginal effects."
  • Mean absolute error: The average absolute difference between predicted and true values, measuring reconstruction or prediction accuracy. "Continuous Susceptibility is reconstructed with mean absolute error of 0.063 on the 0--1 scale."
  • Mediator: A variable that transmits part of the effect of a treatment to an outcome. "Mediator: Prior Academic Preparation"
  • Ordinary least squares (OLS): A method for estimating linear regression parameters by minimizing the sum of squared residuals. "We estimate three progressive specifications that build from a simple ordinary least-squared (OLS) baseline to our preferred two-way fixed effects (TWFE) design with an explicit COVID-window adjustment."
  • Panel (data): Multi-dimensional data following the same units over time. "in his 2018-2025 panel, which removes complicated seminar courses and accounts for cyclicality in offerings."
  • Quasi-experimental evidence: Empirical evidence from study designs that approximate randomized experiments using observational data. "Quasi-experimental evidence at the scale of a full course catalog remains scarce, but a small literature has begun to form."
  • Residualization: Removing the influence of specified covariates from a variable (e.g., subtracting fitted values) before further analysis. "we residualize each first-term grade against its course-by-term mean before averaging and ranking."
  • Selection bias: Systematic differences between groups being compared that can bias estimated effects. "This allows courses to switch treatment groups from year to year, which introduces selection bias into estimates of GenAI's impact."
  • Stratified sample: A sampling design or evaluation subset constructed by dividing the population into subgroups (strata) and sampling within each. "We validate the pipeline against a stratified human-annotated sample of 525 syllabi (two trained coders per syllabus with consensus resolution) and find strong recovery."
  • Treatment group: The set of units exposed to the intervention or condition of interest in a causal study. "This allows courses to switch treatment groups from year to year, which introduces selection bias into estimates of GenAI's impact."
  • Two-way fixed effects (TWFE): A panel-data model including both unit (e.g., course) and time (e.g., semester) fixed effects to control for invariant unit traits and common shocks. "our preferred two-way fixed effects (TWFE) design with an explicit COVID-window adjustment."
  • Within-cohort rank: A ranking of individuals relative to peers who entered in the same cohort, often used for comparability across time. "we measure students' prior academic preparation with a course-residualized, within-cohort first-term GPA rank (AbilityRank)."

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 5 tweets with 103 likes about this paper.