StudentBench: AI and human tutoring yield equivalent GRE learning gains
Abstract: Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether LLMs produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average. In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations. Together, the two studies clearly separate AI tutors across: (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement. Surprisingly, one AI tutor achieved learning gains equivalent to human tutoring (p = .044) at 918 times lower cost (USD 0.0052 for AI versus USD 4.81 for human, per percentage point gained). For Quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains (all p < .002). The StudentBench platform is freely available at https://studentbench.org.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
The paper studies whether AI tutors can help students learn as well as human tutors.
The researchers created StudentBench, a system for testing AI teaching. They used questions from the GRE, an exam taken by many people applying to graduate school. The study compared three groups:
- Students taught by an AI tutor
- Students taught by an expert human tutor
- Students who received no GRE tutoring
The main question was: After one hour of tutoring, did students improve just as much with AI as with a human?
2. What were the researchers trying to find out?
The researchers asked several important questions:
- Can AI tutoring improve students’ GRE scores?
- Does AI tutoring work about as well as expert human tutoring?
- Which AI systems are best at planning lessons and creating practice questions?
- How much does AI tutoring cost compared with human tutoring?
- Does the speed of an AI tutor’s replies affect how much students participate and learn?
The paper is also connected to a bigger idea called recursive human self-improvement. This means that technology could help people learn more effectively, and better-educated people could then use and improve that technology even more.
3. How did the researchers conduct the study?
The tutoring experiment
The researchers studied 2,383 students and collected information from more than 175,000 student–AI messages.
Students first took a pre-test. This showed what they already knew. The test included either:
- Quantitative questions, involving mathematics and problem-solving, or
- Verbal questions, involving reading and language skills.
After the pre-test, students were randomly placed into one of three groups:
- AI tutoring
- Live tutoring from an expert human
- A control group that watched unrelated educational videos
The tutoring lasted about one hour. The AI tutors could:
- Plan a lesson
- Explain ideas
- Ask questions
- Create practice problems
- Give feedback on student answers
After tutoring, students took a different but similar post-test. The researchers calculated the learning gain by subtracting the pre-test score from the post-test score:
For example, if a student scored 60% before tutoring and 70% afterward, the learning gain would be 10 percentage points.
The researchers used different questions on the two tests so students could not simply memorize the answers. They also checked for low effort, very fast answers, excessive tab switching, and other signs that a session might not provide trustworthy data.
Comparing different AI tutors
The study tested 13 AI tutors based on different LLMs. A LLM is an AI system trained to understand and produce text, such as explanations and answers.
In a second study, 51 expert human tutors compared AI-created:
- Lesson plans
- Practice questions
- Answer keys
The experts judged whether the materials were relevant, well organized, appropriately difficult, and accurate.
The researchers also examined tutoring conversations. They looked for behaviors such as whether the tutor:
- Asked students to explain their thinking
- Let students try before showing the answer
- Gave hints connected to the student’s mistakes
- Asked useful follow-up questions
Understanding the statistics
The researchers used statistical tests to determine whether differences were likely to be real rather than caused by chance.
They used an equivalence test, which asks a slightly different question from a normal test. Instead of asking, “Are AI and human tutoring different?” it asks, “Are they close enough that we can reasonably treat them as equally effective?”
The paper also reports numbers such as p = .015. A p-value is a measure of how surprising the results would be if there were no meaningful effect. Smaller values generally provide stronger evidence that the finding is not just random chance.
4. What did the researchers discover?
AI and human tutors produced similar learning gains
The main result was that AI tutoring and expert human tutoring produced statistically equivalent GRE learning gains when their results were combined across the Quantitative and Verbal sections.
The difference between AI and human tutoring was very small: AI learners improved by about 0.58 percentage points less than human-tutored learners. The researchers judged this difference small enough to count the two types of tutoring as equivalent.
Both AI and human tutoring worked better than the no-tutoring control group. Compared with the control group, AI tutoring improved results by about:
- 6.86 percentage points in Quantitative questions
- 5.47 percentage points in Verbal questions
In simple terms, students who received tutoring answered roughly 1.5 to 2 more questions correctly than students in the control group.
Some AI tutors performed better than human tutors in certain areas
The results were not identical in every subject.
- In Quantitative topics, the strongest AI tutors performed better than human tutors in three of the four areas.
- In Verbal topics, human tutors had the highest overall average, although some AI tutors were close.
- Across the seven GRE topic areas, the best AI tutor performed better than the human average in five areas.
This suggests that AI may be especially strong at explaining certain mathematical ideas, while human tutors may still have an advantage in some language-related areas.
AI tutoring was much cheaper
One of the biggest differences was cost.
The study estimated that a human tutor cost about $75 per hour. AI tutoring was much cheaper because it mainly required computer processing.
One AI tutor, called Gemma 4 31B in the paper, produced learning gains equivalent to those of human tutoring while costing about 918 times less per percentage point of learning gain.
This does not mean every AI tutor is equally good or that every student would receive the same results. However, it suggests that AI tutoring could make individualized help available to many more students, including students who cannot afford private tutoring.
Faster AI replies were linked to more learning activity
The researchers also found that faster AI responses were connected with:
- More student messages
- More correct practice answers
- Larger learning gains
This pattern was especially clear for Quantitative tutoring.
A possible explanation is that students lose interest when they have to wait too long. Faster replies may keep the conversation moving, giving students more chances to practice.
However, these findings show relationships, not definite proof that faster replies directly caused better learning. Other factors might also be involved.
Different AI models had different teaching strengths
The AI tutors did not all behave the same way.
Some models were especially good at:
- Planning lessons
- Choosing which topics to teach
- Creating suitable practice problems
- Showing helpful teaching behaviors in conversations
The expert tutors especially liked several models from the Anthropic family for planning lessons and writing practice problems. However, strong performance on these expert ratings did not always perfectly predict the largest learning gains in the student tests.
This is important because an AI can appear to give good explanations to an expert, but the best test is still whether students actually learn.
5. Why are these findings important?
The study suggests that AI tutors can be more than tools that simply provide answers. With the right interaction, they may help students understand ideas, practice skills, and improve their test performance.
The possible benefits include:
- Lower costs: More students could receive one-on-one help.
- Greater access: Students could get tutoring at any time, even when human tutors are unavailable.
- Personalized lessons: AI can focus on the mistakes a student made.
- Quick feedback: Students can receive explanations without waiting.
- Support for human teachers: AI could handle some practice and explanation tasks while teachers focus on more difficult or personal parts of learning.
The paper’s larger message is that AI might support and strengthen human abilities rather than simply replace people.
6. Important limitations
The results should not be treated as proof that AI is always as good as a human teacher.
The study had several limits:
- It measured learning only immediately after tutoring. We do not know whether students remembered the material months later.
- The participants were adults who could read and write English and had access to the study platform.
- The study focused on GRE-style questions, so the results may not apply to every subject or age group.
- The researchers did not compare AI tutoring with students simply practicing GRE questions on their own.
- AI models change quickly, so newer versions may perform differently.
- AI-generated practice questions can sometimes contain mistakes. Students in the study were able to flag answers they believed were wrong.
- The study was funded by Handshake AI, the organization connected to the research, so independent studies would be useful.
Conclusion
The StudentBench paper found that, in this study, AI tutors helped students improve their GRE scores about as much as expert human tutors after one hour. Some AI tutors even performed better than human tutors in particular subject areas.
The biggest advantage of AI was its very low cost. If these results hold in other subjects and for longer periods, AI tutoring could help make personalized education available to many more people.
Still, AI tutoring needs careful testing. Future research should examine whether students remember what they learn, how AI works with human teachers, and whether the same results appear for younger students, different languages, and subjects beyond the GRE.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- Long-term retention is unresolved: The study measures only immediate post-test gains, so it remains unknown whether AI- or human-tutored learning persists over weeks or months.
- Transfer beyond the assessment is untested: The paper does not establish whether tutoring improves performance on novel problem types, real GRE scores, broader academic tasks, or practical reasoning outside the study’s post-test.
- No practice-only comparison was included: Because the control group watched unrelated educational videos, the study cannot isolate the incremental benefit of AI conversation from independently solving GRE practice problems with equivalent time and materials.
- Generalizability is limited by the sample: Participants were English-speaking adults recruited through Handshake, primarily aged 18–23, and paid for participation; effects may differ for younger students, older learners, nonstudents, lower-literacy populations, or learners with disabilities.
- Language and cultural transferability remain unknown: The study does not test AI tutoring in languages other than English or determine whether its effectiveness depends on English-language proficiency, cultural conventions, or familiarity with U.S.-style standardized testing.
- Educational-context generalization is untested: Results from one-hour, one-to-one GRE sessions do not show whether AI tutoring is effective in classrooms, schools, universities, tutoring centers, home environments, or longer curricula.
- Device and access constraints were not evaluated: The paper does not examine performance under mobile-only access, low bandwidth, limited computing resources, intermittent connectivity, or accessibility technologies.
- The human-tutor benchmark is underspecified: More evidence is needed on human tutors’ qualifications, teaching protocols, adherence to the study design, and variability in tutoring quality to determine which level of human tutoring AI matched.
- AI–human comparisons may be affected by modality differences: AI tutoring was text-based whereas human tutoring occurred through live video calls; the study does not separate tutor capability from differences in speech, visual cues, typing demands, social presence, or interaction modality.
- The effects of tutor identity and disclosure are unknown: Students were not told which AI model they received, but the paper does not test whether disclosure that a tutor is AI changes trust, engagement, persistence, learning, or willingness to follow advice.
- Mechanisms of learning remain uncertain: The observed association between latency, engagement, correct practice, and learning gain is correlational; randomized manipulation of response speed is needed to determine whether latency causes improved learning.
- Engagement measures are incomplete and section-dependent: Student message count may not represent cognitive engagement consistently, especially because Verbal students sent fewer but longer messages; future work should use validated measures of reasoning quality, attention, metacognition, and off-task behavior.
- More practice may not imply deeper learning: The study counts correct practice but does not determine whether students solved problems independently, relied on hints, memorized procedures, or developed transferable understanding.
- AI-generated content quality is not fully verified: Student answer-dispute flags were not independently checked for correctness, leaving the true rates of erroneous explanations, practice problems, and answer keys uncertain.
- Safety and pedagogical failure modes are underexplored: The paper does not systematically evaluate hallucinations, misleading explanations, inappropriate confidence, harmful feedback, privacy risks, manipulation, student dependence, or failures involving atypical misconceptions.
- Student heterogeneity is insufficiently characterized: It remains unclear which students benefit most or least by prior achievement, socioeconomic status, motivation, test anxiety, learning differences, demographic characteristics, or baseline AI familiarity.
- The high-performing-student finding is exploratory: The apparent advantage of Gemini models for students in the top proficiency quartile requires preregistered replication with adequate power and interaction tests across models and proficiency levels.
- Domain-level conclusions may be underpowered or confounded: Claims that particular AI tutors outperform humans in five of seven domains may reflect multiple comparisons, unequal sample sizes, model-by-domain interactions, or domain-specific assessment properties.
- Individual tutor equivalence results need multiplicity control: Individual AI tutors were tested for equivalence without correction across tutors, so some apparent human-equivalent tutors may be false positives.
- The equivalence margins are consequential but not empirically validated here: The choice of pooled standard deviations, and the interpretation of smaller margins, may not correspond to a practically meaningful GRE improvement for students or admissions outcomes.
- Very low effective degrees of freedom weaken pooled inference: The combined equivalence test reports 1.59 degrees of freedom; the robustness of this inference under alternative clustering, weighting, and hierarchical models remains to be established.
- Assessment validity and comparability require further evidence: Although new questions were designed to match prior GRE tests, the paper does not provide enough evidence on item calibration, parallel-form equivalence, item exposure, differential item functioning, or whether pre-/post-test differences reflect learning rather than form difficulty.
- Possible test–retest and content-practice effects remain uncertain: Students reviewed pre-test mistakes and then received related tutoring, so some post-test gains may reflect short-term familiarity or strategy practice rather than durable conceptual learning.
- Selection and attrition effects are not fully resolved: The extensive exclusion criteria may remove disengaged, struggling, or technologically constrained participants and could produce estimates that do not apply to typical users.
- The no-tutoring condition may not be a neutral baseline: Watching unrelated videos may differ from waiting, resting, or completing ordinary study activities in motivation and cognitive load, complicating interpretation of the AI- and human-versus-control effects.
- Cost comparisons are sensitive to assumptions: The $75-per-hour human rate, incomplete coverage of AI API costs, reconstructed costs for 3% of calls, and exclusion of platform, monitoring, development, support, and infrastructure costs may substantially affect the reported$918\times$ advantage.
- Cost-effectiveness over a complete learning program is unknown: Session-level cost per percentage point does not capture repeated sessions, remediation, onboarding, human oversight, model downtime, student dropout, or the cost of achieving a target GRE score.
- Model comparisons are not fully controlled for: Models differed in reasoning settings, prompts, availability dates, latency, context handling, and possibly token limits; factorial experiments are needed to separate model capability from configuration and implementation effects.
- Rapid model changes threaten reproducibility: Because models and prices changed during the two-month study, future researchers may not be able to reproduce the exact tutor conditions or determine whether results reflect stable model properties.
- Minimal scaffolding limits conclusions about deployable tutoring systems: The study intentionally uses low-guidance prompts, but it does not determine how structured curricula, retrieval, knowledge tracing, misconception models, tool use, or adaptive interfaces would change learning and cost.
- The teaching leaderboards are not validated against learning outcomes: Expert preferences for lesson plans and practice problems, as well as six transcript indicators, may not predict student learning, retention, equity, or satisfaction.
- Human-reference conversational scores are not directly comparable: AI sessions were typed chat and human sessions were spoken tutoring, so differences in the six pedagogical indicators may reflect transcription, turn-taking, or modality rather than pedagogy.
- Reviewer reliability and representativeness need further study: The paper does not establish how consistently the 51 expert tutors applied the rubrics, whether ratings generalize beyond GRE experts, or whether expert preferences align with student outcomes.
- Student satisfaction and affective outcomes are largely absent: The study does not report trust, enjoyment, frustration, confidence, motivation, perceived usefulness, social connection, or willingness to continue using AI tutoring.
- Equity effects remain unknown: Lower monetary cost does not guarantee equitable access; the paper does not examine disparities caused by language, disability, digital literacy, broadband access, algorithmic bias, or differential ability to formulate questions.
- The study does not test human–AI collaboration: The conclusions concern autonomous AI or standalone human tutoring, leaving unresolved whether AI assistance can improve human tutor effectiveness, increase tutor capacity, or produce better outcomes than either condition alone.
- The RHSI framework is only partially evaluated: The study tests the first proposed RHSI condition—improved human capability—but does not examine whether improved learners use AI more effectively or whether such use produces cumulative, recursive gains over multiple cycles.
- Security and privacy implications of real deployment are unexplored: The paper does not evaluate data retention, student consent in ongoing use, prompt injection, unauthorized disclosure, model-provider access to educational records, or governance requirements for large-scale deployment.
- Effects on official educational outcomes remain unknown: The study uses research assessments and does not show whether AI tutoring improves admissions, course grades, graduation, persistence, or other consequential educational outcomes.
Practical Applications
Immediate Applications
- Low-cost standardized-test preparation platforms (education technology, consumer software).
- analyze a learner’s pre-test errors;
- generate a prioritized lesson plan;
- create targeted practice problems and answer explanations;
- conduct interactive, one-hour tutoring sessions; and
- administer unaided pre- and post-tests to measure improvement.
- The study reports AI–human-equivalent combined GRE learning gains, with one evaluated tutor achieving comparable gains at roughly 918 times lower cost per percentage point gained than the human-tutoring reference.
- Dependencies: The reported evidence is specific to English-speaking adults, GRE-like tasks, text-based tutoring, and immediate post-test gains. Systems should verify generated answers and avoid using secure or leaked examination content.
- AI tutoring as an access and affordability tool for underserved learners (education policy, nonprofits, universities). Schools, libraries, scholarship programs, and workforce organizations can provide AI tutoring where one-to-one human instruction is unavailable or unaffordable. A practical workflow would combine a short diagnostic test, personalized chat tutoring, targeted practice, and escalation to a human tutor for persistent misconceptions. Dependencies: Access to reliable devices and internet, culturally and linguistically appropriate content, privacy protections for student data, and independent monitoring of learning quality. AI should supplement—not automatically replace—teachers and expert tutors.
- Human-tutor augmentation and preparation (education services). Tutoring companies and academic-support centers can use LLMs to draft lesson plans, prioritize student misconceptions, generate practice questions, and suggest pacing. Human tutors can review and adapt these materials before use, reducing preparation time while retaining human oversight. Dependencies: Expert review remains important because the paper documents student disputes of AI-generated answers and evaluates lesson-plan quality separately from actual learning outcomes. Generated problems require automated and human validation for correctness, difficulty, and curricular alignment.
- AI-assisted GRE and SAT preparation in universities (higher education). Graduate-school advising offices, writing and quantitative-skills centers, and career services can integrate StudentBench-like tutoring into existing preparation programs. Quantitative domains appear particularly suitable: the best AI tutors exceeded the human mean in several quantitative domains, and the AI–human equivalence result was statistically supported for Quantitative tutoring. Dependencies: The paper did not test official examination outcomes, long-term retention, or the SAT directly. Institutions should run local validation studies before using AI performance as evidence of admissions readiness.
- Latency-aware tutoring product design (software and learning-platform engineering). AI tutoring interfaces should prioritize fast responses, streaming output, concise turn-taking, and low-latency infrastructure. In Quantitative sessions, faster replies were associated with more student messages, more correct practice, and larger learning gains. Product teams can therefore treat response time as an educational metric rather than merely a user-experience metric. Dependencies: The latency findings are correlational, not proof that reducing latency alone causes learning. Excessively brief or superficial responses could reduce explanation quality, and the relationship was not statistically significant for Verbal tutoring.
- Learning-gain benchmarking for AI model selection (AI industry and procurement).
- unaided pre-to-post learning gain;
- cost per percentage point of gain;
- response latency;
- student engagement;
- lesson-plan quality;
- practice-problem quality; and
- pedagogical behaviors such as eliciting explanations and allowing students to attempt solutions.
- This is more actionable than selecting a model solely by general language or reasoning benchmarks. StudentBench’s open data, code, and platform can support internal replication and model comparisons.
- Dependencies: Model rankings may vary by subject, student population, prompting, inference settings, and software scaffolding. Cost estimates also depend on provider pricing and token usage.
- Research methods for evaluating educational AI (academic research). Researchers can immediately reuse the paper’s experimental design: randomized assignment to AI tutoring, human tutoring, and control conditions; counterbalanced pre- and post-tests; ANCOVA adjustment; equivalence testing; and analysis of cost per learning gain. The platform also supports large-scale collection of real student–AI conversations rather than relying only on simulated users or LLM judges. Dependencies: Valid assessments must be independent of model training data, and studies need safeguards against practice contamination, answer leakage, low-effort participation, and test-retest effects.
- Public and institutional monitoring of AI tutoring quality (policy and educational governance). Regulators, accrediting bodies, and schools can require vendors to report learning outcomes and cost-effectiveness, rather than only model accuracy or satisfaction scores. A practical procurement requirement could be evidence that students perform better on an unaided assessment after tutoring, with subgroup and error-rate reporting. Dependencies: Equivalence to human tutoring in one domain should not be generalized to all subjects or populations. Governance should also address data retention, transparency, accessibility, academic integrity, and the risk that AI-generated explanations reinforce misconceptions.
- Personal study assistance in daily life (consumer use). Students can use an AI tutor to review missed questions, request explanations, generate analogous problems, and practice until they can solve problems independently. The study supports a workflow centered on active student responses rather than passive answer consumption. Dependencies: Users should independently verify uncertain answers, avoid copying solutions, and periodically test themselves without AI assistance. Immediate score gains may not persist without spaced practice and later reassessment.
Long-Term Applications
- General-purpose AI tutors across academic subjects and professional exams (education, healthcare, finance, and workforce training). The demonstrated workflow could be adapted to medical entrance examinations, nursing certification, accounting, coding, legal education, language learning, and technical workforce training. Domain-specific tutors could combine diagnostic assessment, concept sequencing, generated practice, and adaptive dialogue. Dependencies: The paper directly studies only GRE Quantitative and Verbal reasoning. High-stakes domains require subject-matter validation, professional oversight, robust factuality testing, and evidence of long-term retention and transfer to real tasks.
- Longitudinal personalized learning systems (education technology). Future platforms could maintain a learner model over months or years, using knowledge tracing and repeated assessments to schedule review, detect forgetting, and personalize difficulty. StudentBench provides the immediate learning-gain component; longitudinal systems would add retention, transfer, and progression measures. Dependencies: The paper explicitly does not measure retention beyond the immediate post-test. Long-term deployment requires consent, secure student profiles, interoperability with learning-management systems, and safeguards against inaccurate or overconfident learner profiling.
- AI–human tutoring triage systems (schools, universities, and tutoring companies). A scalable model could assign routine explanation and practice tasks to AI, while routing students to human educators when they show persistent errors, emotional distress, accessibility needs, or unusually low progress. Human tutors could review dashboards containing learning gains, engagement, disputed answers, and unresolved misconceptions. Dependencies: Escalation criteria must be validated and should not systematically deprioritize students with disabilities, limited English proficiency, low connectivity, or atypical communication styles. Student engagement metrics cannot be treated as a direct proxy for learning.
- Adaptive tutoring infrastructures with latency–learning optimization (AI systems and cloud computing). Educational platforms could jointly optimize model selection, inference cost, response speed, and learning outcomes. For example, a fast low-cost model might handle routine dialogue, while a more capable model or human tutor handles difficult misconceptions. Dependencies: The observed associations do not establish a universal causal latency threshold. Routing systems must preserve pedagogical continuity, prevent inconsistent explanations, and verify that cost reductions do not lower learning quality.
- Automated generation and validation of educational assessments (publishing and assessment technology). LLMs could generate large pools of practice problems matched to diagnosed misconceptions and calibrated by difficulty. StudentBench’s practice-problem rubric—alignment, difficulty, accuracy, and answer-key quality—provides a foundation for such pipelines. Dependencies: Generated content must undergo expert review, psychometric calibration, bias testing, and contamination checks. Problems used for high-stakes decisions require stronger validation than those used for informal practice.
- New standards for measuring AI’s effect on human capability (academia, industry, and policy). StudentBench suggests a broader benchmark paradigm in which systems are evaluated by how much they improve unaided human performance. Future benchmarks could measure learning gain, retention, transfer, decision quality, creativity, and workplace productivity across populations and domains. Dependencies: Benchmarks must distinguish genuine capability improvement from short-term test familiarity, strategic test-taking, or dependence on AI hints. Equivalence margins, control conditions, and outcome measures should be preregistered and justified for each domain.
- Evidence-based public funding and education policy (government and philanthropy). Governments could fund AI tutoring in schools, adult education, libraries, and workforce reskilling when independent trials show favorable cost per learning gain. Subsidized access could target students who cannot afford private tutoring, potentially reducing disparities in preparation for gatekeeping examinations. Dependencies: The 918-fold cost comparison uses a specific inference-cost estimate and a $75-per-hour human-tutoring reference; it should not be interpreted as a full deployment-cost estimate. Public programs must include teacher training, infrastructure, accessibility, privacy, and independent auditing.
- Recursive human self-improvement through AI-supported learning (long-term societal application). If AI tutoring improves human capabilities, and more capable users subsequently use AI more effectively, integrated learning systems could support continuous skill development in education and employment. This is the paper’s broader RHSI hypothesis, but the study establishes only the initial condition: short-term improvement after AI tutoring. Dependencies: Demonstrating recursive improvement requires longitudinal evidence that gains persist, transfer to new tasks, improve users’ ability to use AI, and generate further measurable gains. It also requires equitable access so that augmentation does not widen educational and economic inequalities.
Glossary
- Analysis of covariance (ANCOVA): A statistical method that compares group outcomes while controlling for a covariate, such as pre-test performance. “To account for differences in average pre-test score across conditions when comparing learning gains, we use the analysis of covariance (ANCOVA) approach”
- Autonomous AI tutor: An AI tutoring system that operates without direct human intervention during instruction. “StudentBench tests autonomous, general-purpose LLMs with minimal software scaffolding against live human tutors”
- Bradley–Terry model: A statistical model for estimating preferences or rankings from pairwise comparisons. “Expert preferences for lesson planning and practice-problem creation, fitted with separate centered Bradley-Terry models”
- Cognitive engagement: The degree to which a learner actively thinks, reasons, explains, or solves problems during learning. “Each action elicits different levels of cognitive engagement from the student”
- Cognitive tutor: An intelligent tutoring system designed to model learners’ knowledge and provide adaptive instruction. “including cognitive tutors (Anderson et al., 1995), ACT-R (Anderson et al., 2004), and knowledge tracing”
- Confidence interval: A statistical range estimating where a population parameter is likely to fall. “The combined AI human difference was percentage points (90% CI ).”
- Conversational pedagogy: The instructional strategies expressed through dialogue between a tutor and a learner. “We also evaluate pedagogical characteristics of tutoring conversations to create the leaderboard in Figure 3C.”
- Covariance: A quantity describing how two variables vary together, often used in statistical estimation. “account for shared tutors and repeated students using CR2 covariance with Satterthwaite degrees of freedom”
- CR2 covariance: A small-sample-adjusted, cluster-robust covariance estimator used for statistical inference with dependent observations. “using CR2 covariance with Satterthwaite degrees of freedom”
- Degrees of freedom: The number of independent values available for estimating statistical quantities. “using CR2 covariance with Satterthwaite degrees of freedom”
- Equivalence bounds: Predefined limits within which two treatments are considered practically equivalent. “We use equivalence bounds of 0.25 pooled standard deviations of learning gains”
- Equivalence test: A statistical test determining whether an observed difference lies within a prespecified range of practically negligible differences. “We test equivalence using standard two one-sided tests at =0.05”
- Expert rubric evaluation: A structured assessment in which specialists judge outputs according to predefined criteria. “expert human tutors to compare LLM-generated lesson plans and practice problems through pairwise rubric evaluations”
- Frontier model: A highly capable model operating near the current leading edge of performance. “Our pooled AI result spans 13 capability-varied AI tutors, including non-frontier and open-weight models.”
- Harness: Software infrastructure that controls, evaluates, or supports an AI system’s execution. “changes in harnesses and software scaffolding led to substantial performance gains”
- Heteroskedasticity-consistent: Designed to provide valid statistical estimates when the variance of errors differs across observations. “Figure 1C shows learning gains for 12 AI tutors, adjusted for pre-test score and section, with HC3 intervals.”
- Inference cost: The computational or monetary cost of generating outputs from a trained AI model. “Its mean inference cost was $0.067 for the entire tutoring session”
- Intelligent tutoring system: A computer-based instructional system that adapts teaching to a learner’s knowledge or responses. “Earlier intelligent tutoring systems have produced learning outcomes comparable to human tutoring”
- Item response theory (IRT): A family of psychometric models that estimates a learner’s ability from responses to test items. “we used item response theory (IRT) to estimate proficiency on the Quantitative pre-test”
- Latency: The time between a request and the system’s response. “Across the 12 AI tutors, lower AI latency was strongly associated with higher student engagement”
- Learning gain: The change in assessment performance between a post-test and a pre-test. “We define learning gain as post-test score minus pre-test score”
- Longitudinal study: A study that follows participants or outcomes over an extended period. “One limitation of this study is that we measured immediate learning gains without evaluating whether they persist over months.”
- Metacognitive strategy: A learning approach involving awareness and regulation of one’s own thinking and learning processes. “An effective metacognitive strategy: learning by doing and explaining with a computer-based Cognitive Tutor.”
- Minimal software scaffolding: Limited auxiliary software support that structures or guides an AI system’s behavior. “We use minimal software scaffolding and prompt adjustments”
- Mixed-initiative dialogue: An interaction in which both participants can independently initiate contributions. “AutoTutor: An Intelligent Tutoring System With Mixed-Initiative Dialogue”
- Open-weight model: An AI model whose trained parameters are publicly available for use or modification. “including non-frontier and open-weight models”
- Pareto frontier: The set of options that cannot be improved on one objective without worsening another. “we explore the Pareto frontier of learning gain, cost, and latency”
- Pareto-dominate: To be superior to another option on at least one objective while being no worse on the others. “Gemini 3.1. Pro had higher mean learning gain, lower inference cost, and faster replies than GPT-5.5. Pro”
- Pedagogy: The theory and practice of teaching. “We also evaluate pedagogical characteristics of tutoring conversations”
- Psychometric: Relating to the measurement of psychological traits or abilities, such as knowledge or proficiency. “Its expert ratings of lesson plans and practice design complement psychometric evaluations of AI-generated exam questions”
- Recursive human self-improvement (RHSI): A proposed process in which improved human abilities enable better use and further development of teaching technologies. “We use the term recursive human self-improvement (RHSI) for this process”
- Reward hacking: Exploiting the specification of an objective to obtain a high measured reward without achieving the intended goal. “To avoid reward hacking in AI tutors”
- Scaffolding: Temporary instructional support that helps learners perform tasks they could not yet complete independently. “StudentBench examines related aspects of tutoring dialogue: scaffolding cues”
- Satterthwaite degrees of freedom: An approximation used to estimate degrees of freedom for inference with unequal or complex variance structures. “using CR2 covariance with Satterthwaite degrees of freedom”
- Singularity: A hypothetical point at which an AI system’s intelligence or capability undergoes an irreversible, transformative change. “A singularity in intelligence can occur when an AI technology recursively self-improves”
- Spearman correlation: A rank-based measure of the strength and direction of a monotonic relationship between two variables. “lower AI latency was strongly associated with higher student engagement (Spearman ”
- Statistical equivalence: A conclusion that two conditions differ by no more than a prespecified practically meaningful margin. “AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains”
- Student- confidence interval: A confidence interval based on the Student’s distribution, commonly used when estimating a mean with limited or unknown population variance. “Error bars depict Student- confidence intervals.”
- Two one-sided tests (TOST): An equivalence-testing procedure that tests whether an effect is both above a lower bound and below an upper bound. “We test equivalence using standard two one-sided tests at =0.05”
- Zone of proximal development: The range of tasks a learner can perform with appropriate assistance but not yet independently. “Students' performance with and without guidance can reveal their zone of proximal development”