Papers
Topics
Authors
Recent
Search
2000 character limit reached

StudentBench: AI and human tutoring yield equivalent GRE learning gains

Published 23 Sep 2026 in cs.AI and cs.CY | (2609.28470v1)

Abstract: Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether LLMs produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average. In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations. Together, the two studies clearly separate AI tutors across: (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement. Surprisingly, one AI tutor achieved learning gains equivalent to human tutoring (p = .044) at 918 times lower cost (USD 0.0052 for AI versus USD 4.81 for human, per percentage point gained). For Quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains (all p < .002). The StudentBench platform is freely available at https://studentbench.org.

Summary

  • The paper compares AI tutoring, expert human tutoring, and a no-tutoring control across 2,383 participants and 2,469 analyzed sessions, finding statistically equivalent learning gains for AI and human tutoring.
  • AI tutoring achieved significant learning gains in both Quantitative and Verbal sections, with a pooled difference of -0.58 percentage points compared to human tutoring and a 90% confidence interval of [-2.18, 1.03].
  • Gemma 4 31B, an AI tutor, demonstrated equivalent learning gains to human tutoring with $918 times$ lower cost per percentage point gained

Research question and contribution

“StudentBench: AI and human tutoring yield equivalent GRE learning gains” (2609.28470) evaluates whether general-purpose LLMs can produce learning gains comparable to expert human tutoring under controlled, one-hour instructional sessions. Rather than treating model competence or response quality as proxies for educational value, the study measures unaided performance on post-tests designed to parallel the GRE. Its central claim is not merely that AI can answer educational questions, but that AI tutoring can improve students’ independent performance at a level statistically equivalent to expert tutoring.

The paper frames this result as evidence for the first condition of “recursive human self-improvement” (RHSI): a technology must improve human capabilities. This framing is conceptually broader than the experiment itself. The data establish immediate tutoring-related gains on newly authored GRE assessments; they do not establish the subsequent RHSI conditions, such as whether improved students use AI more effectively or whether that interaction produces continued gains.

StudentBench contributes three principal components. First, it presents a randomized comparison of AI tutoring, expert human tutoring, and a no-tutoring control across 2,383 participants and 2,469 analyzed sessions. Second, it introduces expert-review and transcript-based evaluations that separate lesson planning, practice-problem generation, and conversational pedagogy. Third, it reports cost, latency, engagement, and practice metrics alongside learning outcomes, producing a benchmark intended to evaluate AI as a system for augmenting human performance rather than as an isolated LLM.

Experimental design and measurement

Participants were recruited through Handshake and were predominantly young adults, although the full age range was 18–67. They completed either a Quantitative or Verbal GRE session. Each session consisted of a 27-question Quantitative or 41-minute Verbal pre-test, approximately five minutes reviewing pre-test errors, one hour in an assigned tutoring condition, and a matched post-test. The assessment forms were counterbalanced: approximately half of participants received form P followed by Q, and the remainder received Q followed by P. The questions were newly written by former ETS and Kaplan exam creators and were designed to match prior GRE assessments in difficulty, content coverage, format, and sequencing. Avoiding publicly available legacy GRE questions was intended to reduce contamination from LLM pretraining.

The AI condition included 13 tutors corresponding to 12 models, with some models evaluated at different reasoning settings. The systems received the student’s complete pre-test record, including questions, answers, correctness, response times, and expert difficulty ratings. They generated a lesson plan, selected concepts and time allocations, created practice problems and answer keys, and conducted the interactive tutoring session. The prompts deliberately supplied minimal scaffolding, leaving the models responsible for diagnosis, sequencing, problem generation, and conversational instruction.

Human tutors were former ETS or Kaplan employees or had at least five years of GRE tutoring experience. They taught through live video calls and received the same pre-test information, but neither human nor AI tutors saw the post-test. The control group reviewed its graded pre-test and watched educational videos unrelated to the GRE. The absence of a practice-only control means that the experiment estimates the effect of the assigned tutoring conditions relative to this control, not the incremental effect of tutoring interaction relative to independent GRE practice.

Learning gain was defined as post-test percentage correct minus pre-test percentage correct. The primary analyses used ANCOVA, adjusting for baseline score, its quadratic term, assessment form, and section where appropriate. AI–human comparisons used CR2 cluster-robust covariance estimates and Satterthwaite degrees of freedom to account for repeated participants and students sharing human tutors. Equivalence was evaluated using TOST with bounds of ±0.25\pm 0.25 pooled standard deviations, corresponding to approximately ±4.09\pm 4.09 percentage points in the combined analysis.

Learning outcomes

The primary result is that pooled AI tutoring was statistically equivalent to expert human tutoring across Quantitative and Verbal sessions. The adjusted AI-minus-human difference was −0.58-0.58 percentage points, with a 90% confidence interval of [−2.18,1.03][-2.18, 1.03]. This interval fell within the prespecified equivalence bounds, yielding p=.015p=.015. Equivalence also held under a stricter ±0.20\pm 0.20-standard-deviation margin, with p=.023p=.023.

The result is substantively important because AI tutoring did not merely outperform the no-tutoring condition; it reached the performance level of a strong human benchmark. Relative to control, AI tutoring increased learning gain by 6.15 percentage points in the combined sample, with a 95% confidence interval of [4.08,8.21][4.08, 8.21] and p<.001p<.001. The section-specific adjusted differences were 6.86 percentage points in Quantitative and 5.47 percentage points in Verbal. Human tutoring also exceeded control by 7.04 percentage points overall. Thus, the equivalence conclusion is consistent with both interventions producing substantial gains over the control condition, rather than with neither intervention being effective.

The section-level pattern is asymmetric. AI and human tutoring were equivalent in Quantitative, with an AI-minus-human difference of 0.88 percentage points and TOST p=.028p=.028. Equivalence was not established in Verbal: the AI-minus-human estimate was ±4.09\pm 4.090 percentage points, with a 90% confidence interval of ±4.09\pm 4.091 and ±4.09\pm 4.092 under the stated margin. Human tutoring therefore retained a possible advantage in Verbal, although the study could not establish a statistically meaningful difference between sections in the human–AI gap.

The domain-level results show specialization rather than a uniformly dominant tutor. The best-performing AI tutor exceeded the human mean in five of seven domains, while human tutoring remained higher in algebra and sentence equivalence. Different models led different domains: Google models led data analysis, geometry, and reading comprehension; GPT-5.5 Pro led arithmetic; Kimi K2.6 led algebra; and Anthropic models led sentence equivalence and text completion. These domain leaderboards are descriptive and should not be interpreted as definitive model rankings for subgroups, since many domain-specific cells contain relatively small samples and the selected frontier model differs across domains.

The combined AI-tutor comparison did not detect statistically significant differences among the 12 shared tutors (±4.09\pm 4.093). This does not imply practical equality among models. Adjusted mean learning gains ranged from 12.1 to 16.0 percentage points, compared with 14.9 for human tutoring and 7.8 for control. Rather, the study had insufficient evidence to distinguish the tutors under its omnibus specification. The distinction matters because the paper simultaneously reports strong separation on other dimensions, particularly cost, latency, expert preferences, and engagement.

Teaching capability evaluations

StudentBench separates several capabilities that are often conflated in evaluations of AI tutors. In the second study, 51 expert GRE tutors conducted 2,028 pairwise evaluations of AI-generated lesson plans and practice problems derived from 381 pre-tests. Bradley–Terry models aggregated preferences over concept relevance, grouping, prioritization, time allocation, test-taking strategy, practice alignment, difficulty, and example accuracy.

Anthropic models, particularly Opus variants, were preferred by experts for lesson planning and practice-problem creation. These rankings demonstrate that expert judgments distinguish model families and model configurations. They also show that high-quality instructional artifacts are not equivalent to high measured learning gains. The paper explicitly treats these leaderboards as evaluations of teaching materials and observed behaviors, not as direct estimates of student learning. Indeed, the model with the strongest expert-preference profile need not be the model with the largest post-test gain.

The practice-problem analyses provide a limited external check on expert judgments. In Quantitative tutoring, models preferred by experts for practice design tended to receive fewer student answer disputes; the combined Spearman correlation was ±4.09\pm 4.094, ±4.09\pm 4.095. For example accuracy, the combined correlation was ±4.09\pm 4.096, ±4.09\pm 4.097. These associations are exploratory, based on a small number of model-level observations, and the dispute button recorded student disagreement rather than independently verified errors. They therefore support alignment between some expert and behavioral indicators without validating either as a substitute for learning outcomes.

The conversational-pedagogy leaderboard operationalized six indicators: scaffolding cues, requests for explanations, interactive questions, early requests to attempt problems, long solutions after student replies, and reasoning checks. AI tutors clustered by provider family, with distinct behavioral profiles across model families. This suggests that provider-level training or prompting practices may influence tutoring style. However, the authors acknowledge two substantial limitations: the indicators were not validated as predictors of learning, and human tutoring was spoken video interaction whereas AI tutoring was typed chat. Consequently, differences in indicator prevalence cannot be interpreted straightforwardly as differences in pedagogical quality.

Cost, latency, and engagement

The cost analysis is one of the paper’s strongest quantitative claims. AI inference costs included lesson planning, practice-problem generation, and interactive tutoring. Human tutoring was valued at a reference price of ±4.09\pm 4.0980.067 for Gemma 4 31B to $21.24 for GPT-5.5 Pro.

Gemma 4 31B produced a mean learning gain of 12.742 percentage points across 159 sessions, compared with 15.608 percentage points for human tutoring. Its AI-minus-human difference was $\pm 4.09$9 percentage points, with a 90% confidence interval of $-0.58$0 and an equivalence-test $-0.58$1. Dividing mean session cost by mean learning gain yielded $-0.58$24.80508 for the human reference. The resulting ratio was approximately 918 times lower cost per percentage point of learning gain.

This comparison should be interpreted narrowly. The numerator for humans is a market-rate assumption rather than an experimentally observed labor cost, while AI costs reflect inference and reconstructed token charges. The denominator is an immediate post-test gain, not retained learning, credential attainment, or broader educational benefit. In addition, the individual-tutor equivalence tests were conducted without correction for multiple comparisons across tutors. The result is therefore a striking cost-efficiency estimate within the study’s operational definition, but not a general estimate of the total cost of educational deployment.

The Pareto analysis further indicates that cost and latency are not monotonically related to learning gain. Gemini 3.5 Flash achieved a similar learning gain to GPT-5.5 Pro at approximately 20 times lower cost. Gemini 3.1 Pro had higher mean gain, lower cost, and faster replies than GPT-5.5 Pro, thereby Pareto-dominating it on the sample estimates. Because frontier membership was based on observed means without uncertainty intervals, these comparisons describe the empirical trade-off surface rather than statistically confirmed dominance.

Reply latency was strongly associated with engagement at the tutor level. Across 12 shared tutors, faster replies correlated with more student messages, with Spearman −0.58-0.583 and −0.58-0.584. Reply times ranged from 1.88 seconds for GPT-5.4 mini to 30.99 seconds for GPT-5.5 Pro. The finding implies that responsiveness may affect learning indirectly by changing the amount of student activity that a tutoring session elicits.

Practice as a potential mechanism

The session-level exploratory analyses connect latency, engagement, practice, and learning in Quantitative tutoring. Among 1,137 Quantitative AI sessions, every link in the reported chain was statistically significant after the paper’s multiple-comparison procedure. Ten additional seconds of reply time were associated with 5.19 fewer student messages, with a 95% confidence interval of −0.58-0.585 and −0.58-0.586. Ten additional student messages were associated with 0.52 more correct practice problems, with a 95% confidence interval of −0.58-0.587 and −0.58-0.588. Ten additional correct practice problems were associated with a 5.62-percentage-point larger learning gain, with a 95% confidence interval of −0.58-0.589 and [−2.18,1.03][-2.18, 1.03]0.

The implication is that response latency may matter less as a direct property of model quality than as a determinant of interaction volume and productive practice. This mechanism could help explain why models with weaker expert-evaluated pedagogical profiles nonetheless achieved human-level learning gains: a fast tutor may induce more attempts and more feedback cycles within the fixed one-hour treatment.

The mechanism did not generalize cleanly to Verbal tutoring. The corresponding associations were not statistically significant under the reported analysis: reply time and student messages had [−2.18,1.03][-2.18, 1.03]1, student messages and correct practice had [−2.18,1.03][-2.18, 1.03]2, and correct practice and learning gain had [−2.18,1.03][-2.18, 1.03]3. Verbal participants sent fewer but longer messages, so message count may be a poorer engagement measure in that section. The contrast emphasizes that engagement metrics are construct-dependent and that the Quantitative pathway should not be treated as a general theory of AI tutoring effectiveness.

Methodological strengths and limitations

The design has several notable strengths. It measures unaided transfer rather than performance during AI assistance, uses counterbalanced assessment forms, withholds post-test content from tutors, recruits expert human comparators, and reports both effectiveness and operational cost. The scale of the study is also substantial for a live tutoring experiment, with 176,614 student–AI messages and 45,463 answered practice problems. The public platform, released data, and reproducibility code make the benchmark unusually inspectable for this class of study.

The principal limitation is that learning was measured immediately after a single one-hour session. The results do not establish retention over weeks or months, cumulative effects from repeated tutoring, or transfer beyond GRE-like items. This is consequential for the cost-per-gain claim: a low cost for an immediate gain may not remain low when durable learning requires spaced practice or repeated intervention.

The population was restricted to adults who could read and write in English, had access to Handshake, and participated in a paid study. The findings therefore do not directly establish performance across languages, younger learners, educational institutions, different devices, or students with different motivational and socioeconomic profiles. The study also did not include a self-study practice condition, leaving the added value of interactive AI tutoring relative to independent problem solving unresolved.

Inference about human equivalence is particularly sensitive to the human comparison structure. The combined analysis retained equivalence after several checks, including excluding participants who completed both sections and omitting human tutors individually. However, the effective degrees of freedom for the principal pooled comparison were only 1.59 because human sessions were concentrated among a small number of tutor clusters. The authors also report that a gain-only sensitivity analysis without the principal covariate and assessment-form adjustments did not establish equivalence in either section. Thus, the central conclusion is supported by the prespecified adjusted model but depends materially on its handling of baseline score, assessment form, clustering, and repeated observations.

Finally, the frontier and individual-model claims should be distinguished from the pooled result. Pooled equivalence supports the claim that the evaluated AI-tutor mixture was comparable to the human benchmark. Six individual AI tutors passed uncorrected equivalence tests, but the study did not correct these tests for multiple comparisons. The evidence is consequently strongest for the pooled AI condition and weaker for claims about any particular model, domain, or subgroup.

Conclusion

StudentBench provides evidence that minimally scaffolded LLM tutors can produce immediate GRE learning gains statistically equivalent to expert human tutoring in a large controlled study. The pooled AI-minus-human difference was [−2.18,1.03][-2.18, 1.03]4 percentage points, and AI exceeded the no-tutoring control by 6.15 percentage points overall. The result was stronger for Quantitative than for Verbal, while domain-level analyses indicated substantial specialization among models.

The paper’s second contribution is methodological: it evaluates tutoring as a multidimensional system involving learning outcomes, planning, practice generation, dialogue, cost, latency, and engagement. The most consequential economic result is that Gemma 4 31B achieved human-equivalent gains under the study’s equivalence test at approximately 918 times lower cost per percentage point gained. This estimate is bounded by immediate post-test measurement, assumed human pricing, model-specific inference accounting, and uncorrected individual-tutor comparisons.

The principal unresolved question is whether these short-term, GRE-specific gains persist and accumulate under repeated use, particularly across Verbal reasoning, independent practice conditions, languages, and educational settings not represented in the sample.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

Explain it Like I'm 14

1. What is this paper about?

The paper studies whether AI tutors can help students learn as well as human tutors.

The researchers created StudentBench, a system for testing AI teaching. They used questions from the GRE, an exam taken by many people applying to graduate school. The study compared three groups:

  • Students taught by an AI tutor
  • Students taught by an expert human tutor
  • Students who received no GRE tutoring

The main question was: After one hour of tutoring, did students improve just as much with AI as with a human?

2. What were the researchers trying to find out?

The researchers asked several important questions:

  1. Can AI tutoring improve students’ GRE scores?
  2. Does AI tutoring work about as well as expert human tutoring?
  3. Which AI systems are best at planning lessons and creating practice questions?
  4. How much does AI tutoring cost compared with human tutoring?
  5. Does the speed of an AI tutor’s replies affect how much students participate and learn?

The paper is also connected to a bigger idea called recursive human self-improvement. This means that technology could help people learn more effectively, and better-educated people could then use and improve that technology even more.

3. How did the researchers conduct the study?

The tutoring experiment

The researchers studied 2,383 students and collected information from more than 175,000 student–AI messages.

Students first took a pre-test. This showed what they already knew. The test included either:

  • Quantitative questions, involving mathematics and problem-solving, or
  • Verbal questions, involving reading and language skills.

After the pre-test, students were randomly placed into one of three groups:

  • AI tutoring
  • Live tutoring from an expert human
  • A control group that watched unrelated educational videos

The tutoring lasted about one hour. The AI tutors could:

  • Plan a lesson
  • Explain ideas
  • Ask questions
  • Create practice problems
  • Give feedback on student answers

After tutoring, students took a different but similar post-test. The researchers calculated the learning gain by subtracting the pre-test score from the post-test score:

Learning gain=post-test score−pre-test score\text{Learning gain} = \text{post-test score} - \text{pre-test score}

For example, if a student scored 60% before tutoring and 70% afterward, the learning gain would be 10 percentage points.

The researchers used different questions on the two tests so students could not simply memorize the answers. They also checked for low effort, very fast answers, excessive tab switching, and other signs that a session might not provide trustworthy data.

Comparing different AI tutors

The study tested 13 AI tutors based on different LLMs. A LLM is an AI system trained to understand and produce text, such as explanations and answers.

In a second study, 51 expert human tutors compared AI-created:

  • Lesson plans
  • Practice questions
  • Answer keys

The experts judged whether the materials were relevant, well organized, appropriately difficult, and accurate.

The researchers also examined tutoring conversations. They looked for behaviors such as whether the tutor:

  • Asked students to explain their thinking
  • Let students try before showing the answer
  • Gave hints connected to the student’s mistakes
  • Asked useful follow-up questions

Understanding the statistics

The researchers used statistical tests to determine whether differences were likely to be real rather than caused by chance.

They used an equivalence test, which asks a slightly different question from a normal test. Instead of asking, “Are AI and human tutoring different?” it asks, “Are they close enough that we can reasonably treat them as equally effective?”

The paper also reports numbers such as p = .015. A p-value is a measure of how surprising the results would be if there were no meaningful effect. Smaller values generally provide stronger evidence that the finding is not just random chance.

4. What did the researchers discover?

AI and human tutors produced similar learning gains

The main result was that AI tutoring and expert human tutoring produced statistically equivalent GRE learning gains when their results were combined across the Quantitative and Verbal sections.

The difference between AI and human tutoring was very small: AI learners improved by about 0.58 percentage points less than human-tutored learners. The researchers judged this difference small enough to count the two types of tutoring as equivalent.

Both AI and human tutoring worked better than the no-tutoring control group. Compared with the control group, AI tutoring improved results by about:

  • 6.86 percentage points in Quantitative questions
  • 5.47 percentage points in Verbal questions

In simple terms, students who received tutoring answered roughly 1.5 to 2 more questions correctly than students in the control group.

Some AI tutors performed better than human tutors in certain areas

The results were not identical in every subject.

  • In Quantitative topics, the strongest AI tutors performed better than human tutors in three of the four areas.
  • In Verbal topics, human tutors had the highest overall average, although some AI tutors were close.
  • Across the seven GRE topic areas, the best AI tutor performed better than the human average in five areas.

This suggests that AI may be especially strong at explaining certain mathematical ideas, while human tutors may still have an advantage in some language-related areas.

AI tutoring was much cheaper

One of the biggest differences was cost.

The study estimated that a human tutor cost about $75 per hour. AI tutoring was much cheaper because it mainly required computer processing.

One AI tutor, called Gemma 4 31B in the paper, produced learning gains equivalent to those of human tutoring while costing about 918 times less per percentage point of learning gain.

This does not mean every AI tutor is equally good or that every student would receive the same results. However, it suggests that AI tutoring could make individualized help available to many more students, including students who cannot afford private tutoring.

Faster AI replies were linked to more learning activity

The researchers also found that faster AI responses were connected with:

  1. More student messages
  2. More correct practice answers
  3. Larger learning gains

This pattern was especially clear for Quantitative tutoring.

A possible explanation is that students lose interest when they have to wait too long. Faster replies may keep the conversation moving, giving students more chances to practice.

However, these findings show relationships, not definite proof that faster replies directly caused better learning. Other factors might also be involved.

Different AI models had different teaching strengths

The AI tutors did not all behave the same way.

Some models were especially good at:

  • Planning lessons
  • Choosing which topics to teach
  • Creating suitable practice problems
  • Showing helpful teaching behaviors in conversations

The expert tutors especially liked several models from the Anthropic family for planning lessons and writing practice problems. However, strong performance on these expert ratings did not always perfectly predict the largest learning gains in the student tests.

This is important because an AI can appear to give good explanations to an expert, but the best test is still whether students actually learn.

5. Why are these findings important?

The study suggests that AI tutors can be more than tools that simply provide answers. With the right interaction, they may help students understand ideas, practice skills, and improve their test performance.

The possible benefits include:

  • Lower costs: More students could receive one-on-one help.
  • Greater access: Students could get tutoring at any time, even when human tutors are unavailable.
  • Personalized lessons: AI can focus on the mistakes a student made.
  • Quick feedback: Students can receive explanations without waiting.
  • Support for human teachers: AI could handle some practice and explanation tasks while teachers focus on more difficult or personal parts of learning.

The paper’s larger message is that AI might support and strengthen human abilities rather than simply replace people.

6. Important limitations

The results should not be treated as proof that AI is always as good as a human teacher.

The study had several limits:

  • It measured learning only immediately after tutoring. We do not know whether students remembered the material months later.
  • The participants were adults who could read and write English and had access to the study platform.
  • The study focused on GRE-style questions, so the results may not apply to every subject or age group.
  • The researchers did not compare AI tutoring with students simply practicing GRE questions on their own.
  • AI models change quickly, so newer versions may perform differently.
  • AI-generated practice questions can sometimes contain mistakes. Students in the study were able to flag answers they believed were wrong.
  • The study was funded by Handshake AI, the organization connected to the research, so independent studies would be useful.

Conclusion

The StudentBench paper found that, in this study, AI tutors helped students improve their GRE scores about as much as expert human tutors after one hour. Some AI tutors even performed better than human tutors in particular subject areas.

The biggest advantage of AI was its very low cost. If these results hold in other subjects and for longer periods, AI tutoring could help make personalized education available to many more people.

Still, AI tutoring needs careful testing. Future research should examine whether students remember what they learn, how AI works with human teachers, and whether the same results appear for younger students, different languages, and subjects beyond the GRE.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

  • Long-term retention is unresolved: The study measures only immediate post-test gains, so it remains unknown whether AI- or human-tutored learning persists over weeks or months.
  • Transfer beyond the assessment is untested: The paper does not establish whether tutoring improves performance on novel problem types, real GRE scores, broader academic tasks, or practical reasoning outside the study’s post-test.
  • No practice-only comparison was included: Because the control group watched unrelated educational videos, the study cannot isolate the incremental benefit of AI conversation from independently solving GRE practice problems with equivalent time and materials.
  • Generalizability is limited by the sample: Participants were English-speaking adults recruited through Handshake, primarily aged 18–23, and paid for participation; effects may differ for younger students, older learners, nonstudents, lower-literacy populations, or learners with disabilities.
  • Language and cultural transferability remain unknown: The study does not test AI tutoring in languages other than English or determine whether its effectiveness depends on English-language proficiency, cultural conventions, or familiarity with U.S.-style standardized testing.
  • Educational-context generalization is untested: Results from one-hour, one-to-one GRE sessions do not show whether AI tutoring is effective in classrooms, schools, universities, tutoring centers, home environments, or longer curricula.
  • Device and access constraints were not evaluated: The paper does not examine performance under mobile-only access, low bandwidth, limited computing resources, intermittent connectivity, or accessibility technologies.
  • The human-tutor benchmark is underspecified: More evidence is needed on human tutors’ qualifications, teaching protocols, adherence to the study design, and variability in tutoring quality to determine which level of human tutoring AI matched.
  • AI–human comparisons may be affected by modality differences: AI tutoring was text-based whereas human tutoring occurred through live video calls; the study does not separate tutor capability from differences in speech, visual cues, typing demands, social presence, or interaction modality.
  • The effects of tutor identity and disclosure are unknown: Students were not told which AI model they received, but the paper does not test whether disclosure that a tutor is AI changes trust, engagement, persistence, learning, or willingness to follow advice.
  • Mechanisms of learning remain uncertain: The observed association between latency, engagement, correct practice, and learning gain is correlational; randomized manipulation of response speed is needed to determine whether latency causes improved learning.
  • Engagement measures are incomplete and section-dependent: Student message count may not represent cognitive engagement consistently, especially because Verbal students sent fewer but longer messages; future work should use validated measures of reasoning quality, attention, metacognition, and off-task behavior.
  • More practice may not imply deeper learning: The study counts correct practice but does not determine whether students solved problems independently, relied on hints, memorized procedures, or developed transferable understanding.
  • AI-generated content quality is not fully verified: Student answer-dispute flags were not independently checked for correctness, leaving the true rates of erroneous explanations, practice problems, and answer keys uncertain.
  • Safety and pedagogical failure modes are underexplored: The paper does not systematically evaluate hallucinations, misleading explanations, inappropriate confidence, harmful feedback, privacy risks, manipulation, student dependence, or failures involving atypical misconceptions.
  • Student heterogeneity is insufficiently characterized: It remains unclear which students benefit most or least by prior achievement, socioeconomic status, motivation, test anxiety, learning differences, demographic characteristics, or baseline AI familiarity.
  • The high-performing-student finding is exploratory: The apparent advantage of Gemini models for students in the top proficiency quartile requires preregistered replication with adequate power and interaction tests across models and proficiency levels.
  • Domain-level conclusions may be underpowered or confounded: Claims that particular AI tutors outperform humans in five of seven domains may reflect multiple comparisons, unequal sample sizes, model-by-domain interactions, or domain-specific assessment properties.
  • Individual tutor equivalence results need multiplicity control: Individual AI tutors were tested for equivalence without correction across tutors, so some apparent human-equivalent tutors may be false positives.
  • The equivalence margins are consequential but not empirically validated here: The choice of ±0.25\pm0.25 pooled standard deviations, and the interpretation of smaller margins, may not correspond to a practically meaningful GRE improvement for students or admissions outcomes.
  • Very low effective degrees of freedom weaken pooled inference: The combined equivalence test reports 1.59 degrees of freedom; the robustness of this inference under alternative clustering, weighting, and hierarchical models remains to be established.
  • Assessment validity and comparability require further evidence: Although new questions were designed to match prior GRE tests, the paper does not provide enough evidence on item calibration, parallel-form equivalence, item exposure, differential item functioning, or whether pre-/post-test differences reflect learning rather than form difficulty.
  • Possible test–retest and content-practice effects remain uncertain: Students reviewed pre-test mistakes and then received related tutoring, so some post-test gains may reflect short-term familiarity or strategy practice rather than durable conceptual learning.
  • Selection and attrition effects are not fully resolved: The extensive exclusion criteria may remove disengaged, struggling, or technologically constrained participants and could produce estimates that do not apply to typical users.
  • The no-tutoring condition may not be a neutral baseline: Watching unrelated videos may differ from waiting, resting, or completing ordinary study activities in motivation and cognitive load, complicating interpretation of the AI- and human-versus-control effects.
  • Cost comparisons are sensitive to assumptions: The $75-per-hour human rate, incomplete coverage of AI API costs, reconstructed costs for 3% of calls, and exclusion of platform, monitoring, development, support, and infrastructure costs may substantially affect the reported$918\times$ advantage.
  • Cost-effectiveness over a complete learning program is unknown: Session-level cost per percentage point does not capture repeated sessions, remediation, onboarding, human oversight, model downtime, student dropout, or the cost of achieving a target GRE score.
  • Model comparisons are not fully controlled for: Models differed in reasoning settings, prompts, availability dates, latency, context handling, and possibly token limits; factorial experiments are needed to separate model capability from configuration and implementation effects.
  • Rapid model changes threaten reproducibility: Because models and prices changed during the two-month study, future researchers may not be able to reproduce the exact tutor conditions or determine whether results reflect stable model properties.
  • Minimal scaffolding limits conclusions about deployable tutoring systems: The study intentionally uses low-guidance prompts, but it does not determine how structured curricula, retrieval, knowledge tracing, misconception models, tool use, or adaptive interfaces would change learning and cost.
  • The teaching leaderboards are not validated against learning outcomes: Expert preferences for lesson plans and practice problems, as well as six transcript indicators, may not predict student learning, retention, equity, or satisfaction.
  • Human-reference conversational scores are not directly comparable: AI sessions were typed chat and human sessions were spoken tutoring, so differences in the six pedagogical indicators may reflect transcription, turn-taking, or modality rather than pedagogy.
  • Reviewer reliability and representativeness need further study: The paper does not establish how consistently the 51 expert tutors applied the rubrics, whether ratings generalize beyond GRE experts, or whether expert preferences align with student outcomes.
  • Student satisfaction and affective outcomes are largely absent: The study does not report trust, enjoyment, frustration, confidence, motivation, perceived usefulness, social connection, or willingness to continue using AI tutoring.
  • Equity effects remain unknown: Lower monetary cost does not guarantee equitable access; the paper does not examine disparities caused by language, disability, digital literacy, broadband access, algorithmic bias, or differential ability to formulate questions.
  • The study does not test human–AI collaboration: The conclusions concern autonomous AI or standalone human tutoring, leaving unresolved whether AI assistance can improve human tutor effectiveness, increase tutor capacity, or produce better outcomes than either condition alone.
  • The RHSI framework is only partially evaluated: The study tests the first proposed RHSI condition—improved human capability—but does not examine whether improved learners use AI more effectively or whether such use produces cumulative, recursive gains over multiple cycles.
  • Security and privacy implications of real deployment are unexplored: The paper does not evaluate data retention, student consent in ongoing use, prompt injection, unauthorized disclosure, model-provider access to educational records, or governance requirements for large-scale deployment.
  • Effects on official educational outcomes remain unknown: The study uses research assessments and does not show whether AI tutoring improves admissions, course grades, graduation, persistence, or other consequential educational outcomes.

Practical Applications

Immediate Applications

  • Low-cost standardized-test preparation platforms (education technology, consumer software).
    • analyze a learner’s pre-test errors;
    • generate a prioritized lesson plan;
    • create targeted practice problems and answer explanations;
    • conduct interactive, one-hour tutoring sessions; and
    • administer unaided pre- and post-tests to measure improvement.
    • The study reports AI–human-equivalent combined GRE learning gains, with one evaluated tutor achieving comparable gains at roughly 918 times lower cost per percentage point gained than the human-tutoring reference.
    • Dependencies: The reported evidence is specific to English-speaking adults, GRE-like tasks, text-based tutoring, and immediate post-test gains. Systems should verify generated answers and avoid using secure or leaked examination content.
  • AI tutoring as an access and affordability tool for underserved learners (education policy, nonprofits, universities). Schools, libraries, scholarship programs, and workforce organizations can provide AI tutoring where one-to-one human instruction is unavailable or unaffordable. A practical workflow would combine a short diagnostic test, personalized chat tutoring, targeted practice, and escalation to a human tutor for persistent misconceptions. Dependencies: Access to reliable devices and internet, culturally and linguistically appropriate content, privacy protections for student data, and independent monitoring of learning quality. AI should supplement—not automatically replace—teachers and expert tutors.
  • Human-tutor augmentation and preparation (education services). Tutoring companies and academic-support centers can use LLMs to draft lesson plans, prioritize student misconceptions, generate practice questions, and suggest pacing. Human tutors can review and adapt these materials before use, reducing preparation time while retaining human oversight. Dependencies: Expert review remains important because the paper documents student disputes of AI-generated answers and evaluates lesson-plan quality separately from actual learning outcomes. Generated problems require automated and human validation for correctness, difficulty, and curricular alignment.
  • AI-assisted GRE and SAT preparation in universities (higher education). Graduate-school advising offices, writing and quantitative-skills centers, and career services can integrate StudentBench-like tutoring into existing preparation programs. Quantitative domains appear particularly suitable: the best AI tutors exceeded the human mean in several quantitative domains, and the AI–human equivalence result was statistically supported for Quantitative tutoring. Dependencies: The paper did not test official examination outcomes, long-term retention, or the SAT directly. Institutions should run local validation studies before using AI performance as evidence of admissions readiness.
  • Latency-aware tutoring product design (software and learning-platform engineering). AI tutoring interfaces should prioritize fast responses, streaming output, concise turn-taking, and low-latency infrastructure. In Quantitative sessions, faster replies were associated with more student messages, more correct practice, and larger learning gains. Product teams can therefore treat response time as an educational metric rather than merely a user-experience metric. Dependencies: The latency findings are correlational, not proof that reducing latency alone causes learning. Excessively brief or superficial responses could reduce explanation quality, and the relationship was not statistically significant for Verbal tutoring.
  • Learning-gain benchmarking for AI model selection (AI industry and procurement).
    • unaided pre-to-post learning gain;
    • cost per percentage point of gain;
    • response latency;
    • student engagement;
    • lesson-plan quality;
    • practice-problem quality; and
    • pedagogical behaviors such as eliciting explanations and allowing students to attempt solutions.
    • This is more actionable than selecting a model solely by general language or reasoning benchmarks. StudentBench’s open data, code, and platform can support internal replication and model comparisons.
    • Dependencies: Model rankings may vary by subject, student population, prompting, inference settings, and software scaffolding. Cost estimates also depend on provider pricing and token usage.
  • Research methods for evaluating educational AI (academic research). Researchers can immediately reuse the paper’s experimental design: randomized assignment to AI tutoring, human tutoring, and control conditions; counterbalanced pre- and post-tests; ANCOVA adjustment; equivalence testing; and analysis of cost per learning gain. The platform also supports large-scale collection of real student–AI conversations rather than relying only on simulated users or LLM judges. Dependencies: Valid assessments must be independent of model training data, and studies need safeguards against practice contamination, answer leakage, low-effort participation, and test-retest effects.
  • Public and institutional monitoring of AI tutoring quality (policy and educational governance). Regulators, accrediting bodies, and schools can require vendors to report learning outcomes and cost-effectiveness, rather than only model accuracy or satisfaction scores. A practical procurement requirement could be evidence that students perform better on an unaided assessment after tutoring, with subgroup and error-rate reporting. Dependencies: Equivalence to human tutoring in one domain should not be generalized to all subjects or populations. Governance should also address data retention, transparency, accessibility, academic integrity, and the risk that AI-generated explanations reinforce misconceptions.
  • Personal study assistance in daily life (consumer use). Students can use an AI tutor to review missed questions, request explanations, generate analogous problems, and practice until they can solve problems independently. The study supports a workflow centered on active student responses rather than passive answer consumption. Dependencies: Users should independently verify uncertain answers, avoid copying solutions, and periodically test themselves without AI assistance. Immediate score gains may not persist without spaced practice and later reassessment.

Long-Term Applications

  • General-purpose AI tutors across academic subjects and professional exams (education, healthcare, finance, and workforce training). The demonstrated workflow could be adapted to medical entrance examinations, nursing certification, accounting, coding, legal education, language learning, and technical workforce training. Domain-specific tutors could combine diagnostic assessment, concept sequencing, generated practice, and adaptive dialogue. Dependencies: The paper directly studies only GRE Quantitative and Verbal reasoning. High-stakes domains require subject-matter validation, professional oversight, robust factuality testing, and evidence of long-term retention and transfer to real tasks.
  • Longitudinal personalized learning systems (education technology). Future platforms could maintain a learner model over months or years, using knowledge tracing and repeated assessments to schedule review, detect forgetting, and personalize difficulty. StudentBench provides the immediate learning-gain component; longitudinal systems would add retention, transfer, and progression measures. Dependencies: The paper explicitly does not measure retention beyond the immediate post-test. Long-term deployment requires consent, secure student profiles, interoperability with learning-management systems, and safeguards against inaccurate or overconfident learner profiling.
  • AI–human tutoring triage systems (schools, universities, and tutoring companies). A scalable model could assign routine explanation and practice tasks to AI, while routing students to human educators when they show persistent errors, emotional distress, accessibility needs, or unusually low progress. Human tutors could review dashboards containing learning gains, engagement, disputed answers, and unresolved misconceptions. Dependencies: Escalation criteria must be validated and should not systematically deprioritize students with disabilities, limited English proficiency, low connectivity, or atypical communication styles. Student engagement metrics cannot be treated as a direct proxy for learning.
  • Adaptive tutoring infrastructures with latency–learning optimization (AI systems and cloud computing). Educational platforms could jointly optimize model selection, inference cost, response speed, and learning outcomes. For example, a fast low-cost model might handle routine dialogue, while a more capable model or human tutor handles difficult misconceptions. Dependencies: The observed associations do not establish a universal causal latency threshold. Routing systems must preserve pedagogical continuity, prevent inconsistent explanations, and verify that cost reductions do not lower learning quality.
  • Automated generation and validation of educational assessments (publishing and assessment technology). LLMs could generate large pools of practice problems matched to diagnosed misconceptions and calibrated by difficulty. StudentBench’s practice-problem rubric—alignment, difficulty, accuracy, and answer-key quality—provides a foundation for such pipelines. Dependencies: Generated content must undergo expert review, psychometric calibration, bias testing, and contamination checks. Problems used for high-stakes decisions require stronger validation than those used for informal practice.
  • New standards for measuring AI’s effect on human capability (academia, industry, and policy). StudentBench suggests a broader benchmark paradigm in which systems are evaluated by how much they improve unaided human performance. Future benchmarks could measure learning gain, retention, transfer, decision quality, creativity, and workplace productivity across populations and domains. Dependencies: Benchmarks must distinguish genuine capability improvement from short-term test familiarity, strategic test-taking, or dependence on AI hints. Equivalence margins, control conditions, and outcome measures should be preregistered and justified for each domain.
  • Evidence-based public funding and education policy (government and philanthropy). Governments could fund AI tutoring in schools, adult education, libraries, and workforce reskilling when independent trials show favorable cost per learning gain. Subsidized access could target students who cannot afford private tutoring, potentially reducing disparities in preparation for gatekeeping examinations. Dependencies: The 918-fold cost comparison uses a specific inference-cost estimate and a $75-per-hour human-tutoring reference; it should not be interpreted as a full deployment-cost estimate. Public programs must include teacher training, infrastructure, accessibility, privacy, and independent auditing.
  • Recursive human self-improvement through AI-supported learning (long-term societal application). If AI tutoring improves human capabilities, and more capable users subsequently use AI more effectively, integrated learning systems could support continuous skill development in education and employment. This is the paper’s broader RHSI hypothesis, but the study establishes only the initial condition: short-term improvement after AI tutoring. Dependencies: Demonstrating recursive improvement requires longitudinal evidence that gains persist, transfer to new tasks, improve users’ ability to use AI, and generate further measurable gains. It also requires equitable access so that augmentation does not widen educational and economic inequalities.

Glossary

  • Analysis of covariance (ANCOVA): A statistical method that compares group outcomes while controlling for a covariate, such as pre-test performance. “To account for differences in average pre-test score across conditions when comparing learning gains, we use the analysis of covariance (ANCOVA) approach”
  • Autonomous AI tutor: An AI tutoring system that operates without direct human intervention during instruction. “StudentBench tests autonomous, general-purpose LLMs with minimal software scaffolding against live human tutors”
  • Bradley–Terry model: A statistical model for estimating preferences or rankings from pairwise comparisons. “Expert preferences for lesson planning and practice-problem creation, fitted with separate centered Bradley-Terry models”
  • Cognitive engagement: The degree to which a learner actively thinks, reasons, explains, or solves problems during learning. “Each action elicits different levels of cognitive engagement from the student”
  • Cognitive tutor: An intelligent tutoring system designed to model learners’ knowledge and provide adaptive instruction. “including cognitive tutors (Anderson et al., 1995), ACT-R (Anderson et al., 2004), and knowledge tracing”
  • Confidence interval: A statistical range estimating where a population parameter is likely to fall. “The combined AI −- human difference was −0.58-0.58 percentage points (90% CI [−2.18,1.03][-2.18,1.03]).”
  • Conversational pedagogy: The instructional strategies expressed through dialogue between a tutor and a learner. “We also evaluate pedagogical characteristics of tutoring conversations to create the leaderboard in Figure 3C.”
  • Covariance: A quantity describing how two variables vary together, often used in statistical estimation. “account for shared tutors and repeated students using CR2 covariance with Satterthwaite degrees of freedom”
  • CR2 covariance: A small-sample-adjusted, cluster-robust covariance estimator used for statistical inference with dependent observations. “using CR2 covariance with Satterthwaite degrees of freedom”
  • Degrees of freedom: The number of independent values available for estimating statistical quantities. “using CR2 covariance with Satterthwaite degrees of freedom”
  • Equivalence bounds: Predefined limits within which two treatments are considered practically equivalent. “We use equivalence bounds of ±\pm0.25 pooled standard deviations of learning gains”
  • Equivalence test: A statistical test determining whether an observed difference lies within a prespecified range of practically negligible differences. “We test equivalence using standard two one-sided tests at α\alpha=0.05”
  • Expert rubric evaluation: A structured assessment in which specialists judge outputs according to predefined criteria. “expert human tutors to compare LLM-generated lesson plans and practice problems through pairwise rubric evaluations”
  • Frontier model: A highly capable model operating near the current leading edge of performance. “Our pooled AI result spans 13 capability-varied AI tutors, including non-frontier and open-weight models.”
  • Harness: Software infrastructure that controls, evaluates, or supports an AI system’s execution. “changes in harnesses and software scaffolding led to substantial performance gains”
  • Heteroskedasticity-consistent: Designed to provide valid statistical estimates when the variance of errors differs across observations. “Figure 1C shows learning gains for 12 AI tutors, adjusted for pre-test score and section, with 95%95\% HC3 intervals.”
  • Inference cost: The computational or monetary cost of generating outputs from a trained AI model. “Its mean inference cost was $0.067 for the entire tutoring session”
  • Intelligent tutoring system: A computer-based instructional system that adapts teaching to a learner’s knowledge or responses. “Earlier intelligent tutoring systems have produced learning outcomes comparable to human tutoring”
  • Item response theory (IRT): A family of psychometric models that estimates a learner’s ability from responses to test items. “we used item response theory (IRT) to estimate proficiency on the Quantitative pre-test”
  • Latency: The time between a request and the system’s response. “Across the 12 AI tutors, lower AI latency was strongly associated with higher student engagement”
  • Learning gain: The change in assessment performance between a post-test and a pre-test. “We define learning gain as post-test score minus pre-test score”
  • Longitudinal study: A study that follows participants or outcomes over an extended period. “One limitation of this study is that we measured immediate learning gains without evaluating whether they persist over months.”
  • Metacognitive strategy: A learning approach involving awareness and regulation of one’s own thinking and learning processes. “An effective metacognitive strategy: learning by doing and explaining with a computer-based Cognitive Tutor.”
  • Minimal software scaffolding: Limited auxiliary software support that structures or guides an AI system’s behavior. “We use minimal software scaffolding and prompt adjustments”
  • Mixed-initiative dialogue: An interaction in which both participants can independently initiate contributions. “AutoTutor: An Intelligent Tutoring System With Mixed-Initiative Dialogue”
  • Open-weight model: An AI model whose trained parameters are publicly available for use or modification. “including non-frontier and open-weight models”
  • Pareto frontier: The set of options that cannot be improved on one objective without worsening another. “we explore the Pareto frontier of learning gain, cost, and latency”
  • Pareto-dominate: To be superior to another option on at least one objective while being no worse on the others. “Gemini 3.1. Pro had higher mean learning gain, lower inference cost, and faster replies than GPT-5.5. Pro”
  • Pedagogy: The theory and practice of teaching. “We also evaluate pedagogical characteristics of tutoring conversations”
  • Psychometric: Relating to the measurement of psychological traits or abilities, such as knowledge or proficiency. “Its expert ratings of lesson plans and practice design complement psychometric evaluations of AI-generated exam questions”
  • Recursive human self-improvement (RHSI): A proposed process in which improved human abilities enable better use and further development of teaching technologies. “We use the term recursive human self-improvement (RHSI) for this process”
  • Reward hacking: Exploiting the specification of an objective to obtain a high measured reward without achieving the intended goal. “To avoid reward hacking in AI tutors”
  • Scaffolding: Temporary instructional support that helps learners perform tasks they could not yet complete independently. “StudentBench examines related aspects of tutoring dialogue: scaffolding cues”
  • Satterthwaite degrees of freedom: An approximation used to estimate degrees of freedom for inference with unequal or complex variance structures. “using CR2 covariance with Satterthwaite degrees of freedom”
  • Singularity: A hypothetical point at which an AI system’s intelligence or capability undergoes an irreversible, transformative change. “A singularity in intelligence can occur when an AI technology recursively self-improves”
  • Spearman correlation: A rank-based measure of the strength and direction of a monotonic relationship between two variables. “lower AI latency was strongly associated with higher student engagement (Spearman ρ=−0.81\rho = -0.81”
  • Statistical equivalence: A conclusion that two conditions differ by no more than a prespecified practically meaningful margin. “AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains”
  • Student-tt confidence interval: A confidence interval based on the Student’s tt distribution, commonly used when estimating a mean with limited or unknown population variance. “Error bars depict 95%95\% Student- tt confidence intervals.”
  • Two one-sided tests (TOST): An equivalence-testing procedure that tests whether an effect is both above a lower bound and below an upper bound. “We test equivalence using standard two one-sided tests at α\alpha=0.05”
  • Zone of proximal development: The range of tasks a learner can perform with appropriate assistance but not yet independently. “Students' performance with and without guidance can reveal their zone of proximal development”

Tweets

Sign up for free to view the 13 tweets with 232 likes about this paper.