StudentBench: Comparative AI Human Tutoring Study
- StudentBench is an evaluation suite for comparing AI tutors, expert human tutors, and a no-tutoring control to measure learning gains through standardized GRE assessments.
- The platform assessed 2383 participants through a rigorous study design, involving more than 2,400 sessions and 175,000 student-AI messages.
- Both AI and human tutoring achieved statistically equivalent immediate learning gains, with AI tutoring being significantly more cost-effective in some cases.
StudentBench is a public platform and evaluation suite for measuring whether artificial-intelligence tutoring improves students’ unaided performance. Its principal benchmark compares AI tutoring, expert human tutoring, and a no-tutoring control using GRE Quantitative and Verbal Reasoning assessments. Rather than evaluating only answer accuracy, StudentBench separates learning gains, lesson planning, practice-problem generation, conversational pedagogy, cost, and engagement. The main study analyzed 2,383 adult participants, 2,469 sessions, more than 175,000 student–AI messages, and 45,463 answered practice problems (Northcutt et al., 23 Sep 2026).
1. Purpose, scope, and platform
StudentBench addresses the question of whether AI tutors produce measurable human learning gains comparable to those produced by expert human tutors. The benchmark’s primary outcome is improvement in unaided post-test performance after a tutoring session, rather than the quality of generated explanations or the tutor’s ability to answer questions correctly.
The platform uses relatively general-purpose LLMs with minimal software scaffolding and two principal prompt templates: one for lesson planning and practice-problem generation, and another for interactive tutoring. Its design is intended to measure teaching ability rather than highly specialized tutoring-system engineering. StudentBench is also presented in relation to recursive human self-improvement: teaching technology may increase human capabilities, which may subsequently enable more effective use and development of technology.
The benchmark evaluates AI tutors as complete configurations comprising a model, prompting and reasoning settings, practice generation, interaction behavior, and response latency. Consequently, its comparisons do not isolate model capability from the surrounding tutoring configuration. The paper evaluates 13 AI tutors per GRE section, representing 12 models, with one model/version combination differing between sections. AI assignment uses least-fill randomization so that tutors with fewer assignments receive subsequent participants.
StudentBench is available as a public platform at https://studentbench.org. Its empirical results concern immediate GRE learning gains and associated tutoring behaviors; they do not establish long-term retention or general educational effectiveness outside the study setting.
2. Learning-gain study design
The principal study involved 2,383 unique adult participants and 2,469 analyzed sessions: 1,291 Quantitative sessions and 1,178 Verbal sessions. Eighty-six participants completed both sections. The median age was 21, and 83.3% of participants were between 18 and 24 years old. Participants were recruited through Handshake from more than 70,000 randomly selected students and were paid $50 independently of performance.
Participants were randomly assigned to AI tutoring, human tutoring, or no-tutoring control conditions. The appendix reports the following condition totals:
| GRE section | AI tutoring | Human tutoring | No-tutoring control |
|---|---|---|---|
| Quantitative | 1,139 | 61 | 91 |
| Verbal | 1,000 | 79 | 99 |
| Total | 2,139 | 140 | 190 |
The apparent difference between the participant, session, and condition totals reflects participants completing both sections and different filtering and linkage conventions in the analyses.
The assessment covered seven GRE domains:
- Quantitative: data analysis, geometry, arithmetic, and algebra.
- Verbal: sentence equivalence, text completion, and reading comprehension.
Each section contained 27 questions. Quantitative sessions allowed 47 minutes, while Verbal sessions allowed 41 minutes. Former GRE question writers from ETS and Kaplan created new questions matching the difficulty, coverage, ordering, categories, and formats of GRE questions. Two assessment forms, P and Q, were counterbalanced: approximately half of participants took P followed by Q, and the remainder took Q followed by P. This design used different pre- and post-test questions covering comparable concepts while reducing confounding from form order.
The session protocol consisted of a pre-test, approximately five minutes reviewing incorrect pre-test answers, one hour in the assigned condition, and a post-test. Unanswered questions counted as incorrect.
AI tutoring
AI tutors received the student’s pre-test answers, correctness, response times, and mistakes, but not the post-test. The AI selected concepts, sequencing, time allocation, practice problems, and explanations. The application displayed practice problems, collected answers, and returned relevant information to the model.
The AI condition used relatively low-guidance prompts. It did not require detailed misconception diagnosis, explicit prerequisite sequencing, or strongly prescribed attempt-before-hint routines. The reported results therefore constitute a low-guidance baseline rather than a ceiling on AI tutoring capability.
Human tutoring
Human tutors conducted live, one-to-one video sessions. They were former ETS or Kaplan employees involved in GRE question writing or tutors with at least five years of GRE tutoring experience. Nine human tutors participated per section. They received students’ pre-test mistakes and were asked to prioritize concepts, allocate time, and focus instruction accordingly. They did not see post-test questions.
Control condition
Control participants reviewed their graded pre-test for approximately five minutes and then watched educational videos unrelated to the GRE. The study did not include an independent-practice condition in which students solved GRE problems without a tutor. Accordingly, the estimated tutoring effects are relative to the no-tutoring control rather than to ordinary self-study practice.
3. Learning-gain measurement and AI–human equivalence
StudentBench defines individual learning gain in percentage points as
where is pre-test percentage correct and is post-test percentage correct.
Because baseline performance could differ between groups, the primary analysis used ANCOVA:
The model included a quadratic baseline term, tutoring condition, section, and assessment form where appropriate. Confidence intervals used HC3 covariance. AI–human comparisons additionally accounted for repeated participants and shared human tutors using CR2 covariance and Satterthwaite degrees of freedom.
Both AI and human tutoring outperformed the no-tutoring control:
| Section | AI minus control | Human minus control |
|---|---|---|
| Quantitative | 6.86 percentage points | 6.27 percentage points |
| Verbal | 5.47 percentage points | 7.52 percentage points |
| Combined | 6.15 percentage points | 7.04 percentage points |
The reported confidence intervals were [4.02, 9.69] for combined AI-minus-control gain and [3.88, 10.20] for combined human-minus-control gain; both comparisons had .
The main AI–human analysis used equivalence testing rather than a conventional test of whether the two means differed. The equivalence margin was pooled standard deviations, corresponding to approximately percentage points in the combined analysis. Equivalence required the entire 90% confidence interval for the AI-minus-human difference to lie inside these bounds.
The pooled adjusted AI-minus-human difference was
percentage points, with
Because the interval lies within the prespecified equivalence interval, the paper reports AI and human tutoring as statistically equivalent for pooled immediate GRE learning gains. The result does not show that AI and human tutoring are identical or that AI tutoring is superior.
The Quantitative-specific comparison also satisfied equivalence:
- AI-minus-human difference: $0.88$ percentage points;
- 90% confidence interval: 0;
- 1.
The Verbal-specific comparison did not meet the paper’s equivalence criterion:
- difference: 2 percentage points;
- 90% confidence interval: 3;
- 4.
Thus, the strongest equivalence claim is pooled across sections rather than unequivocal equivalence within every individual section.
One individual tutor, Gemma 4 31B, also passed the tutor-specific equivalence test:
- AI mean learning gain: 12.742 percentage points;
- human mean learning gain: 15.608 percentage points;
- AI-minus-human difference: 5 percentage points;
- 90% confidence interval: 6;
- equivalence-test result: 7.
Six of the 12 individually tested AI tutors passed the specified equivalence criterion. These model-specific tests were not corrected across the 12 tutors and are therefore more exploratory than the pooled primary analysis.
4. Teaching-material evaluation and conversational pedagogy
StudentBench’s second study assessed AI-generated lesson plans and practice problems through expert pairwise comparisons rather than directly measuring student learning.
The study included 51 expert human tutors, 2,028 pairwise reviews, 381 student pre-tests, and 852 distinct comparison tasks. Reviewers saw a student’s graded pre-test and two anonymized lesson plans generated by different AI tutors. The reviews produced 9,577 lesson-planning ratings and 5,265 practice-problem ratings.
Pairwise preferences were modeled using the Bradley–Terry formulation:
8
Lesson plans were evaluated on:
- relevant concepts;
- concept grouping;
- concept prioritization;
- time allocation;
- test-taking strategies.
Practice problems were evaluated on:
- practice alignment;
- appropriate difficulty;
- example accuracy.
Anthropic models, particularly Opus models, performed strongly in expert preferences for both lesson planning and practice-problem design. These results measure perceived quality of teaching materials rather than directly measured learning. Expert preferences for practice design correlated with fewer student disputes in Quantitative tutoring (9, 0) and in the combined analysis (1, 2). For example accuracy, the corresponding correlations were 3 in Quantitative tutoring and 4 in the combined analysis. These are exploratory associations and do not establish that expert-rated quality caused fewer disputes.
StudentBench also evaluates conversational pedagogy using six deterministic indicators:
- scaffolding cues;
- requests for explanation;
- interactive questions;
- an early request for the student to attempt a problem;
- a long solution following a student reply;
- reasoning checks.
For session 5, the binary indicator for criterion 6 is 7. The session score is
8
and the tutor-level score is 9.
This index is an automatically detected behavioral measure, not a validated causal measure of teaching effectiveness. It favors longer or more turn-rich conversations. The corpus contained 2,139 AI sessions and 135 human sessions with usable recordings. Direct comparison is imperfect because human tutoring was spoken video tutoring whereas AI tutoring was typed chat.
The paper reports model-family clustering in conversational behavior, suggesting common provider-level training or system-design effects. It does not establish that higher conversational-pedagogy scores caused larger learning gains.
5. Engagement, latency, cost, and domain performance
Student engagement was primarily measured through student-authored chat activity, including the number of student messages, messages containing at least five tokens, total student tokens, attempted practice problems, correct or credited practice, and helpfulness ratings from 1 to 5.
Across 12 shared AI tutors, faster replies were associated with more student messages:
0
The negative sign reflects the fact that longer response times corresponded to fewer student messages. AI reply times ranged from 1.9 seconds for GPT-5.4 mini to 31.0 seconds for GPT-5.5 Pro. The latency measure was model-provider call duration rather than necessarily the full browser-perceived delay or time to first streamed token.
For Quantitative tutoring, adjusted observational associations formed the following chain:
1
The reported effects were:
| Relationship | Effect | 2 |
|---|---|---|
| Reply time to student messages | 3 messages per 10 seconds | 4 |
| Student messages to correct practice | 5 correct problems per 10 messages | 6 |
| Correct practice to learning gain | 7 percentage points per 10 correct problems | 8 |
These are adjusted observational associations, not randomized mediation effects. They do not prove that reducing latency would necessarily produce the estimated learning gain. The corresponding Verbal relationships were not statistically significant under the reported analyses.
The abstract reports that the best AI tutor surpassed the human tutor in five of the seven GRE domains. The domains in which the best AI mean exceeded the human mean were data analysis, geometry, arithmetic, text completion, and reading comprehension. Human tutoring had the higher mean in algebra and sentence equivalence. These are domain-level frontier comparisons, not evidence that one AI tutor was significantly superior to human tutoring in every listed domain.
StudentBench also reports aggregate cost per percentage point gained:
9
where 0 is mean session cost and 1 is mean learning gain in percentage points.
For Gemma 4 31B:
2
yielding approximately 375 reference cost and a 15.608-point mean gain, the corresponding value was approximately $4.81. The ratio was reported as approximately 918 times lower cost for Gemma 4 31B.
This metric is an aggregate ratio rather than the marginal price of causing an individual student to improve by exactly one percentage point. It excludes personnel, platform, development, monitoring, safety, and other deployment costs, and concerns immediate post-test improvement rather than long-term learning.
6. Methodological significance and limitations
StudentBench separates several dimensions that are often conflated in AI-tutoring evaluation: learning outcomes, planning quality, practice generation, interactional behavior, engagement, latency, and cost. Its central methodological feature is the use of an outcome-based comparison with expert human tutoring rather than reliance on answer accuracy or subjective fluency alone.
The study also demonstrates that AI tutors can be evaluated at multiple levels:
- Student outcomes: pre-test to post-test learning gain.
- Instructional planning: relevance, grouping, prioritization, time allocation, and test-taking strategy.
- Practice generation: alignment, difficulty, and answer accuracy.
- Conversation: scaffolding, explanation requests, questioning, reasoning checks, and early student attempts.
- Engagement: messages, practice attempts, correct practice, and helpfulness.
- Operational deployment: latency and cost.
Several limitations constrain interpretation. Learning was measured immediately after a one-hour GRE session, so the study does not establish retention over weeks or months. Participants were primarily young adult English readers and writers recruited through Handshake and may not represent children, older learners, non-English learners, students without reliable devices, or ordinary classroom populations.
The domain is also narrow. The findings concern structured GRE preparation and may not generalize to open-ended writing, laboratory learning, conceptual science education, long-term curricula, collaborative learning, or professional training. Claims about other examinations such as the SAT or MCAT are implications rather than results established by the study.
The human-tutoring condition was comparatively small, comprising 61 Quantitative sessions and 79 Verbal sessions. Combined analyses accounted for shared tutors and repeated students, producing 1.59 effective degrees of freedom in the combined comparison. Sensitivity analyses preserved some equivalence results, but a gain-only sensitivity analysis using an alternative clustering specification did not establish equivalence in either section.
AI sessions were required to include at least three student messages, five practice-answer attempts, and three distinct attempted practice problems. These filters exclude incomplete, low-effort, rapid, or protocol-deviating sessions. The findings therefore apply to participants who completed a minimally engaged session rather than necessarily to everyone who opens an AI tutor.
The absence of an independent-practice control prevents separation of the effects of AI dialogue, structured practice, explanations, and simply spending an hour on GRE-related activity. AI tutors also differed in model family, reasoning settings, response speed, output length, practice generation, and conversational behavior, so the study compares complete tutor configurations rather than pure model capability.
The conversational-pedagogy indicators are deterministic proxies and were not validated against learning lift in the dataset. Spoken human tutoring and typed AI tutoring differ in modality. Likewise, expert evaluations of lesson plans and practice problems measure perceived quality rather than directly measured learning.
Finally, numerous exploratory comparisons were performed across domains, tutors, rubrics, latency, practice, and subgroups. Holm correction was applied within specified comparison families, but the individual AI equivalence tests were not corrected across tutors.
StudentBench’s principal conclusion is therefore bounded but specific: in the reported randomized GRE experiment, pooled AI tutoring produced immediate learning gains statistically equivalent to expert human tutoring under prespecified equivalence bounds. The platform’s broader contribution is an evaluation architecture that treats tutoring as a multidimensional educational intervention, combining outcome measurement with analysis of pedagogy, engagement, latency, cost, and instructional materials.