AI Tutors Match Expert Humans at 1/900th the Cost
StudentBench measures whether large language models can produce real learning gains comparable to expert human tutoring. In a randomized study of 2,383 participants taking one-hour GRE prep sessions, AI tutoring improved post-test scores by 6.15 percentage points over control—statistically equivalent to expert human tutors who gained 7.04 points. The study goes beyond evaluating AI responses to measure unaided student performance, revealing that the lowest-cost AI tutor achieved human-level gains at roughly $0.005 per percentage point compared to $4.81 for humans, a 918-fold cost advantage that raises fundamental questions about scalable education.Script
Can artificial intelligence teach as well as an expert human tutor? StudentBench put that question to a direct test: 2,383 people, randomized GRE tutoring sessions, and a startling answer.
The researchers assigned participants to one of three conditions: AI tutoring through a chat interface, live expert human tutoring over video, or a control group watching unrelated educational content. Everyone took matching pre-tests and post-tests with newly written GRE questions, and the AI tutors received minimal scaffolding, forcing them to diagnose weaknesses, plan lessons, generate practice problems, and teach interactively.
AI tutoring increased learning gains by 6.15 percentage points over control, and human tutoring gained 7.04 points. The difference between AI and human was just 0.58 percentage points, falling well within the prespecified equivalence margin with a p-value of point zero one five.
But here is where the story turns radical. The lowest-cost AI tutor, Gemma 4 31B, achieved human-equivalent gains at roughly 1 over 918th the cost per percentage point, spending about half a cent per session compared to 75 dollars for human tutoring.
Why did these systems work so well? In quantitative tutoring, the researchers traced a mechanism: faster AI replies led to more student messages, more messages led to more correct practice problems, and more practice correlated strongly with learning gains. Latency mattered not as a quality signal but as a determinant of how much productive work students could accomplish in one hour.
The caveat is temporal: these gains were measured immediately after one hour of tutoring, leaving retention, transfer beyond GRE-style problems, and cumulative effects of repeated sessions unresolved. Still, StudentBench establishes that minimally prompted AI can match expert human instruction on immediate learning, at a cost structure that could redefine access to personalized education. Explore the full benchmark and create your own educational video at EmergentMind.com.