Ground Truth.
AI, checked against the source.

News · 2026-09-25

StudentBench finds AI and human GRE tutors produced equivalent immediate gains

StudentBench found that AI tutoring and expert human tutoring produced statistically equivalent immediate learning gains for adults preparing for the GRE in a one-hour study. The result is a serious evidence point for structured, feedback-rich exam practice, but it does not demonstrate that chatbots can replace schooling, child supervision or the long-term human work of keeping a learner engaged. Its narrowness is exactly what makes it useful.

Key facts

The design is refreshingly concrete. Participants first completed either Quantitative or Verbal GRE-style questions, reviewed their mistakes briefly, then spent an hour in their assigned condition. AI tutors saw the pre-test and mistakes, planned lessons and generated practice problems in text conversation. Human tutors taught live one-to-one video sessions. The post-test used different questions over the same skills and was completed without AI. The study’s title, “AI and human tutoring yield equivalent GRE learning gains,” says exactly what it tested.

The AI arm was not one chatbot. It used 13 configurations across 12 models, including Gemini, GPT, Claude, Kimi and Gemma variants. The paper pools the AI systems for the principal comparison. Its cheapest individual tutor to pass the authors’ equivalence test, Gemma 4 31B, had an estimated inference cost of about $0.067 per session, versus the study’s $75-per-hour human-tutor reference. The paper’s roughly 918-fold calculation compares inference spending with a labor reference; it is not the retail price of a complete tutoring service, with customer support, curriculum work, safety and marketing included.

A good analogy is a driving lesson on a closed course. If the question is whether a learner can practice turns, get immediate correction and improve on the same kind of road soon afterward, the experiment is strong. If the question is whether the instructor can make a reluctant teenager arrive every week, notice distress, coordinate with parents, build habits or prepare someone to drive unaided in every weather condition, the study is not set up to answer it.

The control condition is another important limit. It watched educational video unrelated to the GRE; it was not a diligent student doing matched independent practice with a book or worksheet. The human comparator was much smaller—140 analyzed sessions—than the 2,139 AI sessions. Participants were paid and screened for effort, the post-test followed immediately, and the research does not measure retention over weeks or months. Equivalence held for pooled results and Quantitative under stated bounds, but not for Verbal alone under every stated analysis.

The paper also separates learning from preference. Experts reviewed 2,028 pairs of AI-generated lesson plans and practice problems, but expert preference for a lesson plan is not the same outcome as a student’s unaided performance. Conversely, a post-test score does not report whether the student felt encouraged, trusted the explanation or would return voluntarily. Those distinctions should shape product claims.

The timing is commercially resonant. Dymocks Education’s closure FAQ says its last trading day is 27 September and suggests families make a modest investment in Gemini or ChatGPT. The company’s decision is a market judgement about value and affordability, not a randomized trial of its tutors. A hybrid-tutoring response from competitors makes the strongest counterargument: explanation is only one component of education; structure, accountability and exam technique are often the product families buy.

The result matters because it identifies where substitution may begin. Motivated adults tackling a bounded syllabus, with immediate feedback and plentiful practice, are a favorable setting for AI. The study’s 6.15-point improvement over control shows a meaningful result, not merely a pleasant conversation. But the appropriate conclusion is conditional substitution. AI may absorb some tutoring explanation and drill at extraordinarily low marginal cost; it has not been shown here to replace human relationships or the environments that make learning happen. For the surrounding economics, see inference cost and token economics.


Primary source, verified: read the paper → (arXiv 2609.28470)

Key questions

Did AI beat human tutors in StudentBench?

No. The pooled result found statistical equivalence for immediate GRE learning gains, with an adjusted AI-minus-human difference of -0.58 percentage points.

Who was studied?

The analysis covered 2,469 sessions from 2,383 English-speaking adults, mostly young adults, rather than schoolchildren.

Does the study show children no longer need tutors?

No. It measured one-hour GRE preparation and an immediate post-test, not motivation, safeguarding, long-term retention or broader schooling.
Cite this

APA

Ground Truth. (2026, September 25). StudentBench finds AI and human GRE tutors produced equivalent immediate gains. Ground Truth. https://groundtruth.day/news/studentbench-ai-human-gre-tutoring.html

BibTeX

@misc{groundtruth:studentbench-ai-human-gre-tutoring,
  title  = {StudentBench finds AI and human GRE tutors produced equivalent immediate gains},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {sep},
  url    = {https://groundtruth.day/news/studentbench-ai-human-gre-tutoring.html}
}

Topics: education · evaluation · tutoring · agents · research

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.