In minutes, not weeks.

Learning Science

AI Tutors vs Human Tutors: What the Research Says

Four decades of tutoring research, from Bloom's two sigma to the first randomized trials of human-AI teams, and what the effect sizes really mean.

Upstack AI ResearchApril 2, 202613 min read
  • Home
  • Insights
  • AI Tutors vs Human Tutors: What the Research Says
Back to all insights
Overview

The question is older than the technology

The debate over AI versus human tutors tends to be framed as new. It is not. Researchers have spent forty years measuring what individual tutoring does to learning, and roughly thirty of those years measuring what computers can do in the tutor's place. The result is one of the better-quantified questions in education, and the honest answer is more interesting than either the boosters or the skeptics usually allow.

The short version: expert human tutoring is very effective, the best computer-based tutoring systems come surprisingly close, and the frontier in 2026 is not "which one wins" but "what happens when you combine them." This article walks through the evidence in the order it accumulated.

0.79σ
Human tutoring vs. no tutoring
VanLehn, 2011
0.76σ
Best step-based AI tutoring systems
VanLehn, 2011
0.66σ
Median ITS effect on achievement
Kulik & Fletcher, 2016
+4 pts
Extra topic mastery from human+AI teams
Tutor CoPilot RCT, 2024
Part 1

Bloom's two sigma, and why it isn't the number to quote

Any discussion of tutoring starts with Benjamin Bloom's 1984 paper. Bloom reported that students tutored one-to-one with mastery learning scored about two standard deviations higher than conventionally taught peers — placing the average tutored student above 98% of the control group. The figure is real, but it is routinely misused.

Two sigma came from small, tightly controlled studies with expert tutors, mastery conditions, and favorable measurement. It is the ceiling, not the norm, and later research has never reliably reproduced it at scale. Treating 2.0 as the benchmark AI tutors must hit is a category error. The more useful benchmarks come from the meta-analyses that followed, which measured tutoring under more realistic conditions and consistently landed lower — while still confirming Bloom's central claim that individual attention is powerful.

Part 2

The landmark comparison: VanLehn, 2011

The single most cited head-to-head is Kurt VanLehn's 2011 review in Educational Psychologist, which synthesized decades of studies comparing human tutoring, computer tutoring, and no tutoring. VanLehn's numbers reset expectations that had been anchored to Bloom's two sigma.

Against a no-tutoring baseline, VanLehn found:

  • Human tutoring: an effect of about 0.79 standard deviations — large, but well short of two sigma.
  • Step-based intelligent tutoring systems: about 0.76 — nearly indistinguishable from human tutors.
  • Substep-based (finer-grained) systems: about 0.40 — meaningful, but a clear tier below.

The finding that mattered was the near-parity between human tutors and well-built step-based systems. When VanLehn compared them directly, the advantage of human tutors over step-based ITS was small (roughly d = 0.21) and in some comparisons vanished. His interpretation was that the effectiveness of tutoring comes less from the tutor being human and more from the granularity of interaction — frequent steps, feedback, and adaptive difficulty — which a well-designed system can supply.

VanLehn's central finding: step-based AI tutoring (0.76σ) came within a rounding error of human tutoring (0.79σ). The mechanism, not the medium, drives the effect.

Part 3

The meta-analyses that followed

Two later meta-analyses tested whether VanLehn's result held across larger bodies of evidence.

Ma, Adesope, Nesbit, and Liu (2014), in the Journal of Educational Psychology, pooled 107 comparisons and found that intelligent tutoring systems produced an effect of about 0.41 standard deviations relative to conventional classroom instruction, and performed comparably to small-group and individual human tutoring in most contrasts. Kulik and Fletcher (2016), in Review of Educational Research, examined 50 controlled evaluations and reported a median effect of roughly 0.66, noting that ITS outperformed conventional classes in the large majority of studies and, in many, matched human tutoring.

The precise numbers vary with the comparison baseline — ITS versus no tutoring reads higher than ITS versus an already-good classroom — but the pattern across three independent syntheses is stable. Well-engineered tutoring systems land in the same broad range as human tutors and comfortably above ordinary classroom instruction. By Kraft's benchmark, where 0.20 is "large" for education, effects in the 0.4 to 0.8 range are exceptional.

Tutoring effects across the major syntheses

StudyWhat it measuredEffect (SD)
Bloom (1984)1:1 mastery tutoring vs. conventional (upper bound)~2.0
VanLehn (2011)Human tutoring vs. no tutoring0.79
VanLehn (2011)Step-based ITS vs. no tutoring0.76
Kulik & Fletcher (2016)ITS vs. conventional (median)0.66
Ma et al. (2014)ITS vs. conventional instruction0.41
Nickow et al. (2020)Structured human tutoring, PreK-120.37
Part 4

What AI tutors do well

Reading across the evidence, the strengths of AI-based tutoring are consistent and specific:

  • Availability and dosage. The Nickow meta-analysis found tutoring effects scale with frequency; the binding constraint on human tutoring is hours. A system that is available at 11 p.m. before an exam delivers dosage a human schedule cannot.
  • Immediate, step-level feedback. VanLehn's analysis attributes much of tutoring's power to fine-grained feedback loops — precisely what well-built systems automate reliably and tirelessly.
  • Adaptive practice and mastery pacing. Systems can hold a student on a concept until mastery without the social cost of doing so in front of peers.
  • Consistency. A system does not have bad days, and its quality does not vary with the tutor's own subject expertise.

The 2025 Harvard physics RCT by Kestin and Miller is the strongest recent demonstration: a purpose-built tutor produced more than twice the learning gains of an already-strong active-learning class, in less time. That result depended entirely on the tutor being engineered around sound pedagogy — a caveat the next section makes central.

Part 5

What AI tutors do poorly

The same literature is clear about the limits. Human tutors retain advantages that current systems approximate weakly or not at all:

  • Motivation and relationship. A large part of what a human tutor supplies is encouragement, accountability, and the read on when a student is discouraged versus confused. VanLehn's parity finding was about cognitive outcomes measured over short spans, not the motivational scaffolding that sustains a struggling student across a term.
  • Open-ended and ill-structured domains. The strongest ITS results come from well-structured subjects — algebra, physics, programming — where correct steps are definable. Essay revision, historical argument, and design critique are harder, and effect sizes there are thinner.
  • Factual reliability. Rule-based tutoring systems did not invent facts; generative chat tutors can. This is a genuine regression that older ITS research does not account for, and it is the subject of the third article in this series.
  • Reading the room. A skilled tutor detects misconception from a facial expression or a hesitant pause. Systems infer only from what a student types.
Part 6

The most promising result: humans and AI together

The framing of "AI versus human" may already be the wrong one. The strongest field evidence in 2024 came from pairing them.

Stanford's Tutor CoPilot study ran the first randomized controlled trial of a human-AI tutoring system in live sessions, involving roughly 900 tutors and about 1,800 K-12 students in under-resourced communities. Tutors were given real-time AI-generated suggestions — expert-style probing questions and next moves — which they could accept, edit, or ignore. Students of tutors using the tool were about 4 percentage points more likely to master the topic, and the gain rose to roughly 9 points for students working with lower-rated tutors. The tool cost about $20 per tutor per year.

The mechanism is telling. Analysis of more than 350,000 messages showed the AI shifted tutor behavior toward better pedagogy: more probing questions, less generic praise, fewer instances of simply handing over the answer. In other words, the AI did not replace the human tutor — it made the median human tutor behave more like an expert one, and it helped the weakest tutors most. That is a different and more defensible value proposition than substitution.

The best evidence does not favor AI over humans or humans over AI. It favors human tutors equipped with AI — which raised topic mastery most for the students working with the weakest tutors.

The honest summary

Four decades of research converge on a defensible position. Human tutoring is one of the most effective interventions education has (roughly 0.79 SD in VanLehn's synthesis). The best computer tutoring systems come remarkably close (0.66 to 0.76 SD across studies), because the mechanism that makes tutoring work — frequent, adaptive, step-level feedback — is something software can deliver. Neither reliably reaches Bloom's two sigma, which was always a ceiling rather than a target.

What AI adds is scale and availability; what it lacks is motivational judgment and, in generative form, factual reliability. The pattern in the newest and most rigorous studies is that the highest-value deployment is not one replacing the other, but AI amplifying human tutors — closing the gap between average and expert practice, and doing the most for the students who have the least support. For institutions designing tutoring at scale, that is the finding to build around.

Transform Your Classroom

Ready to Make Learning
Measurable?

See how LearnLab turns coursework into a measurable loop — automatic grading, conversational tutors, and per-student analytics built for educators who want to teach, not grade.

Our Compliance & Security Standards

GDPR
CCPA
EU AI Act Aligned
Regional Hosting
AES-256
TLS 1.3

Hosted on AWS infrastructure with SOC 2 Type II & ISO 27001 certified data centers — with data residency available across EU, Middle East, and other regions

Across WorkLab and LearnLab, our AI assists — it does not replace — human hiring, grading, and academic decisions. The "EU AI Act Aligned" badge reflects alignment with the Act's principles by design — transparency, human oversight, and documentation — not a certification. Learn more.

Upstack.AIUpstack.AI

Filter the noise. Interview real candidates. One link works anywhere—no ATS migration needed.

Upstack AI FZ-LLC
FOAM2471, Compass Building
Al Shohada Road, AL Hamra Industrial Zone-FZ
Ras Al Khaimah, United Arab Emirates

Microsoft Store
Publisher: UPSTACK AI
Store ID: 9NT2GR4TDZ0G

© 2022 - 2025 Upstack.AI • All Rights Reserved

Powered By

React
TypeScript
Tailwind
Python
AWS
SSL

Last updated: 21/1/2026