In minutes, not weeks.
Four decades of tutoring research, from Bloom's two sigma to the first randomized trials of human-AI teams, and what the effect sizes really mean.
The debate over AI versus human tutors tends to be framed as new. It is not. Researchers have spent forty years measuring what individual tutoring does to learning, and roughly thirty of those years measuring what computers can do in the tutor's place. The result is one of the better-quantified questions in education, and the honest answer is more interesting than either the boosters or the skeptics usually allow.
The short version: expert human tutoring is very effective, the best computer-based tutoring systems come surprisingly close, and the frontier in 2026 is not "which one wins" but "what happens when you combine them." This article walks through the evidence in the order it accumulated.
Any discussion of tutoring starts with Benjamin Bloom's 1984 paper. Bloom reported that students tutored one-to-one with mastery learning scored about two standard deviations higher than conventionally taught peers — placing the average tutored student above 98% of the control group. The figure is real, but it is routinely misused.
Two sigma came from small, tightly controlled studies with expert tutors, mastery conditions, and favorable measurement. It is the ceiling, not the norm, and later research has never reliably reproduced it at scale. Treating 2.0 as the benchmark AI tutors must hit is a category error. The more useful benchmarks come from the meta-analyses that followed, which measured tutoring under more realistic conditions and consistently landed lower — while still confirming Bloom's central claim that individual attention is powerful.
The single most cited head-to-head is Kurt VanLehn's 2011 review in Educational Psychologist, which synthesized decades of studies comparing human tutoring, computer tutoring, and no tutoring. VanLehn's numbers reset expectations that had been anchored to Bloom's two sigma.
Against a no-tutoring baseline, VanLehn found:
The finding that mattered was the near-parity between human tutors and well-built step-based systems. When VanLehn compared them directly, the advantage of human tutors over step-based ITS was small (roughly d = 0.21) and in some comparisons vanished. His interpretation was that the effectiveness of tutoring comes less from the tutor being human and more from the granularity of interaction — frequent steps, feedback, and adaptive difficulty — which a well-designed system can supply.
VanLehn's central finding: step-based AI tutoring (0.76σ) came within a rounding error of human tutoring (0.79σ). The mechanism, not the medium, drives the effect.
Two later meta-analyses tested whether VanLehn's result held across larger bodies of evidence.
Ma, Adesope, Nesbit, and Liu (2014), in the Journal of Educational Psychology, pooled 107 comparisons and found that intelligent tutoring systems produced an effect of about 0.41 standard deviations relative to conventional classroom instruction, and performed comparably to small-group and individual human tutoring in most contrasts. Kulik and Fletcher (2016), in Review of Educational Research, examined 50 controlled evaluations and reported a median effect of roughly 0.66, noting that ITS outperformed conventional classes in the large majority of studies and, in many, matched human tutoring.
The precise numbers vary with the comparison baseline — ITS versus no tutoring reads higher than ITS versus an already-good classroom — but the pattern across three independent syntheses is stable. Well-engineered tutoring systems land in the same broad range as human tutors and comfortably above ordinary classroom instruction. By Kraft's benchmark, where 0.20 is "large" for education, effects in the 0.4 to 0.8 range are exceptional.
| Study | What it measured | Effect (SD) |
|---|---|---|
| Bloom (1984) | 1:1 mastery tutoring vs. conventional (upper bound) | ~2.0 |
| VanLehn (2011) | Human tutoring vs. no tutoring | 0.79 |
| VanLehn (2011) | Step-based ITS vs. no tutoring | 0.76 |
| Kulik & Fletcher (2016) | ITS vs. conventional (median) | 0.66 |
| Ma et al. (2014) | ITS vs. conventional instruction | 0.41 |
| Nickow et al. (2020) | Structured human tutoring, PreK-12 | 0.37 |
Reading across the evidence, the strengths of AI-based tutoring are consistent and specific:
The 2025 Harvard physics RCT by Kestin and Miller is the strongest recent demonstration: a purpose-built tutor produced more than twice the learning gains of an already-strong active-learning class, in less time. That result depended entirely on the tutor being engineered around sound pedagogy — a caveat the next section makes central.
The same literature is clear about the limits. Human tutors retain advantages that current systems approximate weakly or not at all:
The framing of "AI versus human" may already be the wrong one. The strongest field evidence in 2024 came from pairing them.
Stanford's Tutor CoPilot study ran the first randomized controlled trial of a human-AI tutoring system in live sessions, involving roughly 900 tutors and about 1,800 K-12 students in under-resourced communities. Tutors were given real-time AI-generated suggestions — expert-style probing questions and next moves — which they could accept, edit, or ignore. Students of tutors using the tool were about 4 percentage points more likely to master the topic, and the gain rose to roughly 9 points for students working with lower-rated tutors. The tool cost about $20 per tutor per year.
The mechanism is telling. Analysis of more than 350,000 messages showed the AI shifted tutor behavior toward better pedagogy: more probing questions, less generic praise, fewer instances of simply handing over the answer. In other words, the AI did not replace the human tutor — it made the median human tutor behave more like an expert one, and it helped the weakest tutors most. That is a different and more defensible value proposition than substitution.
The best evidence does not favor AI over humans or humans over AI. It favors human tutors equipped with AI — which raised topic mastery most for the students working with the weakest tutors.
Four decades of research converge on a defensible position. Human tutoring is one of the most effective interventions education has (roughly 0.79 SD in VanLehn's synthesis). The best computer tutoring systems come remarkably close (0.66 to 0.76 SD across studies), because the mechanism that makes tutoring work — frequent, adaptive, step-level feedback — is something software can deliver. Neither reliably reaches Bloom's two sigma, which was always a ceiling rather than a target.
What AI adds is scale and availability; what it lacks is motivational judgment and, in generative form, factual reliability. The pattern in the newest and most rigorous studies is that the highest-value deployment is not one replacing the other, but AI amplifying human tutors — closing the gap between average and expert practice, and doing the most for the students who have the least support. For institutions designing tutoring at scale, that is the finding to build around.
See how LearnLab turns coursework into a measurable loop — automatic grading, conversational tutors, and per-student analytics built for educators who want to teach, not grade.
Our Compliance & Security Standards
Hosted on AWS infrastructure with SOC 2 Type II & ISO 27001 certified data centers — with data residency available across EU, Middle East, and other regions
Across WorkLab and LearnLab, our AI assists — it does not replace — human hiring, grading, and academic decisions. The "EU AI Act Aligned" badge reflects alignment with the Act's principles by design — transparency, human oversight, and documentation — not a certification. Learn more.
Upstack.AIFilter the noise. Interview real candidates. One link works anywhere—no ATS migration needed.
Upstack AI FZ-LLC
FOAM2471, Compass Building
Al Shohada Road, AL Hamra Industrial Zone-FZ
Ras Al Khaimah, United Arab Emirates
Microsoft Store
Publisher: UPSTACK AI
Store ID: 9NT2GR4TDZ0G
Powered By
Last updated: 21/1/2026