In minutes, not weeks.
Meta-analyses, effect sizes, and the gap between what adaptive learning can do in principle and what it delivers in practice.
The idea that instruction should adapt to the individual learner predates the computer by decades. Its most influential statement came in 1984, when Benjamin Bloom published what he called the 2 sigma problem: students tutored one-to-one and taught to mastery performed about two standard deviations better than students in a conventional class — a gap so large that the average tutored student outscored 98 percent of the comparison group.
Bloom did not present this as a solution. He presented it as a challenge. One-to-one tutoring was, he wrote, "too costly for most societies to bear on a large scale." The question he set for the field was whether group methods, later including software, could approach the same effect at a fraction of the cost. Four decades of adaptive learning research is, in large part, an attempt to answer that question. The honest summary is: partly, and with conditions.
The two-sigma figure is often quoted as though it were a single intervention. It was not. Bloom's studies separated two ingredients. Mastery learning alone — requiring students to demonstrate competence on each unit before moving on, in a group setting — produced roughly a one-sigma improvement over conventional instruction. Adding one-to-one tutoring on top of mastery learning produced the full two sigma.
That decomposition matters for anyone building adaptive software, because it isolates what is doing the work. A large share of the benefit comes from the mastery structure itself: not advancing until a concept is secure, and giving corrective feedback in between. Personalisation of pace and path is the second layer. Software can reproduce the mastery structure fairly faithfully; reproducing the responsiveness of a skilled human tutor is much harder, which is where most of the remaining gap lives.
Three well-known syntheses give a consistent, if sobering, picture of computerised adaptive tutoring. VanLehn's 2011 review compared human tutoring, intelligent tutoring systems and no tutoring, and organised systems by how finely they interact with the learner. Step-based tutors — those that give feedback on each step of a solution rather than only the final answer — reached an effect of about 0.76 over no tutoring, close to the effect of human tutors, while coarser answer-based systems delivered much less. The granularity of interaction, not the label "adaptive," predicted the benefit.
Kulik and Fletcher's 2016 meta-analysis of 50 controlled evaluations put the median effect of intelligent tutoring systems at about 0.66 standard deviations, moving a median student from the 50th to the 75th percentile — but found that measured gains were much larger on assessments aligned to the instruction than on standardised tests. Ma and colleagues, in a 2014 meta-analysis, reported a more conservative pooled effect of around g = 0.35 for both mathematics and language systems.
Across meta-analyses, the strongest predictor of an adaptive system's effect is not the branding but how finely it interacts with the learner — step-by-step feedback approaches human-tutor levels; answer-only feedback does not.
| Study | What was measured | Effect size | Plain-language gain |
|---|---|---|---|
| Bloom, 1984 | 1:1 tutoring + mastery vs. class | ~2.0 σ | Average tutored student beats ~98% of the class |
| Bloom, 1984 | Group mastery learning alone | ~1.0 σ | Beats ~84% of a conventional class |
| VanLehn, 2011 | Step-based tutoring vs. none | ~0.76 | Approaches human-tutor effectiveness |
| Kulik & Fletcher, 2016 | Intelligent tutoring systems (median) | ~0.66 SD | 50th → 75th percentile |
| Ma et al., 2014 | Intelligent tutoring systems (pooled) | ~0.35 (g) | Modest but reliable improvement |
| Pane et al., 2014 | Cognitive Tutor at scale, year 2 | ~0.20 SD | 50th → 58th percentile |
Several patterns in the literature should temper any all-purpose personalisation pitch. First, the largest effects come from narrow, well-structured domains — algebra, physics, programming — where a correct next step can be defined precisely. The evidence thins considerably for open-ended, ill-structured learning, where "the right next problem" is not computable in the same way.
Second, "learning styles" personalisation — matching content to a supposed visual, auditory or kinaesthetic preference — has repeatedly failed to find empirical support, and should not be confused with the mastery-and-feedback personalisation that does work. Third, effect sizes shrink as studies move from tightly controlled trials to real deployments. The RAND evaluation of Cognitive Tutor Algebra I found no significant effect in year one and roughly a 0.20 standard-deviation gain only in year two, once teachers had adapted their practice. Personalisation at scale is as much an implementation problem as a modelling one.
Adapting to a learner's demonstrated mastery has a strong evidence base. Adapting to a learner's supposed 'learning style' does not — the two are often marketed together and should not be.
Adaptive sequencing is not the only lever with a strong evidence base, and the systems that perform best tend to combine several. Two in particular compound with mastery structure. The first is retrieval practice: having learners actively recall material, rather than re-read it, produces durable gains that a well-designed system can schedule automatically. The second is spacing: distributing practice over time rather than massing it improves long-term retention, and an adaptive engine is well placed to time reviews of concepts a learner has not touched recently.
The significance for platform design is that "personalisation" should not be read narrowly as choosing the next difficulty level. Its higher-value form is deciding what to bring back and when — surfacing a shaky concept for retrieval at the moment forgetting is likely. This reframes the adaptive engine from a difficulty dial into a memory scheduler, and it is a use of learner data that the meta-analytic evidence on retrieval and spacing supports independently of the tutoring-systems literature.
Taken together, the research supports a specific, buildable version of adaptive learning rather than a blanket promise. The through-line from Bloom to the modern meta-analyses is that the mechanism is mastery plus responsive, fine-grained feedback — not novelty, not media richness, and not preference-matching.
Personalised and adaptive learning is not a myth, and it is not magic. It is a well-characterised effect with known boundaries. Institutions that build to those boundaries — mastery structure, granular feedback, aligned assessment, patient implementation — can capture a meaningful and real share of Bloom's original promise. Those that buy the two-sigma headline without its conditions will be disappointed by the 0.20 they get in year one.
See how LearnLab turns coursework into a measurable loop — automatic grading, conversational tutors, and per-student analytics built for educators who want to teach, not grade.
Our Compliance & Security Standards
Hosted on AWS infrastructure with SOC 2 Type II & ISO 27001 certified data centers — with data residency available across EU, Middle East, and other regions
Across WorkLab and LearnLab, our AI assists — it does not replace — human hiring, grading, and academic decisions. The "EU AI Act Aligned" badge reflects alignment with the Act's principles by design — transparency, human oversight, and documentation — not a certification. Learn more.
Upstack.AIFilter the noise. Interview real candidates. One link works anywhere—no ATS migration needed.
Upstack AI FZ-LLC
FOAM2471, Compass Building
Al Shohada Road, AL Hamra Industrial Zone-FZ
Ras Al Khaimah, United Arab Emirates
Microsoft Store
Publisher: UPSTACK AI
Store ID: 9NT2GR4TDZ0G
Powered By
Last updated: 21/1/2026