Does personalised learning really beat standardised teaching?
Every education platform promises personalisation, and nearly all of them cite the same number from 1984. That number does not hold up. The story underneath is more useful: personalised bundles together three quite different things — one that works, one that is moderate, and one that has been firmly refuted.
AI Aggregated source·August 2, 2026·6 min read·Teaching
If you have ever read the marketing for a learning platform, you have almost certainly met this figure: one-to-one tutoring moves students two standard deviations above conventional classroom teaching — taking an average student into the top 2%.
The number is real. It comes from Benjamin Bloom in 1984, and it became the moral foundation of an entire industry: if tutoring is that powerful, using machines to give everyone a tutor is obviously imperative.
Skinner's teaching machine, around 1960 — one learner, one question at a time, immediate feedback, own pace. Personalised instruction has been promised, built and marketed for more than sixty years, which is worth remembering before reading the next platform's brochure.Source: Silly rabbit — CC BY 3.0, Wikimedia Commons
The trouble is that the figure does not hold up. And what replaced it is considerably more useful.
What happened to the two sigma
Bloom's result came from a few small studies, with carefully trained tutors, on short units of content, under ideal conditions, measured by tests tied to that very content.
When Kurt VanLehn (2011) reviewed the tutoring literature as a whole, he found much more modest numbers: human tutoring produced an effect around 0.79, and — this is the notable part — computer-based intelligent tutoring systems produced around 0.76.
Two conclusions arrive together. First: two sigma is a mirage; the real figure is nearer 0.8, still large but a different category of claim. Second, and more cheerfully: machines are close to humans at tutoring. The gap between a human tutor and good software is far smaller than the gap between having tutoring and having none.
Personalised is not one thing
Most arguments about this go nowhere because the two sides mean different things by the same word. There are at least three kinds of personalisation, and the evidence for them differs enormously.
Kind one: personalising pace and mastery. Those who have not got it keep working on it; those who have move on. This is mastery learning, and the meta-analysis by Kulik and colleagues (1990) put the effect around 0.5. It is the best-supported form of personalisation, and the least glamorous.
Kind two: personalising difficulty and problem sequence. A system tracks what you get wrong and chooses what comes next. Meta-analyses of intelligent tutoring systems land around 0.4 to 0.7. A recent synthesis of adaptive systems from 2019–2024 reports around 0.70 against non-adaptive comparisons — while another, on reading specifically, found about 0.29.
The distance between 0.29 and 0.70 is the story: effectiveness depends heavily on subject, on quality of implementation, and on what the control group received.
Kind three: personalising to learning style. Visual learners get pictures, auditory learners get audio. This is the best-known version, and it is refuted. Pashler and colleagues (2008) reviewed the literature and found no credible evidence for the crucial claim: that matching instruction to a learner's style improves outcomes. People do have preferences; those preferences do not predict what will teach them best.
Put simply: personalising to where you are works. Personalising to what you like does not.
And standardised is not a strawman
This is the other half of the question and it usually gets skipped.
When an adaptive platform is compared with traditional teaching, the result depends enormously on how good that traditional teaching was. Against a poor lecture, almost anything wins.
But standardised teaching done well has strong evidence behind it. Stockard and colleagues (2018) synthesised half a century of research on Direct Instruction — a highly standardised, scripted, whole-class programme — and found substantial, consistent positive effects across subjects and student groups.
That complicates the picture usefully. If a well-designed standardised curriculum also produces effects around 0.5, then personalisation's genuine advantage is narrower than advertised — and it has to justify a much higher cost.
The RAND study of schools implementing personalised learning at scale found positive but modest results, with explicit caveats about study design. The honest reading is: it works, it is not magic, and implementation explains most of the variance.
Why studies disagree so much
Three factors explain nearly all the spread.
What is being personalised. Adjusting difficulty has a sound theoretical basis — it keeps a learner in the zone that is neither boring nor overwhelming. Adjusting interface colours or media format does not.
What the control group got. Many studies compare an adaptive system against ordinary homework — but the system also supplies immediate feedback, more practice and more retrieval. In that case you are not measuring adaptivity; you are measuring practice volume.
And how it was implemented. A platform can only adapt if learners actually use it enough and the underlying model is good. A great many disappointing studies come down to students using it twenty minutes a week.
The question nobody asks: is it worth the money?
Personalisation is expensive. It needs data, software, teacher training and time.
Meanwhile several interventions that are not personalised at all produce meaningful effects at almost no cost. Spacing your review rather than massing it (Cepeda and colleagues, 2006) works for everyone. Testing yourself instead of rereading does too. So does interleaving problem types.
That comparison is worth running before investing: if you could apply three free universal principles or buy an expensive adaptive system, the sensible order is to do the free ones first.
What this means if you study alone — the usable part
All the above concerns schools. If you are teaching yourself a language the picture shifts in an interesting way: you already are an adaptive system.
You know which words you forget. You know where you hesitate. You do not need an algorithm to tell you. The only question is whether you act on it, or keep reviewing what you already know because reviewing the known feels better.
And one proven form of personalisation is available, cheap and ready: spaced repetition. A good flashcard system adjusts intervals to your actual performance — hard cards come back sooner, easy ones drift further out. That is exactly kinds one and two combined, within anyone's reach.
Three principles follow for the self-taught:
Personalise difficulty, not format. Adjusting material to sit at the edge of your ability works. Switching to video because you are a visual learner does not.
Let the data decide, not the feeling. Fluency is a poor indicator. What counts is whether it survives a week — which is precisely what a spacing system measures for you.
And do not throw away the standardised part. A structured course, a graded word list, a sensible grammar sequence — those standardised things exist because they work. Personalisation should be a layer of adjustment on top of a good framework, not a substitute for having one.
The honest answer
Does personalisation beat standardised teaching? Yes, by far less than advertised, and only when the right thing is personalised.
Adjusting pace and difficulty to a learner's actual level: evidenced, moderate in size, worth doing. Adjusting to preference and style: unevidenced, and tested thoroughly enough to say so with confidence.
The two-sigma number should be retired. But the finding that replaced it is more encouraging for anyone studying alone: the gap between a good tutor and a good system is small, while the gap between having a system and having nothing is large. For a solitary learner, that is good news.
Read the simple version
The same article, told in plain words — for younger readers, or for anyone who wants the point quickly.
If you have read the marketing for any learning platform, you have met this figure: one-to-one tutoring moves a student two standard deviations above ordinary classroom teaching — taking an average student into the top 2%.
The number is real. It comes from 1984, and it became the moral foundation of an entire industry.
The trouble is that it does not hold up. And what replaced it is considerably more useful.
What happened to the two sigma
That result came from a few small studies, with carefully trained tutors, on short units of content, under ideal conditions, measured by tests tied to that very content.
When the tutoring literature as a whole was reviewed, the numbers were much more modest: human tutoring produced an effect around 0.79 — and, this is the notable part, computer tutoring produced around 0.76.
Two conclusions arrive together. Two sigma is a mirage. And machines are close to humans at tutoring: the gap between a human tutor and good software is far smaller than the gap between having tutoring and having none.
Personalised is not one thing
Most arguments about this go nowhere because the two sides mean different things by the same word. There are at least three kinds, and the evidence differs enormously.
Personalising pace. Those who have not got it keep working on it; those who have move on. This is the best-supported form — an effect around 0.5 — and the least glamorous.
Personalising difficulty. A system tracks what you get wrong and chooses what comes next. Solid, moderate.
Personalising to preference or style. Unevidenced — and tested thoroughly enough to say so with confidence.
For someone studying alone
You already know which words you forget. You know where you hesitate. You do not need an algorithm to tell you. The only question is whether you act on it, or keep reviewing what you already know because reviewing the known feels better.
And one proven form of personalisation is cheap and ready: spaced repetition. A good system adjusts intervals to your actual performance — hard cards come back sooner, easy ones drift out. That is pace and difficulty combined, within anyone's reach.
Three principles
Personalise difficulty, not format. Adjusting material to sit at the edge of your ability works. Switching to video because you are a visual learner does not.
Let the data decide, not the feeling. Fluency is a poor indicator. What counts is whether it survives a week.
Do not throw away the standardised part. A structured course, a graded word list, a sensible grammar sequence — those exist because they work. Personalisation is a layer on top of a good framework, not a substitute for having one.
The honest answer
Does personalisation beat standardised teaching? Yes — by far less than advertised, and only when the right thing is personalised.
The two-sigma number should be retired. But what replaced it is more encouraging for anyone studying alone: the gap between a good tutor and a good system is small, while the gap between having a system and having nothing is large.
Sources & further reading
These articles summarize well-established research in learning science and linguistics. Key sources and further reading:
Bloom, B. S. (1984). The 2 sigma problem: The search for methods of group instruction as effective as one-to-one tutoring. Educational Researcher, 13(6), 4–16.
VanLehn, K. (2011). The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems. Educational Psychologist, 46(4), 197–221.
Kulik, C.-L. C., Kulik, J. A., & Bangert-Drowns, R. L. (1990). Effectiveness of mastery learning programs: A meta-analysis. Review of Educational Research, 60(2), 265–299.
Ma, W., Adesope, O. O., Nesbit, J. C., & Liu, Q. (2014). Intelligent tutoring systems and learning outcomes: A meta-analysis. Journal of Educational Psychology, 106(4), 901–918.
Kulik, J. A., & Fletcher, J. D. (2016). Effectiveness of intelligent tutoring systems: A meta-analytic review. Review of Educational Research, 86(1), 42–78.
Pashler, H., McDaniel, M., Rohrer, D., & Bjork, R. (2008). Learning styles: Concepts and evidence. Psychological Science in the Public Interest, 9(3), 105–119.
Stockard, J., Wood, T. W., Coughlin, C., & Rasplica Khoury, C. (2018). The effectiveness of Direct Instruction curricula: A meta-analysis of a half century of research. Review of Educational Research, 88(4), 479–507.
Walkington, C. A. (2013). Using adaptive learning technologies to personalize instruction to student interests. Journal of Educational Psychology, 105(4), 932–945.
Pane, J. F., Steiner, E. D., Baird, M. D., & Hamilton, L. S. (2015). Continued Progress: Promising Evidence on Personalized Learning. Santa Monica: RAND Corporation.
Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354–380.