6 min read

The 2 Sigma Problem

reading listlearning sciencetutoring

In 1984, Benjamin Bloom put a number on something teachers had always suspected: a good tutor working with one student is worth far more than the same teacher working with thirty. The number was large enough to be uncomfortable. And he turned it into a challenge rather than a boast.

What the paper actually claims

Bloom (1984) reported that students taught one-to-one, using mastery learning, performed about two standard deviations better than students taught the same material in a conventional class. Two standard deviations ("2 sigma") is a large distance. It means the average tutored student scored above roughly 98% of the students in the ordinary classroom.

He did not stop at the headline. The finding sat on a ladder. A conventional class was the baseline. The same class run with mastery learning (teaching to a standard, testing often, and giving corrective feedback until each student reached it) lifted the average student by about one standard deviation, to around the 84th percentile of the original class. Only when mastery learning was delivered one-to-one did the effect reach two sigma.

That gap is the paper's real subject. One-to-one tutoring is the best instruction we know how to give and the least affordable to give at scale. Bloom's question was not "is tutoring good" (it obviously is) but whether group methods could be found that come close to it. He called this the 2 sigma problem, and spent the rest of the paper cataloguing "alterable variables": the levers a teacher or a system can actually change (feedback and correction, reinforcement, the quality of cues and explanations, time on task, student participation), ranked by how much each moves the outcome. The bet was that a well-chosen combination might approach tutoring without needing one tutor per child.

Why it endures

Bloom was not a bystander to this field; he largely built it. This is the same Benjamin Bloom of Bloom's taxonomy and of mastery learning, writing near the end of a long career about the thing he had spent it studying. The paper carries that authority lightly.

It is also short: a dozen pages in Educational Researcher, readable in a sitting. Landmark papers are usually dense; this one is almost conversational, and its influence is out of all proportion to its length. It named a problem crisply enough that four decades of researchers have organised their work around it, which is most of what a great paper does.

What it changes in how you think

Before Bloom, the gap between a tutored student and a classroom student was easy to file under talent or effort. After Bloom, it becomes a design problem. The tutored student and the classroom student are, on average, the same student. What changed was the instruction: someone was watching this particular learner, noticing what they had not yet grasped, and not moving on until they had.

That reframes the classroom's weakness precisely. A class is not worse at teaching because teachers are worse. It is worse because one teacher cannot hold thirty separate models of thirty separate students in mind, correct each error as it appears, and pace each learner independently. The things a tutor does (continuous feedback, correction before the gap compounds, attention to the individual) are exactly the things that do not survive being divided by thirty.

The average tutored student and the average classroom student are the same student. Only the instruction is different.

After Bloom, 1984

The honest part

The 2 sigma figure should be held with care, and the paper is better for saying where it came from. Bloom drew it from small, short studies: a few weeks each, run by his doctoral students with schoolchildren on tightly defined topics. It was never a general law, and it has not been reproduced at full strength. When VanLehn (2011) reviewed decades of tutoring research, human tutors came out around 0.8 standard deviations ahead, and intelligent tutoring systems close behind: real, valuable, and roughly a third of Bloom's number.

So the honest reading is this: two sigma is the ceiling Bloom observed under ideal conditions, not the effect you should expect in the wild. The direction is solid. Individual attention with rapid correction beats one-size teaching. The magnitude is contested. Both things are true at once, and a careful reader keeps both.

Some of the individual levers Bloom pointed to have since been pinned down on their own. Frequent low-stakes testing, one of his feedback-corrective mechanisms, is now among the best-evidenced techniques in the field (Roediger and Karpicke, 2006; Dunlosky et al., 2013). The 2 sigma problem was never going to be solved by one idea. It was always a search for the right combination.

Where we stand

Two sigma is our target, not our claim. Building tutoring-style attention that reaches many people is the whole ambition; asserting we have matched Bloom's ceiling would be exactly the kind of thing this paper teaches you to distrust.

Who should read it, and who can skip

Read it if you build or choose learning tools. Every "personalised tutoring at scale" pitch (ours included) is a promise to make progress on Bloom's problem, and you cannot judge the promise without the paper that framed it. Read it if you teach, especially if you are drawn to mastery learning. And read it if you are an ambitious self-learner: it explains, better than any list of study hacks, why finding someone, or something, that gives you real feedback beats grinding alone.

Almost no one should skip it. It is short, it is foundational, and it is a pleasure to read. The one exception is the casual reader who only wants the headline, and the headline is a single sentence: one-to-one tutoring beat the classroom by about two standard deviations, and the enduring challenge is getting group instruction anywhere near that. If that is all you need, you now have it. Everyone else should spend the hour.

For where this sits in the wider evidence, see the bookshelf behind the method.

Sources

Read more