Anyone who has watched a child freeze over a fractions worksheet knows the fantasy. A patient tutor sits beside them and spots the exact step where things went wrong. Then, instead of handing over the answer, they ask a question. For most of history, only some families could pay for that. Now chatbots offer it to anyone with a phone. That makes AI tutoring effectiveness one of the most consequential questions in education right now.
The honest answer is messier than either the vendors or the doomsayers admit. Some AI tutors have produced striking gains in randomized trials. A generic chatbot, used the lazy way, left high schoolers measurably worse at math once researchers took it away.
Same underlying technology. Opposite outcomes.
What separates them is one design decision: whether the tool makes the learner do the thinking. To see why that matters so much, it helps to go back four decades, to a number education researchers still argue about.
Bloom’s 2 sigma problem, and why the number shrank
Every serious argument about AI tutoring effectiveness eventually runs into 1984.
Where “two sigma” came from
That year, Benjamin Bloom reported a startling result. Students who got one-to-one tutoring plus mastery learning performed two standard deviations above classroom peers. Put plainly, the average tutored student scored higher than 98% of the comparison class.
The “problem” in the name is cost. One tutor per child doesn’t scale, so the challenge became finding something cheaper that works as well. Edtech companies have pitched software as that something for decades. AI tutoring pitches still quote Bloom’s figure today.
What happened when researchers looked closer
In a 2024 reanalysis for Education Next, Paul von Hippel went back to Bloom’s sources. The foundation turned out to be thin. Bloom leaned mainly on two PhD dissertations, by Anania and Burke, and those studies had features that flatter results:
- Short runs. The experiments lasted three weeks.
- Unfamiliar material. Students tackled topics they hadn’t met before: probability in grades 4 and 5, cartography in grade 8.
- A bundle, not a single treatment. The “tutoring” came packaged with extra testing, feedback, retesting and tutor training.
Then there’s the test problem. Von Hippel found that tutoring effects averaged 0.84 standard deviations on narrow tests the study authors built themselves. On broader standardized tests, the average was just 0.27.
He also pointed to a 2020 meta-analysis by Nickow and colleagues covering 96 randomized tutoring studies. Their average effect came to 0.37 standard deviations, roughly 14 percentile points. Not one of the 96 reached two sigma.
Von Hippel puts the original claim in perspective: “Two sigmas is an enormous effect size… If a tutor could raise, say, SAT scores by that amount, they could turn an average student into a potential Rhodes Scholar.”
Years before that reanalysis, VanLehn’s 2011 review had already pulled the human-tutoring figure down to about 0.79. Its more surprising finding concerned software. Step-based intelligent tutoring systems, the pre-chatbot programs that walk students through problems one step at a time, came in at about 0.76.
So a realistic benchmark for strong human tutoring sits around 0.3 to 0.4 standard deviations on broad tests. That’s a meaningful gain by education standards. It just isn’t a miracle. Any AI tool promising “Bloom’s two sigma” is quoting a number human tutors never reliably hit either.
What actually makes human tutoring work
Any judgment of AI tutoring effectiveness needs a human baseline first. Pull Bloom’s bundle apart and the active ingredients look less like a magical person and more like a routine:
- Frequent checks on what the student actually knows, the job extra testing did in Bloom’s studies
- Fast feedback while a mistake is still fresh
- Retesting until mastery before moving on
- Withholding the answer long enough for the student to work at it
Even with all that in place, the payoff is real but modest. The Education Endowment Foundation in England estimates that one-to-one tutoring averages “five additional months’ progress,” according to figures cited by tutoring company Third Space Learning.
The last ingredient on that list matters most for AI. Learning scientists call it desirable difficulty: effort during practice that feels worse in the moment but makes learning stick. General-purpose chatbots exist to be helpful, and helpful usually means answering. If you want the mechanics behind that instinct, our explainer on how large language models work walks through it.
A good tutor often shouldn’t answer. That tension runs through every study below.
What the AI tutoring studies show
Here’s where AI tutoring effectiveness gets interesting. The recent evidence points in a fairly consistent direction, though every study carries caveats.
The Harvard physics trial
The strongest positive result comes from Harvard. Kestin and colleagues published a randomized crossover trial in Scientific Reports in June 2025. Stanford’s SCALE initiative summarizes the headline plainly: AI tutoring outperformed active learning.
The setup was clever. In autumn 2023, 194 students in Physical Sciences 2, an introductory physics course, took part. Each student learned one topic, surface tension or fluid flow, through in-class active learning. They learned the other at home with a custom AI tutor.
The results:
- Students learned more than twice as much from the AI tutor.
- They spent less time doing it: a median of about 49 minutes with the tutor, against 75 in class. That’s roughly 35% less time.
The design is the real story. The researchers told the tutor to reveal one step at a time and to make students attempt each problem before getting help. In effect, they wrote Bloom’s routine into the tutor’s instructions. Raw model capability didn’t drive the result. The pedagogy did.
Keep one detail in mind for later. The comparison group didn’t get one-to-one human tutoring. It got a single active-learning class session.
The Eedi and Google DeepMind trial
A smaller trial from Eedi and Google DeepMind tested a different setup: AI with a human in the loop. In summer 2025, 165 students across five UK secondary classrooms worked with either a human tutor alone or a human tutor supervising Google’s LearnLM.
- On novel problems, the supervised-AI group succeeded 66.2% of the time, against 60.7% for human tutors alone.
- Supervising tutors approved 76.4% of the AI’s drafted messages with zero or minimal edits.
That second number is the quietly important one. If one tutor can sign off on most of what an AI drafts, that tutor might reach far more students. The authors call the trial exploratory, though, and the AI never worked unsupervised. Treat it as a signal worth watching.
Before chatbots: intelligent tutoring systems
Chatbot hype tends to skip this history. VanLehn’s 0.76 for step-based tutoring systems means carefully structured software came close to human tutors years before large language models went mainstream.
The lesson carries forward. Structure did the heavy lifting then, and it still does.
What vendor data can and can’t tell you
Vendor pages are where AI tutoring effectiveness claims get loudest. Third Space Learning reports that students using its Skye tutor rose from 34% to 92% accuracy between check-in and check-out questions within a session. It also says 63.8% of students reported more confidence.
Those are within-session gains. They tell you a student can do the thing right now. They say nothing about whether that student can still do it next month.
What the evidence doesn’t show yet
This is the part most coverage of AI tutoring effectiveness skips. Here’s what the current studies leave open:
- Short time horizons. Every strong positive result so far is short-term. Kestin used short-term tests, and the Skye figures come from single sessions.
- Narrow samples. The Harvard trial covered fewer than 200 introductory physics undergraduates at one institution, studying two topics.
- Soft comparisons. Kestin’s AI beat one active-learning session, not a skilled personal tutor. Eedi’s AI never ran without a human watching.
- Narrow outcome measures. Remember the 0.84 versus 0.27 gap. The SCALE summary of Kestin’s “more than twice as much” doesn’t report a formal effect size.
- Who built it. Google DeepMind co-ran a trial of its own model, and the Skye numbers come from the company that sells Skye. That doesn’t make the findings wrong. It does mean independent replication matters.
So here’s the honest state of AI tutoring effectiveness. The evidence supports well-designed, closely supervised AI tutors on specific, structured material over short periods. It doesn’t yet support handing a child a general chatbot and walking away.
When AI hurts learning: the cognitive offloading problem
The sharpest warning comes from a field experiment by Hamsa Bastani and colleagues, which appeared in PNAS in July 2025 (study record).
Nearly 1,000 high school math students took part. Some practiced with GPT Base, a plain GPT-4 assistant. Others used GPT Tutor, a version that offered hints and teacher guidance but no direct answers. Then the researchers removed access and compared everyone on an exam.
- GPT Base: practice grades rose 48% with access. On the exam, these students scored 17% lower than the control group.
- GPT Tutor: practice grades rose 127% with access. On the exam, these students did about as well as the control group. The design largely cancelled the harm.
Bastani’s own read is blunt. “If we use it sort of lazily and kind of outsource the work… then that’s when we could be in trouble,” she told Knowledge at Wharton. Her deeper worry is about skills: “We’re really worried that if humans don’t learn, if they start using these tools as a crutch and rely on it, then they won’t actually build those fundamental skills…”
That’s cognitive offloading in its plainest form. Performance with the tool went up while ability without it went down. The students handed the thinking to the machine, and the practice stopped building anything. We’ve looked at this trade-off more broadly in Cognitive Offloading to AI: Skill or Crutch?
Same model, opposite results
Now put the two versions side by side. Same underlying model. One produced a 48% practice bump and a 17% exam penalty. The other produced a 127% bump and roughly no penalty.
The Harvard tutor and GPT Tutor share one trait. Both refused to do the student’s work for them.
That’s the hinge of the whole AI tutoring effectiveness debate. “Does AI tutoring work?” is the wrong question. The better one is whether the design forces the learner to do the thinking.
How to judge an AI tutoring effectiveness claim
New studies and product claims arrive constantly. Run each one through five questions before you believe the headline:
- What was the comparison group? No instruction, a lecture, active learning or a human tutor? Effect sizes tend to shrink as the comparison gets stronger.
- What did they measure? A narrow test the authors wrote, or a broad standardized one? The 0.84 versus 0.27 gap shows how much this matters.
- When did they measure it? Within a session, at the end of a unit, or months later? Bastani’s trial shows that gains with the tool can hide losses without it.
- Who funded or built it? Vendor and developer studies aren’t worthless. They do need independent replication before you lean on them.
- Could the tool give away answers? If it could, check whether anyone tested students after taking it away.
What this means if you’re learning with AI
The research turns into a handful of practical habits for anyone studying with a chatbot:
- Ask for hints, not solutions. Tell the AI to reveal one step at a time and to wait for your attempt first. That mirrors both the Harvard tutor and GPT Tutor.
- Test yourself cold. Close the chat and try a fresh problem. If you can’t do it unaided, the practice didn’t stick.
- Distrust easy fluency. Practice that feels smooth with AI open might be the 48% gain that becomes a 17% loss.
- Parents and teachers: ask one question. Does the product withhold answers by default, or only when the student asks it to?
Schools weighing classroom technology face a version of this question with robots, too. Our look at humanoid robots in education asks the same thing: helper or substitute?
So, can AI tutors really teach?
Yes, sometimes, and under conditions the evidence keeps spelling out. The tutor has to withhold answers and pace the student one step at a time. A human supervisor helps. Strip those conditions away and you get a homework machine that makes practice look better while the learning gets worse.
My bet is that the technology isn’t the bottleneck. Defaults are. Most students meet AI through general-purpose chatbots designed to answer, not tutors designed to make them struggle.
If you study with AI, try one experiment this week. Ask it to tutor you without ever giving the final answer. Then close it and solve a new problem alone. You’ll learn more about AI tutoring effectiveness from that hour than from any vendor page.
There’s also an equity question nobody has settled. Well-resourced schools can vet and supervise carefully designed tutors. Everyone else may get the answer machine. We’ve raised a similar worry about brain chips and the class divide, and it applies here too.
If AI finally puts a tutor beside every child, will it be the kind that asks the question, or the kind that hands over the answer?