One of the most ambitious real-world tests of AI tutoring has delivered a sobering verdict. A two-year cluster randomized trial of Khan Academy's Khanmigo tutor across 18 Tennessee middle schools found that access to the AI tutor produced measurable but modest math gains, and that the results were held back by a problem technologists rarely advertise: most students barely used it. The study, titled "One Click Away: AI Tutoring with Khanmigo in a Two-Year School Experiment," was conducted by Philip Oreopoulos and Nina Low and published through the Annenberg Institute at Brown University's EdWorkingPapers series.
The findings matter well beyond a single product, because Khanmigo is one of the most widely deployed AI tutors in American schools and the study is among the first large-scale experimental evaluations of generative AI tutoring in normal classroom conditions. For anyone tracking AI research news and its collision with real-world institutions, the paper is a rare piece of hard evidence in a debate dominated by marketing claims, and it landed on the front page of Hacker News this week as educators and technologists parsed its implications.
What the Study Found
The trial randomly assigned students in 18 Tennessee middle schools to use Khan Academy with its AI tutor Khanmigo during existing daily remedial mathematics sessions. Crucially, the tutor was configured to coach rather than give answers, reflecting Khan Academy's stated pedagogy. The experiment ran for two school years, making it one of the longest field experiments on AI tutoring to date.
The headline results: assignment to the AI tutoring condition raised math achievement by 1.3 national percentile ranks per term, which the authors calculate as roughly 0.06 to 0.08 standard deviations over a school year. The implied effect of a full year of active participation reaches 0.14 standard deviations. Those are real, statistically detectable gains, and for a low-cost software intervention they are respectable.
But the comparison that deflates the hype is this: the gains resemble those from Khan Academy practice without AI assistance. Students who used the platform's existing practice tools, with no AI tutor attached, appear to have done about as well as students who had access to Khanmigo.
The Engagement Problem in the Data
The paper's most revealing section is its forensic account of usage. On the surface, adoption looks strong: 96 percent of students tried Khanmigo at least once. Underneath, engagement collapses. The median student messaged the tutor on only about a third of the days they practiced math, and in only 17 percent of the exercise sessions in which they made a mistake. When students did message the tutor, the authors found the messages were mostly bare answers or clicks on suggested prompts, rather than substantive mathematical dialogue.
In other words, the tutor was configured to Socratically coach students through mistakes, but students mostly avoided the coaching conversation entirely. The authors conclude that the binding constraint appears to be engagement: realizing the promise of AI tutoring will require getting students to use it, not just giving them access.
Why It Matters for Schools Buying AI
Districts across the United States are under pressure to adopt AI tools, and tutoring is the flagship use case in almost every pitch. Generative AI has been promoted as the technology that could finally deliver what education researchers call the two-sigma dream: a personal tutor for every student. The Bloom's two-sigma framing comes from older research on human tutoring, and it is exactly the promise AI tutors are marketed against.
This study suggests the bottleneck is not model quality or access. Every student in the treatment group had a state-of-the-art LLM tutor one click away, at no cost, inside their daily math block. The constraint was whether adolescents would voluntarily sustain a tutoring dialogue, day after day, in a remedial class. Human tutoring works partly because a person notices when you disengage. Software, so far, does not reliably replicate that pull.
Caveats Worth Keeping in Mind
The paper is an EdWorkingPaper, distributed through Brown University's Annenberg Institute, and readers should treat it as working research rather than peer-reviewed findings. The study covers middle school remedial math in one state, so generalizing to other subjects, grade levels, or implementation models is speculative. The authors' own framing is measured: they report what happened when a widely available AI tutor was layered onto existing practice sessions, not what might happen under different curricula or stronger incentives to engage.
It is also worth noting what the study does not say. It does not find that AI tutoring is useless; a 0.14 standard deviation effect from a full year of active participation would be meaningful if engagement could be sustained. The finding is closer to this: the technology worked about as well as the non-AI practice tools it was bolted onto, because students did not use the AI differently from the buttons they already ignored.
The Takeaway
The education technology sector has spent two years promising that generative AI tutors will transform classrooms. This trial, one of the first to follow real students over two full years, suggests the transformation is gated by an ancient problem: getting kids to want to do the work. For school leaders, the lesson is to demand engagement data, not just license counts, from AI vendors. For the AI industry, it is a reminder that the hardest problems in AI deployment are rarely about the model.
Keep Up With AI Research
Field experiments like this one cut through the hype cycle faster than any benchmark. For continuous coverage of AI research with real-world consequences, follow AI news on aibuzzwire.news.
Read more AI news →