A new study tracking roughly 27,000 students in China has produced some of the most sobering evidence yet on generative AI in education: pupils who used AI tools saw their homework scores jump 18 percent and finished assignments in far less time — then scored 20 percent below their classmates on exams taken without AI assistance. The research, led by David Strömberg of Stockholm University together with Victor Lei and Wu Yanhui of the University of Hong Kong, was highlighted by The Economist on August 18 and is circulating as a working paper on SSRN. For readers following breaking AI news, it is a data point educators and parents will not easily dismiss.
The setup and the shock
The researchers set out to answer a question most schools have been guessing at: does AI help students learn, or does it help them finish? By following tens of thousands of pupils across subjects for six months, the team could compare homework performance against subsequent exam results for AI users and non-users alike.
The headline findings, as summarized by The Economist: after six months, pupils using AI saw their average homework score rise by 18 percent across all subjects, while the time spent on each assignment fell from an average of 64 minutes to 45. But when exam time came, the same students scored 20 percent below classmates who had never used AI for their homework.
The deeper finding is subtler and more troubling. "Homework scores once predicted exam performance; now those who score highest are, perversely, more likely to do worse in exams," the Economist's analysis noted. In other words, the traditional signal teachers rely on — homework grades as a proxy for understanding — has been quietly broken. The students whose assignments look best may be the ones who understand the material least.
Outsourcing, not learning
Why did homework scores soar while exam scores collapsed? The data point to what the researchers describe as outsourcing: students handing the work itself to the model rather than using it as a tutor. Figures from the paper cited in public discussion suggest that around 81 percent of AI users in the study were effectively delegating their homework to the language models — and that the outsourcing rate climbed the longer students had been exposed to the tools.
The pattern emerged gradually. Early in the study, some AI-using students still spent more than 65 minutes per assignment — roughly the same time as non-users — and their homework and exam results looked similar to peers who avoided AI. The researchers inferred these students were barely using the models at all. But six months after adoption, no AI-exposed student still spent more than 65 minutes on homework. What begins as occasional assistance appears to slide, assignment by assignment, into full delegation.
That escalation is what makes the findings so difficult to dismiss. It is not that AI tutoring cannot work — used deliberately, as a Socratic partner that explains and quizzes, the technology demonstrably can help people learn. It is that in naturalistic conditions, with teenagers and overdue assignments, the path of least resistance runs straight through the answer.
Why China makes a clean test case
The study's setting strengthens its conclusions. In China, university admission runs through high-stakes standardized examinations, and homework grades do not inflate course results the way they can elsewhere. Exams are externally benchmarked and brutally consequential, which makes the 20 percent gap harder to explain away with easy grading or teacher leniency. The study effectively isolated the channel researchers care about: AI changed how homework was produced, and exams measured what had actually been learned.
Discussion of the paper on Hacker News, where it gathered hundreds of points, wrestled with the obvious follow-up question — whether AI is an "amplifier" that makes diligent students better and disengaged students worse. The paper's own analysis pushes back on that comforting frame: the time-use data suggest that even students who began using AI responsibly drifted toward delegation over months, implying the risk is not confined to those who never intended to learn in the first place.
The policy dilemma for schools
The findings land at an awkward moment. Governments and edtech firms are racing to put AI tutors in front of students, betting that personalized AI assistance will close achievement gaps. This study suggests the opposite outcome is possible at scale: homework becomes frictionless, grades improve, and the measured learning beneath them quietly erodes — discovered only when it is too late to fix.
For schools, the implications are concrete. If homework can no longer serve as evidence of learning, assessment must move in-class and proctored. If AI use is inevitable, assignments may need to be redesigned around the technology — oral defenses, in-class problem sessions, and process-based grading that rewards the struggle the study found students skipping. And if outsourcing intensifies with exposure, delay is not neutral: every month of unstructured AI access appears to deepen the habit.
The researchers' contribution is to replace the anecdotal panic of the ChatGPT-in-the-classroom era with longitudinal numbers, and the numbers are bleak. An 18 percent homework improvement purchased at the cost of a 20 percent exam decline is not a learning gain — it is a transfer of work from student to machine, with interest due on exam day.
Stay Ahead of AI
For research breakthroughs, policy fallout, and honest analysis of AI in education, bookmark AI Buzz Wire.
Read more AI news →