Ground Truth.
AI, checked against the source.

News · 2026-08-21

Students using generative AI got better homework grades and worse exam scores

A study following 26,811 students in grades seven through twelve across one Chinese county for 30 months found that adopting generative AI raised homework scores by roughly 18% and cut homework completion time by roughly 30%, while closed-book monthly exam scores fell about 20% within six months. The paper, "The Generative AI Learning Penalty: Evidence from Chinese Secondary Education," is published as CEPR discussion paper DP21577 and was selected for the National Bureau of Economic Research's 2026 Summer Institute session on the economics of education.

Key facts

The reason this study is worth taking seriously when most AI-and-education research is not comes down to what it measured and how long it watched.

The design matters. Most studies in this area compare one assignment, one class, or one semester, and typically compare different groups of students. This one follows the same students over two and a half years, combining homework scores and completion times, monthly closed-book exams, and eventually the high school and college entrance exams that determine where those students end up in life. Nine subjects. The method is staggered difference-in-differences, which uses the fact that students adopted AI at different times to separate the effect of adoption from the general trend of getting older and covering more material.

The core result is a gap, and it is the gap rather than either number that carries the meaning. Work done with the tool available got faster and looked better. Work done with the tool taken away got worse. The World Bank's commentary on the paper called it a warning shot for human capital, and the framing is apt: the visible output improved while the thing the output was supposed to be building did not.

The mechanism the authors point to is the study's most useful contribution, because it identifies who is affected and who is not. The learning losses concentrate in roughly 80% of AI users, the ones whose homework completion times became unusually short while their homework scores stayed high. That combination, faster and better at the same time, is the behavioral signature of outsourcing rather than assistance. Students who kept spending comparable time on homework showed little or no meaningful loss. Same tool, different use, opposite outcome.

The analogy that fits is a navigation app. Following turn-by-turn directions gets you to the address reliably every time, and it is genuinely better than a paper map. It also means that after a year in a new city you still cannot draw it from memory, because building the map in your head requires the effort the app removed. The route is the homework. The mental map is the exam.

There is an important limit on what the mechanism claim can bear. The paper does not experimentally isolate cognitive offloading. It infers it from the behavioral pattern, which is reasonable but is inference. And the World Bank's own writeup acknowledges a residual confounding story: some other shock could plausibly both increase AI use and decrease performance, though the authors argue such a shock is unlikely to account for the whole effect.

Why it matters: this is the first study at this scale with outcomes that people actually care about, entrance exams rather than quiz scores, and it lands in the middle of a policy conversation that has largely been conducted on vibes. The finding is not "AI is bad for students." It is closer to "the way most students are currently using it is bad for students," which is a considerably more actionable statement. The distinction between the fast group and the normal-pace group suggests the intervention is about how the tool is used rather than whether it is present.

The honest caveat: one county, one country, one 30-month window, one education system with an unusually high-stakes exam culture. A closed-book exam regime is precisely the setting where offloading would show up most sharply, and results may not transfer to systems that assess differently. The study is a serious warning signal with a specific mechanism attached. It is not a universal law, and the authors do not claim it is.


Primary source, verified: read the paper →

Key questions

Is this a randomized experiment?

No. It is a staggered difference-in-differences design on a longitudinal panel, which compares the same students before and after they adopted AI, using variation in when different students adopted it. That is much stronger than a simple correlation but still not a randomized trial.

Did every student who used AI get worse at exams?

No, and this is the study's most useful finding. The learning losses concentrate among roughly 80% of AI users whose homework completion times became unusually short while homework scores stayed high. Students who kept spending similar time on homework showed little or no meaningful loss.

Which subjects were hit hardest?

Social science showed the largest losses on entrance exams, followed by science and mathematics subjects, with languages least affected. The reported drops on the two entrance exams were about 18% and 24%.
Cite this

APA

Ground Truth. (2026, August 21). Students using generative AI got better homework grades and worse exam scores. Ground Truth. https://groundtruth.day/news/students-who-used-generative-ai-scored-lower-on-closed-book-exams.html

BibTeX

@misc{groundtruth:students-who-used-generative-ai-scored-lower-on-closed-book-exams,
  title  = {Students using generative AI got better homework grades and worse exam scores},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {aug},
  url    = {https://groundtruth.day/news/students-who-used-generative-ai-scored-lower-on-closed-book-exams.html}
}

Topics: education · society · research · cognitive-offloading · policy