News · 2026-09-08
The AI 2027 authors graded their own forecast at about 75% speed
The authors of AI 2027 — the forecasting scenario that became the most-argued-about document in AI safety — have graded their own predictions against reality and found the world moving at roughly three-quarters of the speed they projected. Daniel Kokotajlo and Eli Lifland took every quantitative prediction in the scenario that has resolved, measured how much of the distance to each target reality had actually covered, and published the result. One of the authors described the outcome, on camera, as having "surprised me in a bad way."
Key facts
- Individual resolved predictions score a mean of 75% and a median of 84%; across quantitative metrics the overall pace is closer to 65%.
- The largest miss was coding: the scenario expected about 85% on a standard software-engineering benchmark by mid-2025 from a 72% baseline; reality reached about 74.5%.
- Revenue ran slightly ahead of the scenario — about $20B against a predicted $18B.
- Primary sources: the AI Futures Project's Grading AI 2027's 2025 Predictions, and a 90-minute interview published September 8, 2026 on Machine Learning Street Talk.
AI 2027 was published as a narrative scenario describing a fast takeoff — a world where AI systems begin meaningfully accelerating AI research, and the resulting loop compresses a decade of progress into a couple of years. It was widely mocked as science fiction and widely cited as a warning, often by the same people. What makes this week's material unusual is that its authors built a scoring rubric for themselves and then ran it.
The method is simple enough to explain and harder to game than a yes/no scorecard. For each quantitative prediction, they ask how far reality has travelled toward the predicted value compared with how far the scenario said it would travel by now. Aggregate that across predictions and you get a single "how fast are things going compared to the scenario" number. As one author put it in the interview: "the topline number is something like 75% speed. So basically things are on track but just going like a little bit slower."
The most interesting result is a miss the authors handled honestly rather than explaining away. Their metric for AI-driven uplift to software research looked far below the scenario's prediction. The reason turned out not to be that progress stalled. "At the time that we wrote AI27 we had a bad estimate of what the uplift was at the time that we published. We thought it was higher than it actually was," the author explains. "What actually happened is that there was actually significant increase in uplift due to coding agents and so forth, but it was increasing from a lower level than we thought." In other words, the growth rate was roughly right; the starting line was wrong. That is a correction that makes their forecast look worse on the scorecard and the underlying trend look unchanged — the opposite of the direction motivated reasoning pushes.
The analogy is a road trip where you predicted arrival at 6pm and got there at 7:30. You were wrong about the arrival time. You were not wrong about the direction, the route, or the fact that you were driving.
The interview also contains the sharpest thing either author has said publicly about what they want. Asked whether they actually think development should stop, Kokotajlo answers directly: "I think that just stopping everything now, plan S, would be better than the default. Like I would rather just stop everything now than continue going on our current trajectory." He immediately clarifies that this is not the recommendation — their actual proposal is a temporary inference-only pause, used to build transparency and verification infrastructure before continuing "in this more distributed, transparent, cautious way," advancing only as far as can be reliably controlled. The distinction between "what I would prefer to the status quo" and "what I am asking for" is one that forecasting documents usually blur, and he does not.
There is also a genuinely useful reframing buried in the conversation. Rather than arguing about the word AGI, the authors propose a concrete milestone: "the point at which an AI company would rather fire their humans than fire their AIs." That is observable, dated, and does not require anyone to agree on a definition of intelligence. It is the kind of operational threshold that measuring AI by task length tried to supply and mostly did not.
Why it matters: the credibility of AI forecasting has real policy consequences, and almost nobody grades themselves. A scenario running at 65-75% of predicted pace is neither the vindication its supporters will claim nor the debunking its critics will. It says the shape was approximately right and the clock was fast — which, for a document written to argue that timelines are shorter than people think, is a mixed result its authors have chosen to publish rather than bury. It also sits pointedly beside OpenAI's own statement that recursive self-improvement and alignment outrank math research in its priorities.
The caveats are real. This is self-grading, and the authors chose both the predictions and the scoring rubric. The headline number moves depending on which predictions have resolved and how they are weighted — the same work yields 65%, 75% or 84% depending on the aggregation. And a scenario's most consequential claims, about what happens after systems start improving themselves, have not resolved at all and cannot be scored yet.
Key questions
How accurate has the AI 2027 scenario turned out to be?
What was AI 2027's biggest miss?
What do the authors recommend now?
Cite this
APA
Ground Truth. (2026, September 8). The AI 2027 authors graded their own forecast at about 75% speed. Ground Truth. https://groundtruth.day/news/the-ai-2027-authors-graded-their-own-forecast-at-about-75-percent-speed.html
BibTeX
@misc{groundtruth:the-ai-2027-authors-graded-their-own-forecast-at-about-75-percent-speed,
title = {The AI 2027 authors graded their own forecast at about 75% speed},
author = {{Ground Truth}},
year = {2026},
month = {sep},
url = {https://groundtruth.day/news/the-ai-2027-authors-graded-their-own-forecast-at-about-75-percent-speed.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.