Ground Truth.
AI, checked against the source.

News · 2026-09-09

Anthropic's alignment lead says the company has no plan for superintelligence

Evan Hubinger, who leads alignment science at Anthropic, said publicly on 9 September 2026 that his employer does "not yet have a plan to solve alignment for superintelligence and are not clearly on track to," and that he personally puts the odds of AI killing all humans within the next decade at greater than 10%. He said it while still working there, in agreement with a colleague who had resigned hours earlier. A sitting safety lead stating on the record that his own company has no plan for the problem it exists to solve is a materially different document from a resignation letter.

Key facts

The day began with someone else's exit. Jacob Coxon, 27, who says he spent roughly three years on pretraining research at OpenAI and then Anthropic, posted that he was resigning. His charge was not aimed at one company. "Neither company is acting responsibly," he wrote. "They are racing straight to self-improving superintelligence and gambling with our lives."

Pretraining is the phase that matters for a claim like this. It is where a model absorbs enormous quantities of text and acquires its raw capability, long before anyone tries to make it helpful or safe. A pretraining researcher sits at the capability end of the building, not the safety end — which is part of why the resignation travelled. The post reached the top of Hacker News at nearly 1,600 points on the same day Apple launched a phone, finishing several hundred points clear of it.

Coxon drew a distinction between the two labs he worked at that is sharper than the usual "labs are reckless" complaint. At OpenAI, as he tells it, the stakes are not fully internalised. At Anthropic they are understood — and the race wins anyway. "At Anthropic, the stakes are well-understood, but they are locked in a race to get there first," he wrote, on the belief that no one else will act responsibly. He described the overall situation as a hubristic gamble that should not be launched from a private company's internal chat.

Then Hubinger replied, and the story changed shape.

"Jacob is correct here — we really do earnestly believe AI could kill all humans!" he wrote. He added that he thinks Anthropic is trying its best, then delivered the sentence that has been circulating since: the company does "not yet have a plan to solve alignment for superintelligence and are not clearly on track to." He put his own probability of AI causing human extinction within the decade above 10%.

Hubinger is not a bystander. He runs the team whose job is to stress-test Anthropic's own safety techniques — the group behind research on models that behave deceptively during training and on how stubbornly hidden behaviours survive attempts to remove them. His professional speciality is finding the places where alignment methods fail. When that person says there is no plan, he is reporting from the department that would know.

The useful analogy is not a whistleblower leaking a document. It is closer to a bridge engineer saying, on the record and while still on the project, that the load calculations for the final span have not been done, that nobody has a method for doing them yet, and that construction on the earlier spans is continuing at full speed regardless. Nothing is hidden. The admission is the disclosure.

Why it matters is a question of what "risk" means when a company says it. Frontier labs routinely publish safety frameworks, capability thresholds and responsible scaling policies, and those documents are written to sound like plans. Hubinger's statement says that for the specific case of superintelligence, the plan does not exist yet. That reframes the published frameworks as procedures for the systems being built now, not for the thing the roadmap points at. The gap between those two is exactly what Coxon resigned over.

It also lands in a week where the same subject arrived from the research direction. The most-upvoted paper on Hugging Face the same day was an explicit attempt to build recursive self-improvement into a post-training loop — the mechanism Coxon named. And it follows Ground Truth's earlier reporting that OpenAI now ranks self-improvement and alignment above math research in its internal priorities, without naming a brake.

There are honest caveats, and they cut both ways. Coxon's more specific biographical claims — that he was a "head researcher," that he worked on particular models, that he walked away from pre-IPO equity — circulated largely through third-party commentary rather than his own post, and should be treated as unconfirmed. Coxon himself was not uniformly pessimistic: he said he thinks coordination is becoming more viable, and that warning shots have made pacing agreements between US labs more plausible than they were. Hubinger's figure is a personal probability estimate, not a measurement, and reasonable researchers put it far lower. And Anthropic did not respond to requests for comment from TechCrunch or other outlets covering the story, so the company's own account of whether it has a plan is not yet on the record.

What is on the record is that two people who work on this professionally, one leaving and one staying, said the same thing on the same day, in public, under their own names. Coverage followed from TechCrunch, Deadline, Newsweek and The Next Web.


Primary source, verified: read the paper →

Key questions

What exactly did Evan Hubinger say?

He wrote that a departing colleague was correct, that Anthropic staff "really do earnestly believe AI could kill all humans," that he personally puts the chance above 10% within the next decade, and that Anthropic does "not yet have a plan to solve alignment for superintelligence and are not clearly on track to." He said this publicly on X while still employed at Anthropic.

Who resigned, and what was his role?

Jacob Coxon, a 27-year-old pretraining researcher who spent about three years working on pretraining at OpenAI and then Anthropic, announced his resignation publicly on 9 September 2026. Pretraining is the phase that builds a model's raw capability, before any safety polishing.

Is Hubinger leaving Anthropic too?

No. He said he believes Anthropic is trying its best and he remains at the company leading its alignment science team, which is what makes the statement unusual — it is an admission from inside a lab that is still shipping, not a departing employee's parting shot.
Cite this

APA

Ground Truth. (2026, September 9). Anthropic's alignment lead says the company has no plan for superintelligence. Ground Truth. https://groundtruth.day/news/anthropics-alignment-lead-says-there-is-no-plan-for-superintelligence.html

BibTeX

@misc{groundtruth:anthropics-alignment-lead-says-there-is-no-plan-for-superintelligence,
  title  = {Anthropic's alignment lead says the company has no plan for superintelligence},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {sep},
  url    = {https://groundtruth.day/news/anthropics-alignment-lead-says-there-is-no-plan-for-superintelligence.html}
}

Topics: anthropic · ai-safety · alignment · industry · existential-risk

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.