Ground Truth.
AI, checked against the source.

News · 2026-08-17

Anthropic still will not ship the model that found ten thousand vulnerabilities

Anthropic's most cyber-capable model is still not generally available, four months after the company started handing it to a vetted group of defenders. In a May 22, 2026 update on Project Glasswing, Anthropic said that it and roughly 50 partners had used Claude Mythos Preview to find more than ten thousand high- or critical-severity vulnerabilities in the world's most systemically important software. The company's stated reason for withholding Mythos-class models from the public is not that the model failed. It is that Anthropic does not think its safeguards are good enough yet.

Key facts

The setup is unusual enough to be worth spelling out. Most frontier labs ship one model to everyone and hold back capabilities through refusals. Anthropic built a two-track deployment instead. Project Glasswing started in April as a program to point an unusually strong offensive-security model at critical infrastructure before comparable capability reached attackers, and access went to a small group of partners under agreement. By June the arrangement had a name on both sides: Fable 5 as the general-use model wrapped in conservative safety classifiers, and Mythos 5 as the identical base model with some of those classifiers removed, released only through Glasswing to defenders who pass vetting and accept a 30-day data-retention requirement.

The number that makes the program hard to dismiss is the ten thousand. That is not a benchmark score. Those are real vulnerabilities in real software, rated high or critical, found in a few weeks by a group of roughly fifty organizations pointing one model at codebases the internet runs on. And the interesting part is what Anthropic says happened next. In the company's own words: "Progress on software security used to be limited by how quickly we could find new vulnerabilities. Now it's limited by how quickly we can verify, disclose, and patch the large numbers of vulnerabilities found by AI."

That sentence is the actual news. For thirty years, software security has been a discovery-constrained field. Finding a serious bug in a widely used library was hard, slow, specialist work, and the whole apparatus of coordinated disclosure -- embargo windows, maintainer notification, staged patch releases -- was built around the assumption that bugs arrive at a rate humans can process. If a model can find them faster than volunteer maintainers can fix them, the queue is the vulnerability. Anthropic's own numbers make the contrast concrete: enterprise customers using Claude Security patched 2,100 issues in three weeks, which the company notes is much faster than the open-source side, "in large part because enterprises are fixing their own code, whereas open-source fixes usually require volunteer maintainers who work through coordinated disclosure."

Picture a city that has just invented a machine that can inspect every building for structural defects in an afternoon. The inspection is solved. What is now broken is that the city still has the same forty structural engineers who have to certify each repair.

Anthropic has also started shipping the surrounding apparatus rather than just the model. The Glasswing update describes a Cyber Verification Program that lets security professionals doing legitimate vulnerability research, penetration testing and red-teaming operate without certain misuse safeguards, plus a release of the tooling the partners built: reusable skills, a harness that maps a codebase and spins up scanning subagents to triage findings and write reports, and a threat-model builder that prioritizes which parts of a codebase to attack first. That last piece matters more than it sounds. As anyone building agents has learned, the harness around a model frequently determines how capable it looks.

The honest caveat is that all of the load-bearing evidence here is Anthropic's, published by Anthropic, about a program Anthropic runs. There is no independent audit of the ten-thousand figure, no public breakdown of how many of those findings survived triage, and no external evaluation of whether the withheld model is meaningfully more dangerous than the shipped one. The company's public line -- that Mythos-class models stay restricted because current safeguards remain insufficient to prevent severe misuse -- is a judgment call the public cannot check. It is also worth noting that "restricted" has been getting steadily less restrictive: the partner group grew from about 50 in May to roughly 150 more organizations in June, and OpenAI has been running a comparable program with its own offensive cyber models. The trend line is toward wider access to offensive capability under contract, not toward a permanent hold.


Primary source, verified: read the paper →

Key questions

What is Project Glasswing?

Project Glasswing is Anthropic's program for giving a restricted, unusually cyber-capable model to a vetted group of defenders rather than to the public. It launched on April 7, 2026 around Claude Mythos Preview, and its stated purpose is to secure critical software before comparable capability reaches attackers.

What is the difference between Claude Fable 5 and Claude Mythos 5?

They are the same underlying model with different safeguards. Fable 5 is the generally available version with conservative safety classifiers that Anthropic says fire in under 5 percent of sessions, while Mythos 5 is that same base model with some of those safeguards lifted, available only to vetted cybersecurity professionals and infrastructure providers under a 30-day retention requirement.

Why is finding ten thousand vulnerabilities a problem rather than a win?

Because finding them turned out to be the easy part. Anthropic says the bottleneck has moved from discovery to verification, disclosure and patching, which still depend on human maintainers working at human speed.
Cite this

APA

Ground Truth. (2026, August 17). Anthropic still will not ship the model that found ten thousand vulnerabilities. Ground Truth. https://groundtruth.day/news/anthropic-still-wont-ship-the-model-that-found-ten-thousand-bugs.html

BibTeX

@misc{groundtruth:anthropic-still-wont-ship-the-model-that-found-ten-thousand-bugs,
  title  = {Anthropic still will not ship the model that found ten thousand vulnerabilities},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {aug},
  url    = {https://groundtruth.day/news/anthropic-still-wont-ship-the-model-that-found-ten-thousand-bugs.html}
}

Topics: cybersecurity · ai-security · vulnerabilities · red-teaming · frontier-models · anthropic

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.