News · 2026-08-17
Anthropic still will not ship the model that found ten thousand vulnerabilities
Anthropic's most cyber-capable model is still not generally available, four months after the company started handing it to a vetted group of defenders. In a May 22, 2026 update on Project Glasswing, Anthropic said that it and roughly 50 partners had used Claude Mythos Preview to find more than ten thousand high- or critical-severity vulnerabilities in the world's most systemically important software. The company's stated reason for withholding Mythos-class models from the public is not that the model failed. It is that Anthropic does not think its safeguards are good enough yet.
Key facts
- Approximately 50 partners found more than ten thousand high- or critical-severity vulnerabilities using Claude Mythos Preview, per Anthropic's Project Glasswing update of May 22, 2026.
- Claude Security, the public-beta scanning tool for Enterprise customers, was used to patch over 2,100 vulnerabilities in its first three weeks.
- On June 9, 2026 Anthropic split the release: Claude Fable 5 went out generally, while Mythos 5 -- the same base model with safeguards lifted -- went only to vetted security professionals.
- Project Glasswing launched April 7, 2026; primary source is Anthropic's own research blog.
The setup is unusual enough to be worth spelling out. Most frontier labs ship one model to everyone and hold back capabilities through refusals. Anthropic built a two-track deployment instead. Project Glasswing started in April as a program to point an unusually strong offensive-security model at critical infrastructure before comparable capability reached attackers, and access went to a small group of partners under agreement. By June the arrangement had a name on both sides: Fable 5 as the general-use model wrapped in conservative safety classifiers, and Mythos 5 as the identical base model with some of those classifiers removed, released only through Glasswing to defenders who pass vetting and accept a 30-day data-retention requirement.
The number that makes the program hard to dismiss is the ten thousand. That is not a benchmark score. Those are real vulnerabilities in real software, rated high or critical, found in a few weeks by a group of roughly fifty organizations pointing one model at codebases the internet runs on. And the interesting part is what Anthropic says happened next. In the company's own words: "Progress on software security used to be limited by how quickly we could find new vulnerabilities. Now it's limited by how quickly we can verify, disclose, and patch the large numbers of vulnerabilities found by AI."
That sentence is the actual news. For thirty years, software security has been a discovery-constrained field. Finding a serious bug in a widely used library was hard, slow, specialist work, and the whole apparatus of coordinated disclosure -- embargo windows, maintainer notification, staged patch releases -- was built around the assumption that bugs arrive at a rate humans can process. If a model can find them faster than volunteer maintainers can fix them, the queue is the vulnerability. Anthropic's own numbers make the contrast concrete: enterprise customers using Claude Security patched 2,100 issues in three weeks, which the company notes is much faster than the open-source side, "in large part because enterprises are fixing their own code, whereas open-source fixes usually require volunteer maintainers who work through coordinated disclosure."
Picture a city that has just invented a machine that can inspect every building for structural defects in an afternoon. The inspection is solved. What is now broken is that the city still has the same forty structural engineers who have to certify each repair.
Anthropic has also started shipping the surrounding apparatus rather than just the model. The Glasswing update describes a Cyber Verification Program that lets security professionals doing legitimate vulnerability research, penetration testing and red-teaming operate without certain misuse safeguards, plus a release of the tooling the partners built: reusable skills, a harness that maps a codebase and spins up scanning subagents to triage findings and write reports, and a threat-model builder that prioritizes which parts of a codebase to attack first. That last piece matters more than it sounds. As anyone building agents has learned, the harness around a model frequently determines how capable it looks.
The honest caveat is that all of the load-bearing evidence here is Anthropic's, published by Anthropic, about a program Anthropic runs. There is no independent audit of the ten-thousand figure, no public breakdown of how many of those findings survived triage, and no external evaluation of whether the withheld model is meaningfully more dangerous than the shipped one. The company's public line -- that Mythos-class models stay restricted because current safeguards remain insufficient to prevent severe misuse -- is a judgment call the public cannot check. It is also worth noting that "restricted" has been getting steadily less restrictive: the partner group grew from about 50 in May to roughly 150 more organizations in June, and OpenAI has been running a comparable program with its own offensive cyber models. The trend line is toward wider access to offensive capability under contract, not toward a permanent hold.
Key questions
What is Project Glasswing?
What is the difference between Claude Fable 5 and Claude Mythos 5?
Why is finding ten thousand vulnerabilities a problem rather than a win?
Cite this
APA
Ground Truth. (2026, August 17). Anthropic still will not ship the model that found ten thousand vulnerabilities. Ground Truth. https://groundtruth.day/news/anthropic-still-wont-ship-the-model-that-found-ten-thousand-bugs.html
BibTeX
@misc{groundtruth:anthropic-still-wont-ship-the-model-that-found-ten-thousand-bugs,
title = {Anthropic still will not ship the model that found ten thousand vulnerabilities},
author = {{Ground Truth}},
year = {2026},
month = {aug},
url = {https://groundtruth.day/news/anthropic-still-wont-ship-the-model-that-found-ten-thousand-bugs.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.