Ground Truth.
AI, checked against the source.

News · 2026-09-11

Anthropic says AI-run hacking has spread to every kind of attacker it tracks

Anthropic's threat intelligence team says the autonomous style of hacking it first documented in a single suspected state-sponsored campaign in November 2025 "has now proliferated across every class of actors we investigated," from Russian state espionage to opportunistic data thieves. The report, published on 10 September 2026, covers misuse Anthropic disrupted between December 2025 and August 2026. It also describes a cell in northern Yemen that used Claude Code to write guidance software for rockets and missiles.

Key facts

The attacker's assembly line

For most of the history of hacking, two things capped how much damage an attacker could do: the supply of working exploits and the supply of skilled people to use them. Anthropic's report argues that AI has loosened both. "The operators behind observed cases range from state services to lone individuals," it says, and publicly available offensive agent frameworks such as PentAGI "reproduce much of the same scaffolding for anyone who downloads them." Google made a similar case earlier this week when it said attackers have moved from prompting to autonomous agents; Anthropic puts named operations behind the claim.

The report's most quotable conclusion is aimed at investigators: "For threat intelligence investigators, sophistication has stopped being a reliable signal of who is behind an operation."

Three operations

Russian espionage. Anthropic says its attribution of the group it tracks as GTG-20006 "is consistent with public reporting linking the actor to Midnight Blizzard," a Russian state-linked espionage group. The operators built an AI workflow that noticed when security products flagged their malware, then modified and redeployed it. "The result of the above is that AI has inverted the cost back onto defenders," the report says. A new detection used to buy defenders days; now attackers can "close the loop."

Smash-and-grab data theft. A cluster Anthropic labels "ShinyHunters smash-and-grab opportunists" used Claude to speed up mass scanning and exploitation. One operator ran a pipeline across ten cloud servers that downloaded 1.8 million distinct Android apps, unpacked them and scanned them for secrets left in the code. In a separate intrusion, after compromising a software provider, the actor dumped more than 2,100 sets of login tokens for Microsoft's corporate sign-in service, spanning more than 40 companies, in about 34 hours. "AI agents performed nearly all of the work." The group's playbook also included prompt injection against deployments of LiteLLM, an open-source gateway many companies put in front of their AI providers. The same week, the cloud security firm Wiz reported that nearly one in ten internet-facing LiteLLM instances it scanned accepted a default master key.

Exploit foundries. Chinese-speaking operators likely based in Changsha, two of whom Anthropic identified as undergraduate students, used Claude as "the engineering and orchestration layer" of an espionage programme, with autonomous workflows running vulnerability and exploit research around the clock.

AI keys are now loot

A newer pattern runs through several cases: attackers stealing AI API keys from victims and running their own operations on someone else's bill. "In every instance, the API keys involved were stolen from Anthropic customers' environments. Anthropic's own systems were not compromised by this actor," the report says. One group, after breaking into an AI vendor's evaluation sandbox, "took its production keys first." Anthropic's advice is blunt: "Organizations should treat AI keys and agent integrations with the same level of seriousness as they do production credentials."

The Yemen case

The report's weapons section details six cases, "three in China, two in Russia, and one in Yemen." The Yemen case describes "a cell of threat actors based in northern Yemen running three weapons development programs," including a guided rocket and a multi-stage ballistic missile. The actors used Claude Code "in place of human software engineers" to build guidance, navigation and control software, running several Claude sessions like a small team: one writing code, one researching, one reviewing the first one's work. They hid their goals and split tasks across sessions so no single conversation revealed the whole programme.

Anthropic is careful about what it knows. "We do not have evidence the actors succeeded in fielding an operational device; but they did test-fire a guided rocket. This field test appears to have failed: within hours, the actors returned to Claude to work out why it failed." It banned the accounts and shared threat information with public- and private-sector partners, while noting the cell had already built an offline simulation toolkit that does not rely on Claude.

Why it matters, and the caveat

For defenders, the report's message is that the gap between detection and redeployment is closing, and that AI keys and agent integrations are now credentials worth stealing. The same report's distillation section is covered separately in our story on Moonshot and DeepSeek quietly routing their users to Claude, and this week's PaperCut campaign shows the same agent-driven pattern from outside Anthropic's platform.

The caveat is that this is Anthropic's own account of its own platform. Outsiders cannot inspect the logs, several attributions are hedged, and a model provider has reasons to show both that misuse is real and that its safeguards work. Anthropic also notes that none of the cases involved its most restricted Fable or Mythos-class models, with one distillation exception. Coverage from The Record, BleepingComputer and Al Jazeera adds context; the Hacker News thread is where practitioners are arguing about it.


Primary source, verified: read the paper →

Key questions

Did the Houthis use Claude to build missiles?

Anthropic's report does not name the Houthis. It describes a cell in northern Yemen that used Claude Code to develop weapons guidance software, and says it has no evidence the group fielded an operational weapon; the one guided-rocket test it knows of appears to have failed.

What does Anthropic mean by attackers closing the loop?

An AI workflow that notices when security software detects an attacker's tool, then rewrites and redeploys it automatically, which removes the delay defenders used to win by shipping a new detection.

Were Anthropic's own systems breached?

No. Anthropic says the AI keys attackers used to power their operations were stolen from its customers' environments and that its own systems were not compromised.
Cite this

APA

Ground Truth. (2026, September 11). Anthropic says AI-run hacking has spread to every kind of attacker it tracks. Ground Truth. https://groundtruth.day/news/anthropic-says-ai-run-hacking-has-spread-to-every-kind-of-attacker-it-tracks.html

BibTeX

@misc{groundtruth:anthropic-says-ai-run-hacking-has-spread-to-every-kind-of-attacker-it-tracks,
  title  = {Anthropic says AI-run hacking has spread to every kind of attacker it tracks},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {sep},
  url    = {https://groundtruth.day/news/anthropic-says-ai-run-hacking-has-spread-to-every-kind-of-attacker-it-tracks.html}
}

Topics: cybersecurity · ai-security · threat-intelligence · anthropic · agents · prompt-injection

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.