Ground Truth.
AI, checked against the source.

News · 2026-07-31

Twenty-three frontier models were handed a hacked server to clean up and none finished the job

A benchmark released July 29 by Alibaba's language-technology group handed 23 frontier language models a forensic disk image of a genuinely compromised cloud server, along with the alerts and scans a security product would have generated, and asked them to investigate the intrusion and write a remediation plan. No model achieved complete detection and remediation on any single one of the ten test ranges - not one clean result out of 230 attempts.

Key facts

The gap this fills is worth stating precisely, because it explains why so many optimistic AI-security results coexist with so much unease. As the authors put it, "existing cybersecurity benchmarks focus on pre-compromise settings where agents are placed in a clean and idealized environment before an attack occurs." Those benchmarks measure whether an agent can find a vulnerability, or exploit one, in a tidy sandbox. Real security operations centres spend most of their time on the other side of that line: the attacker is already in, the evidence is messy, and the question is what happened and how to make it stop.

SecRespond builds ten cyber ranges, each from a different genuinely compromised cloud host, covering four kinds of entry point, twenty-one attacker techniques catalogued in the standard MITRE ATT&CK framework, and five operating systems. The agent gets a forensic disk snapshot plus the alerts, vulnerability scans and baseline checks a host security product reported, and must produce three forensic reports and a remediation plan. Evaluation ran on the OpenCode agent harness.

The failure pattern is consistent and diagnostic. Agents "can reliably uncover the problems exposed by alerts" but "struggle to proactively investigate the disk for silent intrusions and to produce comprehensive, verified remediation plans." Both halves matter. An intruder who triggered an alert is the easy case; the dangerous one is the intruder who did not, and finding that one requires forming a hypothesis about where to look with no prompt telling you to look there. And a remediation plan that closes four of five holes has, operationally, closed none of them.

The analogy is a burglary investigation where the detective interviews every witness who came forward, writes an accurate report about what those witnesses saw, and never checks the back window nobody mentioned. Every individual step is competent. The case is still open.

The scale of the test matters for how much weight the result carries. Twenty-three frontier models is close to the full field of serious contenders, not a convenient sample, and ten ranges built from genuinely compromised hosts is a lot of expensive environment construction. A benchmark where one or two models fall short invites the reply that better models exist. A benchmark where every model tested fails on every range points at something structural about the task rather than about any model's quality.

This is a useful corrective to the two directions the AI-security conversation currently runs in. On one side, the past month's incidents - Anthropic's disclosure that its own models reached three real organisations during evaluations and the Hugging Face intrusion reconstruction - show agents can be relentlessly effective at offence, where thousands of cheap attempts across many surfaces is a winning strategy. On the other, defensive work is exactly the job that punishes that strategy: it requires knowing when you have found everything, and nothing about a language model's training teaches it what absence of evidence looks like.

The caveat is that this is one benchmark on one harness, days old, with no independent replication, and "complete detection and remediation" is a strict bar that human responders also frequently miss on a first pass. The authors do not claim agents are useless here - they are explicit that current systems accelerate known workflows well. The claim is narrower and better supported: as an autonomous responder, none of the twenty-three models tested is ready to be trusted with the all-clear.


Primary source, verified: read the paper → (arXiv 2607.26791)

Key questions

What is post-compromise incident response?

It is the work that starts after an attacker is already inside: figure out how they got in, what they touched, what they left behind, and how to close it all off. Existing security benchmarks almost all test the earlier phase - finding vulnerabilities in a clean system before anything has happened.

What were the agents actually given?

A forensic disk snapshot of a compromised cloud host, plus the alerts, vulnerability scans and baseline checks that a host security product would have produced. They had to return forensic reports on the intrusion, the baseline risks and the vulnerability risks, along with a remediation plan.

Where did the agents fail specifically?

They handled anything an alert pointed at, but did not proactively dig through the disk for intrusions that generated no alert, and could not produce comprehensive, verified remediation plans. In other words they were good at following the breadcrumbs and poor at finding the ones nobody dropped.
Cite this

APA

Ground Truth. (2026, July 31). Twenty-three frontier models were handed a hacked server to clean up and none finished the job. Ground Truth. https://groundtruth.day/news/no-model-fully-cleaned-up-a-single-hacked-machine.html

BibTeX

@misc{groundtruth:no-model-fully-cleaned-up-a-single-hacked-machine,
  title  = {Twenty-three frontier models were handed a hacked server to clean up and none finished the job},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {jul},
  url    = {https://groundtruth.day/news/no-model-fully-cleaned-up-a-single-hacked-machine.html}
}

Topics: cybersecurity · ai-security · agents · benchmarks · incident-response · evaluation

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.