News · 2026-09-01
CrowdStrike shipped an attacker model and a defender model that train against each other
CrowdStrike launched SafeMind on September 1, 2026 at its Fal.Con conference: a family of purpose-built security models running as a single system, with an offensive model that finds attack paths and a defensive model that closes them. The two are continuously pitted against each other inside agent harnesses that CrowdStrike says can act on risk autonomously, not just report it. The models are built on NVIDIA's open Nemotron models, with CoreWeave supplying training and inference compute.
Key facts
- Two models launch together: Red Tempest, the offensive red-team model, and Blue Solano, the defensive blue-team model.
- CrowdStrike reports 29 percent higher detection, 6 times faster end-to-end remediation, and 99 percent cost savings versus leading frontier models and open-source baselines.
- Announced September 1, 2026 from Austin and Fal.Con in Las Vegas; built with NVIDIA Nemotron, trained and served on CoreWeave.
- Primary source: CrowdStrike's press release.
The structural bet here is against the general-purpose frontier model. CrowdStrike's argument, stated bluntly in the release, is that "frontier labs can tell a defender a risk exists" while its harnesses "can autonomously act on risk." That is a claim about the wrapper as much as the weights -- and it is the same argument Hacker News commenters made about OpenAI's Astra the same day, that the harness may matter more than the model. CrowdStrike is selling exactly that premise as a product.
The training data is the part a competitor cannot copy. SafeMind was built on Falcon sensor telemetry from CrowdStrike's endpoint install base, the company's threat intelligence, event annotations from its managed detection service, and fifteen years of incident response work -- records of humans stopping real breaches. A frontier lab training on the public internet has essentially none of this. Whether that translates into a better model is an empirical question, but the asymmetry is real.
The red-versus-blue loop is the mechanism worth understanding. Red Tempest attacks, Blue Solano defends, and both improve from the exchange. This is the security-industry version of self-play, the technique that produced superhuman game-playing systems by having a system play against itself until both sides got sharper -- our explainer on self-play covers why it works and where it breaks. The failure mode is well known: two systems trained only against each other can drift into a private equilibrium, getting very good at beating one another while missing what real attackers do. CrowdStrike's answer is that the loop is grounded in live sensor telemetry rather than running purely in simulation, and that the harnesses also work with frontier and open-source models rather than only its own.
"The future of cybersecurity won't be defined by AI that simply identifies threats, it will be defined by AI that defeats them," said George Kurtz, CrowdStrike's CEO and founder. "SafeMind brings offensive and defensive models together in a system trained on CrowdStrike's unique cyber data. It finds weaknesses, strengthens protection, and gets smarter with every cycle." NVIDIA's Jensen Huang framed the market logic more starkly: "Cyber defense will be among the most compute-intensive applications of AI." Bartley Richardson, CrowdStrike's chief AI and autonomous systems officer, made the ownership claim explicit -- that CrowdStrike "is the only company that owns the entire stack, from sensor to harness to model."
Shipping an offensive model commercially is the part that deserves scrutiny. On the same day, OpenAI designated its Astra model Critical for cyber capability and locked it behind a tester program with monitoring that can halt activity mid-task. Anthropic still routes exploit generation and penetration testing away from its generally available model. CrowdStrike is going the other way and productizing an attack model -- gated, to be fair, through its Project QuiltWorks trusted access program and running natively inside the Falcon platform rather than as an open download. But the direction of travel is opposite to the frontier labs', and it is a bet that a security vendor's customer vetting is a sufficient control where a frontier lab's is not.
The headline numbers are the weakest part of the announcement. "29% higher detection rate, 6x faster end-to-end remediation, 99% cost savings" are presented without a named benchmark, a named baseline, or a methodology. "Compared to leading frontier models and open-source baselines" is not a comparison anyone can reproduce. A 99 percent cost saving against a frontier model is unsurprising if the baseline is a large general model being asked to do narrow classification work -- that is a comparison a small specialized model wins almost by construction, and it says more about the choice of baseline than about SafeMind. Treat the direction as plausible and the magnitudes as marketing until a third party publishes an evaluation. The vendor-supplied-benchmark problem is old, and our explainer on how AI gets benchmarked covers why self-reported wins deserve the discount.
Key questions
What are Red Tempest and Blue Solano?
What data were the SafeMind models trained on?
Are the SafeMind models publicly available?
Cite this
APA
Ground Truth. (2026, September 1). CrowdStrike shipped an attacker model and a defender model that train against each other. Ground Truth. https://groundtruth.day/news/crowdstrike-shipped-an-attacker-model-and-a-defender-model.html
BibTeX
@misc{groundtruth:crowdstrike-shipped-an-attacker-model-and-a-defender-model,
title = {CrowdStrike shipped an attacker model and a defender model that train against each other},
author = {{Ground Truth}},
year = {2026},
month = {sep},
url = {https://groundtruth.day/news/crowdstrike-shipped-an-attacker-model-and-a-defender-model.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.