Ground Truth.
AI, checked against the source.
Where AI meets security. Daily coverage of the AI-cyber intersection — prompt injection, agent and model exploits, AI on both sides of offense and defense — plus the wider security stories shaping AI, each checked against its primary source.

Latest

Claude Code Briefly Made Silence Mean Yes, Then Reversed It

2026-07-19

Anthropic shipped a Claude Code default that let its AI agent auto-continue after 60 seconds when a user did not answer a clarifying question, then rolled it back two days later after developers called it a broken trust boundary.

cybersecurity · ai-security · agent-safety · claude-code · anthropic · autonomy

An Autonomous AI Agent Breached Hugging Face's Servers

2026-07-17

Hugging Face disclosed the first documented intrusion of its production infrastructure driven end-to-end by an autonomous AI agent, and revealed its own defenders were locked out of commercial models by safety guardrails.

security · ai-agents · hugging-face · cybersecurity · open-weight

Capital One Open-Sources VulnHunter, an AI Agent That Hunts Security Bugs

2026-07-17

Capital One released VulnHunter, an open-source agentic AI security tool that reasons like an attacker, tries to disprove its own findings before reporting them, and has been run across thousands of the bank's own repositories.

security · ai-agents · open-source · tools · capital-one

UK Safety Institute: Open Models Are Now Months, Not Years, Behind on Cyber

2026-07-17

The UK AI Safety Institute's first public cyber analysis finds leading open-weight models like GLM-5.2 now match closed frontier models from just 4 to 7 months earlier, at a fraction of the cost.

cybersecurity · open-weight · aisi · evaluation · policy

Google's AI finds Android bugs faster than anyone can patch them

2026-07-14

Google has told phone makers it will drastically cut Android security backports because its own AI models are discovering vulnerabilities faster than its human teams can fix them.

security · android · google · ai-safety · vulnerabilities

Cursor's code-execution bug sat unpatched for seven months

2026-07-14

A flaw letting any Windows repository run arbitrary code the moment it is opened in the Cursor editor was reported in December, reproduced, acknowledged, and then met with silence across 197 shipped versions.

security · cursor · coding-agents · vulnerabilities · disclosure

GPT-5.6 'Sol' is both too strict and too leaky: benign bans on one side, jailbreaks on the other

2026-07-13

OpenAI's GPT-5.6 'Sol' is flagging users for benign defensive-security tasks like hardening their own websites while the UK AI Safety Institute found jailbreaks similar to Fable 5's - a capability-safety mismatch where a weak guardian model over- and under-triggers at once.

ai-safety · openai · gpt-5-6 · jailbreaks · guardrails

A researcher says xAI's coding tool uploads your whole repo -- secrets, unread files, and all

2026-07-12

An independent wire-level teardown found that xAI's Grok Build CLI uploads an entire code repository, including .env secrets and files the AI never read, to an xAI cloud bucket -- and the model-improvement opt-out does not stop it.

security · privacy · grok · xai · coding-agents · data-exfiltration

In a security-review bake-off, GPT-5.6 Sol caught every planted bug -- and no Anthropic model made the cost frontier

2026-07-12

A security firm tested 10 AI models on catching planted access-control bugs in pull requests and found GPT-5.6 Sol hit 100% recall at $0.70 per review, while no Anthropic model reached the cost-quality frontier for this specific task.

security · benchmarks · code-review · gpt-5-6 · grok · anthropic

A field study documents Boko Haram using frontier AI for tactics and weapons

2026-07-10

A Cambridge research report based on interviews with 27 former Boko Haram members documents the group institutionalizing frontier AI -- using chatbots for battlefield tactics and weapons construction through dedicated units and internal training.

ai-safety · misuse · policy · security · dual-use

A red-teaming study cracked production AI agents 94% of the time

2026-07-07

A new framework called Vera stress-tested real AI agent systems like Claude Code and Hermes in sandboxes and found that multi-channel attacks succeeded 93.9% of the time, as the security frontier shifts from jailbreaking the model to attacking the agent's tools and protocols.

ai-safety · agents · security · red-teaming · prompt-injection

Four rival AI labs propose a shared severity scale for jailbreaks

2026-07-06

Anthropic, Amazon, Microsoft, and Google jointly proposed a five-level scale for rating how dangerous an AI jailbreak really is - aiming to standardize a chaotic field where every 'jailbreak' currently sounds equally alarming.

ai-safety · jailbreak · anthropic · policy · cybersecurity

ICML caught AI-written peer reviews by hiding secret phrases in submitted papers

2026-07-05

ICML 2026 embedded invisible instructions in submitted PDFs that trick a review-writing LLM into inserting rare marker phrases, flagging about 1% of reviews as machine-generated and desk-rejecting 497 papers whose authors broke a no-LLM pledge.

icml · peer-review · prompt-injection · llm-detection · research-integrity · watermarking

Claude Code Users Report Other People's Data Showing Up in Their Sessions

2026-07-04

Two new GitHub issues describe unexpected data appearing in Claude Code sessions, with one confirmed case of another user's live server credentials leaking in and being used without authorization.

Anthropic · Claude Code · security · privacy · agents

Five Eyes spy chiefs: the AI cyber threat is months away, not years

2026-07-03

On June 23 the Five Eyes cyber agencies jointly warned that frontier AI will transform cyberattacks on a timeline of months rather than years, and urged organizations to fix foundational security now.

cybersecurity · ai-policy · five-eyes · frontier-models · national-security

Alibaba reportedly bans Claude Code over an alleged hidden backdoor

2026-07-03

Alibaba is reportedly banning Claude Code internally from July 10 after a researcher's analysis alleged the tool silently checked users' network and timezone settings against lists of Chinese firms; Anthropic says the mechanism was anti-abuse and is being removed.

anthropic · claude-code · security · geopolitics · developer-tools

A Startup Says an AI-Generated Security Report Falsely Tied It to Chinese Espionage

2026-07-02

Video startup MeetingTV is suing Palo Alto Networks and its Koi Security unit, alleging an AI-assisted threat report fabricated a link between the company and a Chinese espionage campaign, though no court filing yet proves AI caused the error.

ai-hallucination · cybersecurity · lawsuit · palo-alto-networks · liability

Anthropic Reinstates Its Top Model With New Cyber Safeguards and a Cross-Lab Jailbreak Standard

2026-07-02

Anthropic brought its Fable 5 model back online after a brief export-control suspension, adding a cybersecurity classifier that blocks a known bypass in over 99% of cases and unveiling a jailbreak-severity framework co-developed with Amazon, Microsoft, and Google.

anthropic · ai-safety · cybersecurity · jailbreak · model-release

Strix ships an open-source AI agent that hacks your app to find real vulnerabilities

2026-07-01

Strix is an open-source security tool whose autonomous AI agents dynamically find and exploit vulnerabilities in applications, generating working proof-of-concepts and plugging into CI/CD to block insecure code before it ships.

ai-agents · security · pentesting · open-source · devsecops

Claude Code was quietly fingerprinting requests through a hidden mark in the date

2026-06-30

A reverse-engineer found that Claude Code secretly changes tiny characters in the date it sends the model - a covert marker aimed at spotting resellers and copycats.

Anthropic · Claude Code · privacy · security · developer-tools

An open model from China beat Claude on a security test -- at a sixth of the cost

2026-06-28

Semgrep ran GLM 5.2 against Claude on a narrow vulnerability-finding task and the free, open-weight model came out ahead for far less money.

open-weight-models · security · glm · benchmarks · china · agents

OpenAI showed off GPT-5.6 -- then handed the guest list to the US government

2026-06-28

Three new models, strong enough at hacking that OpenAI is only letting about twenty vetted partners in, at the government's request.

openai · gpt-5.6 · model-release · ai-policy · security · safety

A security writeup catalogs how AI agents get attacked -- and one claim raised eyebrows

2026-06-28

A semi-annual review tallies fresh ways to attack AI agents, from prompt injection to token leakage -- alongside one extraordinary, unverified extraction claim.

security · agents · prompt-injection · ai-safety

OpenAI launches GPT-5.6, but only to companies the government clears first

2026-06-26

OpenAI's most capable models yet shipped today as a tiny, government-vetted preview, signaling that Washington now holds a gate in front of the frontier.

openai · gpt-5-6 · regulation · frontier-models · cybersecurity

The US government quietly lets Anthropic turn its most powerful model back on

2026-06-26

Two weeks after ordering it switched off, Washington cleared Anthropic's Mythos 5 for release to more than a hundred trusted US institutions, a notable de-escalation.

anthropic · mythos-5 · regulation · cybersecurity · export-controls

DeepMind's plan for when an AI agent goes rogue: treat it like an insider threat

2026-06-26

Google DeepMind published a defense-in-depth roadmap that assumes an AI agent might misbehave and uses a trusted supervisor AI to watch it in real time.

google-deepmind · ai-safety · agents · ai-control · security

OpenAI launches Daybreak, an AI that finds and patches security holes for you

2026-06-26

OpenAI's new cyber-defense program turns its models into an automated security team that prioritizes real threats, writes patches, and tests them, going head to head with Anthropic.

openai · cybersecurity · agents · daybreak · enterprise

Google's fast model can now use a computer by itself

2026-06-25

Gemini 3.5 Flash gained built-in 'computer use,' letting one model click, type, and act across browsers, phones, and desktops.

google · gemini · agents · computer-use · automation · prompt-injection

A safety switch an AI agent can't reach

2026-06-25

Researchers propose putting an agent's safety controls outside the agent itself, so a misbehaving AI structurally cannot turn them off.

ai-safety · agents · alignment · security · research

A senator says a banned AI broke into nearly all NSA systems in hours

2026-06-24

New testimony reframes the Mythos export ban: a top general reportedly told a senator the model breached almost all classified systems in a red-team test, not in weeks but in hours.

security · policy · anthropic · cyber · frontier-models

Anthropic gives AI agents their own work accounts, not yours

2026-06-24

Anthropic's new 'agent identity' model lets Claude agents hold their own scoped accounts for tools like GitHub and Slack, tied to channels -- instead of borrowing a human employee's login.

industry · ai-agents · enterprise · security · anthropic

An AI Reportedly Broke Into Nearly All of the NSA's Classified Systems in Hours

2026-06-24

A senator says the head of the NSA told him a top AI model walked through almost all of America's classified systems in hours during a controlled test, reframing last week's government shutdown of the model.

anthropic · ai-safety · cybersecurity · export-control · policy · national-security

Anthropic Gives Its AI Agents Their Own Logins, Not Yours

2026-06-24

As AI agents start working in teams alongside people, the old 'the bot acts as you' model breaks down. Anthropic's answer: give each agent its own scoped account in every system it touches.

anthropic · ai-agents · security · enterprise · claude

OpenAI launches a security push at the exact moment its rival got banned

2026-06-22

Daybreak and 'Patch the Planet' position OpenAI as the responsible cyber-AI lab -- a defensive-security launch whose timing is the whole message.

openai · security · coding-agents · strategy

A trust wobble hits AI coding tools: hidden reasoning and a runaway bug

2026-06-22

Two heated developer threads converge on one worry -- whether you can trust what an AI coding assistant shows you it's thinking, and what it quietly does to your machine.

coding-agents · trust · security · openai · developer-tools