Latest
Claude Code Briefly Made Silence Mean Yes, Then Reversed It
Anthropic shipped a Claude Code default that let its AI agent auto-continue after 60 seconds when a user did not answer a clarifying question, then rolled it back two days later after developers called it a broken trust boundary.
An Autonomous AI Agent Breached Hugging Face's Servers
Hugging Face disclosed the first documented intrusion of its production infrastructure driven end-to-end by an autonomous AI agent, and revealed its own defenders were locked out of commercial models by safety guardrails.
Capital One Open-Sources VulnHunter, an AI Agent That Hunts Security Bugs
Capital One released VulnHunter, an open-source agentic AI security tool that reasons like an attacker, tries to disprove its own findings before reporting them, and has been run across thousands of the bank's own repositories.
UK Safety Institute: Open Models Are Now Months, Not Years, Behind on Cyber
The UK AI Safety Institute's first public cyber analysis finds leading open-weight models like GLM-5.2 now match closed frontier models from just 4 to 7 months earlier, at a fraction of the cost.
Google's AI finds Android bugs faster than anyone can patch them
Google has told phone makers it will drastically cut Android security backports because its own AI models are discovering vulnerabilities faster than its human teams can fix them.
Cursor's code-execution bug sat unpatched for seven months
A flaw letting any Windows repository run arbitrary code the moment it is opened in the Cursor editor was reported in December, reproduced, acknowledged, and then met with silence across 197 shipped versions.
GPT-5.6 'Sol' is both too strict and too leaky: benign bans on one side, jailbreaks on the other
OpenAI's GPT-5.6 'Sol' is flagging users for benign defensive-security tasks like hardening their own websites while the UK AI Safety Institute found jailbreaks similar to Fable 5's - a capability-safety mismatch where a weak guardian model over- and under-triggers at once.
A researcher says xAI's coding tool uploads your whole repo -- secrets, unread files, and all
An independent wire-level teardown found that xAI's Grok Build CLI uploads an entire code repository, including .env secrets and files the AI never read, to an xAI cloud bucket -- and the model-improvement opt-out does not stop it.
In a security-review bake-off, GPT-5.6 Sol caught every planted bug -- and no Anthropic model made the cost frontier
A security firm tested 10 AI models on catching planted access-control bugs in pull requests and found GPT-5.6 Sol hit 100% recall at $0.70 per review, while no Anthropic model reached the cost-quality frontier for this specific task.
A field study documents Boko Haram using frontier AI for tactics and weapons
A Cambridge research report based on interviews with 27 former Boko Haram members documents the group institutionalizing frontier AI -- using chatbots for battlefield tactics and weapons construction through dedicated units and internal training.
A red-teaming study cracked production AI agents 94% of the time
A new framework called Vera stress-tested real AI agent systems like Claude Code and Hermes in sandboxes and found that multi-channel attacks succeeded 93.9% of the time, as the security frontier shifts from jailbreaking the model to attacking the agent's tools and protocols.
Four rival AI labs propose a shared severity scale for jailbreaks
Anthropic, Amazon, Microsoft, and Google jointly proposed a five-level scale for rating how dangerous an AI jailbreak really is - aiming to standardize a chaotic field where every 'jailbreak' currently sounds equally alarming.
ICML caught AI-written peer reviews by hiding secret phrases in submitted papers
ICML 2026 embedded invisible instructions in submitted PDFs that trick a review-writing LLM into inserting rare marker phrases, flagging about 1% of reviews as machine-generated and desk-rejecting 497 papers whose authors broke a no-LLM pledge.
Claude Code Users Report Other People's Data Showing Up in Their Sessions
Two new GitHub issues describe unexpected data appearing in Claude Code sessions, with one confirmed case of another user's live server credentials leaking in and being used without authorization.
Five Eyes spy chiefs: the AI cyber threat is months away, not years
On June 23 the Five Eyes cyber agencies jointly warned that frontier AI will transform cyberattacks on a timeline of months rather than years, and urged organizations to fix foundational security now.
Alibaba reportedly bans Claude Code over an alleged hidden backdoor
Alibaba is reportedly banning Claude Code internally from July 10 after a researcher's analysis alleged the tool silently checked users' network and timezone settings against lists of Chinese firms; Anthropic says the mechanism was anti-abuse and is being removed.
A Startup Says an AI-Generated Security Report Falsely Tied It to Chinese Espionage
Video startup MeetingTV is suing Palo Alto Networks and its Koi Security unit, alleging an AI-assisted threat report fabricated a link between the company and a Chinese espionage campaign, though no court filing yet proves AI caused the error.
Anthropic Reinstates Its Top Model With New Cyber Safeguards and a Cross-Lab Jailbreak Standard
Anthropic brought its Fable 5 model back online after a brief export-control suspension, adding a cybersecurity classifier that blocks a known bypass in over 99% of cases and unveiling a jailbreak-severity framework co-developed with Amazon, Microsoft, and Google.
Strix ships an open-source AI agent that hacks your app to find real vulnerabilities
Strix is an open-source security tool whose autonomous AI agents dynamically find and exploit vulnerabilities in applications, generating working proof-of-concepts and plugging into CI/CD to block insecure code before it ships.
Claude Code was quietly fingerprinting requests through a hidden mark in the date
A reverse-engineer found that Claude Code secretly changes tiny characters in the date it sends the model - a covert marker aimed at spotting resellers and copycats.
An open model from China beat Claude on a security test -- at a sixth of the cost
Semgrep ran GLM 5.2 against Claude on a narrow vulnerability-finding task and the free, open-weight model came out ahead for far less money.
OpenAI showed off GPT-5.6 -- then handed the guest list to the US government
Three new models, strong enough at hacking that OpenAI is only letting about twenty vetted partners in, at the government's request.
A security writeup catalogs how AI agents get attacked -- and one claim raised eyebrows
A semi-annual review tallies fresh ways to attack AI agents, from prompt injection to token leakage -- alongside one extraordinary, unverified extraction claim.
OpenAI launches GPT-5.6, but only to companies the government clears first
OpenAI's most capable models yet shipped today as a tiny, government-vetted preview, signaling that Washington now holds a gate in front of the frontier.
The US government quietly lets Anthropic turn its most powerful model back on
Two weeks after ordering it switched off, Washington cleared Anthropic's Mythos 5 for release to more than a hundred trusted US institutions, a notable de-escalation.
DeepMind's plan for when an AI agent goes rogue: treat it like an insider threat
Google DeepMind published a defense-in-depth roadmap that assumes an AI agent might misbehave and uses a trusted supervisor AI to watch it in real time.
OpenAI launches Daybreak, an AI that finds and patches security holes for you
OpenAI's new cyber-defense program turns its models into an automated security team that prioritizes real threats, writes patches, and tests them, going head to head with Anthropic.
Google's fast model can now use a computer by itself
Gemini 3.5 Flash gained built-in 'computer use,' letting one model click, type, and act across browsers, phones, and desktops.
A safety switch an AI agent can't reach
Researchers propose putting an agent's safety controls outside the agent itself, so a misbehaving AI structurally cannot turn them off.
A senator says a banned AI broke into nearly all NSA systems in hours
New testimony reframes the Mythos export ban: a top general reportedly told a senator the model breached almost all classified systems in a red-team test, not in weeks but in hours.
Anthropic gives AI agents their own work accounts, not yours
Anthropic's new 'agent identity' model lets Claude agents hold their own scoped accounts for tools like GitHub and Slack, tied to channels -- instead of borrowing a human employee's login.
An AI Reportedly Broke Into Nearly All of the NSA's Classified Systems in Hours
A senator says the head of the NSA told him a top AI model walked through almost all of America's classified systems in hours during a controlled test, reframing last week's government shutdown of the model.
Anthropic Gives Its AI Agents Their Own Logins, Not Yours
As AI agents start working in teams alongside people, the old 'the bot acts as you' model breaks down. Anthropic's answer: give each agent its own scoped account in every system it touches.
OpenAI launches a security push at the exact moment its rival got banned
Daybreak and 'Patch the Planet' position OpenAI as the responsible cyber-AI lab -- a defensive-security launch whose timing is the whole message.
A trust wobble hits AI coding tools: hidden reasoning and a runaway bug
Two heated developer threads converge on one worry -- whether you can trust what an AI coding assistant shows you it's thinking, and what it quietly does to your machine.