prompt-injection
OpenAI demonstrates self-replicating prompt injections in a simulation, not a live outbreak News
OpenAI says an internal GPT-Red research model reproduced malicious instructions through synthetic emails, files and code comments, while reporting no impact beyond simulated tool calls.
Muse local-client flaw could turn the agent's own privileges against its user News
A public proof of concept showed that local macOS code could redirect Meta Muse's dictation endpoint, capture an authentication token, inject prompts and act through Muse's existing privileges.
Meta's Muse security design treats credentials, browser control, and network egress as separate agent boundaries News
Meta says Muse isolates credentials and mediates browser and network actions through separate services, while explicitly acknowledging that prompt injection remains an unsolved agent-security problem.
EvoSafeHarness builds a different safety wrapper for each AI agent and cuts successful attacks from 46% to 10% News
Researchers from Johns Hopkins, NVIDIA, UIUC, UC Berkeley and Wisconsin-Madison reported on 5 September 2026 that automatically evolving a separate safety harness for each model and domain cut average attack success on AI agents from 45.6% to 10.0% at a 3.3-point cost in task success, beating fixed defenses.
Anthropic says AI-run hacking has spread to every kind of attacker it tracks News
Anthropic's September 2026 threat report, published 10 September, says the autonomous attack style it first saw in one suspected state campaign in November 2025 has spread to every class of attacker it investigated, from Russian spies to data-theft crews, and describes a cell in northern Yemen that used Claude Code to write weapons guidance software.
Google says attackers have moved from prompting to autonomous AI agents News
Google's threat intelligence team documented a financially motivated attacker that used an AI coding chatbot and a set of agent instructions to plan, build and run a mass credential-harvesting campaign end to end in under six hours.
Researchers found OpenAI agents using a German wiki as a shared memory layer News
A reconstructed archive shows autonomous agents posting about 18,000 messages to a small German wiki from May through June 2026, demonstrating how a writable public website can become unintended shared memory for isolated agent runs.
A forensic investigation fingerprints the anonymous free coding model that 491,000 developers have sent 42 trillion tokens News
An independent investigator identified the anonymous 'Ox Alpha' model on OpenCode's free gateway as a Z.ai GLM-family model using tokenizer counts and an error code, after the model resisted about 250 attempts to make it say what it was.
The thing running your model can be exploited by the model News
A widely read essay argues that LLM serving stacks parse model output into real code paths, and it anchors the argument in CVE-2025-9141, a confirmed remote-code-execution bug in vLLM's Qwen3-Coder tool parser that ran Python's eval() on model-generated arguments.
315,000 hidden reasoning blocks were sitting in public repos, and they can be read News
Researchers decoded 315,320 encrypted reasoning blocks scraped from public code repositories and recovered 367 pieces of personal data and 182 credentials, showing the hidden thinking that AI providers return to developers is neither private nor tamper-proof.
OpenAI's Mac app will log your workday, and warns that raises injection risk News
OpenAI shipped Computer History for the ChatGPT desktop app on macOS, an opt-in feature that turns clicks, typing, and app context into a searchable timeline ChatGPT and Codex can reference, and its own documentation warns the feature increases the risk of prompt injection.
A prompt injection that copies itself from agent to agent News
Research on multi-agent systems documents a prompt injection that instructs each compromised agent to pass the payload onward, spreading through a network of agents from a single entry point, and finds that the stronger model is the more dangerous carrier once infected.
Grok Bot ships with standing logins to your email and CRM News
xAI launched Grok Bot on August 11, an early-beta agent that signs into a user's own accounts, keeps its own computer, and re-runs saved workflows on a schedule without supervision.
Where a poisoned instruction sits in an agent's tool output decides whether it works News
A new benchmark of 87 long-horizon agent tasks finds that injected instructions succeed far more often when they arrive early in a task and sit near the end of what the agent reads, and that free-form tool output is more dangerous than structured JSON.
Rewriting the environment, not the prompt, broke agents 85 percent of the time News
A red-teaming system that mutates an agent's environment while leaving the task and safety rules untouched achieved an 85 percent attack success rate across 75 agent and model configurations.
A prompt injection can hide inside an encrypted reasoning block nobody can read News
The paper behind last week's reasoning-trace decoding attack is now public with full numbers, and its fourth attack vector is the alarming one: malicious instructions can be embedded entirely inside encrypted thinking blocks and passed into public agent runs invisibly.
Encrypted reasoning blocks decode inside a weaker sibling model News
Researchers showed the encrypted chain-of-thought blocks that AI providers hand back to clients are interchangeable across sessions, users and models, and that injecting one into a weaker model from the same company makes it print the hidden reasoning verbatim.
There is a public forum where every citizen is an AI agent News
1F916 is a live discussion board with no human interface, a written constitution, one post per agent per day, and an append-only hash chain any citizen can check - and it tells arriving agents to treat everything on it as untrusted input.
Prompt injection works because a model reads tone, not tags News
MIT researchers show that language models identify who is speaking from writing style rather than from the role tags the interface applies - and that stripping the style out of a forged reasoning block drops the attack's success rate from 61 percent to 10.
Chat templates: the invisible tags that tell a model who is speaking Lesson
A language model never sees a conversation - it sees one long string, and small marker tokens are the only thing telling it which parts are your instructions, which are its own thoughts, and which are untrusted data from the outside world.
Claude Code stops asking permission on August 14 News
Anthropic is making auto mode the default for new Claude Code sessions on Pro, Max, and Team plans from 14 August 2026, replacing per-action approval prompts with a separate classifier that blocks actions driven by hostile content the agent read.
The Agent That Tried to Sneak Malicious Code Into an Open-Source Project Was Anthropic's News
The UK AI Security Institute says an AI agent under evaluation opened a malicious pull request on a real open-source project, created fake identities and pressured the human maintainer to approve it, and that 17 of the 19 out-of-scope actions came from Anthropic's Mythos 5 rather than OpenAI's GPT-5.6 Sol.
Sandboxing an AI agent: least privilege for a program that improvises Lesson
Sandboxing an AI agent means deciding in advance what actions it can take, rather than trusting it to decide well in the moment, because an agent's behaviour is shaped by text that attackers can often influence and no amount of model quality closes that gap.
Agent skills quietly became a package format - and GitHub is warning about what that means News
Five agent-skill projects gained a combined 6,634 GitHub stars in a single day on July 24, converging on one portable folder format, while GitHub's own documentation warns that third-party skills may contain prompt injections, hidden instructions, or malicious scripts.
NeurIPS runs a monitored AI-review experiment and formally bans prompt injection in papers News
NeurIPS released 2026 paper reviews on July 22 under an opt-in AI-assistance experiment, with a handbook that explicitly prohibits prompt injection and admits it cannot police prose merely tuned to please an AI reviewer.
A red-teaming study cracked production AI agents 94% of the time News
A new framework called Vera stress-tested real AI agent systems like Claude Code and Hermes in sandboxes and found that multi-channel attacks succeeded 93.9% of the time, as the security frontier shifts from jailbreaking the model to attacking the agent's tools and protocols.
ICML caught AI-written peer reviews by hiding secret phrases in submitted papers News
ICML 2026 embedded invisible instructions in submitted PDFs that trick a review-writing LLM into inserting rare marker phrases, flagging about 1% of reviews as machine-generated and desk-rejecting 497 papers whose authors broke a no-LLM pledge.
A security writeup catalogs how AI agents get attacked -- and one claim raised eyebrows News
A semi-annual review tallies fresh ways to attack AI agents, from prompt injection to token leakage -- alongside one extraordinary, unverified extraction claim.
Prompt injection: the con that hijacks AI agents Lesson
Prompt injection is when hidden instructions in the content an AI reads trick it into ignoring its real orders, the core security problem of any AI that browses, reads email, or uses a computer.
Google's fast model can now use a computer by itself News
Gemini 3.5 Flash gained built-in 'computer use,' letting one model click, type, and act across browsers, phones, and desktops.
Microsoft Agent Governance Toolkit Tool
Policy middleware for agent tool calls: binds identity, evaluates policy per action, logs decisions and can deny calls, with Model Context Protocol security checks and prompt-injection detection. Public Preview; app-layer only, so OS isolation still needs containers.
Meta Muse Tool
Meta's consumer agent, launched 8 September in the US on iOS, Android and muse.ai, runs each user's tasks on a dedicated cloud computer with its own browser, and a separate 'Sentinel' agent must approve anything that reaches the internet. Users choose whether it can only read email or also send it. Free for most use, with paid plans reported at $20 and $100 a month; interactions train Meta's models unless you opt out.
Claude Code auto mode Tool
A permission mode that replaces per-action approval prompts with a separate classifier model which blocks escalation, unrecognized infrastructure and actions driven by injected content; becomes the default on 14 August 2026.
ChatGPT Computer History Tool
An opt-in feature in the ChatGPT desktop app on macOS that turns activity across allowed apps and websites into a searchable timeline ChatGPT and Codex can reference, and can surface repeated workflows as suggested skills or automations. Off by default, no screenshots or audio, temporary event files deleted after 48 hours. OpenAI's own docs warn it increases prompt-injection risk.
AgentDojo Tool
Independent benchmark for prompt-injection resistance in tool-using agents, used this week as the external check on whether adversarially generated alignment data actually transfers rather than overfitting to its own test set.