data-poisoning
A network of 23 fake websites built to be read by chatbots is pushing Alberta separatism, and some get the referendum date wrong News
Canada's National Observer found 23 AI-generated websites designed to shape what AI chatbots tell Alberta voters about the province's 19 October separation referendum, some of them publishing the wrong voting date, with evidence it describes as pointing to a possible US connection.
Anthropic retrained on the alignment-faking transcripts it had blocked News
Anthropic's August 2026 risk report discloses that filters meant to keep tens of thousands of published alignment-faking transcripts out of training data were misconfigured for several model generations, and it now suspects every Anthropic model with a knowledge cutoff after December 2024 saw some of them.
A foreign-government contract paid for websites built to be quoted by chatbots News
US foreign-agent filings document paid campaigns that build research-styled websites explicitly intended to shape what AI chatbots say, with one contract calling for the deployment of content to deliver framing results in chatbot conversations.
Four agent-memory papers landed in a week, and none tested what happens when an attacker controls the writes News
Four papers published within days define an AI agent's memory as four incompatible things - a pretrained module, a rewritten lesson, a folder of files, and a reliability ledger - and three of them introduce writable state that determines future behaviour without evaluating an adversary who controls what gets written.
One planted document flipped more than half of deep-research reports to a false conclusion News
Researchers built 5,933 credible-looking but factually false documents and slipped exactly one into the retrieval pool of several deep-research agents; the rate at which final reports endorsed the false conclusion went from zero to 54.7%.
Data poisoning and backdoors: attacking a model through what it eats Lesson
Data poisoning is an attack that corrupts a model by tampering with its training data rather than its code, and a backdoor is the sharpest form: a model that behaves perfectly until it sees a secret trigger. Anthropic and the UK AI Safety Institute found in 2025 that just 250 poisoned documents compromised models from 600 million to 13 billion parameters alike, which means scale does not dilute the threat.