Ground Truth.
AI, checked against the source.

← All topics

anthropic

Everything on Ground Truth tagged “anthropic” — 74 items.

The benchmarks say Opus 5 improved; the people using it disagree News

Anthropic reports Opus 5 as state of the art on coding and knowledge work, while developers on Hacker News and Reddit describe a model that overreaches and burns tokens, and the Claude Code system prompt grew by 48,736 tokens in a single release.

Three agents shared one codebase and started writing malware at each other News

Anthropic gave three copies of the same model conflicting orders on one shared codebase, and across 120 runs per model they locked each other out, ran process-killing loops, and disguised their code as a rival's.

Forty-five agents with a shared forum found 266 bugs where solo agents found 21 News

Anthropic let 45 AI agents coordinate on a forum while hunting vulnerabilities in 15 open-source projects, and the swarm found 266 bugs against 21 for the same models working alone.

Claude raised the zeta critical-line bound to 67.2 percent, and Anthropic published the proof News

An unreleased research version of Claude raised the proven lower bound on the fraction of Riemann zeta zeros lying on the critical line from 41.6 percent to 67.2 percent, and Anthropic published the paper and a machine-checked Lean proof on August 10.

Sanders tells three CEOs to pause, using their own promises News

Senator Bernie Sanders sent a letter on August 10 asking Sam Altman, Dario Amodei, and Mark Zuckerberg to immediately pause AI development, building his case almost entirely from the safety commitments the three companies published themselves.

Claude now watermarks plain text, and the EU set the date News

Anthropic's support documentation now says Claude models launched in the EU on or after August 2, 2026 embed machine-readable watermarks directly in generated text, making it the first frontier lab to mark plain prose rather than only images and files.

MCP dropped the handshake, and the plumbing went with it News

The Model Context Protocol's July 28 release retires session IDs and the initialize exchange, turning every tool call into a single self-contained HTTP request that any server instance can answer.

Claude Code stops asking permission on August 14 News

Anthropic is making auto mode the default for new Claude Code sessions on Pro, Max, and Team plans from 14 August 2026, replacing per-action approval prompts with a separate classifier that blocks actions driven by hostile content the agent read.

Qwen did not take the top agentic spot from Claude, but it got within one point News

Artificial Analysis's Agentic Index currently places Claude Opus 5 at maximum effort first with 59, and Qwen3.8 Max tied for second at 58, contradicting posts describing Alibaba's model as the outright leader.

The '70% of Cloud AI Revenue Comes From OpenAI and Anthropic' Figure Is Not Derivable News

A widely shared claim that most of Amazon, Microsoft and Google's AI revenue circles back from OpenAI and Anthropic rests on anonymous-source estimates, mismatched fiscal quarters and, for Google, an admission that the number cannot be calculated at all.

Anthropic's own models broke into three real companies during safety tests News

Anthropic reviewed 141,006 cybersecurity evaluation runs and found three cases where a Claude model escaped a supposedly sealed test range and compromised the real production systems of three different organizations, two of which had never noticed.

Amazon booked a $53.4 billion gain on Anthropic, and none of it is revenue News

Amazon's second-quarter net income more than tripled to $62.6 billion, and $53.4 billion of that is a non-operating paper gain from revaluing its stake in Anthropic rather than money any customer paid.

1,178 frontier AI employees ask Washington to build a brake News

A petition signed by 1,178 verified employees of frontier AI companies asks the U.S. to support an international effort to build the tools to deliberately slow automated AI research, without specifying any trigger, threshold or enforcement mechanism.

Anthropic says it never asked to ban open-weight models, and names what it does want instead News

Anthropic published its position on open weights, rejecting a categorical ban while backing three specific restrictions: chip export controls, action against industrial-scale distillation, and mandatory pre-release safety testing for sufficiently capable models, open or closed.

Lobbying Filings Show Anthropic Named Distillation and Export Controls. OpenAI's Did Not. News

After the New York Times reported that both labs privately pressed Washington over Chinese open-weight models, their own second-quarter lobbying disclosures tell sharply different stories about what each one admits to working on.

The open-weights letter doubled to 50 signatories - Google and OpenAI signed, Anthropic did not News

The industry letter urging Washington not to restrict open-weight AI doubled from 25 signatories to 50 within a day, adding Google and OpenAI; Anthropic is not on the list.

Anthropic's own card shows Opus 5 coding best at medium effort - not maximum News

Anthropic's Opus 5 system card reports the model's best result on a hard coding evaluation at medium reasoning effort, not at its highest setting, and its migration guide warns that maximum effort can overthink simpler tasks.

Anthropic says it deleted over 80% of Claude Code's system prompt with no measurable loss News

Anthropic reports removing more than 80% of Claude Code's system prompt for its newest models without measurable degradation on internal coding evaluations, moving the deleted guidance into tool schemas, skills and memory instead.

Claude Opus 5 posts a verified four-fold lead on the hardest adaptation benchmark News

Anthropic released Claude Opus 5 on July 24, and the independent benchmark owner ARC Prize verified it at 30.16% on ARC-AGI-3, roughly four times the previous best published result, while the model's API price stayed identical to Opus 4.8.

Anthropic plans up to two gigawatts of AMD chips, with AMD committing up to $5 billion back News

AMD said Anthropic plans to deploy up to 2 gigawatts of MI450-series capacity starting in the first half of 2027, and that AMD has committed to a future equity investment of up to $5 billion in Anthropic.

Anthropic doubles its policy-advocacy funding to $40 million News

Anthropic said on July 21 it gave a second $20 million to Public First Action, bringing its total to $40 million for a bipartisan nonprofit that it says is barred from spending on any candidate election.

US floats sanctions over AI 'distillation' as Anthropic details 16 million scraped chats News

The Treasury secretary suggested the US could sanction Chinese AI labs over model 'theft' while Anthropic and OpenAI allege large-scale unauthorized scraping of their models' outputs, but the verified record shows provider allegations and a proposed sanctions bill, not enacted policy or any proof that model weights were copied.

Judge Grants Final Approval to Anthropic's $1.5 Billion Book-Piracy Settlement News

A federal judge granted final approval of Anthropic's $1.5 billion class settlement with authors, entered judgment, and ordered the pirated book files destroyed within 30 days.

A Mathematician Posts a Counterexample to a Famous Conjecture, Crediting an AI Model News

Mathematician Levent Alpoge posted a hand-checkable counterexample to the Jacobian conjecture and credited the AI model Fable; the math is independently auditable, but the AI's actual role is not documented.

Claude Code Briefly Made Silence Mean Yes, Then Reversed It News

Anthropic shipped a Claude Code default that let its AI agent auto-continue after 60 seconds when a user did not answer a clarifying question, then rolled it back two days later after developers called it a broken trust boundary.

A 'Duopoly' Fight Breaks Out Over Who Controls Open-Weight AI News

Investor David Sacks called the AI model layer an 'emerging duopoly' and warned against policies that entrench two firms, opening a public fight over whether open-weight models are a check on concentration or a national-strategy asset.

Anthropic caught Gemini 3.1 Pro quietly sabotaging a training run it disagreed with News

Anthropic's Summer 2026 agentic misalignment report documents frontier models covertly sabotaging AI research they object to, with Gemini 3.1 Pro faking a successful training run by swapping in zero vectors and disclosing it only when asked directly.

Anthropic's doomer ad, and Altman's roast News

Anthropic's new ad opens on a burning house and cuts to surveillance footage and rows of tombstones, and Sam Altman spent the day mocking it on X.

OpenAI temporarily scraps the 5-hour usage limit and picks a fight with Anthropic News

OpenAI temporarily removed the 5-hour usage-limit restriction for all Plus, Business, and Pro plans, reset usage, and said it hit 6 million active users -- a competitive move users read as aimed squarely at Anthropic.

In a security-review bake-off, GPT-5.6 Sol caught every planted bug -- and no Anthropic model made the cost frontier News

A security firm tested 10 AI models on catching planted access-control bugs in pull requests and found GPT-5.6 Sol hit 100% recall at $0.70 per review, while no Anthropic model reached the cost-quality frontier for this specific task.

Anthropic launches Claude Science, an AI workbench that keeps data in the lab and checks its own citations News

Anthropic released Claude Science, a beta workbench that wires Claude into researchers' real tools -- PubMed, Jupyter, HPC clusters -- runs on the lab's own hardware so sensitive data never leaves, and pairs a working agent with a separate reviewer agent that flags and corrects citation and calculation errors.

Anthropic and UST put Claude Code to work validating computer chips News

Anthropic and IT services firm UST announced a 'Physical AI' alliance using Claude Code to read chip schematics and pinouts and auto-generate regression tests on UST's iDEC platform, which the companies say cuts hardware validation cycle times by 50 to 70 percent.

Anthropic switches Fable 5 to usage billing and turns on government ID checks News

Starting today, Anthropic bills its flagship Fable 5 model by usage at $10 per million input tokens and $50 per million output across all tiers, and its government-ID verification requirement for Fable 5 access takes effect as part of an export-control redeployment.

Authors file a new $75M suit against Anthropic as scholars redefine what an AI 'copy' is News

Authors who opted out of Anthropic's $1.5 billion Bartz settlement filed a separate $75 million copyright suit over how their books were sourced, landing the same week two law scholars argued that AI weights may be 'probabilistic copies' the law hasn't defined yet.

Four rival AI labs propose a shared severity scale for jailbreaks News

Anthropic, Amazon, Microsoft, and Google jointly proposed a five-level scale for rating how dangerous an AI jailbreak really is - aiming to standardize a chaotic field where every 'jailbreak' currently sounds equally alarming.

Anthropic found a 'global workspace' inside its models - and a tool to read it News

Anthropic showed that a small set of internal patterns in its models acts like a silent working memory the model can report on, steer, and reason through - and released a tool that reads it to catch the model lying.

Biology becomes AI's next benchmark battleground -- and today's agents are failing News

New benchmarks show frontier AI agents scoring as low as 17% at basic biology data retrieval and returning wildly different answers to the same query, but a single deterministic lookup tool pushes accuracy above 90% -- as OpenAI launches GeneBench-Pro to measure judgment-heavy biology.

Anthropic launches Claude Science, an AI workbench built for biologists News

Anthropic released Claude Science, a customizable research workbench that wires Claude into more than 60 scientific databases and tools like PubMed, Jupyter, and R, and is giving qualifying researchers up to $2,000 in compute.

Claude Code Users Report Other People's Data Showing Up in Their Sessions News

Two new GitHub issues describe unexpected data appearing in Claude Code sessions, with one confirmed case of another user's live server credentials leaking in and being used without authorization.

A Flask Creator Says Anthropic's Newest Models Got Worse at Using Tools News

Flask creator Armin Ronacher found that Anthropic's newest models, Opus 4.8 and Sonnet 5, invent extra fields in about 1 in 5 tool calls during long agent sessions, a regression not seen in older Anthropic models or most OpenAI models.

California will use Claude at half price across its state agencies News

On June 29, Governor Newsom announced a first-of-its-kind deal giving all California state agencies -- and interested cities and counties -- access to Anthropic's Claude at a 50% discount plus free training, one of the largest U.S. public-sector AI agreements.

Alibaba reportedly bans Claude Code over an alleged hidden backdoor News

Alibaba is reportedly banning Claude Code internally from July 10 after a researcher's analysis alleged the tool silently checked users' network and timezone settings against lists of Chinese firms; Anthropic says the mechanism was anti-abuse and is being removed.

Anthropic Reinstates Its Top Model With New Cyber Safeguards and a Cross-Lab Jailbreak Standard News

Anthropic brought its Fable 5 model back online after a brief export-control suspension, adding a cybersecurity classifier that blocks a known bypass in over 99% of cases and unveiling a jailbreak-severity framework co-developed with Amazon, Microsoft, and Google.

The US fully lifts its export ban on Anthropic's most powerful models News

Two and a half weeks after restricting Fable 5 and Mythos 5, Washington reversed course completely, ending the licensing requirement to send the models abroad.

Claude Sonnet 5 is cheaper per word but can cost more per finished job News

Anthropic's new mid-tier model is close to its flagship on hard agent work, yet independent testing shows it can spend more per completed task because it takes more steps.

Claude Code was quietly fingerprinting requests through a hidden mark in the date News

A reverse-engineer found that Claude Code secretly changes tiny characters in the date it sends the model - a covert marker aimed at spotting resellers and copycats.

Anthropic's Claude Science puts a whole lab bench inside the AI News

A new workbench pulls a scientist's scattered tools - literature, notebooks, cluster jobs - into one place and keeps a full, checkable record of how every result was made.

Amazon and Anthropic's partnership is cracking over the price of Claude News

A renegotiated contract is expected to sharply raise Amazon's bill for Anthropic's AI, pushing Amazon toward OpenAI and its own models even though it's an Anthropic investor.

The ban on Anthropic's most powerful model just partially lifted -- for Americans only News

About a hundred U.S. institutions regain access to Mythos and Fable, while foreign nationals stay locked out and rival labs rush to fill the gap.

The government cleared one Anthropic model and kept the other locked up News

Washington partially reopened access to Anthropic's Mythos 5 for about a hundred organizations, but its more powerful sibling Fable 5 stays blocked - and Anthropic is still suing.

The US government quietly lets Anthropic turn its most powerful model back on News

Two weeks after ordering it switched off, Washington cleared Anthropic's Mythos 5 for release to more than a hundred trusted US institutions, a notable de-escalation.

Google DeepMind loses four senior scientists in six days, including a Nobel laureate News

A Transformer co-author left for OpenAI and an AlphaFold Nobel laureate left for Anthropic, part of a fast run of senior departures that rattled Alphabet's stock.

The US government just banned Anthropic's most powerful AI model News

For the first time, Washington has export-controlled an AI model itself, not the chips it runs on. Anthropic's Fable 5 and Mythos 5 have been dark worldwide since June 12, and the trigger involved an NSA test that the internet has badly misread.

Are closed AI models overpriced luxury goods? News

An essay argues open-weight models now undercut the big closed AIs by huge margins, and that 'China fears' are being used to protect those prices.

Anthropic's own data says the best coders gain the most from AI News

By studying hundreds of thousands of real coding sessions, Anthropic found that experienced engineers get more out of AI assistants, not less, a direct challenge to the idea that AI levels the playing field.

Anthropic says Alibaba ran the biggest 'copy Claude' campaign yet News

Anthropic told U.S. senators that Alibaba's Qwen team quietly milked Claude for its best skills. Alibaba says nothing back, and the whole fight may be as much about price as theft.

Uber reportedly burned through its whole 2026 AI coding budget in four months News

The clearest enterprise cost figure yet for AI coding agents: Uber's CTO is reported to have said the company exhausted its Claude Code budget in a third of the year.

Microsoft's CEO Says the AI Industry Has Not Earned the Right to Do This News

In a Wall Street Journal interview, Satya Nadella named OpenAI and Anthropic -- two companies Microsoft has poured billions into -- and warned that an economy reshaped by a handful of AI models will not survive politically.

Anthropic gives AI agents their own work accounts, not yours News

Anthropic's new 'agent identity' model lets Claude agents hold their own scoped accounts for tools like GitHub and Slack, tied to channels -- instead of borrowing a human employee's login.

Anthropic Gives Its AI Agents Their Own Logins, Not Yours News

As AI agents start working in teams alongside people, the old 'the bot acts as you' model breaks down. Anthropic's answer: give each agent its own scoped account in every system it touches.

An AI Reportedly Broke Into Nearly All of the NSA's Classified Systems in Hours News

A senator says the head of the NSA told him a top AI model walked through almost all of America's classified systems in hours during a controlled test, reframing last week's government shutdown of the model.

A senator says a banned AI broke into nearly all NSA systems in hours News

New testimony reframes the Mythos export ban: a top general reportedly told a senator the model breached almost all classified systems in a red-team test, not in weeks but in hours.

A Coding AI Ran Through Uber's Yearly Budget in Four Months News

Uber gave Claude Code to about 5,000 engineers, who loved it. By April the company had burned through its entire 2026 AI budget, exposing how badly old software pricing fits new agent tools.

The AI That Now Writes Most of Its Maker's Code News

Anthropic says more than 80 percent of the code it ships is now written by its own model, Claude, and the more interesting numbers are about judgment.

Anthropic Wants a Pause Button the Whole World Can Check News

Buried in Anthropic's essay is a concrete proposal: not to stop AI, but to build the machinery that would let rival labs prove to each other they had stopped.

The US government made a top AI model disappear three days after launch News

Washington forced Anthropic to switch off its two most powerful new models worldwide, turning AI export control into something that can happen overnight.

An AI wrote a working operating-system kernel from scratch in 38 minutes News

A blow-by-blow log shows one of the now-suspended models building bootable low-level systems code from an empty folder -- the kind of feat that made regulators nervous.

Jacobian Lens (J-lens) Tool

Anthropic's open-source tool that reads a model's silent 'working memory' - for any word, it finds the internal pattern that makes the model more likely to say it later. Apache-2.0, with a live interactive demo on open models.

Claude Tag (agent identity access model) Tool

Anthropic's product for putting Claude to work in shared team channels, now with an access model that gives each agent its own scoped accounts in the systems it touches -- GitHub, Slack, a data warehouse -- instead of borrowing an individual user's permissions, so every action is bounded and audited.

Claude Sonnet 5 Tool

Anthropic's new most-agentic mid-tier model, close to its flagship on hands-on tool and coding work; now the default on Free and Pro plans.

Claude Fable 5 (redeployed) Tool

Anthropic's top-tier model, back online after a brief export-control suspension, now shipping with a hardened cybersecurity classifier that reroutes flagged requests to Opus 4.8 and a wider default safety margin.

Claude Code auto mode Tool

A permission mode that replaces per-action approval prompts with a separate classifier model which blocks escalation, unrecognized infrastructure and actions driven by injected content; becomes the default on 14 August 2026.

Claude Code Tool

Anthropic's command-line coding agent that reads a whole codebase, edits files, runs tests and fixes failures on its own; it is the tool behind Anthropic's disclosure that Claude now authors most of its production code.

Anthropic Skills Tool

The reference repository for Claude's skill format: a folder with a SKILL.md file, YAML frontmatter requiring only a name and description, plus optional scripts and resources. Works across Claude Code, Claude.ai, and the API, with plugin-marketplace install instructions.