News · 2026-07-24
Claude Opus 5 posts a verified four-fold lead on the hardest adaptation benchmark
Anthropic released Claude Opus 5 on July 24, and the independent benchmark owner ARC Prize verified it at 30.16% on ARC-AGI-3, its test of how well an AI adapts inside unfamiliar interactive environments. The previous best published result on that test was 7.78%. The model ships with a one-million-token context window, reasoning switched on by default, and standard API pricing identical to its predecessor, Opus 4.8.
Key facts
- The headline number: 30.16% on ARC-AGI-3 at high effort, verified by ARC Prize, against a previous listed best of 7.78% for GPT-5.6 Sol.
- When: launched July 24, 2026.
- Who: Anthropic, with the benchmark result independently confirmed by ARC Prize, the non-profit that owns and runs the test.
- Primary source: Anthropic's Opus 5 launch post and its developer release notes.
Most model launches arrive with a chart of the company's own numbers and no way for an outsider to check them. This one is unusual because a third party got there first and published the receipts.
ARC-AGI-3 is not a quiz. Older benchmarks hand a model a question and grade the answer. ARC-AGI-3 drops the model into an interactive environment it has never seen, gives it no instructions in words, and watches whether it can work out what the goal even is and then get better at reaching it across repeated attempts. A perfect score means matching human action efficiency across every game. ARC Prize publishes the configurations, costs, outputs, and task-level replays for every run, which is why its numbers carry weight that a vendor slide does not.
The useful way to read 30.16% is not as an accuracy grade. It is a jump of roughly four times on a test that had barely moved, on a skill - figuring out an unfamiliar system by poking at it - that has been one of the most stubborn gaps between people and machines. ARC Prize also records that Opus 5 cleared five public-demo environments no previous model had beaten. Two caveats belong in the same breath: ARC Prize tested Opus 5 only at high effort, not its maximum setting, because the testing window before launch was short, and the 7.78% comparison figure was produced at max effort. So the lead is real but not effort-matched. And ARC Prize's academic panel reviews methodology rather than re-scoring each individual result.
The mechanism behind the gain is not a disclosed architectural breakthrough. Anthropic's developer notes describe something plainer: the model thinks by default, developers pick a depth from low through max, and higher settings buy it more room to reason, act, and revise its own work through tool loops. Think of it less as a smarter engine and more as a car that has been given permission to circle the block a few times before committing to a turn. That is the same lever behind test-time compute results across the industry, and it comes with the same bill: thought tokens count against your max_tokens budget alongside the visible answer.
Which brings us to the money, where the loudest claim needs the most care. Anthropic's launch post says Opus 5 lands within half a percentage point of Fable 5 on a coding-agent test "at half its per-task cost". The per-token arithmetic behind that is exact: every standard rate is precisely half of Fable's, including cache writes. But Opus 5 is not cheaper than Opus 4.8 - it is priced identically, at $5 per million input tokens and $25 per million output. Anyone already running Opus saw no price change at all. Worse for the cost story, Opus 5 thinks by default while 4.8 did not unless configured, so migrating without setting an effort level can quietly raise real spend. Cost per finished task depends on output length, tool calls, effort, and retries - none of which a price table settles.
The rest of Anthropic's benchmark roster - Frontier-Bench, CursorBench, OSWorld, Humanity's Last Exam and others - remains company-run. Its own Frontier-Bench footnote discloses an internal harness and Opus 4.8 as the fallback on safety refusals. That is launch evidence, not leaderboard fact.
There is also a product wrinkle worth knowing: picking "Opus 5" in a consumer app does not always get you Opus 5. Anthropic documents an automatic, visible fallback to Opus 4.8 on certain higher-risk cyber requests, with the classifier reading uploaded files, connector context, and prior conversation rather than just your last message. Anthropic says these switches happen 85% less often than with Fable 5 - its own estimate. Anyone running informal head-to-head tests should know they may not be testing the model they selected.
The launch hit number one on Hacker News with well over a thousand points. The strongest dissent in that thread is not that the ARC score is fake - ARC Prize verified it - but that per-token pricing tells you almost nothing about per-task economics once effort levels, thinking tokens, tool calls, and fallbacks vary. That objection is correct, and it is the honest caveat here: the benchmark win is independently confirmed, and the universal cost win is not.
Key questions
What is ARC-AGI-3 and why does Opus 5's score matter?
Is Claude Opus 5 cheaper than Claude Opus 4.8?
Does selecting Opus 5 always get you Opus 5?
Cite this
APA
Ground Truth. (2026, July 24). Claude Opus 5 posts a verified four-fold lead on the hardest adaptation benchmark. Ground Truth. https://groundtruth.day/news/claude-opus-5-arc-agi-3-lead.html
BibTeX
@misc{groundtruth:claude-opus-5-arc-agi-3-lead,
title = {Claude Opus 5 posts a verified four-fold lead on the hardest adaptation benchmark},
author = {{Ground Truth}},
year = {2026},
month = {jul},
url = {https://groundtruth.day/news/claude-opus-5-arc-agi-3-lead.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.