News · 2026-07-25
Anthropic's own card shows Opus 5 coding best at medium effort - not maximum
Anthropic's own system card for Claude Opus 5 reports that the model's best result on Cognition's FrontierCode coding evaluation came at medium reasoning effort - 53.4 on the main set and 63.6 on the extended set - not at any of the higher settings above it. That single line undercuts the most common assumption about the new effort dial: that turning it up buys more intelligence. It buys more thinking, and on some workloads more thinking is worse.
Key facts
- Opus 5's best FrontierCode 1.1 score in Anthropic's official system card is at medium effort.
- Anthropic's migration guide says the
maxsetting "can deliver gains on the most demanding tasks" but "may show diminishing returns" and can be "prone to overthinking on simpler ones." - ARC Prize evaluated Opus 5 on ARC-AGI-3 at high effort only, scoring 30.16%, because of a short testing window.
- Both facts surfaced on July 25, 2026, driving a one-day flip in community sentiment from launch enthusiasm to skepticism.
The dial is a workload setting, not an IQ slider
Reasoning effort controls how many tokens the model spends across thinking, tool calls and visible output. Anthropic's general documentation describes the maximum setting as the unconstrained-capability configuration, and its Opus-specific prompting guide recommends the xhigh tier as a starting point for coding and agentic work. The interesting tension is that the vendor's workload-class recommendation and the vendor's own measured peak on a coding benchmark do not point at the same setting.
The mechanism is not mysterious. A model given a larger thinking budget on a task that does not need one will keep going: re-checking work that was already right, exploring alternative approaches, and in agentic runs, burning tool calls on re-exploration rather than execution. That is the same failure shape our explainer on test-time compute describes - extra deliberation has a cost curve, and the top of that curve is not always the top of the score curve.
The practical instruction in Anthropic's guide is the useful one: run a fresh effort sweep against your own evaluations rather than inheriting a setting from an older model. Effort is now part of what you are configuring, which means a benchmark claim about "Opus 5" without the effort level attached is an incomplete claim.
The ARC-AGI-3 record, and the "benchmaxxed" allegation
A prominent thread on r/singularity alleged the model's record ARC-AGI-3 result was gamed. The primary source narrows what is actually in dispute. ARC Prize's Opus 5 result card lists the score at 30.16% at high effort, and states explicitly that the short testing window meant ARC-AGI-3 was not evaluated at maximum. The maximum-effort figures circulating alongside the accusation - 97.5% on ARC-AGI-1, 90.4% on ARC-AGI-2 - belong to the older, static benchmarks.
ARC's testing policy is unusually candid about what the organization can and cannot certify. It publishes model configurations, runs evaluations through open-source harnesses, and records task replays so a viewer can inspect an individual run. The headline number comes from a semi-private task set, not from the public demos. What ARC does not claim is that it can rule out training on ARC-like puzzle distributions, or guarantee that the score transfers to unfamiliar game genres. It calls the frontier set "semi-private" precisely because tasks are sent to external APIs, acknowledges residual leakage risk, and relies on zero-data-retention agreements plus periodic benchmark replacement as its defenses.
That produces a cleaner three-way split than the argument happening online. That Anthropic may have trained on similar interactive puzzle distributions is plausible, and not against the rules - ARC publishes public tasks for development. That ARC's policy does not police distribution-level tuning is partly true. That the reported score is wrong or fraudulent is not supported: ARC lists it as verified, and no invalidation or evidence of exact-task leakage has appeared.
Why this matters beyond one model
Both stories are instances of the same problem, which our guide to how AI gets benchmarked covers at length: a benchmark can verify a score without being able to certify what the model learned, or how far that learning travels. A semi-private evaluation is meaningfully stronger than a public one - it demonstrates generalization to unseen instances within a benchmark family. It is much weaker than proof of general reasoning, and ARC says so itself.
The effort finding pushes in the same direction. When a single model ships with several reasoning budgets, and its best coding result lands in the middle of that range, a leaderboard row stops being a property of the model and becomes a property of a configuration.
The honest caveat
Anthropic has not published the FrontierCode methodology behind the system card table, and the community threads driving the skepticism offer no primary citation for their causal explanations - the claim that maximum effort triggers "unnecessary refactors" is an inference, not a measurement. The non-monotonic result is one benchmark on one model. It is a reason to sweep your settings, not a rule about reasoning effort in general.
Key questions
Does more reasoning effort always make Claude Opus 5 better at coding?
Was Opus 5's ARC-AGI-3 record set at maximum effort?
Is there evidence Anthropic cheated on the ARC benchmark?
Cite this
APA
Ground Truth. (2026, July 25). Anthropic's own card shows Opus 5 coding best at medium effort - not maximum. Ground Truth. https://groundtruth.day/news/opus-5-effort-dial-peaks-at-medium.html
BibTeX
@misc{groundtruth:opus-5-effort-dial-peaks-at-medium,
title = {Anthropic's own card shows Opus 5 coding best at medium effort - not maximum},
author = {{Ground Truth}},
year = {2026},
month = {jul},
url = {https://groundtruth.day/news/opus-5-effort-dial-peaks-at-medium.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.