News · 2026-08-07
Qwen did not take the top agentic spot from Claude, but it got within one point
The current Agentic Index published by Artificial Analysis places Claude Opus 5 at maximum reasoning effort in first position with a score of 59, and Alibaba's Qwen3.8 Max at 58, tied for second with Claude Opus 5 at extra-high effort. Social posts describing Qwen as the new overall leader are describing something the board does not say. The real result is narrower and more interesting: an API-only Chinese model is now within one point of the top Western agentic score.
Key facts
- Live standings: Claude Opus 5 (maximum effort) 59; Claude Opus 5 (extra-high effort) 58; Qwen3.8 Max 58.
- The Agentic Index is an equal-weighted average of Artificial Analysis's agentic benchmarks, not the general Intelligence Index that most coverage cites.
- The public page exposes no per-model run date, so there is no timestamp on any of these scores.
- Primary source: Artificial Analysis Agentic Index and its benchmarking methodology.
The error that produced the wrong headline is a specific and repeatable one, so it is worth naming. Artificial Analysis lists reasoning effort as part of the entry, not as a footnote. Claude Opus 5 appears twice on this board with two different scores, because how hard you let a model think is a deployment setting that materially changes what it can do. If you collapse those rows into a single model name - taking the lower one, as it happens - Qwen ties for the lead. Keep them separate, as the board does, and Opus 5 at full effort is alone at the top. Same data, opposite headline, and the difference is entirely in whether you read the configuration column. We have made this exact point about a different leaderboard before.
What the index measures also deserves more attention than it gets, because "agentic" is doing a lot of work in that name. Artificial Analysis's methodology says the index averages GDPval-AA v2, a task benchmark where models get shell and web access through a harness called Stirrup and outputs are judged by blind pairwise comparison anchored to human experts, and a banking-workflow benchmark that scores multi-step tasks by inspecting the resulting database state rather than reading the model's summary of what it did. That second design choice is the good one. Checking the database is the difference between grading what happened and grading what the model says happened - which, as the same week's work on lenient model judges shows, is not a small distinction.
One real gap in the published board: there is no run date. The page does not expose when any given model's score was measured. For a leaderboard whose entries change as vendors ship updates behind the same API name, that is a meaningful omission - a score with no date is a claim with no shelf life. Anyone quoting these numbers should note the day they read them.
The genuinely newsworthy fact underneath the bad headline concerns openness, not ranking. Qwen3.8 Max is not a model you can download. Alibaba's own Qwen3-Max announcement presents the family as available through Qwen Chat and the Alibaba Cloud API, and Artificial Analysis's own analysis of the family describes it as proprietary with unreleased weights. That is a notable position for the lab that built much of its reputation on open weights, and we covered it when the model shipped as a paid API. A one-point gap at the frontier is a story about capability convergence. It is not a story about anything you can run.
The caveat, as always with composite indices: an equal-weighted average of two benchmarks is a choice, not a measurement. Change the weighting or add a third benchmark and the order can flip without any model changing at all. One point on a board like this is well inside the range where methodology decisions dominate. The right reading is not "Opus 5 beats Qwen3.8 Max" - it is that at this level of measurement precision, these three entries are not distinguishable, and anyone who tells you otherwise in either direction is over-reading the board.
Key questions
What does the Agentic Index measure?
Why does the same model appear twice with different scores?
Are Qwen3.8 Max's weights available to download?
Cite this
APA
Ground Truth. (2026, August 7). Qwen did not take the top agentic spot from Claude, but it got within one point. Ground Truth. https://groundtruth.day/news/opus-5-still-leads-the-agentic-index-and-qwen-is-one-point-back.html
BibTeX
@misc{groundtruth:opus-5-still-leads-the-agentic-index-and-qwen-is-one-point-back,
title = {Qwen did not take the top agentic spot from Claude, but it got within one point},
author = {{Ground Truth}},
year = {2026},
month = {aug},
url = {https://groundtruth.day/news/opus-5-still-leads-the-agentic-index-and-qwen-is-one-point-back.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.