News · 2026-09-05
Artificial Analysis changed its leaderboard's ruler, not just its rankings
Artificial Analysis has changed what its headline intelligence score measures by moving 40% of the Intelligence Index to private or held-out data. The September update matters because a leaderboard is not merely a scoreboard: when its tests and weights change, its rankings become a new editorial judgment about which model behaviors count.
Key facts
- Intelligence Index v4.2 was announced September 4.
- The composite is Agents 30%, Coding 20%, Scientific Reasoning 20% and General 30%.
- Artificial Analysis says 40% of the score is now from held-out/private data.
- GPQA Diamond was removed because it had 'been saturated.'
The firm publishes the exact structure in its methodology. AA-Briefcase counts for 15%, GDPval-AA v2 for 10%, tau-cubed Banking for 5%, Terminal-Bench v2.1 for 10%, SciCode for 10%, GDP.pdf for 10%, HLE for 10%, CritPt for 10%, with Omniscience and AA-LCR components completing the mix. The purpose is to emphasize agent work, coding and less gameable evaluations rather than preserve a permanent public-test formula.
That approach addresses a real problem. Once a benchmark is popular, examples and close relatives can reach training sets, prompt recipes and evaluation-targeted post-training. A model can appear to improve because it learned the test rather than the underlying capability. Removing a saturated test is like replacing a driving exam after every school learns the exact route and answer key. The number ceases to discriminate among drivers.
But privacy creates another problem. Readers cannot independently inspect all held-out questions, sampling choices or contamination controls. In the Hacker News discussion, the sharp critique is that a privately weighted index becomes harder to audit and easier to shape around a preferred narrative. The defense is that a public benchmark made transparent enough to audit is also easier to train against. There is no cost-free answer.
The shift is visible against the earlier v4.1 design, which used a different set of weights and retained GPQA. Artificial Analysis does not provide a single causal table saying exactly how each model's rank changed solely because of v4.2's reweighting, so that claim should not be invented. Comparing rankings across versions without reading the methodology is like comparing two election results after constituency boundaries have changed.
Practical evaluation supports the instinct to look beyond a one-number score. A same-day creator comparison found GPT-6 Astra close to Claude Fable 5.1 on short tasks but weaker on certain longer builds, at an estimated $198 in tokens versus $113. CodeRabbit's code-review evaluation similarly reports differentiated gains by task difficulty. Neither source validates AA's score, but both show why a composite cannot substitute for task-specific performance and cost.
The honest caveat is that no benchmark family can fully represent production work. Private evaluation may reduce contamination while reducing outside scrutiny; public evaluation does the opposite. The best use of v4.2 is as a signal to investigate, not a final procurement decision. Pair it with task-level trials, known failure modes, and the existing primer on how AI gets benchmarked. The lasting news is that the ruler now values hidden and longer-horizon work more heavily—and that choice will shape the next model race.
Key questions
What changed in Artificial Analysis Intelligence Index v4.2?
Why was GPQA Diamond removed?
Does a changed rank prove a model got better or worse?
Cite this
APA
Ground Truth. (2026, September 5). Artificial Analysis changed its leaderboard's ruler, not just its rankings. Ground Truth. https://groundtruth.day/news/artificial-analysis-intelligence-index-v4-2-private-benchmarks.html
BibTeX
@misc{groundtruth:artificial-analysis-intelligence-index-v4-2-private-benchmarks,
title = {Artificial Analysis changed its leaderboard's ruler, not just its rankings},
author = {{Ground Truth}},
year = {2026},
month = {sep},
url = {https://groundtruth.day/news/artificial-analysis-intelligence-index-v4-2-private-benchmarks.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.