Ground Truth.
AI, checked against the source.

← All topics

leaderboards

Everything on Ground Truth tagged “leaderboards” — 4 items.

Qwen did not take the top agentic spot from Claude, but it got within one point News

Artificial Analysis's Agentic Index currently places Claude Opus 5 at maximum effort first with 59, and Qwen3.8 Max tied for second at 58, contradicting posts describing Alibaba's model as the outright leader.

Google Falls Off One Leaderboard's Top 15, as a Report Describes a Gemini-Specific Chip News

Google dropped out of the top 15 on LLM Stats' composite leaderboard while remaining its fastest model, and Reuters separately reported an unannounced Gemini-specific inference chip.

Vals AI Tool

Independent evaluator that scores frontier and open-weight models on professional workloads - finance, tax, legal, medical, public benefits - alongside coding benchmarks, with per-model cost figures. A useful counterweight to vendor-published charts.

Artificial Analysis Agentic Index Tool

A public leaderboard averaging agentic benchmarks that give models shell and web access, including a multi-step banking workflow scored on the resulting database state rather than the model's own summary. Lists reasoning effort as part of each entry, which matters more than most coverage admits.