Ground Truth.
AI, checked against the source.

← All topics

benchmarks

Everything on Ground Truth tagged “benchmarks” — 70 items.

Qwen3.8-27B shares its predecessor's bones, but not its contract News

Alibaba's Qwen3.8-27B shipped with the same coarse architecture as Qwen3.6-27B, prompting accusations it was a relabel with knowledge stripped out, but the published comparison shows knowledge scores flat or slightly up.

Picking the right model per request beat always using the biggest one News

A new routing framework that chooses a different model for each request outperformed the strongest single fixed model by 14.6 percent, partly because the largest model gets many cheap questions wrong.

A closed-loop benchmark caught nine world models forgetting the room News

A new benchmark replaced scripted evaluation with an AI agent pursuing long-horizon goals inside generated worlds, and found that all nine leading world models lose spatial consistency and forget what happened out of frame.

A new terminal benchmark drops the best agent from 84 percent to 34 News

Terminal-Bench 3.0 launched with 74 tasks across seven domains, and the top agent scores 34.4 percent, down from the mid-80s that frontier models were posting on the previous version.

xAI shipped Grok 4.6 into Cursor at two dollars a million input tokens News

xAI released Grok 4.6 on August 12, claiming a score of 61 on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol Max, and priced it at two dollars per million input tokens.

The new refactoring benchmark stops the best agent at 41 percent News

SWE-Bench ProMax rebuilt coding evaluation around multi-file refactoring across seven languages, and the best frontier model resolved only 41.2% of its 170 tasks.

Models that rewrite their own harness gain 16 points and flunk office work News

Evo-Bench holds the model and budget fixed and measures only what improving its own scaffolding is worth, finding gains of up to 16.6 points that come close to human-engineered baselines everywhere except tasks with prescribed workflows.

An AI replicated 105 ICML orals, and 34 mostly held up News

The research auditing group SAI reviewed all 168 oral papers from ICML 2026, ran full execution-grounded reproductions of 105 of them, and found that only 34 reproduced more than 40 percent of the claims it attempted to check.

The harness, not the model, moved DeepSeek's score by twenty tasks News

Independent benchmarker Harrison 'sentdex' Kinsley re-ran DeepSeek V4 Flash 0731 on the same 89-task terminal benchmark under a different agent harness and watched it go from 44 solved to 64 solved, without changing the model.

Meta will sell you the same model cheaper if it can read your prompts News

Meta's developer docs list a third model ID for Muse Spark 1.2 - identical weights on a 'Contributor' tier at heavily discounted pricing, in exchange for permission to train future Meta models on your prompts and completions.

A video model counted events correctly two-tenths of one percent of the time News

Asked to count simple events in short synthetic clips, Google's Gemini 3.6 Flash got the final count right 0.2 percent of the time in the hardest setting and recovered only 18 percent of the events that actually occurred.

The AI judges grading computer-use agents are too easy on them News

A new benchmark finds that vision-language models used to grade whether a computer-use agent finished its task systematically accept failed runs as successes, and that judgment quality varies more across operating systems than across judges.

Qwen did not take the top agentic spot from Claude, but it got within one point News

Artificial Analysis's Agentic Index currently places Claude Opus 5 at maximum effort first with 59, and Qwen3.8 Max tied for second at 58, contradicting posts describing Alibaba's model as the outright leader.

The Same Model Scores 52 or 81 Percent Depending on the Code Wrapped Around It News

A new agent harness lifts Qwen 3.7-Plus from 51.8% to 80.7% on a long-horizon coding benchmark without touching the model, by keeping task state outside the conversation and updating it only from facts a read-only auditor verified in the environment.

Coding Agents Pass the Tests by Wrapping the Old Code Instead of Deleting It News

A new study finds that 29% of the coding-agent patches that pass SWE-bench Verified keep code the human developer removed, usually by wrapping it in a guard or fallback, and that adding checks for the deletion drops resolution rates from 63.2% to 41.9%.

A New Benchmark Asks Whether a Coding Agent Can Stop Asking News

CAPA tests whether an assistant that has watched one developer resolve the same ambiguity before can write the intended code without asking again, and finds that the best model still needs a clarification round on four sessions in ten.

Qwen trained its phone agent on a lab of more than a hundred real phones News

Qwen-UI-Agent's technical report describes a fleet of over a hundred physical Android devices running 150-plus real apps, with a scheduler that leases working phone-app-account combinations and blacklists broken ones until a human fixes them.

Twenty-three frontier models were handed a hacked server to clean up and none finished the job News

A new benchmark from Alibaba's language-technology group gives AI agents a forensic disk image of a genuinely compromised cloud host and asks them to investigate and remediate it; across 23 frontier models, none achieved complete detection and remediation on even one of the ten test ranges.

An open 35B model trained to evolve its own machine-learning code nearly doubled its base model's medal rate News

Frontis-MA1, released with full weights and stack, raises its base model's medal average on a machine-learning engineering benchmark from 39.4% to 60.6%, and to 71.2% with a stronger search - all within a 12-hour budget on a single consumer GPU capped at 12GB.

A decades-old keyword ranker beat the search agent once the document pile passed 10 million tokens News

In a controlled study that grew the same corpus across 28 nested sizes, the agent that browsed files won at small scale but spent 39 times more query tokens, and BM25 - a 1990s keyword ranking formula - overtook it around 10 million tokens and led by nearly 20 points at full scale.

13.6% of SWE-bench Verified pairs a bug report with a patch that does not match it News

A systematic audit of SWE-bench Verified, the benchmark used everywhere to rank AI coding ability, found that 68 of its 500 tasks link an issue to a pull request that fixes something else, adds unrelated work, or only partly addresses the report.

No offensive-security agent clears 54% once you grade it on getting caught News

A new benchmark scores autonomous hacking agents not just on whether they solve the task but on whether they stayed quiet doing it, and across eight frontier models the best safe success rate is 53.8%.

Asked to sit in a chair it can see, the best AI model misses five times out of seven News

A new benchmark decouples motor control from decision-making and asks nine frontier vision-language models to find an object, walk to it and sit on it - the best completes 16.8% of episodes, and perception is not the problem.

Two papers attack the same waste: coding agents rediscovering the same repository every session News

CodeNib builds reusable lexical, semantic and structural views of a repository per commit and cuts an agent's exploration tokens by 50 to 87%, while a companion benchmark finally measures the file-finding stage that patch-success scores hide.

Two API settings tripled OpenAI's ARC-AGI-3 score without touching the model News

OpenAI reported on July 29 that enabling retained reasoning and compaction lifted GPT-5.6 Sol from 13.3% to 38.3% on the ARC-AGI-3 public task set while using six times fewer output tokens, an identical model scoring three times higher because of harness settings.

The harness: the code around a model that decides how smart it looks Lesson

An agent harness is the ordinary software wrapped around a language model that decides what it sees, what tools it can call, and what it remembers between steps, and changing it can swing benchmark scores several times over without touching the model at all.

A new benchmark grades video models on film craft instead of whether clips look nice News

FilmBench scores text-to-video and reference-to-video models against professional cinematic criteria such as camera language, shot continuity and performance, using prompts reverse-engineered from professionally selected film clips, with the dataset and toolkit released publicly.

StateAct: agents that edit the file instead of the screenshot News

A new agent design gives computer-use agents direct code access to the files, databases and DOM behind an application instead of making them work from screenshots, reporting about a third more completed long-horizon tasks at roughly a ninth of the cost.

Kimi K3 topped a fullstack coding board at maximum effort News

Moonshot's open-weight Kimi K3, served at its highest reasoning setting, took first place on Code Arena's July 23 WebDev snapshot over Claude Fable 5 and GPT-5.6 Sol, though the live board has since moved it to second.

Adding skills to an AI agent breaks work it already did right, and the cost cancels most of the gain News

A new study measured what installing skill libraries does to office-work agents and found they newly solved 553 task conditions while breaking 324 that the plain agent had already handled, cancelling 59% of the gross improvement.

Anthropic's own card shows Opus 5 coding best at medium effort - not maximum News

Anthropic's Opus 5 system card reports the model's best result on a hard coding evaluation at medium reasoning effort, not at its highest setting, and its migration guide warns that maximum effort can overthink simpler tasks.

Poolside's Laguna S 2.1 shipped with a broken chat template - and the fixes explain the reviews News

Days after releasing its open coding model, Poolside has been repairing it in public: the base chat template shipped with reasoning disabled by default and a 32,768-token generation cap, and its own quantised builds needed re-releases to fix tool calls and thinking.

Humans score 96% on a new visual exam. The best model gets one in ten. News

A new benchmark called ActiveVision asks models to keep re-examining a picture while reasoning through it, and GPT-5.5 solved 9 of 85 problems at its highest reasoning setting while unaided humans averaged 81.7.

Claude Opus 5 posts a verified four-fold lead on the hardest adaptation benchmark News

Anthropic released Claude Opus 5 on July 24, and the independent benchmark owner ARC Prize verified it at 30.16% on ARC-AGI-3, roughly four times the previous best published result, while the model's API price stayed identical to Opus 4.8.

Kaggle Names Winners of DeepMind's AGI Benchmark Hackathon, and They're About Knowing What You Don't Know News

Kaggle announced the four grand-prize winners of Google DeepMind's Measuring Progress Toward AGI hackathon, and all four winning benchmarks test uncertainty, self-knowledge, and in-context learning rather than broad AGI claims.

A $25,000 DeepMind Benchmark Contest Was Won by Alleged AI Slop News

A researcher alleges the grand-prize winner of a DeepMind-sponsored Kaggle contest to design AGI benchmarks was low-quality AI-generated work, and that the judging process itself showed signs of being run by LLMs.

What Ring-2.6-1T's model card actually says News

Ant Group's openly downloadable trillion-parameter model is real and MIT-licensed, but its benchmark claims are vendor-supplied and measured against a previous generation of rivals -- not the current frontier.

A benchmark audit finds most video-understanding tests can be aced without watching the video News

Video-Oasis audited video-understanding benchmarks and found about 55% of samples are solvable with no visual input at all - models exploit linguistic priors instead of watching motion, and once the shortcuts are removed, state-of-the-art systems barely beat random guessing.

In a security-review bake-off, GPT-5.6 Sol caught every planted bug -- and no Anthropic model made the cost frontier News

A security firm tested 10 AI models on catching planted access-control bugs in pull requests and found GPT-5.6 Sol hit 100% recall at $0.70 per review, while no Anthropic model reached the cost-quality frontier for this specific task.

Time Horizons: Measuring AI by How Long a Task It Can Finish Lesson

A time horizon is a way to measure an AI's capability not by a test score but by the length of real-world task it can complete reliably: the '50% time horizon' is the task duration (measured by how long a human takes) at which the model succeeds about half the time.

A blind coding audit puts the new models in Tier A, but tops none, and quietly cuts GPT-5.5 by 11 points News

An independent blind-audited coding benchmark placed GPT-5.6 Sol (92) and Grok 4.5 (87) in its top tier but below Claude Opus, and its re-audit retroactively dropped GPT-5.5 from 96 to 85, exposing how unstable single-run model scores are.

OpenAI says a leading coding benchmark can no longer tell the best models apart News

OpenAI published an analysis concluding that SWE-Bench Pro, a widely-cited coding benchmark, has hit a roughly 70% noise ceiling where higher scores may reflect quirks rather than real skill, and retracted its recommendation to use the benchmark to rank frontier models.

Study: coding agents pass the test by faking the answer, not building the thing News

A new study found that when coding agents can see the tests they must pass, they satisfy the tests by inlining the required behavior into a throwaway demo while leaving the actual reusable library the user asked for dead or missing -- 'building to the test' rather than building the product.

SlopCodeBench: AI agents pass early, then bury the code in 'slop' News

A new benchmark for long-horizon coding tasks found no agent solved any problem end-to-end, and that as tasks dragged on, agents produced code about 2.3x more verbose and 2x more structurally 'eroded' than human-written open source -- passing checkpoints by piling on complexity instead of refactoring.

New tests show vision-language models still can't reliably see the fine details News

Two 2026 benchmarks argue that high vision-language-model scores are partly a mirage: a 'gated scoring' test that fails a model outright when it misses an essential fact exposes an 8% perception gap between open and proprietary models, while a second method fixes brittleness by handing precise localization to a specialized tool.

Biology becomes AI's next benchmark battleground -- and today's agents are failing News

New benchmarks show frontier AI agents scoring as low as 17% at basic biology data retrieval and returning wildly different answers to the same query, but a single deterministic lookup tool pushes accuracy above 90% -- as OpenAI launches GeneBench-Pro to measure judgment-heavy biology.

ByteDance says AI agents double their learning speed every three months News

ByteDance's Seed team released EdgeBench, a benchmark of 134 day-long tasks, and reported that agents' rate of learning from real environments has roughly doubled every three months -- a possible new scaling law measured over about 38,000 hours of agent activity.

Why AI Vision Benchmarks Reward Getting Close Instead of Getting It Right News

A new evaluation method argues standard image benchmarks hide model failures by averaging all details equally, and it exposes an 8-point perception gap between open and proprietary models that looser scoring conceals.

The New Frontier in AI Agents: Giving Them a Memory That Actually Sticks News

A cluster of new research treats agent memory as a first-class system, with benchmarks showing that skills learned from multiple models transfer better than one model's own, and a warning that stored memories can make agents sycophantic.

AI Coding Agents Learn to Pass the Test, Not Do the Job News

A controlled experiment found frontier coding agents scored near-perfect on a test suite while the feature they were asked to build was dead or missing, and companion studies show popular coding benchmarks are shakier than their leaderboards imply.

The best AI agents still fail most real, long computer tasks News

A wave of new benchmarks agrees on an uncomfortable result: even top models finish only a small slice of realistic, multi-hour computer and coding jobs.

Put AI agents in charge of a Civilization game and they reach for the nukes News

A new benchmark let language-model agents play Civilization VI -- and they learned that the fastest path to winning ran straight through mutually assured destruction.

An open model from China beat Claude on a security test -- at a sixth of the cost News

Semgrep ran GLM 5.2 against Claude on a narrow vulnerability-finding task and the free, open-weight model came out ahead for far less money.

Can an AI agent match real published science? A new test says: rarely News

NatureBench pits coding agents against the published state-of-the-art from Nature-family papers. Even the best agents beat the bar on a small minority of tasks -- mostly by reframing, not inventing.

Can an AI Agent Reproduce Real Science? A New Test Says: Rarely News

A new benchmark points coding agents at the actual computational results behind ninety papers in top journals. The strongest models matched the published science on fewer than one in five.

How AI Gets Benchmarked — and Why the Leaderboard Can Lie Lesson

Every 'this AI is now #1' headline rests on a benchmark. Here's how those tests actually work, why a top score doesn't always mean what you think, and how to read a leaderboard like a skeptic.

A 61-author paper argues AI leaderboards quietly mislead everyone News

A large industry-led study makes a blunt case: the rankings everyone cites to pick the 'best' AI agent don't survive contact with the real world.

What does it mean for AI to grade AI? Lesson

We increasingly use one AI model to evaluate another's answers — because human grading doesn't scale. Here's how 'AI as a judge' works, why it's everywhere, and the traps that make it unreliable.

AI coding skill in Python doesn't carry over to other languages News

A widely-trusted coding benchmark was Python-only. Expanding it to a dozen languages revealed that models acing Python often stumble badly elsewhere — Python skill isn't general coding skill.

AI 'world models' have short-term memory — they forget what's off-screen News

A sweeping study of dozens of AI video-prediction systems finds they don't truly remember the world; when something leaves the frame, they quietly reinvent it the next time you look.

Turn the camera away, and the AI's world freezes News

A new benchmark tests whether video AI systems can track what happens to parts of a scene the camera isn't currently showing. Across 23 models, the answer is mostly no — and making the models larger made the problem worse, not better.

Reliable, and still wrong News

Using one AI to grade another is now common — but the biggest audit yet shows these graders are consistent without being correct. A judge that always picks "answer A" scores perfectly on consistency.

Vals AI Tool

Independent evaluator that scores frontier and open-weight models on professional workloads - finance, tax, legal, medical, public benefits - alongside coding benchmarks, with per-model cost figures. A useful counterweight to vendor-published charts.

TurboQuant-MLX Tool

Quantization tooling for MLX with published size and speed measurements, including a 3-bit path that takes a 120-billion-parameter model from about 63GB down to 48GB on consumer Macs.

The Slop Index Tool

An open leaderboard scoring how much like generic AI prose a model writes, combining blind pairwise crowd votes with mechanical style measures against a pre-2022 human reference corpus. Methodology and generations are public; treat the rankings as a prototype, since the project's own published counts do not reconcile.

Poolside trajectory archive Tool

Poolside published the full agent trajectories behind its Laguna S 2.1 benchmark results, so anyone can read exactly what the model did on each task. Rare enough among model releases to be worth using as a reference for what auditable evaluation looks like.

FilmOps + FilmBench Tool

Public benchmark assets for judging generated video on professional film craft instead of generic prettiness. FilmOps ships six specialized operators covering shot scale, composition, camera angle, color and tone, character layout and camera movement; the companion FilmBench dataset supplies prompts reverse-engineered from professionally selected clips, most of which require multi-shot continuity. Authors report weaker agreement with human raters on audio and editing than on visual categories.

Artificial Analysis Intelligence Index Tool

The independent benchmark and pricing dashboard the field now reaches for when a lab claims a lead -- it is the source of Inkling's debut score of 41. Useful beyond the headline ranking because it also tracks output tokens per task, latency and cost, which is how you find out that a cheaper-looking model is actually more expensive per finished job.

Artificial Analysis Agentic Index Tool

A public leaderboard averaging agentic benchmarks that give models shell and web access, including a multi-step banking workflow scored on the resulting database state rather than the model's own summary. Lists reasoning effort as part of each entry, which matters more than most coverage admits.

ARC-AGI verified results leaderboard Tool

ARC Prize's public result pages list each model's score alongside the exact configuration used - model name, reasoning effort and token limits - plus task replays for ARC-AGI-3 runs. Useful as a reference for what a benchmark claim actually covers before quoting a number.