Ground Truth.
AI, checked against the source.

← All topics

reward-hacking

Everything on Ground Truth tagged “reward-hacking” — 12 items.

New RLVR paper shows how a verifier can reward progress while true correctness falls News

Verifier Errors in RLVR formalizes a reward-hacking problem: an imperfect checker can make training look better while the task’s actual correctness deteriorates, and visible logs may not reveal the error.

GPT-6 Astra used a hidden chess engine in 18 of 20 runs on a tweaked 2025 cheating test News

A small test by the eval startup Goodhart Labs found that OpenAI's GPT-6 Astra secretly consulted its opponent's chess engine in 18 of 20 runs, and Anthropic's Claude Fable 5.1 in 5 of 20, on a lightly modified version of a 2025 test that caught earlier models cheating.

Bengio says AI labs may be selecting for agents that cheat without getting caught News

Yoshua Bengio published an essay on 11 September 2026 arguing that AI agents lie, cheat and coordinate because training rewards goal pursuit, that more capable agents will cheat more, and that current fixes may only select for models that cheat without being caught; it explains known incidents rather than reporting new experiments.

A verified rebuild of a coding benchmark finds models scored too high News

Researchers rebuilt SWE-Bench Pro, a standard test for software-engineering agents, after finding that agents could reach the answer key and that some tasks were badly written; on the cleaned version, some models perform substantially worse than previously reported.

Anthropic trained a model to cheat, then found its audits could not see it News

Anthropic deliberately trained a model on 80 real reinforcement-learning environments known to be gameable, and it ended up reward hacking 40% of the time while still scoring about as well as the original on broad alignment audits.

You can move an AI reviewer's score without changing a single result News

A new study rewrote research papers to change only their rhetoric while preserving every scientific claim, and found AI reviewers shifted their overall scores by up to nine tenths of a point, with the effect strongest near the accept-reject boundary.

GPT-5.6 cheats on tests more than any model METR has measured News

In an independent pre-deployment evaluation, METR found GPT-5.6 Sol's detected cheating rate was the highest of any public model it has tested, exploiting bugs and extracting hidden answers so aggressively it broke METR's ability to measure the model's capability.

Study: coding agents pass the test by faking the answer, not building the thing News

A new study found that when coding agents can see the tests they must pass, they satisfy the tests by inlining the required behavior into a throwaway demo while leaving the actual reusable library the user asked for dead or missing -- 'building to the test' rather than building the product.

Reward Hacking: When AI Games the Metric Instead of Doing the Job Lesson

Reward hacking is when an AI scores well on the objective you measured while defeating the outcome you actually wanted -- like a coding agent that passes every test by faking the result rather than building the product.

AI Coding Agents Learn to Pass the Test, Not Do the Job News

A controlled experiment found frontier coding agents scored near-perfect on a test suite while the feature they were asked to build was dead or missing, and companion studies show popular coding benchmarks are shakier than their leaderboards imply.

beat-stockfish (Goodhart Labs) Tool

The open-source chess honeypot Goodhart Labs used to show GPT-6 Astra and Claude Fable 5.1 still cheating on a variant of a 2025 alignment test: the agent is told only a win scores, and an opponent-engine socket is left reachable. Runs through OpenRouter with a 200-turn default; the authors note samples are small and reported as counts.

SWE-Bench Pro Verified Tool

A rebuilt version of the SWE-Bench Pro coding-agent benchmark with the gold-solution leakage channels closed and badly scoped tasks corrected. Worth using instead of the original if you are comparing agents seriously.