llm-as-a-judge
You can move an AI reviewer's score without changing a single result News
A new study rewrote research papers to change only their rhetoric while preserving every scientific claim, and found AI reviewers shifted their overall scores by up to nine tenths of a point, with the effect strongest near the accept-reject boundary.
NeurIPS papers average six objective mistakes each, up from four News
A study of 2,500 machine learning papers using an automated checker found that the average number of objective mistakes in a NeurIPS paper rose from 3.8 in 2021 to 5.9 in 2025, a 55 percent increase over four years.
The AI judges grading computer-use agents are too easy on them News
A new benchmark finds that vision-language models used to grade whether a computer-use agent finished its task systematically accept failed runs as successes, and that judgment quality varies more across operating systems than across judges.
A $25,000 DeepMind Benchmark Contest Was Won by Alleged AI Slop News
A researcher alleges the grand-prize winner of a DeepMind-sponsored Kaggle contest to design AGI benchmarks was low-quality AI-generated work, and that the judging process itself showed signs of being run by LLMs.
What does it mean for AI to grade AI? Lesson
We increasingly use one AI model to evaluate another's answers — because human grading doesn't scale. Here's how 'AI as a judge' works, why it's everywhere, and the traps that make it unreliable.