News · 2026-09-30
LiveNerf begins measuring Claude drift, with no degradation verdict yet
LiveNerf is collecting a preregistered baseline to test whether Claude Code’s measured performance changes over time, and it has not yet found an Opus 5.5 degradation. The project had six daily runs through September 29, while some users separately reported worse answers after that day’s Claude outage. The new evidence is an auditable measurement plan, not confirmation that the outage changed the model.
Key facts
- LiveNerf tests a frozen panel of 78 Opus 5.5 questions and a 12-question control.
- Baseline collection began September 24, 2026, at 22:10 UTC.
- Its decision rule requires a persistent change across two subsequent ten-day windows.
- Primary source: LiveNerf’s public repository and results.
Complaints that a model feels different are hard to settle with a single repeated prompt. Model outputs vary; tasks differ; tools and product settings change. The author’s contribution is to commit to a panel, a sampling schedule, and a decision rule before enough results accumulate to support a conclusion. That makes a future claim easier to audit and makes an early dramatic interpretation harder to justify.
The preregistration freezes the questions, prompts, grading rules, and harness. The project runs through Claude Code on a Claude Max subscription, with version 2.1.280 pinned and effort explicitly set to high. Calls use one turn without tools, project hooks, memory, or connected-tool servers. It logs outputs, usage, latency, versions, and a harness identifier. It measures that product path, which may differ from the raw model delivered by a direct programming interface.
The panel is deliberately sensitive. Out of 2,336 candidate academic questions, calibration found that 97% were always right or always wrong. Those questions offer little signal about a modest performance shift. LiveNerf selected questions the model sometimes answered correctly, then used fresh confirmation samples to address the selection bias. It is like choosing a scale that moves visibly when a small weight is added, while checking that the sensitivity was not an accident of calibration.
Its primary outcome compares accuracy on the same items over time. Output length provides a secondary indicator of changed thinking effort. A declared change must exceed three percentage points, clear a 99% uncertainty interval in each of two ten-day windows, and not appear in the Opus 5 control. The design’s predicted detectable shift for one window is roughly 7.5 points. The three-point rule is therefore not a promise to reliably detect every three-point change at its current budget.
The strongest caveat comes from the project itself. Its validation report could not distinguish Opus 5 from Opus 5.5 at 99% confidence on accuracy or token count in the tested sample. That result limits claims about detecting subtle model replacement. The panel audit also identified ambiguous questions and suspect answer keys; those remain frozen, with a planned sensitivity analysis. Transparent imperfections are preferable to quietly changing a test after seeing its results.
The timing also prevents a launch-day comparison. Opus 5.5 launched September 22, but collection began roughly two-and-a-half days later. A change before the first run becomes part of the reference rather than something the study can recover. The baseline-update discussion explicitly says the alleged dip will be in the baseline. The repository’s 30-day design anticipates a first comparison around day 20 and a full two-window decision around day 30, approximately October 24.
Anthropic’s September 29 incident record describes 59 minutes of errors and access problems from 14:00 to 14:59 UTC. It names affected products, including Claude Code and the API, but not a downgraded model. A user’s after-outage report establishes an impression and a temporal association. It cannot identify the cause, and LiveNerf’s daily series was not designed to isolate that hour-long incident.
Anthropic’s earlier engineering postmortem offers useful context: infrastructure bugs have previously degraded outputs, even while the company denied intentionally reducing quality with demand. That precedent makes investigation reasonable, but does not explain this outage. Evan Miller’s paper, “Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations,” explains the uncertainty problem, cited by LiveNerf, helps explain why repeated observations and uncertainty matter. The project’s value is turning a disputed perception into a testable prospective claim, with an explicit ceiling on what the instrument can resolve.
The published run history is therefore a starting point for future comparison, with the baseline dates and instrument limits attached to every later interpretation.
Key questions
Has LiveNerf shown that Opus 5.5 was nerfed?
Does LiveNerf test the Claude API directly?
Could the meter detect any model substitution?
Cite this
APA
Ground Truth. (2026, September 30). LiveNerf begins measuring Claude drift, with no degradation verdict yet. Ground Truth. https://groundtruth.day/news/livenerf-baseline-not-a-nerf-verdict.html
BibTeX
@misc{groundtruth:livenerf-baseline-not-a-nerf-verdict,
title = {LiveNerf begins measuring Claude drift, with no degradation verdict yet},
author = {{Ground Truth}},
year = {2026},
month = {sep},
url = {https://groundtruth.day/news/livenerf-baseline-not-a-nerf-verdict.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.