Learn · Intermediate
Preregistration: deciding what counts as evidence before an AI test runs
Preregistration is a time-stamped commitment to an experiment’s question, measurements, and decision rules before the results are known. In AI evaluation, it helps distinguish a genuine prediction from a persuasive story assembled after trying many prompts, metrics, or time windows. Its value is accountability for the analysis, rather than a guarantee that the benchmark measures the right capability.
The simplest analogy is writing the scoring rules before a match. If the winner can decide afterward that only second-half goals count, an apparently precise score loses meaning. An evaluator faces a quieter version of that freedom: remove an awkward question, change the effort setting, choose a more favorable baseline, or stop collecting data on the day a result looks dramatic. Each decision may sound reasonable in isolation. Together they can manufacture a convincing finding from ordinary variation.
Brian Nosek and colleagues’ The preregistration revolution explains the distinction between prediction and post-result explanation. Both are useful, but they deserve different evidential weight. For an AI system, the preregistered plan should state the model or product path, prompts, task selection, grading method, sample budget, main metric, comparison, exclusions, and decision threshold. A public repository commit can preserve such a plan, provided its timing and amendments are inspectable.
Consider a team testing whether an assistant’s quality falls after release. “We will ask it questions each day” is incomplete. The plan should specify which questions, how many samples per question, the requested reasoning effort, whether tools are allowed, how failures are graded, and which dates form the baseline. It should also say whether a claim requires one bad day, a sustained shift, or a difference from a control model. The rules determine what observation can actually count as evidence.
LiveNerf’s preregistration provides a concrete example surfaced in today’s dossier. It freezes a question panel and a Claude Code harness, then compares a baseline with two later ten-day windows. Its registered decision requires a sufficiently large shift, stringent uncertainty bounds, and an absent corresponding shift in the control. That design deliberately gives up immediate verdicts in return for a more interpretable later claim.
A useful plan also identifies the smallest effect it is designed to detect and explains why that effect matters operationally. A tiny score movement may be statistically interesting yet irrelevant to users; a study with too few samples may miss a change users would care about. Preregistration does not fix a weak instrument. LiveNerf’s validation could not confidently distinguish the tested Opus 5 and Opus 5.5 configurations. That limits what a null result can exclude. If a thermometer only resolves whole degrees, agreeing readings cannot rule out a half-degree change. Similarly, a frozen academic panel may miss a deterioration in browser use or writing style even if it measures factual answers carefully.
Uncertainty still needs an appropriate method. Evan Miller’s Adding Error Bars to Evals addresses statistical uncertainty in language-model measurement. Repeated samples of the same problem are not the same as equally many independent problems, and comparisons on matched questions can be more informative than unmatched aggregate scores. Preregistration records the chosen method; it does not make an invalid method valid. This connects to benchmark design and null baselines.
Repeated monitoring adds a stopping problem. Looking at an ordinary significance test every day and stopping on the first favorable result changes the false-alarm behavior. One option is to use fixed evaluation windows and keep collecting through them. Another is to preregister a statistical procedure specifically designed for sequential evidence. Merely announcing a daily dashboard does not make its thresholds safe for continuous peeking. Exploratory plots can remain visible without being treated as final confirmatory decisions.
A plan can change when something breaks. A corrupted log, a service outage, or a demonstrably wrong answer key may justify an amendment. The honest response is to preserve the original plan, state when the problem was found, explain the change, and show how the conclusion depends on it. A sensitivity analysis can report both the frozen panel and a version excluding suspect items. Quietly deleting the inconvenient cases defeats the reason to register the test.
Preregistration concerns how evidence is produced; holdout sets concern which data an evaluation uses. The two complement each other. A good prospective AI study reports its original rules, data and tool versions, uncertainty, deviations, and sensitivity limits. Its conclusion says exactly what the instrument found. A result that survives those commitments is easier to trust, and an inconclusive result remains useful because readers can see what was tested and what still needs a better measurement.
Evan Miller: Adding Error Bars to Evals
Brian Nosek and colleagues: The preregistration revolution
Key questions
What does preregistration prevent in an AI evaluation?
Can a preregistered study change its plan?
Does a preregistered null result prove a model never changed?
Cite this
APA
Ground Truth. (2026, September 30). Preregistration: deciding what counts as evidence before an AI test runs. Ground Truth. https://groundtruth.day/learn/preregistration-for-ai-evaluations.html
BibTeX
@misc{groundtruth:preregistration-for-ai-evaluations,
title = {Preregistration: deciding what counts as evidence before an AI test runs},
author = {{Ground Truth}},
year = {2026},
month = {sep},
url = {https://groundtruth.day/learn/preregistration-for-ai-evaluations.html}
}