News · 2026-10-07
TasteVal reports cheaper experimental search, without demonstrating scientific invention
TasteVal reports that Claude Opus 5.5 reached the benchmark’s expert reference score with a compute multiplier of 2.30 across eight experimental machine-learning tasks. The preprint measures how efficiently systems choose and interpret experiments, using a fixed coding agent and constrained compute budgets. It is evidence of competitive experimental search under those conditions, not a demonstration that AI independently invents better science.
Key facts
- P-Zero Research’s Oliver Jaffe and Dane Sherburn authored the paper; the reviewed version is dated October 6, 2026.
- The study includes eight tasks and 24 experienced human researchers.
- Opus 5.5’s reported compute multiplier is 2.30, with a 95% confidence interval of 1.15 to 4.37.
- Primary source: the TasteVal paper and its full text.
Research skill is often discussed as a mysterious ability to know which experiment to run next. TasteVal makes one component measurable: given a task and a stream of results, how much experimental compute does a researcher spend before reaching a target score? That definition is narrower than the everyday meaning of taste, but the narrowness is what makes the comparison possible.
The paper’s title is “Measuring the Experimental Research Taste of AI Systems Against Human Experts.” Jaffe and Sherburn’s full paper operationalizes the experimental part, rather than claiming to encompass choosing a field, identifying a worthwhile problem, or negotiating with collaborators. The benchmark tasks span data curation, pretraining, fine-tuning, preference modeling, and adversarial-prompt alignment in text and vision settings.
The evaluated model plays the Researcher role. It chooses experiments and interprets their outcomes. A fixed Coder agent carries out implementation and execution. Separating those roles helps hold one part of the machinery steady while comparing experimental decision-making. It also means the headline result comes from a system with a researcher and supporting execution infrastructure, not a bare chatbot independently conducting an entire project.
The analogy is a racing team testing car settings. The researcher decides which adjustment to try and reads the lap results; the mechanic makes the adjustment. A good researcher can avoid wasting runs on unhelpful settings. That does not mean the team invented a new engine or chose a better racing championship. It means experimental search improved inside a defined task with a clear score.
Each run has a cap of 40 H100-hours or 120 wall-clock hours. Human participants received between four and 16 hours of task-specific literature review and could not use AI to design experiments or interpret results. The benchmark uses the best human attempt per task as its reference, with at least two human attempts on each task. That is a demanding comparison, but it is not a population-wide survey of scientific ability.
Against that reference, the authors report Opus 5.5’s 2.30 compute multiplier and roughly one-thirtieth the average per-run cost. These numbers address different denominators. The compute multiplier describes performance under the benchmark’s experimental-efficiency framework. The monetary estimate depends on the study’s run-cost accounting. Neither is a direct estimate of the cost of replacing an employed research team or of accelerating a frontier laboratory by the same factor.
The creativity finding should travel with the performance finding. The authors labeled 540 selected official model-generated submissions and found none classified as invented. They also examined about 1,600 experiments from top-decile runs, including failed and rejected experiments, and found none under that label. The result limits an expansive interpretation of the benchmark. It does not prove that models can never originate ideas, because the task design, sample, and taxonomy determine what the study can observe.
The authors spell out conditions that could inflate apparent real-world readiness. Tasks have clean training, validation, and test signals; scores arrive quickly with relatively low noise; experiments fit on one H100; and no task requires live coordination among agents. Scientific projects often lack those conveniences. A field experiment may take months, a failed result may be ambiguous, and the most useful question may not have a clean benchmark score.
There are limits in the other direction too. The paper notes restricted elicitation and a substantial spending gap between human and model runs, which could understate model potential. Honest analysis keeps both possibilities. A constrained benchmark can simultaneously exaggerate transfer to real research and fail to extract every capability of the evaluated system.
The dossier does not establish independent replication or substantive expert reception. A Reddit feed shows discovery of the topic, but not the discussion needed to infer consensus. The reported trend fits are estimates over the benchmark’s model set, not forecasts of the speed of future AI research. Benchmark interpretation and multiple-comparison discipline help explain why a good measured result still needs a defined scope.
A useful companion is the AI Theorist study, where agents iterated physical hypotheses after humans supplied real measurements. That asks a different question about scientific work. TasteVal’s contribution is a more controlled way to compare experimental choices. Its strongest counterargument is that efficient search on well-scored tasks leaves problem selection, originality, collaboration, and external validation largely untouched. Those are the next tests for any broader claim about research judgment.
Key questions
What does TasteVal mean by research taste?
How were the human experts compared with the models?
Did TasteVal demonstrate models inventing new research ideas?
Cite this
APA
Ground Truth. (2026, October 7). TasteVal reports cheaper experimental search, without demonstrating scientific invention. Ground Truth. https://groundtruth.day/news/tasteval-experimental-research-compute-efficiency.html
BibTeX
@misc{groundtruth:tasteval-experimental-research-compute-efficiency,
title = {TasteVal reports cheaper experimental search, without demonstrating scientific invention},
author = {{Ground Truth}},
year = {2026},
month = {oct},
url = {https://groundtruth.day/news/tasteval-experimental-research-compute-efficiency.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.