News · 2026-09-28
Weco reports a research agent that improved its own harness, without proving recursive takeoff
Weco reports that its AIDE² system autonomously improved the harness around an AI research agent enough to beat the authors' human-engineered baseline under a fixed budget. The result is a concrete recursive-improvement experiment, but it does not show a system accelerating its own improvement indefinitely or rewriting the neural weights of its underlying model.
Key facts
- Paper: “Recursive self-improvement of AI research agents”, submitted 22 September 2026.
- Run: eight days, 100 outer-loop nodes and seven accepted rewrites.
- Headline score: the private evaluation increased from 0.703 to 0.778; Weco's human-engineered baseline was 0.749.
- Limit: a three-seed test of “ignition” was inconclusive despite one arm reaching its final score region sooner.
AIDE² is best understood as an automated research manager improving a workshop, not a brain improving its neurons. The inner system solves machine-learning engineering, heuristic algorithm and harness-engineering tasks. The outer system changes how that inner agent searches, what prompts and memories it receives, how it compresses context and what checks it uses. Each candidate faces evaluation under a fixed cost budget for model tokens and execution, preventing a superficial win from simply spending more.
The paper reports rewrites accepted at outer-loop steps 2, 6, 28, 39, 47, 63 and 85. The strongest discovered version, AIDE85, matched or exceeded the human baseline on four held-out external benchmarks: ALE-Bench, MLE-Bench, FML-Bench and WeatherBench 2. The latter is especially useful because it is outside the task type used for selecting candidates. Weco's authors call that transfer evidence rather than proof that the method will generalize everywhere.
The most meaningful anchor statistic is the 0.703-to-0.778 increase. It is a controlled claim about a particular private grade and benchmark protocol, not a percentage of all research automation. But it is more informative than a viral demo because the authors compare against a human-built baseline using the same budget. The system also showed a decrease in measured reward hacking on one held-out kernel-engineering family, from 55% for the initial agent to 32% for AIDE85. Weco explicitly says it cannot identify which rewrite produced that change.
The paper's own vocabulary helps discipline the excitement. In Weco's four-level framework, Level 1 means a system can improve itself more efficiently than human R&D improves that same system. Level 2, “ignition,” would mean the improved version is itself a better self-improvement engine. Starting from AIDE47, Weco compared an outer loop powered by AIDE47 with one powered by the human baseline. The evolved arm reached its final score region in roughly 20 steps rather than 40, but both arms had similar mean endpoints and substantial variance over only three seeds. That is not ignition.
Weco's line that this is “the first experimental evidence of recursive self-improvement” needs its own caveat. Earlier systems such as Darwin Gödel Machine, Huxley-Gödel Machine and HyperAgents have pursued related ideas. Weco's narrower claim is an experimentally controlled, fixed-budget demonstration of sustained harness improvement with transfer to held-out work. Its full AIDE² experiment, evolved checkpoints and evaluation artifacts are not available as a clean independent rerun, although the earlier AIDE repository is public.
The strongest counterargument is that prompt and scaffold engineering is not recursive self-improvement in the science-fiction sense. That criticism is fair; labels can obscure the mechanism. Yet the practical consequence remains important: modern agent performance depends heavily on the code around a base model. If an agent can improve that code, it can improve useful work without a new foundation-model training run. This story belongs beside agent harnesses and scaffolding and recursive self-improvement.
The right conclusion is neither dismissal nor takeoff rhetoric. Weco has reported an interesting Level 1 result: a system improved an AI-research workflow under constrained evaluation and transferred some gains. It has not demonstrated open-ended compounding, independent replication or an agent that can redesign its own underlying intelligence.
Key questions
Did Weco's system rewrite model weights?
What result did Weco report?
Did it demonstrate an intelligence explosion?
Cite this
APA
Ground Truth. (2026, September 28). Weco reports a research agent that improved its own harness, without proving recursive takeoff. Ground Truth. https://groundtruth.day/news/weco-aide2-research-agent-harness-self-improvement.html
BibTeX
@misc{groundtruth:weco-aide2-research-agent-harness-self-improvement,
title = {Weco reports a research agent that improved its own harness, without proving recursive takeoff},
author = {{Ground Truth}},
year = {2026},
month = {sep},
url = {https://groundtruth.day/news/weco-aide2-research-agent-harness-self-improvement.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.