News · 2026-09-06
ARC-AGI-3 says Astra beat its human baseline on action efficiency
ARC Prize reports that GPT-6 Astra used 51.7% fewer environment-changing actions per level than its human baseline on average in a provider-adapter harness. The result matters because ARC-AGI-3 is trying to measure not only whether a system eventually succeeds, but whether it acts efficiently in a novel environment. ARC Prize explicitly says the finding is not proof of AGI.
Key facts
- ARC-AGI-3's methodology uses Relative Human Action Efficiency.
- The human baseline came from 458 participants in controlled first-run sessions in San Francisco.
- Participants had 90 minutes, were paid about $130 plus $5 per solved environment, and were not told about ARC Prize or AI.
- ARC Prize's Astra analysis reports fewer actions on 96.0% of levels and 51.7% fewer actions on average.
The scoring design is unusually easy to misunderstand. An action is an environment-changing command. Private reasoning, retries, and internal tool calls are deliberately outside the score. The goal is to avoid rewarding a system merely for verbose narration or a giant hidden search tree; the benchmark asks how economically it changes the world once it acts. Imagine two people solving the same escape room. Both can think, consult notes, and try keys, but the score counts only moves that actually alter the room. One opens the door in five meaningful moves, the other in fifteen.
ARC Prize built its comparator rather than relying on a casual online sample. Its human-dataset account says 458 people took part in weekly in-person sessions under first-run conditions, with the same prior information and affordances as the AI. That matters because action efficiency depends on what a participant knows about the interface and how much practice it gets. The team says it had pre-registered the opposite expectation: 'action efficiency would remain a dividing line between humans and AI.'
The result remains bounded by its harness. A provider adapter makes choices about prompting, tools, memory, retry policy, and stopping rules. That surrounding system can change apparent competence, which is why the agent harness deserves as much attention as the model name. Action count also ignores cost, latency, hidden reasoning, and the possibility that a model spends enormous compute choosing one excellent move. Those exclusions are deliberate, but they prevent readers from turning 51.7% into a general measure of intelligence or commercial efficiency.
ARC Prize's own framing is the right one: it asks whether a system can learn like a human and then execute as efficiently as a human. That is narrower and more testable than 'is this AGI?' It also gives the result real significance. A system that consistently needs fewer irreversible actions can be more useful in operations, robotics, and software tasks where every action carries risk, provided its selected actions are correct.
The dossier adds separate signals that should not be fused into this result. François Chollet publicly said AGI could arrive 'Sooner, given progress is happening faster than I expected,' a forecast rather than an ARC score. SRE-Bench contains 19 private programs, 44 anti-analysis primitives, 262 binaries, and 1,572 graded reverse-engineering tasks; a follow-up says Astra reached a nearly 100% solve rate. That is important for security workflows, but it is not a general autonomy result. The durable lesson is to ask what a benchmark counts, what it ignores, and whether its human baseline and harness are credible.
Key questions
What does ARC-AGI-3 count as an action?
How large was Astra's reported efficiency advantage?
Does ARC Prize say this proves AGI?
Cite this
APA
Ground Truth. (2026, September 6). ARC-AGI-3 says Astra beat its human baseline on action efficiency. Ground Truth. https://groundtruth.day/news/arc-agi-3-astra-beats-human-action-efficiency-baseline.html
BibTeX
@misc{groundtruth:arc-agi-3-astra-beats-human-action-efficiency-baseline,
title = {ARC-AGI-3 says Astra beat its human baseline on action efficiency},
author = {{Ground Truth}},
year = {2026},
month = {sep},
url = {https://groundtruth.day/news/arc-agi-3-astra-beats-human-action-efficiency-baseline.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.