Ground Truth.
AI, checked against the source.

News · 2026-09-06

ARC-AGI-3 says Astra beat its human baseline on action efficiency

ARC Prize reports that GPT-6 Astra used 51.7% fewer environment-changing actions per level than its human baseline on average in a provider-adapter harness. The result matters because ARC-AGI-3 is trying to measure not only whether a system eventually succeeds, but whether it acts efficiently in a novel environment. ARC Prize explicitly says the finding is not proof of AGI.

Key facts

The scoring design is unusually easy to misunderstand. An action is an environment-changing command. Private reasoning, retries, and internal tool calls are deliberately outside the score. The goal is to avoid rewarding a system merely for verbose narration or a giant hidden search tree; the benchmark asks how economically it changes the world once it acts. Imagine two people solving the same escape room. Both can think, consult notes, and try keys, but the score counts only moves that actually alter the room. One opens the door in five meaningful moves, the other in fifteen.

ARC Prize built its comparator rather than relying on a casual online sample. Its human-dataset account says 458 people took part in weekly in-person sessions under first-run conditions, with the same prior information and affordances as the AI. That matters because action efficiency depends on what a participant knows about the interface and how much practice it gets. The team says it had pre-registered the opposite expectation: 'action efficiency would remain a dividing line between humans and AI.'

The result remains bounded by its harness. A provider adapter makes choices about prompting, tools, memory, retry policy, and stopping rules. That surrounding system can change apparent competence, which is why the agent harness deserves as much attention as the model name. Action count also ignores cost, latency, hidden reasoning, and the possibility that a model spends enormous compute choosing one excellent move. Those exclusions are deliberate, but they prevent readers from turning 51.7% into a general measure of intelligence or commercial efficiency.

ARC Prize's own framing is the right one: it asks whether a system can learn like a human and then execute as efficiently as a human. That is narrower and more testable than 'is this AGI?' It also gives the result real significance. A system that consistently needs fewer irreversible actions can be more useful in operations, robotics, and software tasks where every action carries risk, provided its selected actions are correct.

The dossier adds separate signals that should not be fused into this result. François Chollet publicly said AGI could arrive 'Sooner, given progress is happening faster than I expected,' a forecast rather than an ARC score. SRE-Bench contains 19 private programs, 44 anti-analysis primitives, 262 binaries, and 1,572 graded reverse-engineering tasks; a follow-up says Astra reached a nearly 100% solve rate. That is important for security workflows, but it is not a general autonomy result. The durable lesson is to ask what a benchmark counts, what it ignores, and whether its human baseline and harness are credible.


Primary source, verified: read the paper →

Key questions

What does ARC-AGI-3 count as an action?

It counts commands that change the environment, while internal reasoning, retries, and tool calls do not count toward the action-efficiency score.

How large was Astra's reported efficiency advantage?

ARC Prize says its provider-adapter Astra harness used 51.7% fewer actions per level on average and fewer actions on 96.0% of levels.

Does ARC Prize say this proves AGI?

No. ARC Prize explicitly says the result is not proof of AGI and frames the benchmark as a test of human-like learning and efficiency.
Cite this

APA

Ground Truth. (2026, September 6). ARC-AGI-3 says Astra beat its human baseline on action efficiency. Ground Truth. https://groundtruth.day/news/arc-agi-3-astra-beats-human-action-efficiency-baseline.html

BibTeX

@misc{groundtruth:arc-agi-3-astra-beats-human-action-efficiency-baseline,
  title  = {ARC-AGI-3 says Astra beat its human baseline on action efficiency},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {sep},
  url    = {https://groundtruth.day/news/arc-agi-3-astra-beats-human-action-efficiency-baseline.html}
}

Topics: benchmarks · agents · openai · evaluation · reasoning

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.