Ground Truth.
AI, checked against the source.

News · 2026-09-21

Remote Labor Index puts Astra at 20.83% client-acceptable project completion

GPT-6 Astra achieved a 20.83% client-acceptable completion rate on Scale’s Remote Labor Index, while Fable 5.1 reached 17.92%. The result is a meaningful advance in end-to-end agent performance, but it also means that even the leading system failed roughly four out of five projects at the benchmark’s professional delivery standard.

Key facts

RLI is designed to measure delivery rather than a model’s ability to produce a persuasive chat answer. Tasks include briefs, input files, professional reference deliverables and economic information across documents, code, data work, design, audio, video, 3D and CAD-like artifacts. Three independent evaluators compare the AI artifact with a human reference. A score of two means a reasonable client would accept it; a score of three means it exceeds the reference. The automation rate is the share scoring two or three.

That definition makes the 20.83% figure more meaningful than a multiple-choice score and less sweeping than a labor-market claim. It says the evaluated model-plus-agent system cleared a particular client-quality bar on about one project in five. It does not say one in five jobs has disappeared, that a human worker is replaceable, or that the model can reliably handle an unmeasured project. The benchmark projects represent over 6,000 hours of human work and $143,991 in stated value, but they are still a finite sample.

The trajectory is substantial. The original frontier was 2.5%, about one accepted project in forty. CAIS’s July update put Fable 5 at 15.8%; Scale’s current figures reach roughly one in five for Astra. CAIS called the earlier increase “significant,” but also emphasized that most projects fail. The difference between “can sometimes finish” and “can take responsibility for the work” remains huge.

The failure mechanism is the story. The paper’s HTML version highlights poor professional quality, incomplete artifacts, malformed or corrupted files and inconsistency across deliverables. A visually impressive architectural render can mask a broken 3D file; a polished web page can omit required files or have incorrect data. This is the last-mile problem: an agent needs not only generate pieces but also verify that its exported artifact opens, matches the brief and survives a customer’s workflow.

The leaderboard also cannot be read as a clean model-only ranking. Scale says it tests agents in both CLI and computer-use environments, reports the single best performance and caps generation spending at $30 per task. The CAIS setup used different scaffolding, computer use, revision loops, time budgets and sometimes higher spending. Comparisons across published evaluations are directionally useful but not perfectly apples-to-apples.

There is a second measurement warning. CAIS reported that an automated judge put GPT-5.5 at 17.9% while human evaluation found 6.25%, and put Opus 4.8 at 18.8% while humans found 8.33%. That is why the human-evaluated RLI number is valuable and why glossy agent demos should not be treated as client acceptance. The counterargument—that benchmark scaffolds and private tasks obscure real-world variation—is valid. It strengthens, rather than weakens, the practical conclusion: companies should run comparable end-to-end acceptance tests on their own files, tools and reviewers.


Primary source, verified: read the paper → (arXiv 2510.26787)

Key questions

What does the 20.83% RLI score mean?

It means human evaluators judged Astra’s deliverable at least as good as a professional reference and acceptable to a reasonable client on 20.83% of evaluated projects.

Does the RLI score mean AI automates 20.83% of all jobs?

No: it is a score on a finite set of 240 end-to-end freelance-style projects and is not an economy-wide job-automation estimate.

Why are RLI scores not pure model comparisons?

The leaderboard evaluates model-plus-agent systems in CLI or computer-use environments and reports each agent’s best result.
Cite this

APA

Ground Truth. (2026, September 21). Remote Labor Index puts Astra at 20.83% client-acceptable project completion. Ground Truth. https://groundtruth.day/news/remote-labor-index-astra-fable-automation-rate.html

BibTeX

@misc{groundtruth:remote-labor-index-astra-fable-automation-rate,
  title  = {Remote Labor Index puts Astra at 20.83% client-acceptable project completion},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {sep},
  url    = {https://groundtruth.day/news/remote-labor-index-astra-fable-automation-rate.html}
}

Topics: agents · benchmarks · automation · evaluation · labor

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.