News · 2026-09-21
Remote Labor Index puts Astra at 20.83% client-acceptable project completion
GPT-6 Astra achieved a 20.83% client-acceptable completion rate on Scale’s Remote Labor Index, while Fable 5.1 reached 17.92%. The result is a meaningful advance in end-to-end agent performance, but it also means that even the leading system failed roughly four out of five projects at the benchmark’s professional delivery standard.
Key facts
- The Scale RLI leaderboard lists Astra at 20.83% and Fable 5.1 at 17.92%.
- RLI contains 240 freelance-style projects across 23 work categories; 230 are private evaluation tasks.
- A pass is a human judgement that the AI output is as good as the professional reference and acceptable to a reasonable client.
- The canonical methodology paper is Remote Labor Index.
RLI is designed to measure delivery rather than a model’s ability to produce a persuasive chat answer. Tasks include briefs, input files, professional reference deliverables and economic information across documents, code, data work, design, audio, video, 3D and CAD-like artifacts. Three independent evaluators compare the AI artifact with a human reference. A score of two means a reasonable client would accept it; a score of three means it exceeds the reference. The automation rate is the share scoring two or three.
That definition makes the 20.83% figure more meaningful than a multiple-choice score and less sweeping than a labor-market claim. It says the evaluated model-plus-agent system cleared a particular client-quality bar on about one project in five. It does not say one in five jobs has disappeared, that a human worker is replaceable, or that the model can reliably handle an unmeasured project. The benchmark projects represent over 6,000 hours of human work and $143,991 in stated value, but they are still a finite sample.
The trajectory is substantial. The original frontier was 2.5%, about one accepted project in forty. CAIS’s July update put Fable 5 at 15.8%; Scale’s current figures reach roughly one in five for Astra. CAIS called the earlier increase “significant,” but also emphasized that most projects fail. The difference between “can sometimes finish” and “can take responsibility for the work” remains huge.
The failure mechanism is the story. The paper’s HTML version highlights poor professional quality, incomplete artifacts, malformed or corrupted files and inconsistency across deliverables. A visually impressive architectural render can mask a broken 3D file; a polished web page can omit required files or have incorrect data. This is the last-mile problem: an agent needs not only generate pieces but also verify that its exported artifact opens, matches the brief and survives a customer’s workflow.
The leaderboard also cannot be read as a clean model-only ranking. Scale says it tests agents in both CLI and computer-use environments, reports the single best performance and caps generation spending at $30 per task. The CAIS setup used different scaffolding, computer use, revision loops, time budgets and sometimes higher spending. Comparisons across published evaluations are directionally useful but not perfectly apples-to-apples.
There is a second measurement warning. CAIS reported that an automated judge put GPT-5.5 at 17.9% while human evaluation found 6.25%, and put Opus 4.8 at 18.8% while humans found 8.33%. That is why the human-evaluated RLI number is valuable and why glossy agent demos should not be treated as client acceptance. The counterargument—that benchmark scaffolds and private tasks obscure real-world variation—is valid. It strengthens, rather than weakens, the practical conclusion: companies should run comparable end-to-end acceptance tests on their own files, tools and reviewers.
Key questions
What does the 20.83% RLI score mean?
Does the RLI score mean AI automates 20.83% of all jobs?
Why are RLI scores not pure model comparisons?
Cite this
APA
Ground Truth. (2026, September 21). Remote Labor Index puts Astra at 20.83% client-acceptable project completion. Ground Truth. https://groundtruth.day/news/remote-labor-index-astra-fable-automation-rate.html
BibTeX
@misc{groundtruth:remote-labor-index-astra-fable-automation-rate,
title = {Remote Labor Index puts Astra at 20.83% client-acceptable project completion},
author = {{Ground Truth}},
year = {2026},
month = {sep},
url = {https://groundtruth.day/news/remote-labor-index-astra-fable-automation-rate.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.