News · 2026-08-08
The data firms behind frontier AI sell judgment, not labels
The companies supplying frontier AI labs have stopped selling data labels and started selling judgment. Mercor, Surge AI and AfterQuery -- three of the largest -- now advertise near-identical product lines: reinforcement-learning environments, expert demonstrations, scoring rubrics, and human evaluations. What is being invoiced is graded professional reasoning, packaged as a service, and it has quietly become one of the more consequential inputs to model capability.
Key facts
- Mercor sells benchmarks, evaluation environments, large-scale human datasets, RL environments and production rubrics, and publicly claims use by "the top 5 AI labs and 6 of the Mag 7."
- Surge AI sells RL environments, rubrics and verifiers, RLHF and supervised fine-tuning data, human evaluation, expert professional domains and multimodal data.
- AfterQuery sells supervised fine-tuning data, reinforcement learning with rubrics, agent environments exposed through APIs, and computer-use trajectories.
- Primary sources: Mercor's research page, Surge AI's products page and AfterQuery.
To see why the product changed, look at what changed in training. The previous generation of data work was annotation: label this image, rank these two responses, mark this answer as toxic. Useful, but static -- a fixed dataset a model reads once.
Post-training by reinforcement learning needs something else entirely. It needs a task the model can attempt, an environment that responds to the attempt, and a rule for scoring what came back. You cannot buy that as a spreadsheet. You have to buy a working environment plus the criteria for grading behaviour inside it -- which is why the vendor pages read like descriptions of examinations rather than descriptions of datasets. Surge describes human evaluation as the gold standard; Mercor frames its evaluations as rigorous, repeatable tests of what an agent can actually do; AfterQuery's pitch is that models trained on outputs plateau while models trained on reasoning keep improving.
The analogy: the old business sold flashcards. The new business builds the practical exam, hires the examiner, and writes the mark scheme. That is a much harder thing to produce, which is why the firms doing it have become significant companies rather than staffing agencies -- and why a rubric written by a working radiologist or a securities lawyer is now a tradeable asset.
This is also the layer where reinforcement learning with verifiable rewards meets its limits and needs people. Maths and code can be checked automatically. Whether a legal memo is competent, whether a diagnosis is defensible, whether a financial recommendation is sound -- those need someone qualified to say so, or a rubric written by someone qualified, or an LLM judge calibrated against people who are. All three routes run through purchased expertise.
Which brings up the claim that has been circulating about this market, and why it does not survive contact with the record. Aggregator write-ups and social posts have carried a figure of roughly $500 million a year in training data sold by US vendors to Tencent, Alibaba and ByteDance, described as coming from company filings. No public filing discloses a buyer-country split or any such number. The revenue figures in the underlying coverage trace to unnamed sources and company-reported run rates -- a private-market estimate, not a disclosure. None of the firms has publicly confirmed or denied Chinese-lab sales either. The number may be roughly right. It is not a documented fact, and it is being repeated as one.
The export-control question is genuinely unsettled rather than obviously answered. The Bureau of Industry and Security's Export Administration Regulations govern commodities, software and technology, along with specific end-user and end-use controls. Human annotation services do not sit naturally in any of those categories -- the deliverable is not a controlled item, and the expertise being sold is not classified. The plausible pressure points are who the customer is and whether any sanctions or entity-list restrictions attach to them, not the nature of the work itself. That is a much narrower legal surface than the chip controls people reach for by analogy.
The strategic point is the one worth carrying away, and it cuts against the dominant narrative. Most US-China AI argument is about theft: accusations of distillation, weight exfiltration, illicit chip transfer. Meanwhile a legal, invoiced, above-board market has grown up around buying the thing everyone claims cannot be copied -- expert human judgment, converted into rubrics and demonstrations that any lab with a purchase order can apply to its own model. Capability that used to be a byproduct of having hired the right people is now a line item.
The honest caveat: nearly everything here comes from the vendors' own marketing pages, and marketing pages describe what a company would like to sell as much as what it does sell. Contract values, customer identities and volumes are private. What can be verified is the shape of the offering -- and that three competitors independently converged on the same shape is itself evidence about where model improvement is currently coming from.
Key questions
What do AI data vendors actually sell now?
Is selling training data to Chinese AI labs illegal for US firms?
Why did labs move from labels to environments?
Cite this
APA
Ground Truth. (2026, August 8). The data firms behind frontier AI sell judgment, not labels. Ground Truth. https://groundtruth.day/news/the-data-firms-behind-frontier-ai-sell-judgment-not-labels.html
BibTeX
@misc{groundtruth:the-data-firms-behind-frontier-ai-sell-judgment-not-labels,
title = {The data firms behind frontier AI sell judgment, not labels},
author = {{Ground Truth}},
year = {2026},
month = {aug},
url = {https://groundtruth.day/news/the-data-firms-behind-frontier-ai-sell-judgment-not-labels.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.