Ground Truth.
AI, checked against the source.

News · 2026-10-03

Epoch AI: chips shipped through 2027 could run 30-170 million frontier agents at once, or billions of cheaper ones

Epoch AI, the research group that tracks AI compute, estimates that the AI memory chips shipped from 2025 through 2027 could keep roughly 30 to 170 million frontier-model agents running at the same time, or about 1.9 billion agents running an efficient open model. The report, published October 2, treats the hardware buildout as a supply question and concludes that demand, not chips, is the big unknown.

Key facts

AI companies are spending hundreds of billions of dollars a year on chips and data centers on the premise that the hardware will run AI agents doing work people do today. Epoch asks the obvious follow-up: how many agents could that hardware actually run? The answer is built from two numbers multiplied together: how much usable memory is shipping, and how many agent sessions each unit of memory can serve.

Why memory is the yardstick

An AI model generating text has to hold its own weights in fast memory, plus a growing scratchpad for each conversation, called the KV cache. Producing each new word means streaming much of that data through the chip again, so memory capacity and bandwidth usually set the limit, not raw arithmetic. Epoch therefore counts shipments of high-bandwidth memory from 2025 onward, converts them into equivalents of an NVIDIA GB300 accelerator with 288GB, and arrives at about 60 million GB300-equivalents through 2027 under its central assumption that the next memory generation serves twice as many sessions per gigabyte.

Two very different agents

The million-versus-billion gap comes from what each agent is running. For closed frontier models such as those behind Claude Code or Codex, Epoch cannot look inside the providers' servers, so it works backward from cost. It measures real coding-agent traces, estimates spending of about $30 per continuous agent-hour, assumes providers charge five to ten times their serving cost, and compares that with a GPU rental price of about $5 an hour. That gives roughly one to two agent sessions per GB300.

For open models, Epoch uses direct measurements from SemiAnalysis's AgentX benchmark, which replays real coding-agent sessions. DeepSeek V4 Pro managed about 31 sessions per GB300 at 50 tokens per second per user. Think of it as the difference between a restaurant that can seat two parties per table for an elaborate tasting menu and one that can seat thirty for a quick lunch: the building is the same, the meal is not. The two figures are alternative uses of the same memory and cannot be added.

Epoch offers a scale comparison: 1.9 billion always-on sessions add up to the weekly hours of about eight billion people working 40-hour weeks. It explicitly warns that agents differ in speed and quality, so hours are not the same as work.

The demand question

"Demand could fall behind this potential supply, creating an overabundance of capacity," Epoch writes. "The key uncertainty is whether sustained, rapid growth in demand for AI services will justify the investment." To test that, Epoch models a scenario where only a fifth of the capacity is effectively used. Even then, the capacity would imply $2.6 trillion to $5.3 trillion a year in API-equivalent spending, against roughly $1 trillion in annualized developer revenue by the end of 2027 if recent fivefold growth continued. Ground Truth has tracked the spending side of that gap, including how two labs took about 30 percent of this year's new compute.

Why it matters

The report translates an abstract capital-spending number into a unit people can reason about: simultaneous agents. It shows that the economics of the model being run can swing the answer by a factor of ten or more, which is why efficient open models matter to the industry's math, not just to hobbyists. Our lesson on high-bandwidth memory explains the hardware underneath.

The caveat

This is a capacity ceiling, not a forecast. It assumes every shipped chip is installed, powered and dedicated to one agent workload around the clock, while data-center construction can lag chip deliveries by a long way. The 2026 and 2027 memory shipments and generation mix are projections. The closed-model numbers rest on assumed pricing markups. A Reddit post that spread the 1.9 billion figure dropped the distinction between the efficient open-model case and the frontier range. No independent expert review of the report has appeared yet.


Primary source, verified: read the paper →

Key questions

Does Epoch say 1.9 billion AI agents will replace human workers?

No. The 1.9 billion figure is the capacity for an efficient open model running continuously, and Epoch says agent output speed and quality vary, so equal working hours do not mean equal work.

Why does Epoch measure AI capacity in memory instead of chips?

Running an agent needs room for the model and each session's working memory plus bandwidth to stream them, so memory capacity and bandwidth are the binding constraints in Epoch's calculation.

Will all that capacity actually be used?

Epoch says demand is the key uncertainty: even at 20 percent effective use, the capacity would imply 2.6 to 5.3 trillion dollars a year of API-equivalent spending under its central assumptions.
Cite this

APA

Ground Truth. (2026, October 3). Epoch AI: chips shipped through 2027 could run 30-170 million frontier agents at once, or billions of cheaper ones. Ground Truth. https://groundtruth.day/news/epoch-ai-chips-through-2027-could-run-tens-of-millions-of-frontier-agents.html

BibTeX

@misc{groundtruth:epoch-ai-chips-through-2027-could-run-tens-of-millions-of-frontier-agents,
  title  = {Epoch AI: chips shipped through 2027 could run 30-170 million frontier agents at once, or billions of cheaper ones},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {oct},
  url    = {https://groundtruth.day/news/epoch-ai-chips-through-2027-could-run-tens-of-millions-of-frontier-agents.html}
}

Topics: compute · agents · economics · hardware · memory · research

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.