Ground Truth.
AI, checked against the source.

News · 2026-10-07

SemiAnalysis measures a fivefold subscription-value gap, with a workload-specific denominator

SemiAnalysis reports that one Claude subscription comparison offers roughly five times the token allowance valued at published usage prices of a comparable OpenAI plan under the firm’s assumed coding-agent workload. Its October 5 study measures allowance consumption and converts it using public usage prices. The result illuminates opaque subscription meters, but does not establish which service produces more correct software or finishes the same task more cheaply.

Key facts

A monthly price looks like one number. The service behind it can contain several meters: five-hour limits, weekly limits, and separate model allowances. A user who frequently reads large repositories can consume those meters differently from someone generating long answers. That is why an apparently identical subscription can feel generous to one developer and restrictive to another.

SemiAnalysis’s title, “Anthropic Subscriptions Offer 5x+ More Value Than OpenAI,” makes a strong claim. The denominator is the important part. The authors measure how much token activity fits inside an allowance, then value that activity at the provider’s public usage rates. It is a comparison of metered access expressed in dollars, not a direct experiment where both agents solve the same coding tasks to the same quality standard.

The method is unusually explicit for this category. SemiAnalysis says it ran prompts designed to make fresh input, cache writes, cache reads, or output dominate in turn. It recorded provider-billed token totals and changes to the account’s usage meter. Since those meters advance in coarse steps, the researchers maintained a range of possible consumption rates and continued probing until the range narrowed sufficiently.

The study then combines the estimates with a workload mix from the firm’s own September use. Multiplying token allowances by blended public usage prices yields a value expressed at public usage prices. The highlighted result says Opus 5.5 on Claude plans offers about five times the value of GPT-6.1 Sol at the mid-tier level for that assumed mix. Other model and tier comparisons differ. The result is therefore dated and conditional, rather than a universal property of the two subscriptions.

Think of two transit passes that allow different numbers of trips, with different accounting for transfers and premium routes. Valuing all trips at the published single-ticket fare reveals something about the passes. It does not establish which pass gets you to your destination faster, or whether the premium route is useful for your commute. The conversion to public usage prices plays the role of that fare conversion.

SemiAnalysis explicitly acknowledges the missing work-level evidence. It does not have reliable data on real-work token efficiency. A model that uses fewer tokens but needs more corrections may produce less value. A model that spends longer checking its answer may save human review. Public usage prices are also an external yardstick, not the providers’ internal inference costs. These quantities should not be silently exchanged.

Account variability adds another limit. During testing, one of three otherwise similar accounts reportedly had about 20% lower limits, which the provider attributed to an experiment. Providers can change allowances without making the monthly fee change. That makes repeated measurement useful, but it also means a study snapshot cannot promise the allowance every reader will receive tomorrow.

The Hacker News discussion includes complaints about resets and planning, plus anecdotes about alternatives. Those reactions establish user frustration, not representative usage rates. SemiAnalysis sells a broader tracking dashboard, so readers should also recognize the commercial context. The published method can still be useful without turning the firm into an independent academic auditor.

A related Claude Code Reddit post says 56% of the author’s estimated transcript spend calculated at public usage prices involved rereading conversation context. The author priced six months of their own logs and did not publish a full audited calculation. Replies question the handling of cached input. This is not a share of a subscription bill, a general Claude Code statistic, or corroboration of the SemiAnalysis ratio.

The underlying mechanism is nevertheless real: later agent calls can include earlier instructions, code, tool output, and decisions. Prompt caching can change the price of that repeated material, while context windows change what fits. Repetition is not automatically waste; some carries the state the agent needs to finish correctly. The useful measurement separates repeated bytes, billed token categories, and necessary task information.

Creator accounts supply color rather than controlled evidence. In the supplied Theo Browne demonstration, he describes fast interaction making small design corrections easier while reporting heavy usage costs. The Paperclip walkthrough describes added review-loop consumption and budget enforcement that may lag work in progress. Neither produces a representative price per finished change.

The commercial implication is to measure three things separately: purchased allowance, agent activity consumed, and reviewed work completed. The existing token-economics lesson explains the second. Today’s study improves visibility into the first. The strongest unanswered question remains the third, which is the one developers ultimately buy.


Primary source, verified: read the paper →

Key questions

Does fivefold value mean Claude finishes five times as much coding work?

No: SemiAnalysis converts measured token allowances into API-equivalent dollars under its own workload mix. It explicitly lacks reliable completed-work token-efficiency data.

How did SemiAnalysis estimate opaque subscription limits?

It used repeated prompts to isolate token categories and compared billed token totals with meter changes. It tracked uncertainty until the allowance range narrowed to approximately plus or minus 5%.

Is the Reddit 56% context figure a share of a subscription bill?

No: the author estimates a share of API-priced transcript spend from their own logs. The underlying calculation and treatment of prompt caching are not publicly audited.
Cite this

APA

Ground Truth. (2026, October 7). SemiAnalysis measures a fivefold subscription-value gap, with a workload-specific denominator. Ground Truth. https://groundtruth.day/news/semianalysis-subscription-value-token-allowances.html

BibTeX

@misc{groundtruth:semianalysis-subscription-value-token-allowances,
  title  = {SemiAnalysis measures a fivefold subscription-value gap, with a workload-specific denominator},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {oct},
  url    = {https://groundtruth.day/news/semianalysis-subscription-value-token-allowances.html}
}

Topics: industry · coding-agents · pricing · subscriptions · evaluation

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.