Ground Truth.
AI, checked against the source.

News · 2026-10-08

Claude Haiku 5.5 makes agent work cheaper at entry rates, with a fivefold long-prompt step-up

Anthropic released Claude Haiku 5.5 on October 7 with adjustable thinking, a one-million-token context window and entry prices of $0.10 per million input tokens and $0.50 per million output tokens. Prompts above 100,000 tokens cost five times those rates. The release makes repeated agent work more accessible, while making context size and effort central to the real bill.

Key facts

The useful way to understand Haiku is as a worker in a larger software process. An application might ask it to extract one fact, route a support request, summarize a document or operate a browser for a bounded task. A more capable lead model can reserve its attention for the final judgment. Anthropic’s launch examples emphasize that division of labor, while continuing to recommend Sonnet or Opus for complex end-to-end coding.

The model documentation says adaptive thinking is enabled by default and medium effort is the default setting. Text and images go in; text comes out. The model can hold far more context than the lowest price band covers. Capacity and price are separate promises: a large notebook may have room for a million words, while the photocopier starts charging a different rate after its first hundred thousand.

That distinction is especially important in agents. A short next action can arrive inside a large prompt containing tools, previous messages and retrieved material. Once the prompt crosses the threshold, a seemingly small request enters the expensive band. Anthropic’s pricing page also lists much cheaper cache reads and a batch discount, so repeated content and execution style change the result further. The relevant comparison is cost per correctly completed job, including repair and retries.

There is a second denominator change. The new tokenizer represents identical text with approximately 30% more tokens than Haiku 4.5, according to Anthropic’s model documentation. The launch footnote calls the increase “slightly more.” Anthropic’s reported average running-cost reduction accounts for that change, but no average guarantees savings for a specific workload. A team should measure its own text, effort settings and distribution of prompt lengths before projecting a budget.

Plotly’s analytics experiment provides an unusually concrete early check. It held the agent’s tools and prompts fixed and asked 43 questions about a synthetic, messy wind-farm warehouse. Haiku 5.5 answered 37 correctly, compared with 27 for its predecessor. The full run cost $0.38, with a median of 14 seconds per question. Luna scored 38 and cost $0.13. This is evidence of an improvement over the old Haiku in that harness, alongside a clear counterexample to assuming the new model wins every economic comparison.

Independent evaluators also warn against a universal ranking. Artificial Analysis reported that maximum-effort Haiku used roughly three times Luna’s output tokens on its selected tasks. Its initial cost accounting did not yet include the long-prompt price tier. Vals AI supplies another task mix, provider and effort setting. These are different instruments, rather than interchangeable measurements of one abstract intelligence score.

Effort changes reliability as well as cost. In its prompting guide, Anthropic says low effort can skip searches, verification or later steps. Moving to medium roughly halved early stopping in its tests while more than doubling output tokens per attempt. It also documents occasional empty visible answers at very high effort and specific refusal handling. Those operational details can erase savings if a harness retries without understanding the failure.

The launch discussion shows enthusiasm for fast delegated work and dissent from users whose established prompts perform worse or consume more tokens. A widely shared pagoda comparison calculated hypothetical API bills from subscription use; it did not report those amounts as money paid. That distinction matters when community examples become pricing claims. The lessons on prompt caching and model routing explain the architecture behind more disciplined comparisons.

Haiku 5.5 is a substantial shipping release, with a plausible role in inexpensive, narrow agent calls. The honest caveat is that vendor benchmarks, customer testimonials and early evaluator suites do not establish a universal cost-per-success winner. Its cheapest rate is real; how often a working agent can use that rate successfully is the deployment question.


Primary source, verified: read the paper →

Key questions

What happens to Haiku 5.5 pricing above 100,000 prompt tokens?

The input and output rates rise fivefold above that threshold: to $0.50 and $2.50 per million tokens. Its one-million-token context capacity does not keep every request in the cheapest band.

Did Haiku 5.5 beat Luna in Plotly’s analytics test?

No: Haiku scored 37 of 43 while Luna scored 38 of 43 in Plotly’s reported experiment. Haiku’s run cost $0.38 against Luna’s $0.13, so this test did not make Haiku the cost-per-task winner.

Can developers tune how much Haiku 5.5 thinks?

Yes: Haiku 5.5 has adjustable effort and adaptive thinking, with medium effort the documented default. Lower effort saves time and output tokens but can increase premature stopping.
Cite this

APA

Ground Truth. (2026, October 8). Claude Haiku 5.5 makes agent work cheaper at entry rates, with a fivefold long-prompt step-up. Ground Truth. https://groundtruth.day/news/claude-haiku-5-5-price-cliff-agent-economics.html

BibTeX

@misc{groundtruth:claude-haiku-5-5-price-cliff-agent-economics,
  title  = {Claude Haiku 5.5 makes agent work cheaper at entry rates, with a fivefold long-prompt step-up},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {oct},
  url    = {https://groundtruth.day/news/claude-haiku-5-5-price-cliff-agent-economics.html}
}

Topics: models · anthropic · agents · pricing · inference

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.