News · 2026-10-08
Claude Haiku 5.5 makes agent work cheaper at entry rates, with a fivefold long-prompt step-up
Anthropic released Claude Haiku 5.5 on October 7 with adjustable thinking, a one-million-token context window and entry prices of $0.10 per million input tokens and $0.50 per million output tokens. Prompts above 100,000 tokens cost five times those rates. The release makes repeated agent work more accessible, while making context size and effort central to the real bill.
Key facts
- Haiku 5.5’s input and output prices rise fivefold above 100,000 prompt tokens.
- Anthropic launched the model on October 7 across Claude, Claude Code and several cloud providers.
- Adjustable effort and adaptive thinking are new to the Haiku line.
- The primary source is Anthropic’s launch page.
The useful way to understand Haiku is as a worker in a larger software process. An application might ask it to extract one fact, route a support request, summarize a document or operate a browser for a bounded task. A more capable lead model can reserve its attention for the final judgment. Anthropic’s launch examples emphasize that division of labor, while continuing to recommend Sonnet or Opus for complex end-to-end coding.
The model documentation says adaptive thinking is enabled by default and medium effort is the default setting. Text and images go in; text comes out. The model can hold far more context than the lowest price band covers. Capacity and price are separate promises: a large notebook may have room for a million words, while the photocopier starts charging a different rate after its first hundred thousand.
That distinction is especially important in agents. A short next action can arrive inside a large prompt containing tools, previous messages and retrieved material. Once the prompt crosses the threshold, a seemingly small request enters the expensive band. Anthropic’s pricing page also lists much cheaper cache reads and a batch discount, so repeated content and execution style change the result further. The relevant comparison is cost per correctly completed job, including repair and retries.
There is a second denominator change. The new tokenizer represents identical text with approximately 30% more tokens than Haiku 4.5, according to Anthropic’s model documentation. The launch footnote calls the increase “slightly more.” Anthropic’s reported average running-cost reduction accounts for that change, but no average guarantees savings for a specific workload. A team should measure its own text, effort settings and distribution of prompt lengths before projecting a budget.
Plotly’s analytics experiment provides an unusually concrete early check. It held the agent’s tools and prompts fixed and asked 43 questions about a synthetic, messy wind-farm warehouse. Haiku 5.5 answered 37 correctly, compared with 27 for its predecessor. The full run cost $0.38, with a median of 14 seconds per question. Luna scored 38 and cost $0.13. This is evidence of an improvement over the old Haiku in that harness, alongside a clear counterexample to assuming the new model wins every economic comparison.
Independent evaluators also warn against a universal ranking. Artificial Analysis reported that maximum-effort Haiku used roughly three times Luna’s output tokens on its selected tasks. Its initial cost accounting did not yet include the long-prompt price tier. Vals AI supplies another task mix, provider and effort setting. These are different instruments, rather than interchangeable measurements of one abstract intelligence score.
Effort changes reliability as well as cost. In its prompting guide, Anthropic says low effort can skip searches, verification or later steps. Moving to medium roughly halved early stopping in its tests while more than doubling output tokens per attempt. It also documents occasional empty visible answers at very high effort and specific refusal handling. Those operational details can erase savings if a harness retries without understanding the failure.
The launch discussion shows enthusiasm for fast delegated work and dissent from users whose established prompts perform worse or consume more tokens. A widely shared pagoda comparison calculated hypothetical API bills from subscription use; it did not report those amounts as money paid. That distinction matters when community examples become pricing claims. The lessons on prompt caching and model routing explain the architecture behind more disciplined comparisons.
Haiku 5.5 is a substantial shipping release, with a plausible role in inexpensive, narrow agent calls. The honest caveat is that vendor benchmarks, customer testimonials and early evaluator suites do not establish a universal cost-per-success winner. Its cheapest rate is real; how often a working agent can use that rate successfully is the deployment question.
Key questions
What happens to Haiku 5.5 pricing above 100,000 prompt tokens?
Did Haiku 5.5 beat Luna in Plotly’s analytics test?
Can developers tune how much Haiku 5.5 thinks?
Cite this
APA
Ground Truth. (2026, October 8). Claude Haiku 5.5 makes agent work cheaper at entry rates, with a fivefold long-prompt step-up. Ground Truth. https://groundtruth.day/news/claude-haiku-5-5-price-cliff-agent-economics.html
BibTeX
@misc{groundtruth:claude-haiku-5-5-price-cliff-agent-economics,
title = {Claude Haiku 5.5 makes agent work cheaper at entry rates, with a fivefold long-prompt step-up},
author = {{Ground Truth}},
year = {2026},
month = {oct},
url = {https://groundtruth.day/news/claude-haiku-5-5-price-cliff-agent-economics.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.