Ground Truth.
AI, checked against the source.

← All topics

inference-cost

Everything on Ground Truth tagged “inference-cost” — 14 items.

GPT-6.1 Sol approaches Astra on an independent index at much lower task cost News

GPT-6.1 Sol scored one point below Astra on Artificial Analysis’s Intelligence Index at about 22% of its estimated task cost.

Fireworks ships Ember-1, an API model built to spend fewer reasoning tokens News

Fireworks launched Ember-1 as a public inference API and says it achieves comparable quality with 35–50% shorter reasoning traces, a vendor-reported claim aimed at reducing the compounding cost of agent work.

RTK claims up to 90% token savings for coding agents. A 1,740-run cost test found about 5%. News

Quesma ran 1,740 coding-agent attempts with and without RTK, a popular open-source tool that compresses terminal output before an AI agent reads it, and found it cut one setup's total bill by about 5% while raising another's by about 5%, even as RTK's own counter reported an 89% reduction.

Google's new Flash model scores higher and costs more to finish a job News

Google released Gemini 3.8 Flash on September 2, 2026, and the per-token price is unchanged, but the model deliberately spends about 30% more output tokens per task, pushing measured cost per task from roughly $0.40 to $0.58.

DeepSeek gave its cheapest model eyes and did not change the price News

DeepSeek shipped an experimental vision version of its V4-Flash model that accepts images by base64, URL, or file upload, and bills it at exactly the same rate as the text-only model.

Anthropic's cheaper model is not cheaper - its cache is News

Claude Fable 5.1 kept the same $10 and $50 per-million sticker price as Fable 5, but cache reads dropped to a quarter of the old rate, which is why one developer's 22,022 API calls got about 31% cheaper per prompt while using 31% more tokens.

A House bill would tax AI tokens and raise the rate when unemployment rises News

H.R. 10044, the AI Tax and Work Protection Act, would place an excise tax on foundation-model usage starting at 2% of token value and escalating automatically as the national unemployment rate climbs above 5%.

Qwen3.8-27B spent 22,276 thinking tokens on one drawing News

Alibaba's new open-weight 27B model ships with its reasoning effort set to the highest level by default, and Simon Willison measured a simple drawing prompt taking 21 minutes instead of two.

Picking the right model per request beat always using the biggest one News

A new routing framework that chooses a different model for each request outperformed the strongest single fixed model by 14.6 percent, partly because the largest model gets many cheap questions wrong.

DeepSeek re-trained V4 Flash without touching the architecture and its coding-agent score went from 7 to 54 News

DeepSeek published new MIT-licensed weights for V4 Flash on July 31 that change only the post-training, lifting the model's score on a real-world software-engineering agent test from 7.3 to 54.4 out of 100.

A decades-old keyword ranker beat the search agent once the document pile passed 10 million tokens News

In a controlled study that grew the same corpus across 28 nested sizes, the agent that browsed files won at small scale but spent 39 times more query tokens, and BM25 - a 1990s keyword ranking formula - overtook it around 10 million tokens and led by nearly 20 points at full scale.

China's GLM-5.2 Ships as the Top Open-Weight Model, Under MIT License News

Z.ai released GLM-5.2, a 753-billion-parameter model, as open weights under an MIT license, and an independent index ranks it the strongest open-weight model available, close behind the leading closed models at a fraction of the price.

OpenRouter Tool

A production gateway to hundreds of models behind one API, with public rankings built from real usage and the ability to sort by price, throughput, latency and popularity.

LLMRouter Tool

A unified framework for building, evaluating and deploying model routers, with a quickstart, single and batch routing calls, and a benchmark that dispatches queries across eighteen candidate models with cost tracking.