inference-cost
GPT-6.1 Sol approaches Astra on an independent index at much lower task cost News
GPT-6.1 Sol scored one point below Astra on Artificial Analysis’s Intelligence Index at about 22% of its estimated task cost.
Fireworks ships Ember-1, an API model built to spend fewer reasoning tokens News
Fireworks launched Ember-1 as a public inference API and says it achieves comparable quality with 35–50% shorter reasoning traces, a vendor-reported claim aimed at reducing the compounding cost of agent work.
RTK claims up to 90% token savings for coding agents. A 1,740-run cost test found about 5%. News
Quesma ran 1,740 coding-agent attempts with and without RTK, a popular open-source tool that compresses terminal output before an AI agent reads it, and found it cut one setup's total bill by about 5% while raising another's by about 5%, even as RTK's own counter reported an 89% reduction.
Google's new Flash model scores higher and costs more to finish a job News
Google released Gemini 3.8 Flash on September 2, 2026, and the per-token price is unchanged, but the model deliberately spends about 30% more output tokens per task, pushing measured cost per task from roughly $0.40 to $0.58.
DeepSeek gave its cheapest model eyes and did not change the price News
DeepSeek shipped an experimental vision version of its V4-Flash model that accepts images by base64, URL, or file upload, and bills it at exactly the same rate as the text-only model.
Anthropic's cheaper model is not cheaper - its cache is News
Claude Fable 5.1 kept the same $10 and $50 per-million sticker price as Fable 5, but cache reads dropped to a quarter of the old rate, which is why one developer's 22,022 API calls got about 31% cheaper per prompt while using 31% more tokens.
A House bill would tax AI tokens and raise the rate when unemployment rises News
H.R. 10044, the AI Tax and Work Protection Act, would place an excise tax on foundation-model usage starting at 2% of token value and escalating automatically as the national unemployment rate climbs above 5%.
Qwen3.8-27B spent 22,276 thinking tokens on one drawing News
Alibaba's new open-weight 27B model ships with its reasoning effort set to the highest level by default, and Simon Willison measured a simple drawing prompt taking 21 minutes instead of two.
Picking the right model per request beat always using the biggest one News
A new routing framework that chooses a different model for each request outperformed the strongest single fixed model by 14.6 percent, partly because the largest model gets many cheap questions wrong.
DeepSeek re-trained V4 Flash without touching the architecture and its coding-agent score went from 7 to 54 News
DeepSeek published new MIT-licensed weights for V4 Flash on July 31 that change only the post-training, lifting the model's score on a real-world software-engineering agent test from 7.3 to 54.4 out of 100.
A decades-old keyword ranker beat the search agent once the document pile passed 10 million tokens News
In a controlled study that grew the same corpus across 28 nested sizes, the agent that browsed files won at small scale but spent 39 times more query tokens, and BM25 - a 1990s keyword ranking formula - overtook it around 10 million tokens and led by nearly 20 points at full scale.
China's GLM-5.2 Ships as the Top Open-Weight Model, Under MIT License News
Z.ai released GLM-5.2, a 753-billion-parameter model, as open weights under an MIT license, and an independent index ranks it the strongest open-weight model available, close behind the leading closed models at a fraction of the price.
OpenRouter Tool
A production gateway to hundreds of models behind one API, with public rankings built from real usage and the ability to sort by price, throughput, latency and popularity.
LLMRouter Tool
A unified framework for building, evaluating and deploying model routers, with a quickstart, single and batch routing calls, and a benchmark that dispatches queries across eighteen candidate models with cost tracking.