pricing
GPT-6 Astra improves computer use sharply, but OpenAI reports a monitoring trade-off News
OpenAI's GPT-6 Astra posts its clearest gains in computer use and coding-agent tasks while costing 2.5 times GPT-5.6 Sol per token, and its system card says chain-of-thought-only monitoring is weaker even as prompt-injection robustness improves.
OpenAI shipped GPT-6 Astra, and its headline benchmark score has two different answers News
OpenAI began a staged rollout of GPT-6 Astra on September 3, 2026 at $10 per million input tokens and $50 per million output, and ARC Prize's own results page shows the model scoring 62.71% on ARC-AGI-3 under one test harness and 99.95% under another.
Anthropic's cheaper model is not cheaper - its cache is News
Claude Fable 5.1 kept the same $10 and $50 per-million sticker price as Fable 5, but cache reads dropped to a quarter of the old rate, which is why one developer's 22,022 API calls got about 31% cheaper per prompt while using 31% more tokens.
Anthropic shipped one model under two names and two safety settings News
Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1, 2026 -- the same underlying model shipped twice, with the only difference being how tightly its cybersecurity and biology safeguards are wound.
OpenAI cut Sol's price, and OpenRouter cut it again News
OpenAI dropped GPT-5.6 Sol to $4 per million input tokens and $20 per million output on August 21, 2026, a 33 percent cut on output, and OpenRouter is separately listing the same model from OpenAI at half that.
DeepSeek is selling a checkpoint it has not published News
DeepSeek's API now serves a model version named DeepSeek-V4-Pro-0813 and at least five commercial hosts resell it by that exact name, but the company has not published a matching dated weights page, and none of the resellers undercuts DeepSeek's own price.
DeepSeek starts charging rush-hour prices on August 17 News
DeepSeek is replacing flat API pricing with peak and off-peak rates on August 17, and the steepest change hits cached input on its Pro model, which goes up twelvefold during Beijing business hours.
xAI shipped Grok 4.6 into Cursor at two dollars a million input tokens News
xAI released Grok 4.6 on August 12, claiming a score of 61 on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol Max, and priced it at two dollars per million input tokens.
DeepSeek warns of a significant API price rise, five days after being called 100 times cheaper News
DeepSeek added a footnote to its official pricing page warning that it plans to raise API prices significantly in the near future with no figure and no date attached, five days after an independent benchmark study priced its model at roughly 100 times less per task than Western frontier models.
Qwen3.8-Max Shipped as a Paid API, Not as Open Weights News
Alibaba put Qwen3.8-Max live as a hosted API at $2 per million input tokens and $6 per million output tokens, a fifth cheaper than the model it replaces, while the open weights it promised for Max and a 27B sibling have not shipped.
OpenAI cut its cheapest model's price 80%, and credits one of its own models for making it possible News
OpenAI dropped GPT-5.6 Luna's API price by 80% and Terra's by 20% effective July 30, and says its Sol model autonomously rewrote production kernels that cut the cost of serving the model by 20%.
Claude Opus 5 posts a verified four-fold lead on the hardest adaptation benchmark News
Anthropic released Claude Opus 5 on July 24, and the independent benchmark owner ARC Prize verified it at 30.16% on ARC-AGI-3, roughly four times the previous best published result, while the model's API price stayed identical to Opus 4.8.
OpenAI temporarily scraps the 5-hour usage limit and picks a fight with Anthropic News
OpenAI temporarily removed the 5-hour usage-limit restriction for all Plus, Business, and Pro plans, reset usage, and said it hit 6 million active users -- a competitive move users read as aimed squarely at Anthropic.
SpaceXAI ships Grok 4.5, trained on trillions of Cursor coding sessions News
SpaceXAI released Grok 4.5 on July 8, its first model as a public SpaceX subsidiary, trained on trillions of Cursor developer-interaction tokens and priced aggressively at $2 per million input tokens, though that rate only holds below 200K context.
Grok 4.5 arrives claiming Opus-class quality at a third the price News
SpaceXAI released Grok 4.5, a 1.5-trillion-parameter model priced at $2 per million input tokens and $6 per million output, undercutting frontier rivals roughly threefold while claiming comparable coding quality.
Anthropic switches Fable 5 to usage billing and turns on government ID checks News
Starting today, Anthropic bills its flagship Fable 5 model by usage at $10 per million input tokens and $50 per million output across all tiers, and its government-ID verification requirement for Fable 5 access takes effect as part of an export-control redeployment.
A startup router is giving away 100 million tokens of Kimi, MiniMax and GLM News
API aggregator Dahl Inference is handing out 100 million free tokens across top open-weight Chinese models like Kimi K2.6 and MiniMax M2.7 - not a price cut from the labs themselves, but a router burning money to win users amid a glut of cheap compute.
Claude Sonnet 5 is cheaper per word but can cost more per finished job News
Anthropic's new mid-tier model is close to its flagship on hard agent work, yet independent testing shows it can spend more per completed task because it takes more steps.
Amazon and Anthropic's partnership is cracking over the price of Claude News
A renegotiated contract is expected to sharply raise Amazon's bill for Anthropic's AI, pushing Amazon toward OpenAI and its own models even though it's an Anthropic investor.
Frontier AI is getting more expensive while open models keep getting cheaper News
Closed frontier models are raising prices and tightening access just as Chinese open-weight models slash theirs, a structural reversal with big consequences for who builds with AI.
Are closed AI models overpriced luxury goods? News
An essay argues open-weight models now undercut the big closed AIs by huge margins, and that 'China fears' are being used to protect those prices.
OpenRouter discounted models Tool
A live collection of models currently carrying provider discounts on OpenRouter. GPT-5.6 Sol from the OpenAI provider is listed at roughly half OpenAI's own promotional rate, against $5 and $30 for the same model via Azure.
Gemini Developer API pricing page Tool
Google's own current rate card, including the context-caching rates that determine whether a long-running agent is cheap or ruinous. Worth reading before assuming a headline per-token price describes what you will actually pay, since caching rates and introductory-period expiry dates do most of the work.
Gemini 3.6 Flash and 3.5 Flash-Lite Tool
Google's economy-tier models went generally available on July 21, with 3.6 Flash keeping a million-token context and 64,000-token output while dropping its output price roughly a sixth versus 3.5 Flash and using about 17 percent fewer output tokens per task. Note the migration-breaking changes: some sampling parameters are deprecated and prefilled model turns are no longer supported.
Dahl Inference Tool
Third-party inference router reselling top open-weight models (Kimi K2.6, MiniMax M2.7, GLM 5.2) at low per-token prices, currently running a 100M-free-token promotion.
Artificial Analysis Intelligence Index Tool
The independent benchmark and pricing dashboard the field now reaches for when a lab claims a lead -- it is the source of Inkling's debut score of 41. Useful beyond the headline ranking because it also tracks output tokens per task, latency and cost, which is how you find out that a cheaper-looking model is actually more expensive per finished job.
Artificial Analysis Tool
A free public dashboard that independently benchmarks and compares AI models on a combined intelligence index alongside price and speed; the source of this week's finding that GLM-5.2 leads the open-weight class.
Anthropic prompt caching pricing reference Tool
Anthropic's documentation of how cached tokens are billed: cache hits at 10% of standard input, five-minute writes at 1.25 times base input, one-hour writes at 2 times. This is the page that explains why Claude Fable 5.1 can be substantially cheaper per prompt for long agent sessions while costing exactly the same for one-shot calls.