Ground Truth.
AI, checked against the source.

News · 2026-09-22

Grok 4.7 improves agentic coding, but its gains consume more reasoning tokens

xAI has released Grok 4.7 as a stronger model for coding and agentic knowledge work, with the standard API price unchanged from Grok 4.6. The important caveat is that the independently measured improvement comes with substantially more reasoning-token use, so a low list price is not the same as low completed-task cost.

Key facts

xAI says Grok 4.7 uses a “new, larger base model,” a longer reinforcement-learning run on harder multi-hour tasks, better self-verification, and improved long-context management. Parameter count and training data are not disclosed. The public API offers low through xhigh reasoning effort, with high the default. These statements describe xAI's intended mechanism, not independently measured causes of the observed gains.

The pricing message needs translation. Standard Grok 4.7 costs $2 per million input tokens, $0.50 for cached input, and $6 output—matching Grok 4.6's rate card. Input prompts above 200,000 tokens are billed at $4 per million input and $12 per million output, so a 500,000-token context does not carry the headline price all the way through. xAI's pricing documentation is the relevant source, not launch shorthand. The model is available through the API and, according to xAI, through Grok Build, Cursor, OpenRouter, Vercel, and Cloudflare.

The phrase “twice as fast” describes a separate Fast variant, not standard 4.7 against 4.6. xAI says standard 4.7 is served at the same price and speed as 4.6. Fast uses quicker infrastructure at twice the standard token price, and the documentation says it is available only in Cursor and Grok Build. That distinction matters to a developer who expects an API switch to double throughput. It will not.

Artificial Analysis's report gives the most useful external reading. At xhigh it finds an overall Intelligence Index score of 46, versus 44 for Grok 4.6. The bigger movement is on long-horizon work: AA-Briefcase rose 111 Elo, and Grok Build with 4.7 rose nine points on its coding-agent index. DeepSWE increased from 65% to 73% and Terminal-Bench 4.0 from 18% to 33% in AA's components. That is a meaningful result, but it measures a model plus its native harness and chosen effort level. It does not show that the bare model beats every rival in every environment.

The efficiency bill is the story behind the score. AA measured roughly 81,000 output tokens per Intelligence Index task for Grok 4.7 xhigh, compared with about 36,000 for Grok 4.6 high in the report summary. A later comparison gives 38,000 for 4.6 xhigh, an internal inconsistency that does not change the direction: 4.7 spends far more output tokens. It is like paying a talented contractor the same hourly rate but discovering that the job now takes a longer deliberative shift. Higher quality can still be worth it; the economics must be measured at task level.

xAI's own table is mixed. It reports wins over 4.6 in several coding and professional-work tests, but trails Fable 5.1 on CursorBench and Terminal-Bench and trails leading models in HealthBench Professional. AA likewise reports regressions on some non-agentic tests. In the dossier snapshot, Epoch's benchmark hub put GPT-6 Astra first on its broader aggregate capabilities index. xAI's release is therefore a credible agentic-work upgrade, not a universal frontier crown.

Discussion focused on exactly these issues: useful coding performance and value on one side, effort-setting comparisons, time-to-completion, and output-token overhead on the other. The strong counterargument is methodological and right: system-level coding-agent scores combine model behavior, tool use, context construction, and harness implementation. The so-what remains commercially important. xAI is competing on distribution and list price while accepting a more compute-hungry reasoning regime; buyers should compare cost per completed task, not cost per token or one leaderboard cell.

A realistic procurement trial should fix the harness, task suite, maximum effort, latency target, and total-output budget before comparing vendors; otherwise an apparent model win may just be permission to spend more inference.

That discipline also prevents a provider from winning an evaluation simply by silently changing its reasoning budget between versions.


Primary source, verified: read the paper →

Key questions

How much does Grok 4.7 cost through the API?

The standard API price is $2 per million input tokens, $0.50 cached input tokens, and $6 output tokens, with higher pricing above 200,000 input tokens.

Is Grok 4.7 Fast available in the public xAI API?

No: xAI's documentation says the faster variant is restricted to Cursor and Grok Build rather than the public API.

Did Grok 4.7 become the overall best model?

No single result supports that: it improved strongly in a coding-agent measure but had only a two-point overall Intelligence Index gain over Grok 4.6.
Cite this

APA

Ground Truth. (2026, September 22). Grok 4.7 improves agentic coding, but its gains consume more reasoning tokens. Ground Truth. https://groundtruth.day/news/grok-4-7-agentic-coding-price-token-tradeoff.html

BibTeX

@misc{groundtruth:grok-4-7-agentic-coding-price-token-tradeoff,
  title  = {Grok 4.7 improves agentic coding, but its gains consume more reasoning tokens},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {sep},
  url    = {https://groundtruth.day/news/grok-4-7-agentic-coding-price-token-tradeoff.html}
}

Topics: models · agents · coding · api-pricing · benchmarks

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.