Ground Truth.
AI, checked against the source.

News · 2026-09-30

GPT-6.1 Sol approaches Astra on an independent index at much lower task cost

OpenAI released GPT-6.1 Sol on September 29, and an independent evaluation places it close to the company’s more expensive Astra model at substantially lower task cost. Artificial Analysis scored Sol 6.1 at 52 against Astra’s 53 on its Intelligence Index, estimating $0.72 versus $3.26 per weighted task at maximum reasoning effort. The result strengthens the case for cheaper capable agents, while leaving real-work reliability and subscription value unresolved.

Key facts

The significant part of a cheaper model is the work it can finish. A low price means little if the assistant needs repeated attempts, overlooks a document, or leaves a human to repair its output. Sol’s launch matters because Artificial Analysis’s evaluation offers an independent check on that tradeoff. Its headline is “near-Astra intelligence,” a useful description of the measured composite rather than a promise that the models behave identically.

The index combines different kinds of work. A one-point gap across that mixture does not mean a one-point gap on every coding, science, or document task. At a different effort setting, Sol’s coding-agent result was stronger: Artificial Analysis reports that xhigh effort scored one point above Astra for less than 15% of Astra’s task cost. Maximum effort did worse in that particular comparison. More thinking is therefore a setting to evaluate, rather than a universal upgrade button.

A concrete analogy is choosing delivery services. The per-mile price is the token rate; the total bill also depends on route length, failed deliveries, and repeat trips. Artificial Analysis’s release page and methodology estimate the token bill for its weighted test mixture. They do not measure the full cost of a customer’s workflow, including human checking, external tools, or the value of elapsed time. That distinction is central to inference economics.

OpenAI’s own tests cover long software projects, professional questions drawn from complex documents, business workflows, computer use, science, and difficult factual questions. The company reports matching Astra on its long-horizon software-engineering comparison at roughly one-fifth the cost. Its science result is less sweeping: Astra still scored highest, and OpenAI recommends it for the hardest scientific work. These are vendor-run comparisons with specific harnesses and effort settings, not independently replicated general rankings.

The official model documentation lists a 1,050,000-token context window and up to 128,000 output tokens. Cached input costs $0.10 per million tokens; cache writes cost $2.50. Requests above 272,000 input tokens incur higher rates across the full request. Those conditions make prompt caching, context management, and the requested reasoning effort material parts of the bill. The model is provided as a hosted service; this announcement does not offer a checkpoint to download.

Efficiency also does not mean less generation. Artificial Analysis reports roughly 10–30% more output tokens than the preceding Sol across effort settings. Better measured performance can still make the extra output economical, but an application that prizes short answers or low latency should test those qualities separately. Anthropic’s Sonnet 5.5 announcement lists the same standard input and output prices. That price equality alone cannot settle which assistant finishes comparable jobs more cheaply.

The community discussion of Artificial Analysis’s chart reveals the divide. Developers paying by token welcome the cost improvement. Subscription users compare plan allowances, integration, speed, and their own unsuccessful attempts. Those anecdotes are useful workload clues, but they do not overturn a benchmark or establish a population-wide failure rate. A separate apparently pro-Sol benchmark post was satire: its graph compared version numbers, not measured intelligence.

There is a security cost to increased capability too. OpenAI’s system-card addendum classifies Sol as Critical in cybersecurity and says it applies Astra’s safeguards. That is a capability classification and the company’s clearance rationale, not an independent guarantee of safe deployment. The practical conclusion is narrower than a winner-takes-all headline: Sol offers a compelling new candidate for workload-specific evaluation, and its strongest verified advantage is benchmark task economics.

A useful deployment comparison records not only the model name but also the effort setting, accepted output, and cost of every unsuccessful attempt. Those records let a team test whether the suite-level advantage survives the particular documents, repositories, and approval steps its work actually requires.


Primary source, verified: read the paper →

Key questions

How much cheaper is GPT-6.1 Sol than Astra?

Its standard token prices are one-fifth of Astra’s; Artificial Analysis estimated $0.72 versus $3.26 per weighted index task at maximum effort. Those are benchmark-suite costs, not subscription prices.

Does Sol 6.1 beat every competing model?

No: Artificial Analysis placed it one point below Astra on its Intelligence Index and said Opus 5.5 still led overall. Results also change with the task and reasoning effort.

Where was Sol 6.1 available at launch?

OpenAI announced API, ChatGPT Work, and Codex availability on September 29; the launch did not include Chat.
Cite this

APA

Ground Truth. (2026, September 30). GPT-6.1 Sol approaches Astra on an independent index at much lower task cost. Ground Truth. https://groundtruth.day/news/gpt-6-1-sol-moves-the-cost-frontier.html

BibTeX

@misc{groundtruth:gpt-6-1-sol-moves-the-cost-frontier,
  title  = {GPT-6.1 Sol approaches Astra on an independent index at much lower task cost},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {sep},
  url    = {https://groundtruth.day/news/gpt-6-1-sol-moves-the-cost-frontier.html}
}

Topics: model-release · openai · inference-cost · coding · benchmarks

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.