Ground Truth.
AI, checked against the source.

← All topics

frontier-models

Everything on Ground Truth tagged “frontier-models” — 16 items.

OpenAI shipped GPT-6 Astra, and its headline benchmark score has two different answers News

OpenAI began a staged rollout of GPT-6 Astra on September 3, 2026 at $10 per million input tokens and $50 per million output, and ARC Prize's own results page shows the model scoring 62.71% on ARC-AGI-3 under one test harness and 99.95% under another.

OpenAI formally designates Astra as its first Critical cyber-capability model News

OpenAI announced on September 1, 2026 that its Astra model meets the Critical cybersecurity threshold under its Preparedness Framework -- the first model the company has ever placed at that level -- after experts used it to find unknown browser and operating-system vulnerabilities and chain two zero-days into a working exploit.

Capability Thresholds and Responsible Scaling Policies Lesson

Capability thresholds are pre-committed lines that AI labs draw in advance -- specific dangerous abilities that, once a model demonstrates them, trigger specific mandatory safeguards. They are the industry's main attempt to make safety decisions before the incentive to fudge them arrives.

Anthropic still will not ship the model that found ten thousand vulnerabilities News

Anthropic says roughly 50 partners used its restricted Claude Mythos Preview model to find more than ten thousand high- or critical-severity software vulnerabilities, and the company still will not release Mythos-class models to the public because its safeguards are not good enough yet.

The White House's Open-Weight Carve-Out Is a Private Briefing, Not a Published Rule News

Reporting says the White House finished an AI framework that covers only closed frontier models and will not publish it, but the only public legal instrument is June's Executive Order 14409, which contains no definition of open-weight, no US-origin condition and no mandatory testing regime to be exempt from.

No, the White House Isn't Licensing AI Models - Here's What EO 14409 Actually Sets Up News

Executive Order 14409 creates a voluntary US government pre-release access framework for "covered frontier models," thresholded by a classified NSA-led cyber-capability benchmark, and explicitly does not authorize mandatory licensing, preclearance, or permitting of AI model releases.

Grok 4.5 arrives claiming Opus-class quality at a third the price News

SpaceXAI released Grok 4.5, a 1.5-trillion-parameter model priced at $2 per million input tokens and $6 per million output, undercutting frontier rivals roughly threefold while claiming comparable coding quality.

OpenAI previews GPT-5.6 -- and shows it to the government first News

OpenAI previewed a three-model GPT-5.6 family on June 26 and released it only to a small set of vetted partners after briefing the U.S. government, making pre-launch government coordination a routine step for a frontier model.

Five Eyes spy chiefs: the AI cyber threat is months away, not years News

On June 23 the Five Eyes cyber agencies jointly warned that frontier AI will transform cyberattacks on a timeline of months rather than years, and urged organizations to fix foundational security now.

The government cleared one Anthropic model and kept the other locked up News

Washington partially reopened access to Anthropic's Mythos 5 for about a hundred organizations, but its more powerful sibling Fable 5 stays blocked - and Anthropic is still suing.

OpenAI launches GPT-5.6, but only to companies the government clears first News

OpenAI's most capable models yet shipped today as a tiny, government-vetted preview, signaling that Washington now holds a gate in front of the frontier.

Google promised Gemini 3.5 Pro in June. June is almost over. News

Google said its next flagship would arrive in June; with days left it's still limited preview. The timing is awkward -- it overlaps a gap where another Western flagship is also unavailable.

A senator says a banned AI broke into nearly all NSA systems in hours News

New testimony reframes the Mythos export ban: a top general reportedly told a senator the model breached almost all classified systems in a red-team test, not in weeks but in hours.

Recursive self-improvement: when AI starts building AI Lesson

The idea that an AI good enough at AI research could improve itself, and the improved version could improve itself again, faster each round. Here's what it actually means, why a major lab now says we're getting close, and why "close" is not the same as "here."

Anthropic Wants a Pause Button the Whole World Can Check News

Buried in Anthropic's essay is a concrete proposal: not to stop AI, but to build the machinery that would let rival labs prove to each other they had stopped.

GPT-6 Astra API Tool

OpenAI's agentic flagship, aimed at computer use, browsing, coding and long multi-step workflows. Five reasoning effort levels from low to max, with no off switch. $10 per million input tokens and $50 per million output, cached input at $1 -- cache discipline is the difference between an affordable agent loop and an unaffordable one.