Ground Truth.
AI, checked against the source.
← 2026-09-292026-09-30later →

GPT-6.1 Sol approaches Astra on an independent index at much lower task cost

2026-09-30

GPT-6.1 Sol scored one point below Astra on Artificial Analysis’s Intelligence Index at about 22% of its estimated task cost.

model-release · openai · inference-cost · coding · benchmarks

OpenAI launches persistent dots, but pausing one does not stop its delegates

2026-09-30

OpenAI’s new dots run persistent tasks on cloud computers, while delegated tasks and scheduled runs require separate stop controls.

agents · openai · product-release · permissions · agent-memory

ChatGPT’s $200 Pro plan returns with a lower allowance

2026-09-30

OpenAI reopened Pro 200 at $200 a month with less included usage and an October 29 transition for eligible existing subscribers.

openai · pricing · subscriptions · industry

Anthropic reports GLM-5.3 built exploits near Mythos’s rate on one test

2026-09-30

Anthropic reports that GLM-5.3 built 50 working exploits in 410 benchmark attempts, close to Mythos Preview’s 56, under controlled conditions.

cybersecurity · ai-security · vulnerabilities · red-teaming · open-weights

LiveNerf begins measuring Claude drift, with no degradation verdict yet

2026-09-30

LiveNerf is collecting a fixed-task Claude Code baseline and has not established that Opus 5.5 became worse after the September 29 outage.

evaluations · reliability · anthropic · open-source · model-drift

Bain says a $6 trillion AI market would be needed to support its 2031 buildout scenario

2026-09-30

Bain derives a $6 trillion annual AI revenue threshold from projected $1.5 trillion annual infrastructure spending and a 25% spending-to-revenue assumption.

industry · infrastructure · economics · market-research

Menlo estimates consumer AI reached $40 billion as existing users spend more

2026-09-30

Menlo Ventures estimates global consumer-AI spending rose to $40 billion in 2026, with payment concentrated among a relatively small high-spending cohort.

consumer-ai · industry · subscriptions · market-research

The Collatz artifact passed two buggy checkers; the conjecture remains open

2026-09-30

A September interview revisits a July AI-assisted Collatz artifact that exploited distinct bugs in Lean and an older independent checker.

formal-verification · mathematics · ai-research · reliability

Raven releases a framework that makes agent handoffs and harness changes inspectable

2026-09-30

Raven’s paper and code organize specialist agents through typed task graphs and separately evaluate controlled changes to their surrounding software.

agents · harnesses · open-source · research · evaluation

Environment Steering tests runtime data-flow checks against agent attacks

2026-09-30

Environment Steering reports safer agent behavior by checking whether data may flow to tool inputs and final answers under task-specific policies.

cybersecurity · ai-security · prompt-injection · agents · data-flow-control · research

← 2026-09-292026-09-30later →