GPT-6.1 Sol approaches Astra on an independent index at much lower task cost
GPT-6.1 Sol scored one point below Astra on Artificial Analysis’s Intelligence Index at about 22% of its estimated task cost.
OpenAI launches persistent dots, but pausing one does not stop its delegates
OpenAI’s new dots run persistent tasks on cloud computers, while delegated tasks and scheduled runs require separate stop controls.
ChatGPT’s $200 Pro plan returns with a lower allowance
OpenAI reopened Pro 200 at $200 a month with less included usage and an October 29 transition for eligible existing subscribers.
Anthropic reports GLM-5.3 built exploits near Mythos’s rate on one test
Anthropic reports that GLM-5.3 built 50 working exploits in 410 benchmark attempts, close to Mythos Preview’s 56, under controlled conditions.
LiveNerf begins measuring Claude drift, with no degradation verdict yet
LiveNerf is collecting a fixed-task Claude Code baseline and has not established that Opus 5.5 became worse after the September 29 outage.
Bain says a $6 trillion AI market would be needed to support its 2031 buildout scenario
Bain derives a $6 trillion annual AI revenue threshold from projected $1.5 trillion annual infrastructure spending and a 25% spending-to-revenue assumption.
Menlo estimates consumer AI reached $40 billion as existing users spend more
Menlo Ventures estimates global consumer-AI spending rose to $40 billion in 2026, with payment concentrated among a relatively small high-spending cohort.
The Collatz artifact passed two buggy checkers; the conjecture remains open
A September interview revisits a July AI-assisted Collatz artifact that exploited distinct bugs in Lean and an older independent checker.
Raven releases a framework that makes agent handoffs and harness changes inspectable
Raven’s paper and code organize specialist agents through typed task graphs and separately evaluate controlled changes to their surrounding software.
Environment Steering tests runtime data-flow checks against agent attacks
Environment Steering reports safer agent behavior by checking whether data may flow to tool inputs and final answers under task-specific policies.