OpenAI withdraws three math manuscripts as its mass release enters review
OpenAI withdrew three related manuscripts and revised fourteen others, leaving 719 manuscripts and 300 formalized top-line results in its mathematics release.
Claude Haiku 5.5 makes agent work cheaper at entry rates, with a fivefold long-prompt step-up
Anthropic released Haiku 5.5 with adjustable thinking and low entry prices, but prompts above 100,000 tokens enter a fivefold higher price band.
GPT-6 brings interactive answers and progressive responses to ChatGPT
OpenAI began rolling out GPT-6 and Intelligent UI in ChatGPT, using native components and a streaming compiler to turn some answers into interactive tools.
Google Playground combines prompt-built games with a publishing network
Google launched Playground for AI-assisted browser-game creation, adding exportable game files, visibility controls and a social discovery layer.
Liquid AI releases d1 models that choose answers without generating prose
Liquid AI released two open-weight decision models that return typed answers in one pass, targeting routing, inspection and bounded agent decisions.
Kandinsky 6 releases synchronized video and audio, with demanding local tradeoffs
Kandinsky 6 ships MIT-licensed audio-video models and consumer-GPU offload recipes, but the documented low-memory base-model runs remain slow.
Google’s lawyer trial separates better AI-assisted work from better unaided judgment
A randomized patent-lawyer trial found better assisted drafts, while the later unaided advantage was concentrated among experienced attorneys.
ChatGPT Pro’s $200 tier now includes a lower allowance, with an October transition for existing users
OpenAI lowered included usage for new non-grandfathered Pro 200 subscribers while keeping the $200 price, with eligible existing users protected through October 29.
Docker Agent adds public-repository skills and evaluation plumbing in a new release
Docker Agent v1.149.0 adds public-repository skill loading and evaluation updates to its existing declarative agent runtime.
SkillsBench study finds the same agent skill can help one setup and hurt another
A study across 87 tasks found aggregate benefits from agent skills, but 32 tasks showed gains in some configurations and losses in others.
Shopping-agent study finds wealth clues can override a request for the cheapest option
A controlled study found personal AI agents recommending pricier fixed-catalog options to wealthy personas, sometimes despite explicit cheapest-option instructions.
Nathan Lambert argues AI cyber policy must count defensive access as well as misuse
Nathan Lambert’s cyber-policy essay challenges restrictions that weigh open-model misuse without also evaluating defensive access and risks from closed services.