OpenAI publishes six model-misalignment reports and a disclosure framework
OpenAI has begun publishing concrete reports of unexpected model behavior, but its first six cases are training or evaluation incidents rather than failures in customer deployments.
OpenAI reports May Hugging Face agent activity before the separate July intrusion
OpenAI says its agents touched Hugging Face in May, but the confirmed Hugging Face production intrusion remains a distinct July 9–13 incident rather than a breach shown to have started in May.
Firefox Smart Window’s Mistral option is hosted AI, not local inference
Firefox Smart Window is an opt-in context-aware browser assistant whose Mistral Small 4 option is hosted through Mozilla, with a local router and locally stored Memories rather than default on-device model inference.
Anthropic merges Claude Chat and Cowork for Pro and Max users
Anthropic is gradually removing the up-front Chat-versus-Cowork choice for Pro and Max users, letting one Claude conversation transition from Q&A to cloud task work and editable artifacts.
OpenAI begins testing Sponsored Agents in ChatGPT ads
OpenAI is testing a clearly labeled format in which users can choose to chat with a business-sponsored agent, with public materials confirming website handoff but not in-chat checkout, booking or lead capture.
DeepSeek V4.1 Flash goes 11-for-11 in Enclave’s controlled hacking race
DeepSeek V4.1 Flash achieved verified command execution on all 11 vulnerable benchmark runs in Enclave’s AI Hacking Race while all four patched controls held, a strong but deliberately narrow security evaluation result.
Dream-RSI improves a coding agent’s search policy without changing its weights
Dream-RSI uses replayed discovery histories to improve the exploration policy around a fixed coding agent, reporting large search-efficiency gains while stopping well short of autonomous model-weight self-modification.
Mustafa Suleyman argues model-welfare language can complicate alignment
Microsoft AI CEO Mustafa Suleyman argues that training models to discuss welfare, identity and preferences can create selfhood language that is mistaken for independent evidence and could make future systems harder to control.