Xiaomi releases MiMo-V2.6, a trillion-parameter open agent family with a 9B distill
Xiaomi released two one-million-context multimodal sparse agent models and a 9B Qwen distill, but its efficient Flash model still requires a roughly 178 GB weight download.
Grok 4.7 improves agentic coding, but its gains consume more reasoning tokens
xAI's Grok 4.7 offers a verified coding-agent improvement at the same standard API list price as 4.6, while independent testing finds it spends far more output tokens per task.
Linear rebuilt its CI after agent-assisted coding made validation the bottleneck
Linear says its test suite nearly quadrupled as agents accelerated code production, and it cut critical-path overhead enough to hold pull-request waits just above five minutes.
Kev brings Jev-like local decision models to Qwen, while audits challenge universal calibration claims
Kev released open local models for typed probability decisions, as published evaluations and an out-of-distribution audit show that confidence calibration varies sharply by task.
Alibaba's verified Apsara news is a Qwen4 architecture preview and the Zhenwu M890 chip
Alibaba has not publicly released Qwen4, but its Qwen3.8-Flash-Next repository previews the planned architecture and its T-Head unit has announced the 144 GB Zhenwu M890 accelerator.
OpenAI says an internal model resolved 100-plus math problems and asks an independent group to advise on disclosure
OpenAI has claimed more than 100 long-standing mathematical results from a new internal model, but has published neither an itemized list nor the proofs for that batch.
Meta's Muse security design treats credentials, browser control, and network egress as separate agent boundaries
Meta says Muse isolates credentials and mediates browser and network actions through separate services, while explicitly acknowledging that prompt injection remains an unsolved agent-security problem.
Apple's M5 Ultra makes large local AI practical by capacity, not by beating an RTX 5090
A 256 GB M5 Ultra Mac Studio ran large sparse models and long contexts locally in a detailed review, but an RTX 5090 remained faster whenever the tested workload fit in its VRAM.