Z.ai changed only the post-training, and the model learned to find exploits
Z.ai released GLM-5.3 on August 14 using the same base model as GLM-5.2, with every gain coming from post-training, and the largest jump was in finding and exploiting software vulnerabilities.
Grok Bot ships with standing logins to your email and CRM
xAI launched Grok Bot on August 11, an early-beta agent that signs into a user's own accounts, keeps its own computer, and re-runs saved workflows on a schedule without supervision.
You can move an AI reviewer's score without changing a single result
A new study rewrote research papers to change only their rhetoric while preserving every scientific claim, and found AI reviewers shifted their overall scores by up to nine tenths of a point, with the effect strongest near the accept-reject boundary.
Qwen3.8-27B shares its predecessor's bones, but not its contract
Alibaba's Qwen3.8-27B shipped with the same coarse architecture as Qwen3.6-27B, prompting accusations it was a relabel with knowledge stripped out, but the published comparison shows knowledge scores flat or slightly up.
The benchmarks say Opus 5 improved; the people using it disagree
Anthropic reports Opus 5 as state of the art on coding and knowledge work, while developers on Hacker News and Reddit describe a model that overreaches and burns tokens, and the Claude Code system prompt grew by 48,736 tokens in a single release.
Google's private AI runs on sealed hardware, not on encrypted math
Google's shipping private inference product runs Gemini inside hardware enclaves on custom chips, which is confidential computing rather than homomorphic encryption, and the company's actual homomorphic work is an unsupported research compiler.
A closed-loop benchmark caught nine world models forgetting the room
A new benchmark replaced scripted evaluation with an AI agent pursuing long-horizon goals inside generated worlds, and found that all nine leading world models lose spatial consistency and forget what happened out of frame.
Picking the right model per request beat always using the biggest one
A new routing framework that chooses a different model for each request outperformed the strongest single fixed model by 14.6 percent, partly because the largest model gets many cheap questions wrong.
Someone compiled a working computer into transformer weights by hand
A team constructed transformer weights analytically rather than training them, producing a model that runs arbitrary C programs through a WebAssembly interpreter encoded entirely in its attention layers at about 30,000 tokens per second.
An AGI-thesis fund fell 67 percent and took a market maker with it
Situational Awareness, the investment firm founded by Leopold Aschenbrenner around an artificial general intelligence thesis, dropped 67 percent in July and sold most of its stock portfolio to meet margin calls, with the Financial Times reporting a roughly 15 billion dollar hit at Jane Street connected to the episode.