Ground Truth.
AI, checked against the source.
← 2026-08-132026-08-14later →

Z.ai changed only the post-training, and the model learned to find exploits

2026-08-14

Z.ai released GLM-5.3 on August 14 using the same base model as GLM-5.2, with every gain coming from post-training, and the largest jump was in finding and exploiting software vulnerabilities.

cybersecurity · ai-security · vulnerabilities · open-weights · coding · post-training · china · glm

Grok Bot ships with standing logins to your email and CRM

2026-08-14

xAI launched Grok Bot on August 11, an early-beta agent that signs into a user's own accounts, keeps its own computer, and re-runs saved workflows on a schedule without supervision.

cybersecurity · prompt-injection · ai-security · agents · product-launch · xai

You can move an AI reviewer's score without changing a single result

2026-08-14

A new study rewrote research papers to change only their rhetoric while preserving every scientific claim, and found AI reviewers shifted their overall scores by up to nine tenths of a point, with the effect strongest near the accept-reject boundary.

peer-review · evaluation · reward-hacking · llm-as-a-judge · research-integrity

Qwen3.8-27B shares its predecessor's bones, but not its contract

2026-08-14

Alibaba's Qwen3.8-27B shipped with the same coarse architecture as Qwen3.6-27B, prompting accusations it was a relabel with knowledge stripped out, but the published comparison shows knowledge scores flat or slightly up.

open-weights · qwen · local-models · benchmarks · china · quantization

The benchmarks say Opus 5 improved; the people using it disagree

2026-08-14

Anthropic reports Opus 5 as state of the art on coding and knowledge work, while developers on Hacker News and Reddit describe a model that overreaches and burns tokens, and the Claude Code system prompt grew by 48,736 tokens in a single release.

anthropic · agents · evaluation · developer-tools · model-behavior · harness

Google's private AI runs on sealed hardware, not on encrypted math

2026-08-14

Google's shipping private inference product runs Gemini inside hardware enclaves on custom chips, which is confidential computing rather than homomorphic encryption, and the company's actual homomorphic work is an unsupported research compiler.

privacy · cybersecurity · ai-security · google · encryption · infrastructure

A closed-loop benchmark caught nine world models forgetting the room

2026-08-14

A new benchmark replaced scripted evaluation with an AI agent pursuing long-horizon goals inside generated worlds, and found that all nine leading world models lose spatial consistency and forget what happened out of frame.

world-models · robotics · benchmarks · video-generation · evaluation

Picking the right model per request beat always using the biggest one

2026-08-14

A new routing framework that chooses a different model for each request outperformed the strongest single fixed model by 14.6 percent, partly because the largest model gets many cheap questions wrong.

routing · inference-cost · infrastructure · benchmarks · open-source

Someone compiled a working computer into transformer weights by hand

2026-08-14

A team constructed transformer weights analytically rather than training them, producing a model that runs arbitrary C programs through a WebAssembly interpreter encoded entirely in its attention layers at about 30,000 tokens per second.

transformers · interpretability · research · open-source · architecture

An AGI-thesis fund fell 67 percent and took a market maker with it

2026-08-14

Situational Awareness, the investment firm founded by Leopold Aschenbrenner around an artificial general intelligence thesis, dropped 67 percent in July and sold most of its stock portfolio to meet margin calls, with the Financial Times reporting a roughly 15 billion dollar hit at Jane Street connected to the episode.

industry · finance · markets · ai-investment · risk

← 2026-08-132026-08-14later →