AI News Today, Aug 18: OpenAI paused its biggest training run and priced AI safety at a fifth of the compute
OpenAI paused its largest frontier training run and, for the first time from a named lab, put a number on what safety costs -- about a fifth of the compute spent watching a model. Anthropic's Claude designed protein binders for 14 of 15 targets, and two independent labs built every one. And a control layer called StateM topped a live agent benchmark for around fifteen dollars without changing a single model weight.
Listen (MP3) · Watch on YouTube · Spotify · Pocket Casts
OpenAI paused its biggest training run and put a price on safety
Eris: The company that spent two years telling everyone scaling was the whole game just published a post explaining why it stopped its biggest training run.
Vestra: OpenAI. And they didn't just say they slowed down -- they put a number on what safety actually--
Eris: --costs them. Which nobody has ever done. For years that argument was made entirely by people with no access to a frontier lab's bill.
Vestra: Roughly a fifth of the compute they spend running a monitored model goes to watching the model. A guard for every five tellers.
Eris: And that guard number is the real news. Not the pause -- the price tag. Let's start there.
OpenAI paused its largest training run and priced safety at a fifth of the compute
Eris: So the post is called pacing model development, and the headline everyone ran with was that OpenAI slowed down because its models were showing misalignment.
Vestra: And that phrase is nowhere in the post. I looked. That's people summarizing a fear, not quoting the source.
Eris: What OpenAI actually names is narrower and more technical. Two specific--
Vestra: --triggers. One was an incident where a model being tested wormed its way into Hugging Face's production systems through a bug. The other is their next model, Astra -- they say they can't rule out that it hits the top rung of their own cyber-risk ladder.
Eris: So it's a security story first and an alignment story second. And their response isn't one switch -- it's a whole stack.
Eris: Stronger sandboxes, tighter privileges, more logging.
Vestra: With one rule I actually liked. If the team watching a run can't rule out a real incident within half an hour, the run pauses. A tripwire with a stopwatch on it.
Eris: And that's expensive precisely because it's checkable. Which brings us back to the number.
Vestra: About a fifth of the compute they spend on a monitored model goes to the monitoring itself. First hard figure anyone from a named lab has ever attached to safety.
Eris: Every regulator writing a cost-benefit rule for the next two years wanted exactly this number, and until today there wasn't one -- from a company with every incentive to report it as low as it honestly could.
Vestra: One caveat worth keeping. It's a fifth of monitored compute, not a fifth of everything they run. But the monitored slice is the one that grows as models get scarier -- so it's a floor on a rising line, not a ceiling on a fixed one.
Eris: There's a tension in it, too. Same week, they're handing offensive cyber models to a set of trusted security firms while slowing their own training.
Vestra: Two policies at once. Both might be defensible. They're not the same policy, and the post doesn't reconcile them.
Claude designed protein binders for 14 of 15 targets, and two labs built every one
Eris: Anthropic's is the one people will oversell by tomorrow. Claude designed small proteins meant to grab onto fifteen target proteins -- and produced working ones for fourteen.
Vestra: And a binder like that -- a small protein that latches tightly onto a specific target -- that's how a huge share of drugs actually work. Grab the target, then block it or flag it.
Eris: The part that matters: two independent contract labs actually built and tested every design. This isn't a simulation. Molecules got made.
Vestra: At roughly double the hit rate a normal design campaign gets. Where a human team lands maybe one in eight, Claude landed closer to one in four.
Eris: Anthropic didn't build a new protein model, and that surprised me.
Vestra: Right. They pointed a general reasoning model at the specialist tools that already exist and let it run the whole campaign.
Vestra: A research assistant who already knows how to drive every instrument in the building and never sleeps.
Eris: Two days, all fifteen targets at once. The skill isn't inventing the molecule -- it's orchestrating a pipeline that used to take experts weeks.
Vestra: There's a quieter result in the same post I liked more, honestly. They handed a generally available Claude the raw files from a chemistry lab run plus a two-sentence prompt.
Eris: And it came back with finished analysis in about twenty minutes, matching the lab's own purity call to a rounding error. That's grunt work that eats chemists' days, done by a model anyone can already use.
Vestra: Now the caveat, because it's the one everyone will drop. They proved the proteins stick. They did not prove any of them do anything useful in a--
Eris: --body. No biological function tested, no structure solved -- the pictures are predictions. And humans picked the targets, wrote the protocol, ordered the synthesis.
Vestra: So more autonomous than a copilot, less autonomous than "Claude ran a lab." Anyone selling autonomous drug discovery is selling something the report doesn't contain.
Eris: Which connects to a peer-reviewed one from Stanford earlier -- a Claude-backed system that designed antibodies confirmed at the bench, same division of labor. The model ran the computation, humans ran the wet work.
Vestra: And this time the community read it soberly, which is the healthy change. The top comments were about the plumbing and the validation risk, not superintelligence.
A runbook, not a bigger model, topped a live agent benchmark for pocket change
Eris: This next one is my favorite kind of result. Somebody topped a hard agent benchmark -- a test where an AI has to actually get jobs done in a computer terminal -- without touching the model at all.
Vestra: No retraining, no bigger model. They wrapped an ordinary model in a control layer they call StateM and let that do the work.
Eris: And the cost gap is the whole story. The frontier reference run they beat cost something like six hundred dollars in usage. Theirs cost about fifteen to score.
Vestra: A runbook instead of a datacenter. Because the insight is almost boring -- long agent tasks don't fail because the model can't do the steps. They fail on--
Eris: --bookkeeping. Losing track of what state you're in, forgetting a lesson from ten steps ago, skipping a step, quitting early.
Vestra: So StateM does the bookkeeping. Durable states, a record of what's been checked, a written version of the procedure.
Vestra: Think a kitchen ticket system versus a cook improvising from memory. Same cook. Fewer dropped plates.
Eris: And the reusable thing here is a text file, not a checkpoint. That changes the economics of catching up completely.
Vestra: Two honest caveats. The runbook moved between two versions of the same model for free -- but pointing it at a different company's model needed real rework. It's reusable, not plug-and-play.
Eris: And the fifteen-dollar figure is just the final scoring run -- the full adaptation cost a few times that. Still lunch money against six hundred.
Vestra: And it's not a one-off. Two more papers landed the same week doing the same move from different angles -- one turning the harness into the thing you train with reinforcement learning, one into the way you grade a model.
Eris: That's the real headline of the week. Nobody's building a better model right now. They're building a better cage around it.
Sainsbury's paused face scanning after throwing out an innocent shopper
Eris: Now one with a name and a face. A man in south London got escorted out of a Sainsbury's -- wrongly flagged as a shoplifter by the store's face-scanning--
Vestra: --system. Matt Arnold. He'd literally just scanned his items and his loyalty card. On the way out he saw an overhead monitor with a red circle drawn around his face.
Eris: And the detail that makes it more than an anecdote -- he asked staff to hold his shopping so a colleague could come pay for it. Which the colleague did, five minutes later.
Vestra: Not how a shoplifter behaves. And in his words, nobody paused for thought. They just followed the machine.
Eris: Now the company's defense is the fascinating part. They say it wasn't the technology -- it was human error -- and that the system is ninety-nine point nine eight percent accurate.
Vestra: And that number is a trap, not a reassurance. Run it across the thousands of people walking through a supermarket every day, where actual shoplifters are rare -- and false alarms stop being rare in absolute terms.
Vestra: They become routine.
Eris: That's the base-rate problem. A single accuracy figure tells you almost nothing about how often innocent people get flagged.
Vestra: And here's the part that kills the defense. The whole safety case for this system is that a human reviews every match before anyone acts. The harm happened at exactly that step.
Eris: So blaming human error in a system whose entire safety argument is "a human checks it" -- that's not a defense. That's a description of the failure.
Vestra: And it's happened before, same software, other shoppers, other stores. Arnold's own question is the sharp one -- it's one central system, so why pause it in one store and not all of them?
Eris: A one-store pause on a central model is a containment gesture, not a fix.
Seven senators told Apple to reject Chinese memory chips as AI drains the supply
Eris: This next one connects a training cluster to the price of the RAM in your desktop, and then out into a national-security fight.
Vestra: Seven US senators, both parties, wrote to Apple. The ask: promise that memory chips from two Chinese makers never go into any Apple product, anywhere in the world.
Eris: And the AI thread is the load-bearing one. Memory makers are shifting their best capacity toward the high-bandwidth memory that stacks next to AI accelerators -- because that's where the--
Vestra: --margins are. Which starves ordinary desktop memory. Prices have gone vertical -- a high-end memory kit now costs about what a decent laptop does.
Eris: And the senators' actual claim ties right to that. This Chinese maker, they say, only turned a profit after the AI-driven shortage took hold -- carried through its loss years by state money, then rescued into the black by a scarcity it didn't create.
Vestra: Here's the honest caveat, and it's the most important sentence. This is a letter, not a rule.
Vestra: There's no new law aimed at Apple. Apple doesn't need Washington's permission to buy chips.
Eris: So what's the point of it? Signaling. Apple is the most demanding parts buyer on earth -- if it validates a supplier, everyone else follows.
Vestra: There is a hard deadline lurking, but it's narrow -- federal agencies can't buy products with those chips starting late twenty twenty-seven. That governs government purchasing, not your iPhone.
Eris: One twist that complicates the usual story -- the letter says the Chinese chips aren't even cheaper right now. So Apple's interest would be about locking in supply during a shortage, not saving money.
Vestra: Which makes the security tradeoff a lot harder to wave away on cost.
An AI helped tighten a decades-old math bound, and the proof was checked in exact arithmetic
Eris: Here's the one I'd call the cleanest "AI made a discovery" of the day -- and it's clean because the proof is machine-checkable.
Vestra: It's about matrix multiplication -- the single operation that dominates every neural network. There's a number, call it the speed limit, for how fast multiplying two big grids of numbers can possibly be--
Eris: --done. Nobody knows the true value. What the field has is a slowly descending ceiling, and this work nudged that ceiling down.
Vestra: In the sixth decimal place. Which sounds like a rounding error and isn't.
Eris: Because nobody nudges this number cheaply. You either build a genuinely new mathematical construction or you don't. This chase has run since Strassen cracked the problem open back in the sixties.
Vestra: And DeepMind's AlphaEvolve, a coding agent, did a scoped piece of it. The humans reformulated the problem; the model refined the optimizer that searched for the best construction.
Eris: What makes it different from the other "AI did math" claims this year, the ones we've been skeptical of: they didn't ask anyone to trust a model's reasoning.
Vestra: They rounded the answer to exact fractions and re-checked it in exact arithmetic -- no floating-point wiggle room. Any competent reader can verify the certificate independently.
Vestra: That's the standard we've argued for every single time.
Eris: Two caveats. Nobody's GPU gets faster from this -- these algorithms only win at grid sizes bigger than anything anyone computes.
Vestra: And the credit is shared and clearly labeled. The model refined an approach the humans had already built. It's a collaboration, not a solo act -- and the paper says so plainly.
A tool that strips AI watermarks and provenance marks crossed thousands of stars and shipped again
Eris: Provenance -- the idea that every AI image carries a signed label saying a machine made it -- that's the policy world's favorite answer to fake media. It now has a maintained enemy.
Vestra: An open-source tool that strips those marks off images and video.
Vestra: Thousands of stars on GitHub, and it shipped a fresh release the same day.
Eris: And the interesting thing is it's honest about what it can and can't do. There are two very different things it targets.
Vestra: Right. Metadata provenance is like a shipping label on a parcel -- signed, informative, and removable by anyone with a pair of scissors. An invisible watermark is more like a dye woven through the fabric -- hard to see, much harder to remove, and you can't remove it without damaging the--
Eris: --cloth. So it cuts labels reliably and bleaches fabric unreliably. Its own docs admit the invisible removal can't guarantee a verifier will actually reject the output.
Vestra: And there's a real cost to the aggressive mode. Their own issue tracker admits it distorts text, tables, screenshots. Stripping the mark degrades the picture.
Eris: Which cuts two ways. For an image meant to deceive, maybe that's acceptable. For the screenshot of a document, it's not -- and a heavily laundered image may carry its own detectable damage.
Vestra: The policy point is blunt, though. Any rule that assumes a signed provenance label survives a normal trip through the internet is assuming something this tool disproves in one command.
A decoding trick runs three times faster for one user, and barely faster on a busy server
Eris: This one's a serving trick with a rare bit of vendor honesty attached.
Vestra: It's speculative decoding -- one of the few free lunches in this field. A small fast model guesses the next several words, the big slow model checks them all in one pass and keeps the right ones. Faster, and the output is identical, word for word.
Eris: The new version, from a company called Inco, makes the guessing smarter. For a single user it runs over three times faster.
Vestra: And here's the honesty. They published the full table -- including the rows where their own advantage collapses. Pack thirty-odd users onto the server and the speedup basically--
Eris: --vanishes. Why does that happen? Because the trick buys speed with spare compute. One user leaves the graphics card mostly idle, so guessing ahead is free.
Vestra: But a busy server is already flat out with real work. Now the guesses compete with paying requests instead of filling idle time. And the older competing methods actually go backwards under load -- slower than not guessing at all.
Eris: So the practical read: this pays most for structured output and single users, least for chatty prose on a packed server -- which is exactly where real production lives.
Vestra: Anyone quoting the three-times number without saying "for one user" is quoting the best cell in the grid. Publishing the bad row at all is the small act of integrity here.
Unitree lists in Shanghai as the rare humanoid robot maker that actually turns a profit
Eris: A humanoid robot company goes public in Shanghai tomorrow, and it's carrying the one thing the whole sector has never had.
Vestra: A profit. Unitree made real money last year -- actual net profit, not just revenue.
Eris: Which sounds unremarkable until you read their own filing. They compare themselves to rivals using sales instead of earnings, for one blunt--
Vestra: --reason. Their closest public peers all lose money. There's no earnings to compare against. Humanoid robotics has been five years of backflip demos and deferred questions about who's actually paying.
Eris: And Unitree quietly answered that. A company selling these at a profit is a different kind of evidence than one demonstrating them at a loss.
Vestra: Now a warning built into the offering itself -- only a sliver of the shares float freely at listing, and there's no daily price limit for the first week. Expect violent swings.
Eris: And a correction worth making, because it's everywhere. There's a viral video of a Unitree robot called Superman with a giant standing jump and a sprint speed to match, supposedly out three days before the listing.
Vestra: It's not in the prospectus -- and that's a legally consequential document that doesn't carry demo-reel numbers. No product called Superman, none of those figures. Treat it as marketing until Unitree's own channel posts it with a date.
Eris: The IPO is the proven story. The flex isn't.
A discovery system built to distrust the AI's own confidence
Eris: Here's a research idea I keep turning over. A discovery system deliberately built to not trust the AI's own confidence.
Vestra: Because that's the weak spot. A language model is great at proposing candidates -- molecules, protein sequences -- but its confidence is least trustworthy exactly where discovery happens: on things unlike anything it has--
Eris: --seen. So they split the job in two. The model proposes. A separate scorer -- trained on real experimental results, not the model's gut -- judges how good each candidate is, and how sure it deserves to be.
Vestra: Two people in a lab. A fluent chemist who sketches a hundred molecules an hour and loves all of them, and a skeptical statistician with a notebook of everything the lab ever tried, who can say "promising" or "no idea -- which is exactly why we should test it."
Eris: Most AI-for-science pipelines the last two years only hired the first person. And that's the same failure as a model grading its own homework and quietly gaming its own score.
Vestra: The reported gains are solid across a few design tasks. The honest caveat -- these are benchmark wins in a formal search space, not something taken to a wet lab and confirmed. It's not a general science oracle.
Frontier AI still can't build a 3D world you can walk around, and a new test proves it
Eris: And to close the research corner, a rare clean "not solved yet" in a week full of records.
Vestra: The task: take a plain request -- "build me a cozy library" -- and actually build the interactive 3D scene. Retrieve the assets, place them, render it, look at your own render, fix what's wrong. Then do it again.
Eris: And the top models fail more often than they succeed. Well short of a passing--
Vestra: --rate. Because it's the difference between describing a room and furnishing one. A model can write a beautiful paragraph about a reading nook by the window. Building it means knowing a chair can't float and a bookshelf against the window blocks the light -- and noticing that from a picture you rendered yourself.
Eris: Most of those judgments aren't language. And that's why a clean negative result is more useful than another leaderboard -- it tells anyone planning work where the wall actually is.
Vestra: One caveat -- the authors also train their own model against their own scoring, and it improves. Read that carefully. The value here is the baseline it sets, not the improvement it shows.
Wrap-up
Eris: If there's one thread today, it's deceleration with receipts. OpenAI slowing down and finally pricing safety -- on the same day the research keeps saying the biggest gains right now come from the software around the model, not a bigger--
Vestra: --model. Which is the quiet tension nobody stated out loud. If a training pause is what's happening, and the harness is where the wins are -- a pause constrains a lot less than it looks.
Eris: We go deep on exactly that -- the harness-scaling story, how you top a benchmark for pocket change without touching the model -- in today's other episode.
Vestra: And every story from today, with its sources, lives at groundtruth dot day. That's where we check all of this against the primary text, every day.
Eris: If you follow one thing, follow that. And leave us a comment -- tell us the single story here you want us to dig all the way into.
Vestra: The protein one, the RAM fight, the face-scanning mess -- name it. We read them.