Ground Truth.
AI, checked against the source.

AI News Today, Aug 23: An open world model that remembers, and a robot outruns Bolt's clock

2026-08-23 · Breach Protocol: Inside the AI Blackbox — full transcript

Alaya Lab's Evoke keeps a generated world consistent for an hour by storing scene geometry outside the model -- open weights, whole training ladder included. Generalist AI says its robot learns new tasks from a single twelve-second demo, with nothing but a blog post as evidence. And Ramp's spend data shows businesses refusing to pay for the market's best model, even as Anthropic's overall adoption keeps climbing.

Listen (MP3) · Watch on YouTube · Spotify · Pocket Casts

An AI world that stops forgetting itself

Eris: So the chair problem might actually be solved. You know the one -- every AI-generated world forgets its own furniture the second you look away.

Vestra: The one where you turn back and it's a different chair. That failure has killed every interactive world model so far.

Eris: A lab called Alaya just shipped one that keeps the world in an external memory. Open weights, the whole training recipe published, and it holds a scene together for an hour.

Vestra: And the catch is in the fine print, because it always is. Let's get into it.

Evoke, an open world model that remembers where it has been

Eris: The release is called Evoke, fourteen billion parameters, from Alaya Lab, and it went straight to the top of Hugging Face's daily paper list. The pitch is a generated world you can walk around in for an hour without it forgetting itself.

Vestra: And how it does that is the actual news, because it's not a bigger model with a longer memory. They gave up on the generator remembering anything. The model draws a chunk of video, about a second and a half at a time, then estimates the depth of what it just drew, turns those pixels into 3D geometry, and files that geometry in a separate store organized by where the camera was standing.

Eris: So when the camera swings back to a spot it has visited, the system looks up the stored geometry for that position, warps it into the current view, and hands the generator a picture of what should be there -- plus a mask marking the parts that are actually unknown.

Vestra: Which shrinks the generator's job down to rendering. It doesn't have to remember the chair. Something else wrote the chair down.

Eris: The analogy I keep coming back to is a film crew's continuity binder. The camera operator doesn't memorize what was on the desk in scene four. Someone wrote it down, and hands the page back when the crew revisits the set.

Vestra: Okay, that image works. Now the fine print, because the launch coverage got a big thing wrong. Posts today called this real-time on a single GPU. The paper's own timing says it takes a bit over two seconds of compute to produce a second and a half of video, on a serious datacenter GPU, at fairly low resolution. That is slower than playback. The authors say it themselves -- further acceleration still needed.

Eris: In fairness, they measured that with none of the usual speed tricks turned on. No caching, no compilation, no quantization. There's real headroom.

Vestra: There's headroom, and there's also a claim about the shipped thing that isn't true of the shipped thing. Both matter. Same with the memory: it's bounded -- roughly ninety seconds of geometry in the long-session setup -- and the paper describes what it recovers as recognizable rather than pixel-faithful. A world with a working short-term memory, not a perfect one.

Eris: What I love anyway is the generosity of the release. They didn't just publish the finished model, they published every checkpoint along the training ladder, so another lab can pick the method up halfway instead of reproducing it from scratch. That almost never happens.

Vestra: Apache licensed, too, with one warning for anyone commercial: the depth components are separate downloads under a non-commercial license. Read that boundary before you build on it.

Eris: And hold onto why this matters, because it threads through half of today. An interactive world that stays consistent for an hour is exactly the training ground embodied AI has been missing, and this is the first open-weights system in that class. The robot half of that same bet landed this week too.

A robot learns a new task from a twelve-second demo

Eris: Generalist AI says its new robot model, GEN-1.5, can watch a single demonstration -- three to twelve seconds of a person doing a task -- and then just attempt it. No fine-tuning, no training run. Across ten household-style manipulation tasks it worked a bit better than a coin flip.

Vestra: Which sounds unimpressive until you remember the baseline for a robot meeting a task it was never trained on is basically zero.

Eris: Right, and the comparison they draw is to language models in 2020. GPT-3's big trick wasn't knowing more, it was that examples in the prompt were enough. Show it a pattern, it generalizes, no retraining. Robotics has chased that for decades and only ever gotten narrow versions of it.

Vestra: The mechanics here: the model takes in video, language, and its own joint sensors across a thirty-second memory window, and the demonstration goes into that window like example text goes into a prompt. They call it physical prompting. The demo is not a training example. It's context.

Eris: And the weird behaviors are honestly the best part. A demonstration recorded in simulation works as a prompt for the real robot, even though there's no simulation anywhere in its training data. A person can demonstrate with their bare hands and the robot maps the motion onto its grippers. Denied the tool it was shown, it grabs a different tool and improvises.

Vestra: And the company insists none of that was engineered. No architecture changes to encourage it, no meta-learning loop, no improvisation objective. It just appeared after eight-plus months of continuous pretraining on physical interaction data. One of their researchers dates the discovery to a specific evening in early August.

Eris: Emergence with a timestamp. I find that genuinely charming.

Vestra: I find it genuinely unverifiable, which is my job here. This is a company blog post, evaluating its own model, on its own hardware. No paper, no weights, no model card, no outside reproduction. When GPT-3 did this for text, anyone with API access reproduced it within weeks. Nobody can reproduce this.

Eris: Agreed -- a claim, not a measurement, for now. But put it next to Evoke and the shape of the week is hard to miss. One lab is manufacturing consistent worlds to train robots in, another says the robot already generalizes if the physical pretraining is big enough. Everyone's converging on the same bet: the environment, not the architecture, is the bottleneck.

Vestra: If the weights ever ship, we'll find out whether the bet paid.

Businesses are not paying for the best model on the market

Eris: The corporate card company Ramp published its July spend data, and one number is getting passed around everywhere. Claude Fable 5 -- scored as the single most capable model you can buy -- made up just six percent of the tokens businesses bought from Anthropic. About one dollar in nine of their Anthropic spend.

Vestra: And before anyone runs with the obvious headline -- the same report says the opposite about the company. Anthropic extended its lead in business adoption in July, to well over four in ten U.S. businesses, still ahead of OpenAI. Customers are not leaving. They are declining the top shelf.

Eris: Ramp's economist put it about as bluntly as an economist can: with this model we've found a new upper bound for how much businesses are willing to spend on AI. More performance is not worth the price tag.

Vestra: The mechanism is worth spelling out, because it isn't about quality. Fable 5 costs roughly twice OpenAI's flagship per token. And when independent benchmarkers work out cost per completed task rather than cost per token, Anthropic's own cheaper model -- Opus 5 -- scores essentially the same for meaningfully less money. The best model on the leaderboard is not the best buy on the invoice, even inside its own vendor's lineup.

Eris: There is an honest confounder, though, and the data sits right on top of it. Fable 5 launched in June, got suspended three days later under a government export directive, and only came back at the start of July with usage caps for its first week. A model that was unavailable for most of a month is going to look under-bought the month after.

Vestra: Which is why the number to actually watch is next month's index. August is the first clean month. If the share stays this low with no availability excuse, the pricing-ceiling story hardens from a data point into a finding.

Eris: And there's a slower signal underneath. The share of businesses paying for model-serving platforms -- the on-ramp for running open-weight models yourself -- has been climbing all year, and the growth is coming from the heaviest, most sophisticated spenders, not the newcomers. That's substitution pressure arriving from the top of the market, which is the dangerous direction if you sell frontier models.

Vestra: So what's under pressure is a pricing theory, not a company. A frontier model now has to beat the cheap alternative by more than the price gap, every single month. That distinction is the whole story.

A humanoid robot ran 100 meters in 9.39 seconds in Beijing

Eris: Beijing, last night, the World Humanoid Robot Games. A robot from the Tianzhuo team ran the hundred meters in 9.39 seconds. Usain Bolt's world record, unbeaten since 2009, is 9.58.

Vestra: So on the stopwatch, a machine is now faster than the fastest human ever recorded. And the honest version of that sentence keeps its asterisks attached: robot-only race, robot-specific rules, a preliminary heat, on a prepared indoor track. Nobody lined up against Bolt.

Eris: The part that impressed me more than the time is the rule change. Last year you could steer your robot down the track. This year the sprint is fully autonomous. The machine has to perceive the track, keep its own balance, and hold a line at sprint pace with no human in the loop.

Vestra: That's the actual engineering story. Making a teleoperated platform fast is a hardware problem. A two-legged machine stabilizing its own gait at the edge of what its motors can do, by itself, is a control problem, and a much harder one. It's the difference between a remote-controlled car and a self-driving one. Both can be quick. Only one has to know where it is.

Eris: The Games themselves are telling too. Fifty-one events, and twenty-one of them aren't sports at all -- they're scenario challenges staged like factories and homes. A Beijing official said the quiet part on state radio: robots are an industry, and a future, not just a technology. This is a shop window for Chinese manufacturers.

Vestra: One warning about the other numbers circulating alongside the sprint. People are comparing a robot's standing jump to the human high jump record, and burst top-speed figures to a full-race average. Those measure different events entirely. Category errors, not records. The 9.39 is real. Most of the garnish around it is not.

Eris: A sprint is about the narrowest capability a robot can have. The scenario events are where we find out whether any of these machines can stack a shelf.

OpenAI argues open models will power unstoppable cyber-attacks

Eris: OpenAI's head of global affairs, Chris Lehane, gave the Guardian the sharpest version yet of the case against open-weight models. His argument: freely downloadable models are only months behind the frontier, and once the weights are on an attacker's machine, nobody can rate-limit them, refuse them, or switch them off. Continuous, persistent, automated attacks, and you'll need superior models just to defend yourself.

Vestra: The distinction he's leaning on is real. A hosted model has a provider who can close your account. A downloaded one is like owning the locksmith's tools instead of renting the locksmith. The tools don't decline the job.

Eris: And his ask is specific: a U.S. law making pre-release safety proof mandatory. You could not release or deploy a model without demonstrating a level of safety first, with a pause mechanism built into the process. He puts the legislative window in early 2027, with a new Congress.

Vestra: Now the pressure points. The core claim is a projection, not a measurement -- nobody has published a study showing open-weight models sustain attack campaigns better than closed ones, and the strongest documented AI-enabled intrusions to date ran on commercial models. His own product category, in other words.

Eris: The timing doesn't help him either. A company that sells closed API access, arguing that free downloadable models are the danger, weeks before an expected market listing at a reported valuation north of eight hundred fifty billion dollars. Mandatory pre-release proof is a compliance burden that lands hardest on the people who give weights away.

Vestra: And the fiercest pushback in the same article came from the safety side, not the open-source side. David Krueger, who helped found the UK's AI Security Institute, says nobody should be building more powerful systems at all because we can't control or understand them, and calls the labs reckless. So this fight has three corners: OpenAI says regulate the open models, open advocates say that's market protection, and the hardline safety researchers say all of you are the problem.

Eris: Buried in the same reporting, and almost bigger than the interview: OpenAI paused development of its most advanced internal models this week over safety concerns. Their safety lead says they are very far from everything running back to normal.

Vestra: Single-sourced through the Guardian so far -- no first-party publication from OpenAI. The pause is also the only evidence Lehane offered that his company is being careful. Watch whether it gets its own write-up. Continued silence would tell us something too.

Iran-linked hackers took a UK power plant offline for four days

Eris: The UK government confirmed that a cyber-attack took a British power plant offline for four days last month. It's the first publicly acknowledged case of an intrusion actually stopping power generation there, and the government blames hackers linked to Iran.

Vestra: The four days is the detail that matters. Most reported intrusions into energy systems are reconnaissance -- someone gets read access, maps the environment, and leaves. Keeping a generator down for four days means the attackers reached systems that control a physical process, and recovery took more than restoring a workstation.

Eris: The official framing is carefully small: a small-scale generator, no risk to the wider grid at any point. Both of those are probably true. Neither answers whether the same access at a larger site would have meant a larger outage.

Vestra: And there's a conflation to resist here, because the timing invites it. Days earlier, five U.S. agencies warned about AI-assisted scripts probing industrial controllers in American infrastructure -- we covered that advisory on the show. Same category of target, same news cycle, and zero public evidence connecting the two. The American advisory describes reconnaissance, not disruption, and the British incident has no published technical detail at all. No entry vector, no malware family, nothing.

Eris: So the attribution to Iran is a government assertion, made in the middle of an active dispute. The UK just extended permission for U.S. operations against Tehran from British bases, and Iran's Revolutionary Guard had already called those bases legitimate targets.

Vestra: A plausible assertion with zero published evidence -- those can both be true at once. What the record does say is that the documented Iran-linked infrastructure campaigns got in through default passwords on internet-exposed equipment. Neglect, not sophistication. That has been the real story for a decade.

Eris: And a decade of advisories never moved a security budget. A confirmed four-day loss of generation might. That's the cynical reason this story matters.

Ornith-1.5 writes its own training problems and grades them

Eris: Ornith released an open-weight model family, Ornith-1.5, with a training loop that feeds itself. The model invents its own practice problems, builds the grader for each one, attempts a solution, and the reward flows back through all three stages at once. On the vendor's own agentic coding evaluations, the flagship lands level with Anthropic's best.

Vestra: The reason this isn't just synthetic data again is the reward design. Ask a model to write practice questions and it will happily write easy ones, broken ones, or duplicates of what it already knows. Ornith makes task-writing itself a rewarded skill, and the reward multiplies three terms together. The task and its grader have to actually work. The difficulty has to sit right at the model's frontier -- they aim for problems the model solves about one time in five. And it has to be novel against everything already generated. Multiply the three, and failing any one zeroes the whole reward. You can't farm points with impossible problems or reworded old ones.

Eris: And the curriculum advances on its own. Difficulty is measured against the current model's attempts, so as it masters a task, the reward for generating that task decays and the proposer moves on to harder material. A tutor who writes tomorrow's problem set from what you got wrong today.

Vestra: Caveats, and they're structural. This family is built on top of other labs' base models -- a Qwen and Gemma lineage with continued pretraining -- which the release doesn't exactly lead with. The benchmark comparisons are vendor-run against a peer set the vendor chose. And the deep one: the same model writes the exam and takes it. If its idea of correctness has a flaw, that flaw flows into the reward signal invisibly. Their anti-gaming terms are a mitigation, not a proof.

Eris: File it next to another paper from today, Co-RL -- the same move at a different layer. When you run out of human-made supervision, you build a machine that manufactures it, and then everyone argues about whether the machine can be trusted.

Vestra: One independent result on a private test suite would settle a lot. Until then, a clever recipe with a self-grading problem.

Self-hosting a frontier open model now takes a whole server node

Eris: imec, the Belgian research institute, did the accounting nobody selling GPUs wants done. They ran open-weight coding models through the same agent harness as the commercial APIs and priced the whole stack. Their verdict on buying your own hardware, in their own words: probably not to save money.

Vestra: The finding that jumped out at me is that the best open model no longer fits on hardware a serious team would plausibly buy. Kimi K3's weights alone are about 1.4 terabytes. An eight-GPU node of the previous generation holds one and a half total, which leaves no room for the working memory every active session needs. So it's the newest, most expensive node class or nothing -- and even there it served sixteen simultaneous users and ran roughly eight times slower per task than the hosted frontier agent they compared it against.

Eris: And the killer isn't the hardware price, it's the idle time. Demand from a developer team is spiky -- dead overnight, peaking mid-afternoon. You buy for the peak and pay around the clock. Measured utilization for internal developer tools runs somewhere around one working hour in five, so a box that idles most of the day costs a multiple of its sticker rate per useful hour.

Vestra: Their spending data explains why anyone is tempted anyway. The median developer costs about a hundred and forty dollars a year in API usage -- trivial. But the heaviest agent users run into the tens of thousands, and that tail arrives suddenly, because agents spend tokens on your behalf without asking. The bill problem was never the average. It's the tail.

Eris: The structural point underneath is the keeper: memory, not compute, now sets the floor for running frontier open models. Which quietly narrows what open weights deliver in practice. The license says anyone can run it. The memory bill says almost nobody can.

Vestra: There's an irony worth naming, given the Lehane interview. The property he calls the threat -- nobody upstream can rate-limit or shut off a downloaded model -- is precisely what imec lists as the real reason to self-host, along with data that never leaves the building. Same fact, opposite signs.

Eris: Sovereignty costs a full node and slower answers. At least someone finally published the price.

Alibaba wants to raise ten billion dollars purely for AI

Eris: Alibaba proposed a share placement worth eighty billion Hong Kong dollars -- call it ten billion U.S. -- and said something companies almost never say. One hundred percent of the net proceeds go to full-stack AI, including infrastructure. Not general corporate purposes. One destination, the whole amount, in a legal disclosure document.

Vestra: Two structural details deserve attention. It's a proposal, not a done deal -- their own release says there can be no assurance it completes. And the shares are offered only to non-U.S. investors, structured entirely outside American markets. That's a Chinese company signaling where it expects to fund its compute buildout from now on, and the answer is not the United States.

Eris: The timing fits a pattern that's been building all month: compute is repricing upward from every direction at once. Hardware vendors raising prices on memory supply, imec showing the real serving floor for big open models, and now ten billion dollars of fresh equity earmarked for datacenters. When the input costs of a datacenter climb, the equity you need to build one climbs with it.

Vestra: And equity is the unfakeable signal. Issuing new shares dilutes every existing holder, and management only accepts that when it believes the thing being bought is worth more than the dilution. Whatever a press release says, this is belief priced in their own stock.

Eris: The catch is that full-stack AI capabilities is broad enough to mean chips, power contracts, model training, or acquisitions -- and an announced intention carries no duty to report back on where the money actually went. Precise about the amount, vague about everything else.

Vestra: So track two things: whether the placement closes, and whether any later disclosure breaks out the spending. Until then it's a very expensive statement of belief.

Gallup is testing AI agents that answer surveys for real people

Eris: Gallup -- the polling institution -- has built AI stand-ins for about a thousand of its real panel members. Each one is grounded in a long recorded interview with that specific person plus their past survey answers, and its job is to predict how that person would answer new questions. Gallup is now formally testing where the stand-ins work and where they fail.

Vestra: The research behind this has a famous accuracy figure that's usually quoted wrong, so let me do the honest version. The benchmark isn't truth. The original study re-asked people the same questions two weeks later, and people don't fully agree with their own earlier answers. Measured against that ceiling -- your consistency with yourself -- an agent built from your interview gets nearly as close to predicting you as you do. An agent built from your demographics alone does noticeably worse. The claim is that what you said about your life beats a stereotype assembled from your age, income, and zip code.

Eris: And part of why it works is real inference, not lookup. Strip out every question the agent could answer by searching the transcript, and the interview-grounded agents still beat the demographic ones.

Vestra: Meanwhile the commercial side has run far ahead of the science. The startup involved, Simile, just raised over two hundred million dollars at a two-billion valuation, has Fortune 100 customers, and states its mission as simulating all eight billion people on earth. That sentence should raise your pulse a little.

Eris: Which makes Gallup's posture the actual news for me. They're inside the partnership and drawing the hard line themselves: simulated responses will never appear in Gallup's published population estimates, will never replace direct measurement, and they name the danger out loud -- this technology could erode public trust.

Vestra: There's also a trap built into the whole idea, the validation paradox. The only way to confirm the simulation got it right is to run the human study it was supposed to replace. That's fine for pre-testing questionnaire wording. It's fatal for the cases where real behavior would have surprised you -- which is exactly when polling matters.

Eris: Gallup's own phrase is the one to keep: simulated and human responses are not interchangeable. Hold anyone who quotes an AI poll to that.

SenseNova generates images in raw pixels with no compressor in the middle

Eris: SenseTime open-sourced an image model, SenseNova-U1.5, that removes two components nearly every image generator on earth depends on -- and most people don't know the components exist. Standard image models never actually touch pixels. They generate inside a compressed sketch produced by a separate network called an autoencoder, and a second frozen network reads images for them. Everything the compressor throws away -- tiny text, small faces, fine texture -- is gone before generation even starts.

Vestra: It's a translator working from a summary instead of the original document. Faster, usually fine, but you cannot recover a detail the summarizer dropped, no matter how good you are.

Eris: This model deletes both middlemen. One backbone reads and generates raw pixels directly, building each image up through progressive upsampling stages instead of predicting a compressed sketch. It's Apache licensed, and that pixel decoder is sitting right there in the public code -- so the claim is real engineering, not marketing.

Vestra: Warnings before anyone downloads it. The name says 8B, the download says otherwise -- it's around eighteen billion parameters and fifty gigabytes of weights, with quantized paths documented for a big consumer card. And the vendor's own quality tables put it competitive with, not ahead of, the strongest closed editors. The architecture is the news. The images are just the argument that the architecture is viable.

Eris: The adoption signal I trust more than any comparison table: within days, people filed feature requests to get it running in the popular local tooling. Researchers admire models. Practitioners open issues.

Vestra: And the larger point stands regardless of this one model. The compressor-in-the-middle design has gone unquestioned for years mostly out of convenience, and someone just showed the convenience is optional.

Models can train each other without a single correct answer

Eris: Same family of idea as Ornith's self-written curriculum, one layer down. A method called Co-RL trains reasoning models with no correct answers anywhere in the loop. Two models train side by side on unlabeled problems, and each one gets rewarded for agreeing with the other model's majority vote -- never its own.

Vestra: That never-its-own is the entire trick. Let a model reward its own most confident answers and it reinforces its own biases until training collapses -- the paper shows those runs literally diverging off the chart. A peer trained separately makes different mistakes, so its vote carries information you cannot get from yourself.

Eris: Two students grading each other's homework with no answer key. If they studied from the same textbook, they share the same mistakes and learn nothing. Different textbooks, and each catches what the other missed.

Vestra: And in their tests it mostly matches the identical training recipe run with real ground-truth answers -- same models, same data, same compute, across most of their settings. That's a striking result for a method that never sees a right answer.

Eris: The resource it burns is disagreement itself. Models from different families make usefully different errors. Models from the same lineage agree too much to teach each other anything. Which reframes model monoculture, doesn't it -- if everyone keeps training on the same web data, that's not just an ecosystem worry, it's draining the exact fuel this approach runs on.

Vestra: And the ceiling nobody can wish away: agreement is not truth. Two models wrong in the same direction reward each other just as confidently as two models that are right. This is math-benchmark work, mostly single runs, with a tiny public footprint so far. Promising mechanism, unproven range.

Wrap-up

Eris: If today has one thread, it's this: the field kept replacing the missing human. A world model manufacturing training grounds for robots, Ornith manufacturing its own curriculum, two models manufacturing each other's supervision, Gallup testing manufactured survey respondents. And every one of them then has to answer the same question -- why should we trust the machine that replaced the person?

Vestra: Every one answered with a different mechanism, too, which is exactly why this is worth watching instead of dismissing.

Eris: In today's other episode we go deep on the day's research papers -- including that world model with the external memory, properly cracked open. And every story from this brief lives on Ground Truth, our news site at groundtruth.day, with links to the original sources so you can check our work.

Vestra: If this recap earns its spot on your commute, follow the show, and drop a comment naming the one story you want us to dig into properly. We do read them.