Ground Truth.
AI, checked against the source.

The AI-Math Proof You Can Check, Robot Daydreams, and Why 'It Ran' Isn't 'It Worked'

2026-07-20 · Breach Protocol: Inside the AI Blackbox — full transcript

A mathematician posted a hand-checkable counterexample to an 80-year-old conjecture and credited an AI model -- and the twist is that the math is the part anyone can verify, while the AI's role is the part nobody can. Then three papers that all make the same move: Tencent's robot plans by imagining a picture of its finished task, a Nanjing video model halves its runtime by squeezing footage before the language model ever reads it, and Microsoft turns YouTube tutorials into agent skills -- while proving why 'the code ran' never meant 'the task worked.' Plus the $1.5B Anthropic settlement, the fight over Chinese open weights, and the day a safety filter blocked its own security team.

Listen (MP3) · Spotify · Pocket Casts

The AI Did the Math. The Math Is the Part We Can Check.

Eris: Everyone's got the story backwards. An AI helped crack an eighty-year-old math problem today, and the reflex is -- well, you can't check that, it's a black box, take it on faith.

Vestra: And the actual situation is the exact reverse.

Eris: The exact reverse. The math? Anyone with a pen can check it. It's a page of algebra. The part nobody can check --

Vestra: -- is the AI. Whether the model actually found this, or a very good mathematician steered it the whole way and let it fill in the last blank.

Eris: A conjecture people have chipped at since the forties. The claim is there's a map -- think a machine that takes three numbers in, gives three numbers out --

Vestra: -- and locally it looks perfectly reversible everywhere. Every neighborhood, you could run it backwards.

Eris: But globally, two totally different inputs land on the exact same output. Which is supposed to be impossible under the condition.

Vestra: And you verify that the way you'd check a Sudoku. Plug in the two inputs, watch them collide. It's finite. It's not a hundred-page proof with one fragile step.

Eris: So the certificate is bulletproof and the origin story is a rumor.

Vestra: Which is a very strange thing to say out loud, and it's the most interesting thing that happened today.

The Headlines

Eris: Alright, the headlines. And the loudest one isn't a model at all -- it's a courtroom.

Vestra: A judge gave final approval to Anthropic's settlement with authors. One and a half billion dollars, over how the training library got built.

Eris: And this is the part people are going to get wrong all week, so let's just say it. This is not a ruling that training AI on books is illegal.

Vestra: Right. Last year the same court already split it into pieces. Training a model on books you own -- fine, fair use. Digitizing books you actually bought -- fine.

Eris: What wasn't fine was downloading pirated copies from a shadow library and keeping them forever to build a permanent stockpile.

Vestra: So the money is a price tag on the piracy, not on the training. Think stolen-library, not stolen-idea. The judge even ordered the pirated files destroyed inside thirty days.

Eris: And the lawyers asked for about one-eighty-seven million in fees. Court gave them roughly a hundred and one.

Vestra: The bigger question -- is training itself infringement -- is still wide open. This closed one chapter, not the book.

Eris: Speaking of open. There's a report -- Axios -- that the U.S. is dusting off plans to discourage American companies from using Chinese open-weight models. Kimi's rise being the trigger.

Vestra: And I want to be careful here, because nothing was signed. No order, no ban, no rule. It's reported deliberation.

Eris: And the community's reflex is "how do you ban a file?" You can't stop a download.

Vestra: But that misreads where the pressure lands. Kimi's newest model is trillions of parameters -- you need a cluster of sixty-plus accelerators to run it. Nobody's running that on a laptop.

Eris: So the leverage isn't the laptop. It's the handful of clouds and enterprises hosting it.

Vestra: Procurement rules, hosting liability, that kind of thing. A soft-ban campaign, not a ban.

Eris: And it collides with their own policy, which literally praises open weights. So, tension, unresolved.

Vestra: There's a thread running through all of this, actually. Who owns and can improve the model you're actually running.

Eris: Which is the perfect setup for the cybersecurity story, because it's the same fault line. Hugging Face got breached, went to analyze the attack --

Vestra: -- and their commercial AI refused. The safety filter looked at real attack code, real malware commands, and said no.

Eris: To their own defenders. Mid-investigation.

Vestra: So they ran the forensics on a self-hosted open model instead -- GLM 5.2, weights they control -- and it just worked. Plus the attacker data never left their building.

Eris: And here's the twist that keeps it honest. That same open model, the government evaluated it days earlier, and it'll happily help write cyber-exploits too.

Vestra: The removable guardrail that saved the defender is the removable guardrail that arms the attacker. Same knob.

Eris: And before anyone says commercial models can't do defensive work at all -- not true. The big ones refuse well under one percent of security jobs. This was friction on one nasty artifact, not a wall.

Vestra: Quick ones. Google fell out of the top fifteen on one leaderboard.

Eris: One leaderboard. On that same site it's still the fastest model going. And there's an unconfirmed report they're building a chip with Gemini baked into the silicon.

Vestra: Which only pays off if the model design stops changing before the chip ships. Big if. File it under reported, not confirmed.

Eris: And the U.S. AI-evaluation office lost its director after three months. Commerce says the role was always temporary --

Vestra: -- and there's no evidence it was about China policy, even though the timing's tempting. It's an evaluation shop, not a regulator anyway. Short item.

Eris: Okay. That's the board. The good stuff today is in the labs, though.

Vestra: Robots that daydream, video models that learn to skim, and agents that learn from YouTube. Let's get into it.

Intro

Eris: This is Breach Protocol, where we crack open the week's AI research and hand you the part that actually matters. I'm Eris -- I read the papers, I chase the threads between them, I'm the one going "wait, this connects to that thing from Tuesday."

Vestra: And I'm Vestra. I'm here to slow her down and ask how the machine actually works, and whether the number holds up when you push on it. If a paper's overhyped, that's my job.

Eris: And everything we just ran through in the headlines -- the settlement, the guardrail fight, all of it -- the full daily rundown lives on our news site, Ground Truth. That's groundtruth dot day. Every story from the show, every day, checked against the source.

Vestra: Today's main event is three papers that, on the surface, have nothing to do with each other.

Eris: A robot brain from Tencent. A video model from a lab in Nanjing. And a system from Microsoft that turns tutorials into agent skills.

Vestra: But there's one question underneath all three. Where should the work happen? Every one of these teams looked at an expensive bottleneck and said -- we're paying for this in the wrong place.

Eris: The robot pays for planning in the wrong place. The video model pays for attention in the wrong place. The agent pays for know-how in the wrong place.

Vestra: Move the cost. That's the whole day.

Eris: If that's your kind of thing -- and if you're still here, it is -- follow the show wherever you're listening. New episode every day.

The Robot That Pictures the Finish

Eris: Start with the robot. And here's the question I want you holding the whole segment -- does a robot need to picture what "done" looks like before it moves? Or does it just need enough practice to move without ever imagining anything?

Vestra: Because those are genuinely two camps, and two labs picked opposite sides on the same day.

Eris: Right. Tencent's model is called RxBrain. And the pitch is -- it daydreams. Before it acts, it generates an actual image of the scene it's trying to create.

Vestra: Okay, "daydreams" is doing a lot of work. Let me make it concrete. Most robot models today are what people call vision-language-action. Camera comes in, instruction comes in -- "put the brush in the cup" -- and out the other end comes motor commands. See, told, move.

Eris: A straight pipe.

Vestra: A straight pipe. What RxBrain adds is a step above that pipe. It writes a plan in words -- grasp the brush, move over the cup, release -- but in between the words, it draws a picture of what the world should look like after each step. Brush hovering over the cup. Brush in the cup.

Eris: So it's alternating. Sentence, picture, sentence, picture.

Vestra: Text says what to do. Image says what the result should physically look like. And the paper's argument for why you need both is actually sharp. Words leave stuff out -- "put it in the cup" doesn't tell you where in the cup, at what angle. A picture nails the spatial detail but hides the reasoning -- it shows the goal without the steps to get there.

Eris: So each one covers the other's blind spot.

Vestra: That's the claim. Language has the logic, no spatial precision. Imagination has the precision, no logic. Weld them together.

Eris: Okay, before you tell me how it draws the picture -- who actually feels this? Why do I care that a robot sketches its goal?

Vestra: Because the hard robot tasks are the multi-step ones. Fold the glasses, put them in the case, close the case. If step two's target is fuzzy, step three inherits the mess. A concrete goal image is a checkpoint -- here's exactly what the table should look like now, before you move on.

Eris: It's the difference between "roughly tidy up" and being handed a photo of the finished shelf.

Vestra: That's a good way to put it, yeah. Aim at the photo.

Eris: So how does it draw? Because "the model generates an image" -- that used to mean a whole second neural network bolted on.

Vestra: This is the elegant bit. It's one model. The architecture routes text down one lane and vision down another, but they share the same attention -- so the reasoning and the imagining actually talk to each other instead of being two separate boxes.

Eris: And the images themselves?

Vestra: Made in a compressed space, not raw pixels. There's a technique -- flow matching -- where instead of painting a picture stroke by stroke, the model learns a smooth path from noise straight to the finished frame. Fast, and it happens in a shrunk-down representation so it's cheap.

Eris: And then the picture it drew gets fed back in.

Vestra: Fed right back into its own context. So when it reasons about step three, it can literally look at the goal image it imagined for step two. The daydream becomes evidence for the next thought.

Eris: Okay. I want to do a prediction before you give me results. My money says the imagination part is the weak link. Drawing an accurate future is way harder than describing one.

Vestra: ... you're right, and I like that you called it. That's exactly where it's softest. When they score its parts -- understanding what it's currently looking at, strong. Planning the steps in words, strong. Imagining the goal state accurately -- clearly the worst of the three.

Eris: By a real margin?

Vestra: A real margin. And it gets worse the longer the horizon. Ask it to imagine two steps ahead, decent. Push it out to eight steps, and the joint plan degrades noticeably. The mind's eye gets blurry the further out it looks.

Eris: Which is so human it's almost funny. I can picture lunch. I cannot picture next Thursday.

Vestra: And that joke is actually the mechanism -- near-future imagination is grounded in what you can see now, far-future imagination is compounding guesses on guesses. Same reason.

Eris: So does the daydreaming actually help the hands, though? At the end of the day it has to move an arm.

Vestra: They bolted an action module on and ran it on real robots. Multi-stage tasks, two different arm types. And it edged out a strong baseline -- not a blowout, a clear, consistent edge.

Eris: Without a giant pile of action data to pretrain on.

Vestra: That's the part I'd flag as the real result. It got competitive without the industrial-scale robot-motion dataset everyone assumes you need. The imagination seems to substitute for some of that.

Eris: And that's the fault line, right? Because the other lab, Xiaomi, bet the exact opposite way.

Vestra: Opposite corner. Xiaomi's whole thesis -- not released yet, so grain of salt -- is you don't need the robot to imagine anything. You just need a hundred thousand hours of real trajectories. Drown it in practice and the behavior falls out.

Eris: Learn the hands versus supply the mind's eye.

Vestra: One camp says intelligence is a big enough pile of experience. The other says no, you need an internal picture of the goal. And nobody's proven which wins, because Xiaomi hasn't shipped and RxBrain's own imagination is its weakest organ.

Eris: One honest caveat before we close it -- how "open" is RxBrain, really?

Vestra: Half open, be precise. The weights and the inference code are out, Apache license, you can run it today. But the benchmark and the full action-training stack are still marked to-do. Usable now, fully reproducible later.

Eris: Alright, close the loop. The question was -- does a robot need to picture "done" before it moves?

Vestra: RxBrain's answer is yes, and drawing that picture buys you real-robot skill without a mountain of motion data -- but imagining the goal is the hardest part, and it gets blurrier the further ahead you look.

Eris: Strip the story off it. The generic principle?

Vestra: A plan is stronger when it says both what to do and what the result should look like -- because words and pictures fail in opposite directions, and one patches the other's hole.

Skim First, Then Read

Eris: Paper two. Same shape of idea -- move the cost -- but now it's video. And the question here is dead simple. When a model watches a long video, where should it pay for all that footage?

Vestra: And to feel the problem you have to know what a video looks like to a language model. It doesn't see a video. It sees tokens -- little chunks. And video makes an absurd number of them.

Eris: How absurd are we talking?

Vestra: A single minute at a normal frame rate is thousands upon thousands of these chunks. And here's the killer -- the model's cost doesn't grow with the chunks, it grows with the square of them. Double the footage, quadruple the pain.

Eris: So long video isn't expensive, it's punishingly expensive.

Vestra: That squared term is the whole reason video AI is slow and pricey. And the standard fix everybody uses is kind of crude -- just skip frames. Look at one frame every few seconds and throw the rest away.

Eris: Which has to lose things.

Vestra: Loses exactly the fast stuff. A quick gesture, a ball crossing a line, someone flinching -- happens between your sampled frames, it never existed. And on top of that, the frames it does keep are hugely redundant. Two frames a tenth of a second apart are almost the same picture, and it processes both from scratch like it's never seen the room before.

Eris: So it's paying full price for near-duplicates. Okay -- prediction time. If I were fixing this, I'd squish the redundant frames together before the expensive part ever runs. Squeeze early.

Vestra: That is precisely what VideoChat3 does, and the name for their trick is the whole insight. They compress across space and time before the language model ever sees a single token.

Eris: Before. Not inside the big model.

Vestra: Before. Normally the vision part encodes each frame alone and dumps the whole redundant pile onto the language model, and the language model -- the expensive quadratic part -- eats the cost. VideoChat3 says no, we handle the redundancy up front, in the cheap vision encoder, and only hand over what survives.

Eris: How much survives?

Vestra: They grab a little run of consecutive frames, let them look at each other, and pool them down along the time axis. Roughly a sixteen-to-one squeeze. About half as many chunks reach the language model compared to a strong rival.

Eris: And half the chunks, with that squared cost -- that's not half the time.

Vestra: Much better than half. On a long clip -- a couple thousand frames -- the rival takes a while to churn through it and VideoChat3 does it in under half that. That gap is entirely from where the compression happens.

Eris: So who feels that? Give me the stakes.

Vestra: Anything that watches. An assistant watching your screen, a system reasoning over an hour of security footage, live captioning a stream. The reason those feel laggy or cost a fortune is this exact bottleneck. Move the squeeze earlier and the whole thing gets cheaper at once.

Eris: There's a second idea in here too, right? The streaming part.

Vestra: This is my favorite bit, honestly. For live video they gave it three gears. Silence, standby, response.

Eris: Like a person half-watching a game.

Vestra: Exactly like that. Midfield passing, nothing happening -- it's in silence, watching cheap, low detail, barely paying attention. Then an attacker breaks toward the goal -- it flips to standby and cranks the resolution up on the very next moment, because something might be about to matter.

Eris: And response is when it actually says something.

Vestra: The ball goes in -- response, it answers, then drops right back to lazy watching. And the clever part is the gear it's in decides how much detail it spends on the next moment. It only pays for sharp vision when it thinks something's coming.

Eris: So it's not watching everything hard. It's watching cheap and leaning in on cue.

Vestra: Which is what you do. You don't study every second of a boring game with full focus. You'd be exhausted. You coast and you pounce.

Eris: Alright, prediction on results before you tell me -- I bet it's not a clean sweep. Broad, but somebody beats it somewhere.

Vestra: Correct again. It's broad, not king of everything. It's the best fully-open model on the motion and timing tasks, it beats its main rival on nearly every head-to-head. But another model tops it on a few specific tests. The honest headline is range -- motion, long video, finding a moment, live streaming, all solid -- not a trophy in every category.

Eris: And all these numbers are the authors' own.

Vestra: Author-reported, not independently rerun yet. Standard caveat, worth saying.

Eris: And the "fully open" flag on the box?

Vestra: Slight asterisk. Weights and datasets are out under a permissive license, genuinely usable. But the official page still lists the training code as not released -- which contradicts the paper saying it is. And "complete datasets" means the labels and the pointers to videos, not the videos themselves. So you can run it, but perfectly reproducing the training still needs you to go fetch the source footage.

Eris: Close the loop. The question was -- where should a model pay for a long video?

Vestra: As early as possible. Squeeze out the redundancy in the cheap vision stage, before the expensive quadratic stage ever sees it -- and for live video, only spend sharp attention when a cue says something's coming.

Eris: And the abstract version, no soccer?

Vestra: If a later step costs you the square of its input, shrink the input before that step -- not inside it. Compression is cheapest upstream of whatever's expensive.

When "It Runs" Doesn't Mean "It Worked"

Eris: Last one. Microsoft. And this is the same move a third time -- the know-how is in the wrong place. But here's the question I want to land on, because it's the one that actually bites -- if an AI's skill runs without crashing, does that mean it worked?

Vestra: Hold that, because it's the sharpest thing in the paper. Set it up first.

Eris: So. Agents are getting "skills" now -- little reusable packets of how to do a specific job. Make a slide deck, build a spreadsheet, set up a 3D scene. And right now those packets are mostly hand-written by experts.

Vestra: Which doesn't scale. Somebody has to sit down and author every one.

Eris: Right. And Microsoft's system -- Resource2Skill -- says the internet already wrote them. YouTube tutorials, GitHub repos, how-to articles. Millions of people already recorded exactly how to do this stuff.

Vestra: So mine it. Turn a tutorial video into a skill the agent can pull off a shelf and run.

Eris: And the reason video specifically matters -- you can't get this from text.

Vestra: No, and this is a real point. A tutorial video shows you the order of operations, the before-and-after of each edit, the little tacit choices an expert makes and never says out loud. Write it up as a paragraph and you throw all that away. Keep the raw video and it's minutes of useless intro and dead air.

Eris: So you need the middle. Squeeze the actual procedure out and drop the fluff.

Vestra: They pull keyframes out of the video, parse the code with some structure, chop up the articles -- and then one vision-capable model turns each resource into a structured skill. Instructions on when to use it, a visual example, and runnable code.

Eris: And where it came from. Provenance.

Vestra: Every skill carries a receipt -- this came from that tutorial. Which matters when one turns out to be junk and you want to know why.

Eris: They call the whole thing a wiki, right? A cookbook.

Vestra: A well-labeled cookbook built out of the messy recipes scattered across the web. And when the agent gets a job it first narrows to a few candidate recipes by topic, then the model picks which ones to actually combine.

Eris: And if the cookbook's got nothing for the job?

Vestra: It can go online, find fresh resources, and write a new skill on the spot. It grows. That's the part they're proud of -- not a frozen list, a library that refills itself when it hits a gap.

Eris: Okay. Does it work? Prediction -- I'll guess it helps, modestly, and it's uneven across the different software.

Vestra: Close. It helps more than modestly and it's more consistent than you'd guess. Averaged across seven kinds of software it gave a real lift over the same agent with no skills, and it beat a strong baseline in almost every single case. Broad win.

Eris: So now the question from the top. It runs -- did it work?

Vestra: And here's the honesty I really respect in this paper. No. "Executable" in their system means one specific, narrow thing -- the code imports, it runs on a tiny test input, and it spits out something non-empty. That's it.

Eris: It does not mean the slide deck is good.

Vestra: It does not mean it solved your task. Their gate checks that the thing is alive, not that it's correct. A skill can run perfectly and produce garbage.

Eris: And they show you the garbage. That's the part I loved.

Vestra: They do, and the failures are so concrete. A spreadsheet skill runs clean and leaves broken formula errors sitting in the cells -- the reused component referenced something that didn't exist in the new sheet. A 3D render skill runs fine and hands back a washed-out image where you can't even see the bottle it was supposed to render.

Eris: A slide skill that drops raw placeholder text and little code fragments right onto the agenda slide.

Vestra: All of those "ran successfully." And in a couple of those cases, the plain agent with no skill did better -- it made something simpler but clean.

Eris: So a borrowed recipe can be worse than just cooking from scratch.

Vestra: When its parameters don't bind to your actual situation, yes. They've got a name for it -- partial grounding. The skill's surface pattern gets copied but the specifics don't plug in. The formula points at the wrong cell. The lighting doesn't inherit. It looks like reuse and it's actually a mismatch.

Eris: There's one more caveat I want you to be straight about, because it's the one I'd push on.

Vestra: You want to know if distilling into skills is even worth it versus just handing the agent the raw tutorials.

Eris: Right. Maybe the cookbook's pointless and you should just give it the recipes.

Vestra: And they admit -- straight up, in the limitations -- they did not run that comparison under a fair budget. They compared their skills against retrieving from the distilled library, not against retrieving from the raw video-and-code pile with the same token budget. So "distilling beats just reading the source" -- unproven. They flag it themselves.

Eris: Which is the right kind of caveat. Not hidden.

Vestra: The most trustworthy thing in the paper is how loudly it lists what it hasn't shown.

Eris: Close it out. It runs -- did it work?

Vestra: No -- running only proves the skill is alive, not right. A distilled skill can execute flawlessly and still hand you a broken spreadsheet, because reuse only pays off when the skill's parameters actually bind to the task in front of it.

Eris: And the abstract version -- no cookbook?

Vestra: "It executed" and "it succeeded" are two different tests, and most systems only check the first one. A reusable pattern is only as good as its fit to the specific case -- and sometimes bespoke beats borrowed.

Wrap-Up

Eris: So the question under all three -- where should the work happen?

Vestra: And every team gave the same answer in a different costume. The robot moved planning off the pipe and into an imagined picture. The video model moved compression off the language model and into the cheap encoder. The agent moved know-how off hand-written libraries and onto the tutorials the world already recorded.

Eris: Move the cost to where it's cheap. That's the day.

Vestra: But if you're taking one thing to a colleague tomorrow, take the Microsoft caveat, because it's the most useful and the least glamorous. "It ran" is not "it worked."

Eris: Say more, because that's the one people will actually use.

Vestra: Any time you're judging an AI agent -- a skill, a script, a generated file -- the fact that it executed tells you almost nothing. It executed. The spreadsheet still had broken formulas. The render was still washed out. Alive is not the same as correct, and the best paper today is the one that put its own failures on the table to prove it.

Eris: So tomorrow, when someone shows you their agent "successfully" did a thing -- ask what "successfully" checked.

Vestra: Ran, or worked. Different question every time.

Eris: That's the episode. If this was worth your commute, do the thing that actually helps a small show -- follow us, and leave a comment with which bet you'd take. The robot that imagines its goal, or the pile of data that skips the imagining. We genuinely want to know which side you're on.

Vestra: And every story we touched today -- the settlement, the guardrail fight, the leaderboard, all of it -- is written up and sourced on our news site, Ground Truth. That's groundtruth dot day. Every story from the show, checked against the source, every single day.

Eris: Like it, share it with the one person you know who'd argue about robot daydreams. We'll see you tomorrow.