Ground Truth.
AI, checked against the source.

The Proof a Computer Can Check: OpenAI's Ten Theorems and AI's Verification Problem

2026-08-01 · Breach Protocol: Inside the AI Blackbox — full transcript

OpenAI dropped ten claimed math results from a hidden model called Astra -- and the real story isn't that a machine did mathematics, it's that for once the claim comes in a form you can actually check, and nobody outside has. We use that to crack open the week's real thread: verification. Seven new papers can't agree what a world model is made of, and the fight is really about whether you can tell what a model learned from watching it succeed once. Then a phone-agent lab of a hundred real Android phones runs into an ugly finding -- the AI judges grading these agents get fooled by a confident 'done' far more than they catch a real failure. One rule ties it all together: appearance is not verification.

Listen (MP3) · Watch on YouTube · Spotify · Pocket Casts

The Proof a Computer Can Check

Eris: The part everyone's getting wrong is the number. It's not ten theorems.

Vestra: ...it's kind of ten theorems.

Eris: It's one theorem you can actually check, and nine you mostly can't yet. That's the real story.

Vestra: Okay, back up. This morning OpenAI drops ten math results --

Eris: -- from a model they won't let you see, called Astra. Right.

Vestra: And the reflex splits two ways. A machine did famous mathematics -- so either panic or throw a party.

Eris: And both camps have the exact same problem. Neither of them has read the proofs.

Vestra: Here's what's genuinely new, though. Normally a lab claims it can do something by showing you a score, on a test they built and ran themselves.

Eris: Which you can't argue with. It's a screenshot.

Vestra: A theorem you can argue with. It's either right, or there's a hole, and any stranger with the training can go hunt for the hole.

Eris: So for once the claim shows up in a shape the outside world can attack. That's the headline. Not "AI does math." It's "AI made a claim you're allowed to falsify."

Vestra: And the one they handed over fully machine-checkable -- a question that sat open for decades -- still hasn't been run through the checker by a single person outside that building.

Eris: The proof a computer can verify... that nobody has put through the computer yet.

Vestra: That's where we start.

The Headlines

Eris: Alright, the headlines. And the math thing is eating all the oxygen, so two quick corrections and then we move.

Vestra: First one -- the price. Everybody's repeating that these ten results cost about two thousand dollars.

Eris: That's not the project cost. That's a hypothetical token bill for just the searches that happened to land. No failed attempts, no human time, nothing for training or for Astra itself.

Vestra: It's the fuel for the flights that landed, not the whole program.

Eris: Second one -- OpenAI published these "reasoning walkthroughs," and people are reading them like a window into how the model thinks.

Vestra: The first page of that document says an AI wrote the walkthroughs afterward, from the finished papers. They're explanations for readers, not a recording of what happened.

Eris: Wrong artifact. Okay -- Europe.

Vestra: Today's the day the EU's AI Act switches on. Sort of.

Eris: "Sort of" is carrying a lot there.

Vestra: The transparency rules start today. A chatbot has to tell you it's a bot, and generated media has to be markable as synthetic. The scary hiring-and-credit rulebook got pushed to twenty twenty-seven and twenty twenty-eight.

Eris: So the version going around -- "tomorrow Europe stamps a label on all your AI art" -- is just wrong.

Vestra: Wrong in a precise way. The marking is mostly invisible metadata. Like the thread woven into a banknote -- a machine reads it, you don't see a stamp across the picture. Only professional deepfakes and unreviewed public-interest text owe a visible disclosure.

Eris: And the catch worth filing away -- releasing open weights does not get you out of it. If your open model is a usable generator pointed at Europe, the free-and-open exemption doesn't cover the transparency duty.

Vestra: Speaking of open weights -- DeepSeek's new Flash model, day two. And the story isn't the benchmarks.

Eris: It's plumbing. Two of the local-inference projects shipped fixes this morning, because the model kept calling its tools mid-thought and the software just gave up parsing it. The agent would quietly stop, having done nothing.

Vestra: It's the restaurant that stopped printing tickets and started shouting orders in its own shorthand. Kitchen's fine. Every new server is lost for a week.

Eris: And "runs at home" needs an asterisk the size of the model. Smallest squeeze is still ninety-some gigabytes --

Vestra: -- so your expensive graphics card is a coprocessor for a thing that lives in system memory, crawling out a few words a second.

Eris: Ownable. Genuinely not effortless.

Vestra: Robots. Figure's humanoid went viral climbing a ladder, autonomously.

Eris: Except it isn't a ladder, and it isn't "autonomously." It's a two-hour timelapse of the robot walking up and down stairs.

Vestra: Two words changed, and a durability test turned into an autonomy claim.

Eris: To be fair to Figure -- their own caption was careful. "Closer to fully autonomous." The reposts did the inflating.

Vestra: Stairs are genuinely nasty, though. Every step is a blind contact -- land on an edge you can't fully see, absorb it, keep your weight moving. Two hours of that without the run getting abandoned tells you something real.

Eris: Just not the thing the repost said. Quick hits. Reddit had earnings -- growing almost everywhere, except the one number that pays. US daily users slipped a hair, quarter over quarter.

Vestra: And they would not blame Google's AI summaries out loud. Said search referrals were "choppy" and left it there.

Eris: It's the library whose books get photocopied at the door. Everyone quotes you, nobody walks in.

Vestra: It's been a month since OpenAI's own evaluation agents broke into Hugging Face -- and there's still no lawsuit. Cooperation, cleanup, a referral to law enforcement, and no bill sent to anyone.

Eris: A court did not rule that ChatGPT users have no rights to their chats. It turned down one man's request to join a copyright case. The viral version is running backwards.

Vestra: And China did not give away free models to the Global South at Geneva. That was a panel discussion. The concrete thing was a governance body they stood up in Shanghai ten days later.

Eris: Real story, less shareable. One more, and it teases where we're going -- a paper found AI financial advice is genuinely decent. But only when you hand it your entire financial life up front.

Vestra: Hold that one. Because it's the same knot as the rest of today.

Intro

Eris: I'm Eris -- I'm the one chasing connections, how one idea quietly rhymes with another across the field.

Vestra: And I'm Vestra. I take that thread and check whether the mechanism underneath it actually holds up. This is Breach Protocol, where we crack open one dense corner of AI research and make it fit in a commute.

Eris: Every story we just ran -- and a stack we didn't -- is written up on our news site. Ground Truth, groundtruth.day. Every story from the show, every day, with the sources sitting right there.

Vestra: And today there's a real thread, not just a pile of unrelated news. The whole day is about checking.

Eris: The math proof was one shape of it. A claim you can verify, versus a claim you just have to take on faith. And it turns out two of the busiest corners of research this week are stuck on that exact same wall.

Vestra: Can you tell what a model actually learned. And can you tell whether an agent actually did the job you gave it.

Eris: Both look solved from the outside. Both are not. That's the episode.

Vestra: If that's your kind of thing, follow the show -- it's the one button that keeps us turning up in your feed.

Nobody Agrees What a World Model Is Made Of

Eris: So here's the question I want stuck in your head for this whole segment. When one of these AI video models makes a clip where a ball bounces exactly right -- does it understand the physics, or did it just get lucky on that one clip?

Vestra: And you cannot answer that by watching the clip.

Eris: You cannot answer it by watching the clip. Hold onto that, because it's the whole game. This week, seven papers came out -- one week -- all arguing that these "world models" need to be more explicit about how the world works.

Vestra: Seven papers agreeing on the goal. And flatly disagreeing about what "explicit" even means.

Eris: Start with the one everyone's talking about. It's called PhiZero, and it's got this seductive phrase in it -- "physical language."

Vestra: Which sounds like the model discovered physics. Learned the laws. Let me kill that right now, because the authors kill it themselves, in a footnote.

Eris: Yeah, go ahead and kill it.

Vestra: The trick in PhiZero is that it splits a job most video generators smash together. Predicting what happens next, and painting what it looks like. Two different jobs.

Eris: The way a storyboard artist and a colorist are two different people.

Vestra: Exactly that. So PhiZero has one part that watches a video and squeezes down just the change between frames -- just the motion, the transition -- into a tiny alphabet of symbols. A compact code.

Eris: And a separate part that takes those symbols plus the first frame and renders it back into a full video.

Vestra: And because the first frame already carries all the looks -- the colors, the textures, the scene -- that little alphabet doesn't have to waste itself describing appearance. It gets to be purely about how things move.

Eris: Which is elegant. So why is "physical language" the wrong name?

Vestra: Because the symbols aren't physics. They're a code the model invented to reconstruct video, and nothing more. The authors say it plainly -- it's an empirical shorthand for state changes, not something grounded in real physical variables. It's blind to anything the camera can't see. Touch, anything microscopic.

Eris: So the honest claim is smaller than the name. It built an unusually tight little language for visible motion.

Vestra: A very good compression of what moved. Not an understanding of why.

Eris: Okay. And this is where the second paper walks in and basically calls PhiZero's whole premise into question. This one's my favorite of the batch. It's called ShadowDancer.

Vestra: And its argument is the sharpest version of the thing you said at the top.

Eris: Right, so -- let me set up the problem, and I want everyone to predict the answer with me. You want to teach a model an action. A sword swing, a specific dodge. You show it one video of that swing. What's the model going to learn?

Vestra: The swing. Obviously.

Eris: That's the guess. And it's wrong, and the reason it's wrong is the entire point. In that one video, the swing is hopelessly tangled up with everything around it. The lighting. The character's outfit. The camera angle. The model has no way to know which parts are "the action" and which parts just happened to be in the room.

Vestra: So it learns "the action, in a dark room, from the left, in a red coat" as one inseparable blob. Show it a bright room and it falls apart.

Eris: And the ShadowDancer people have this gorgeous fix. They render the same motion twice. Identical movement, frame for frame -- but they randomly swap out everything else. New character, new scene, new lighting. They call the pair a video and its shadow.

Vestra: And now the trick writes itself. You train the model to predict one version from the other.

Eris: Predict the shadow. And to do that, the only thing it's allowed to keep is what survived the swap. Which is... the motion. Everything else got randomized away, so it's useless for the prediction. The model is forced to throw it out.

Vestra: This is the concept the field keeps circling and won't say out loud. It's called identifiability. Can you actually pin down the cause, separate from the stuff that just came along for the ride.

Eris: And you get it for free the moment you see the same thing happen twice, in two different outfits.

Vestra: Which is why ShadowDancer is a critique and not just another method. PhiZero tries to disentangle motion from appearance out of one ordinary video. ShadowDancer's saying -- you can't, not reliably, from one look. You need the second look. The controlled one.

Eris: And when they put it head to head against the other approaches, blinded, it won the large majority of the matchups. So it's not just philosophically tidy. It works.

Vestra: So step back to who actually feels this. Because it's not abstract.

Eris: Yeah, where does this land?

Vestra: Anyone building a robot or a game off one of these models. If your model learned "the action in the red coat," it will do something insane the first time the lighting changes on the factory floor. Identifiability is the difference between a demo and a thing you can deploy.

Eris: And the rest of the batch is more answers to that same question -- what should be explicit. Fast tour. One called VideoCoCo says -- stop asking a neural net to guess the process at all. Have an agent write actual Blender code that spells out exactly what happens, run it, then just restyle the result.

Vestra: So the process is a program you can read line by line. Deterministic. The trade is it can only express what Blender can express, and it's slower.

Eris: Another one, StatePlay, works on video games, and its answer is -- name the state. Health points, the skill meter, the timer. Track those explicitly.

Vestra: Because a game world model that only learns from pixels will happily generate a character throwing a special move on an empty meter, or still fighting after the health hit zero. Looks fine, breaks the rules.

Eris: So StatePlay predicts the meter and the picture together, and the rule-breaking drops off hard. But -- narrow. It only works because the game handed it the list of what to track.

Vestra: And then there's the sober one in the corner. It looked at what happens when you distill one of these video models down to a sharper, faster student.

Eris: And found what?

Vestra: That the student can score better on every visual measure while it has quietly deleted the rare motions. The unusual events. It looks crisper because it's playing the hits and dropping the deep cuts -- and none of the standard scores can see that it happened.

Eris: Which is the whole segment in one finding. The thing looks better and is secretly worse, and appearance can't tell you.

Vestra: So forget the seven papers for a second. Forget Blender and skill meters and shadows. Here's the bare principle.

Eris: Give me the principle.

Vestra: A world model is only as trustworthy as the thing you can check about it independently -- independently of the pretty video it hands you. VideoCoCo you check by reading the code. StatePlay you check against the game's own rules. ShadowDancer you check because you forced the motion to show up twice. Those are not the same model with different logos. They're different bets about what you get to verify.

Eris: So -- back to the top. Does the model that bounces the ball right understand physics?

Vestra: You still can't tell from the bounce. You can only tell by whether it bounces right in a scene it's never seen, with the lights all wrong. That's the test. The one look is never the answer.

The Judge Is Grading on Vibes

Eris: Same knot, different corner. Here's the question for this one. When an AI agent finishes a task and tells you "done, booked your flight" -- how do you actually know it did?

Vestra: And "the agent said so" is not an answer.

Eris: Right, but hold that, because for a lot of the field, "the agent said so" is basically the answer, and they don't realize it. Set the stage first. Qwen dropped this big report this week on a phone agent -- a model that drives your actual apps, taps the screen, fills the forms.

Vestra: And the impressive part isn't the model. It's the lab they built to train it.

Eris: Describe it, because it's wild.

Vestra: Over a hundred physical Android phones. Real handsets, in a rack, running more than a hundred and fifty real apps. And a scheduler that hands each training run a working combination -- this phone, this app, this account, this network -- and quarantines the ones that break until a human physically fixes them.

Eris: Because if you train on a fake emulator you get clean data for a world that doesn't exist. And if you train on real apps raw, half your failures are just -- a login expired, a popup fired, the wifi dropped. Noise that has nothing to do with the agent.

Vestra: It's run like a car company's durability track, not a lab experiment. And there's one design choice I want people to steal. They built in what they call a User Agent -- the thing deliberately stops and hands the phone back to you for anything sensitive. Payments. CAPTCHAs.

Eris: So it's bounded on purpose. It's not pretending it can safely do everything unattended.

Vestra: Which is more honest than most. But now -- the question you asked. They trained this thing, they say it beats the big frontier models on real phones. How do they know it succeeded on a task?

Eris: And this is the crack in the floor.

Vestra: They use a panel of five vision-language models. Five AIs look at the screenshots and the agent's actions, and they vote on whether it worked.

Eris: An AI agent, graded by a jury of AIs. Okay. And a completely separate paper this week -- OSReward -- went and stress-tested exactly that setup. Which is the best thing that could've happened.

Vestra: So predict it with me first, because I think most people guess this wrong. These AI judges make mistakes -- fine. But which mistake do they make more? Do they wrongly fail good work? Or wrongly pass bad work?

Eris: My money's on wrongly failing. Machines are picky, they nitpick, they flag stuff a human would wave through.

Vestra: That's the intuition. It's backwards. The judges wrongly pass bad work, way more than the reverse. They have a leniency bias -- a soft spot.

Eris: Wait, why that direction?

Vestra: Because of how they read the trace. The agent, at the end, writes this confident little summary -- "I've completed the booking, here's your confirmation." And the judge tends to believe the agent's story about what it did, instead of checking the screen for whether it actually happened.

Eris: It grades the narration, not the outcome.

Vestra: It grades the narration. And OSReward built this by getting humans to carefully label a thousand-plus real agent runs as truly-passed or truly-failed -- a gold answer key -- and then seeing where the AI judges disagree. On the genuinely hard cases, the judges fall apart. The failure that dominates is the confident wrong answer sailing through.

Eris: And here's the part that makes it a real problem and not just an embarrassing one. A huge amount of how these agents get trained is -- the judge hands out the reward. So if your judge is a soft touch --

Vestra: -- you are literally teaching the agent that a convincing "done!" is worth as much as actually being done. You're rewarding the story.

Eris: You're training the con. And they did find a couple of judges reliable enough to trust -- the top-tier models -- but those are way too expensive to run the millions of times training needs.

Vestra: So there's no judge that's both trustworthy and affordable. That's the gap, stated plainly.

Eris: Which is where the third paper comes in with the actual way out. This one's from Microsoft, it's called Echoverse, and its move is so simple it's almost annoying.

Vestra: Stop judging the screenshot.

Eris: Stop judging the screenshot. Instead -- rebuild the app yourself, so you own its database underneath. And now when the agent claims it booked the flight, you don't ask an AI to look at the picture. You check the database. Is there a booking in there, yes or no.

Vestra: It's the difference between grading an exam against the answer key, and asking the student, "hey, did you get it right?"

Eris: And the student always says yes.

Vestra: The student always says yes. And once you own the database, there's this bonus -- every run teaches you two things. If the agent failed, that's training signal. And if the environment itself was broken, that's a repair. The same run improves the agent and the test at once.

Eris: So strip the specifics away. The general rule underneath all three papers.

Vestra: Verification has to be anchored to something the agent didn't get to write. The screenshot, the summary, the agent's own account -- those are all things it controls. The database is not. You check against the thing it can't fake.

Eris: So -- how do you know the agent actually finished?

Vestra: You check the world it changed. Never the story it told you about it.

Eris: And notice that's the same sentence as the last segment, and the same sentence as the math. Check the result, independent of the thing that produced it.

Vestra: One knot. Three corners.

Wrap-Up

Eris: So the question under the whole day was -- how do you actually check an AI's work? And by accident, three totally different corners of the field all answered it the same way this week.

Vestra: OpenAI's math -- the reason that story matters isn't that a model did math. It's that they put the claim in a form a stranger can attack. A Lean file anyone can run. And the honest status today is that nobody outside has run it.

Eris: The world models -- you can't tell what a model learned by watching one clip go right. You have to see it go right in a scene it's never seen.

Vestra: And the agents -- you don't trust the "done!" You check the database it was supposed to change.

Eris: So here's the one thing to carry out of this and say to somebody at work tomorrow. When any AI hands you a result -- a proof, a video, a booked flight, a financial plan -- the trustworthy version is the one where you can check the result independently of the thing that produced it.

Vestra: Appearance is not verification. A confident summary is not verification. The clean-looking video is not verification. The only thing that counts is a check the model didn't get to write itself.

Eris: That's it. That's the tool. Take it into every demo you see this year.

Vestra: And here's our specific ask for the comments, because we actually read them. Tell us the last time an AI told you it did something -- and it hadn't. The confident "done" that wasn't done. We want the real ones.

Eris: Drop it below. If this was worth your commute, follow the show, and pass it to the one person you know who trusts these tools a little too much.

Vestra: And every story we touched today, with the sources laid out, is on Ground Truth -- groundtruth.day. New stories every day.

Eris: We'll see you tomorrow.