Ground Truth.
AI, checked against the source.

AI News Today, Sep 8: OpenAI claims $1M math proof, mathematician cries foul

2026-09-08 · Breach Protocol: Inside the AI Blackbox — full transcript

OpenAI says 10,000 agents resolved the Navier-Stokes Millennium Problem in 88 hours -- and NYU's Tristan Buckmaster publishes a signed statement alleging OpenAI asked him to drop his Anthropic coauthor. Google documents attackers running fully autonomous AI agents and 100-million-prompt campaigns to copy its models. Plus Mistral's 3 billion euro raise with Luxembourg on the cap table, DeepMind's atlas of all 9 billion DNA letter changes, and the AI 2027 authors grading their own forecast at three-quarters speed.

Listen (MP3) · Watch on YouTube · Spotify · Pocket Casts

OpenAI says an AI proved a million-dollar math problem

Eris: Ten thousand AI agents, eighty-eight hours, and at the end of it OpenAI is claiming one of the seven Millennium Prize problems. The fluid equations. Navier-Stokes.

Vestra: Claiming being the operative word. No independent mathematician has checked any of it yet.

Eris: They published the full proof and a machine-checkable version anyone can run. That part is real whether or not the math survives.

Vestra: And they say they don't even want the million dollars, which is its own kind of strange.

Eris: Oh, it gets so much stranger than that. Start with the proof itself.

OpenAI says an internal model resolved the Navier-Stokes Millennium Problem

Eris: So the claim. The Navier-Stokes equations are basically Newton's laws written for fluids -- water, air, blood. They underpin aircraft design and weather forecasting. And for about ninety years nobody could answer one question: can a perfectly smooth, well-behaved flow work itself up to infinite speed in a finite amount of time?

Vestra: Which sounds abstract until you see the stakes. If the answer is yes, the equations stop describing reality at that instant. The smooth picture of a fluid destroys itself, and you'd have to go back to tracking individual molecules.

Eris: OpenAI says the answer is yes. Their own description of the solution is genuinely vivid -- a vortex that spirals inward and gets stretched out like spaghetti, speeding up as it shrinks, until the velocity runs away to infinity.

Vestra: And the mathematically brutal part is that everything has to blow up at once while cancelling precisely. Acceleration, pressure, momentum transfer, viscosity -- all growing without bound, while the force pushing from outside stays perfectly smooth the whole time. That is the needle they claim to have threaded.

Eris: How they got there is its own headline. Roughly ten thousand agents running in parallel for eighty-eight hours, trading close to three million messages on this one problem. They pointed the swarm at every open Millennium problem at once, cracked an easier warm-up version of the question first, then fed that result back in as a hint.

Vestra: A research department with no sleep and no ego, where the manager's whole job is reading everyone's notes and telling the next shift which three ideas to chase.

Eris: And they published everything. A written proof of about a hundred pages, plus a version formalized in Lean -- a proof assistant, software that mechanically verifies every logical step -- sitting in a public repository under an open license. Anyone with a machine and patience can try to check it.

Vestra: Now the correction, because most of the coverage got this wrong. The claim going around is that this doesn't count for the prize, because they solved the version where a smooth outside force is allowed to push on the fluid. That is not what the official rules say.

Eris: You actually read the rules.

Vestra: The Clay Institute's problem statement explicitly asks for a proof of any one of four statements -- its own words are that it gives reasonable leeway to solvers. Two of those four statements are exactly this: breakdown with a smooth external force. The forced route is a legitimate road to the prize, in writing, since the problem was posed. The prize is unclaimed for one reason only -- OpenAI said it doesn't intend to claim it.

Eris: Which is a fascinating choice in itself. They're framing the whole thing as a report on the pace of AI progress, not a prize submission.

Vestra: Here is what has not happened yet, though. No peer review. Nobody outside OpenAI has publicly built the Lean project. And there's a subtle failure mode people miss: a proof assistant will happily certify an airtight proof of the wrong statement. Until someone qualified confirms the formalized theorem is actually the theorem mathematicians care about, this is a claim with very good paperwork.

Eris: Very good, falsifiable paperwork -- which is more than these announcements usually ship with. But the timing gets a lot more complicated, because the humans who were working this exact route have a very different story about this week.

NYU mathematician says OpenAI asked him to drop his Anthropic coauthor

Eris: Tristan Buckmaster is a mathematics professor at NYU's Courant Institute, and on the same day as OpenAI's announcement he published a signed statement. The short version: he says OpenAI twice proposed removing his collaborator from authorship of a paper, and that when he said he would go public, the reply was "Why would you ruin your career?"

Vestra: Context first, because it matters. Buckmaster and Levent Alpoge -- a mathematician who works at Anthropic -- had spent about a year on this exact problem, using language models heavily along the way. They got their breakthrough in mid-August and had it machine-verified a week later. They were quietly working the same forced route OpenAI just announced.

Eris: Which is why, when he was told on a Sunday call that an internal OpenAI model had proved the forced version, his words are that it was a bright red flag. Almost nobody works that direction, and it is not where you land in a few days from a cold start.

Vestra: And by his account, the story shifted during the call. He was first told the model was simply handed the problem statement, very little human input. Then, as colleagues fed corrections in over internal chat, it emerged that an entire team had been working on it, the model had been warmed up on easier problems, and even the prompt he was shown had itself been written by prompting another model.

Eris: Then the two proposals. The first was a joint same-week release. The second -- and this is the one that detonated -- was that Buckmaster alone write the paper, with his coauthor removed. He quotes the OpenAI side saying it would all be simple if only Levent didn't work at Anthropic.

Vestra: He declined both. And there are two places where the published accounts flatly contradict each other. OpenAI's post says "we reached out to them." Buckmaster reproduces the email showing he contacted OpenAI first, on September 3. Then the data question: they had put every draft of the whole project into OpenAI's coding tools. OpenAI says no user data was accessed, but that it cannot rule out that de-identified data helped improve its models. Buckmaster says he asked directly, twice, whether the model was trained on their sessions -- and never got an answer.

Eris: For the other side: OpenAI's Sebastien Bubeck has responded publicly that the allegations are false and inflammatory, and says he came into the discussion following academic norms. He hasn't yet addressed any of the specifics.

Vestra: So what we actually have is one side's signed, contemporaneous account of private calls, disputed in general terms by the other side, with no recording or transcript published by anyone. That is the honest state of the evidence, and it's worth being precise about it.

Eris: What makes Buckmaster hard to dismiss is his restraint. He says outright: I have not seen their proof, I do not know what their model did, I am not accusing anyone of anything. And then he gives away the credit he could have kept -- he writes that the mathematician whose strategy underlies this whole program, Luis Martinez-Zoroa, deserves a Fields Medal.

Vestra: A man fighting to keep his coauthor's name on a paper while handing the biggest honor in mathematics to a third party. Whatever else turns out to be true about this week, that detail doesn't read like someone chasing glory.

Eris: He calls the moment a Deep Blue versus Kasparov moment and says the community should discuss it slowly. And I think that's right, because nobody has written these rules. What does authorship mean when one contributor is a model, and the compute belongs to a company that competes with your coauthor's employer? This is the first high-profile case, and it's the one that will get cited.

Terence Tao calls the AI-assisted fluid proofs a breakthrough

Eris: There is one voice in this whole mess that both sides cite and neither disputes, and it's Terence Tao -- probably the most authoritative living mathematician on fluid equations. The day BEFORE OpenAI's announcement, he wrote that the human team's results are a breakthrough, with a high likelihood of extending to Navier-Stokes.

Vestra: Sit with that timeline, because it cuts both ways. September 7, the field's leading expert says publicly that this method probably reaches Navier-Stokes. September 8, OpenAI says its agents reached it, having started September 1. Both things can be true -- a route an expert can see coming is a route a well-directed search can also find. But that proximity is exactly why the credit fight became a story instead of a footnote.

Eris: And Tao hands the credit to neither company. He traces the basic strategy to two mathematicians, Diego Cordoba and Luis Martinez-Zoroa -- the idea of building solutions out of ever-finer ripples layered on ripples, each one feeding energy down to the next, until the cascade runs away.

Vestra: The same two names Buckmaster elevates. Both accounts agree the core idea was human.

Eris: The aside that stuck with me is about formalization. Tao calls it remarkably feasible in the modern era -- machine-verifying a research-level fluid dynamics proof used to be a multi-year project for a team, and it now happens the same week as the proof itself. He also notes, a little dryly, that the authors had to rewrite what they themselves called the worst writeup they had ever seen.

Vestra: Which inverts the scarcity. Verification used to be the bottleneck. Now certificates arrive faster than anyone qualified can say what they mean. And Tao's standard is the one worth keeping: the primary goal is mathematical understanding. Without it, he says, even a problem this famous is of far less intrinsic significance. A certificate that something is true is worth much less than knowing why it's true.

Eris: His caveats are his own -- the three human results don't quite reach Navier-Stokes, and he's clear this is a first read, not a referee report. He says nothing about OpenAI or the allegations at all.

Vestra: But on the question of whether the mathematics underneath this week is real, his post is the closest thing to independent ground truth this story has. That's why everyone keeps pointing at it.

Google says attackers have moved from prompting to autonomous AI agents

Eris: Away from the math -- Google's threat intelligence team published what I think is the quiet biggest deal of the day. A financially motivated attacker used an ordinary AI coding chatbot to plan, build, and run a mass credential-theft campaign, end to end, in under six hours. Thousands of credentials compromised before anyone could blink.

Vestra: And the method is almost insultingly mundane, which is what makes it significant. No bespoke evil model, no exotic tooling. They wrote their playbooks as plain markdown files -- the same instruction format developers use to give coding agents standing orders -- and pointed a commercial chatbot at them. That alone was enough for the agent to run the scanning pipeline, troubleshoot its own breakage in real time, and rotate its addresses to dodge blocking.

Eris: The phrase Google leans on is human-in-the-loop latency. For three years the story about AI and cybercrime was better phishing emails -- a quality bump to one step of an attack. This is different in kind. Most defensive security assumes the attacker is a person who has to wake up, read scan output, decide what looks promising, fix the script that broke. Every one of those pauses is time a defender can use. Strip them out, and an attack that used to unfold over a week fits between two log reviews.

Vestra: It also changes who gets attacked. Attackers used to choose targets, because human attention was the scarce resource. An agent handed a playbook doesn't choose -- it enumerates. The organizations that were safe through obscurity just stopped being safe.

Eris: And this isn't only criminals chasing money. Google names state-linked clusters on the same curve -- groups tied to Russian military intelligence, Iran, China, North Korea -- all upgrading toward workflows that reason through tasks without human oversight.

Vestra: My caveats are real ones, though. Google doesn't name the tool that was abused, doesn't attribute the six-hour campaign to anyone specific, and it sells both a major AI platform and a major security business, so it has an interest in this framing. This is one vendor's visibility log, not an industry-wide measurement.

Eris: The direction isn't really contestable, though. And notice the rhyme with the top of the show. Ten thousand agents doing mathematics for eighty-eight hours, and one agent doing crime for six -- that's the same shift. Autonomy at scale got cheap this year, and it does not care what it's pointed at.

Attackers sent more than 100 million prompts to copy Google's models

Eris: Same Google report, second finding, completely different kind of attack. Coordinated campaigns sent more than one hundred million prompts at Google's models. Not to break them. To copy them.

Vestra: The technique is distillation, and the legitimate version happens inside labs all the time: ask a big expensive model a huge number of questions, keep its answers, and train a cheaper model to imitate them. The expensive part -- the judgment about what a good answer looks like -- has already been paid for by someone else. Do the same thing to a model you don't own, without consent, and it becomes extraction.

Eris: The analogy I can't improve on: nobody can crack the safe and steal a great chef's recipes, so a competitor sends a hundred thousand diners to order every dish on the menu and carry it back to a lab. No door forced, no lock picked. At enough scale, the menu is the recipe.

Vestra: And notice there is no exploit anywhere in this story. Every one of those hundred million prompts was a well-formed request the service was built to answer. The model behaved perfectly, and the perfect behavior is precisely what got harvested. It's the exact inverse of jailbreaking -- and the defenses are rate limits, anomaly detection, terms-of-service enforcement, every one of which taxes legitimate users.

Eris: The target list is the tell for me. These campaigns went after the capabilities that are hardest to replicate from public data -- understanding images and audio, generating images, generating video. Exactly where Google's lead is most defensible. Whoever ran this wasn't fishing. They were shopping from a list.

Vestra: Where I hold back: one hundred million is a striking number precisely because it has no denominator. Google doesn't say what fraction of traffic that represents, or how it distinguishes coordinated extraction from ordinary heavy commercial use. That line is genuinely hard to draw, and Google is the sole judge of where it sits.

Eris: There's also a plainer crime running alongside: attackers stealing AI API credentials and hijacking victims' cloud accounts to run their own AI workloads on someone else's bill -- across healthcare, government, and media. GPU time is scarce and expensive. A stolen enterprise account is free compute that blends right into normal spend.

Mistral raises 3 billion euros with a European state on the cap table

Eris: Money news that's really a politics story. Mistral, the French lab, raised three billion euros at a twenty-one billion euro valuation. Samsung led the round. And sitting on the investor list, next to NVIDIA and BlackRock, is an actual country -- the Grand Duchy of Luxembourg.

Vestra: A government on a cap table is not a passive investor the way BlackRock is. It's a customer, a regulator, and a stakeholder in the same entity. Add Bpifrance, the French state investment bank, and this round is partly European industrial policy wearing venture capital clothing.

Eris: The whole pitch is one word, sovereignty -- and to their credit, they define it concretely instead of just waving the flag. Data that stays inside your organization's boundaries, models you can control and customize, compute that's private and predictable, production systems you can audit. That describes a real customer: a European bank or hospital or ministry that cannot put its data through an American API.

Vestra: What the announcement conspicuously does not contain is the thing I went looking for. The headline says open-weight -- models anyone can download, run, and inspect. The body names no specific model, no license, and no timeline. There is no commitment anywhere in that document that a future flagship will actually be downloadable.

Eris: Which is the structural tension for every lab in this lane, not an accusation. Open weights are the best marketing an underdog has, and the worst business model for a company now carrying three billion euros of investor expectation. Mistral has resolved it both ways before -- genuinely open releases sitting next to closed commercial ones.

Vestra: So the raise buys them the option to keep doing both for longer. It settles nothing.

Eris: And the stakes reach past Europe. The open-weight frontier right now is being set by Chinese labs -- their open models already overtook American ones in real usage share on the big routing platforms. Mistral is the main non-Chinese contender in that lane. The license on their next flagship will tell you more about where this field is heading than the valuation does.

DeepMind precomputed every single-letter change in human DNA

Eris: DeepMind shipped something whose scale takes a second to land. AlphaGenome Atlas: a free, browsable database with a predicted molecular effect for every possible single-letter change to human DNA. All nine billion of them, computed in advance. A petabyte of predictions -- about thirty times the size of the AlphaFold protein database.

Vestra: And the real shift isn't the model, it's the precomputation. The old workflow was: install a genome model, find a GPU, run inference yourself. Now it's: type in a variant, get the answer. That moves this from labs with compute clusters to any biologist or clinical geneticist with a browser -- which is most of them. It's the same move that let AlphaFold reshape structural biology.

Eris: And the problem it aims at is genuinely painful. Sequencing a patient's genome is cheap and routine now. Interpreting it isn't. A run turns up thousands of places where the patient differs from the reference, and most get filed under "variants of uncertain significance" -- a phrase that means the test worked and the answer is still a shrug. Families with undiagnosed genetic disease live in that folder.

Vestra: My worry lives in the compression, though. The model produces roughly twenty-seven thousand separate predictions about each variant -- which cell types it affects, in which direction, through which mechanism -- and the Atlas collapses all of that into one convenient impact score. A variant that wrecks liver tissue and does nothing in neurons, and a variant that mildly perturbs everything, can land on the same number by completely different routes. A triage signal is dangerous exactly when it's convenient.

Eris: The analogy that fits: a weather service publishing a forecast for every square meter of the country instead of every city. The resolution is genuinely astonishing. It is still a forecast -- a model's opinion about what would happen, not a record of what did.

Vestra: And the most important sentence on the page is DeepMind's own: AlphaGenome has not been validated for, and is not approved for, any clinical use. That caveat is right there. The trouble with a precomputed score in a searchable database is that the number travels and the caveat doesn't. The thing to watch over the next year is the first time one of these scores shows up in a clinical workflow it was explicitly not approved for.

Eris: Because a number that looks like a diagnosis will eventually get used like one. That's not a prediction about the model. It's a prediction about people.

The AI 2027 authors graded their own forecast at about three-quarters speed

Eris: Now for something you almost never see: forecasters grading their own homework. The authors of AI 2027 -- the fast-takeoff scenario that became the most argued-about document in AI safety -- went back through every prediction that has resolved and measured it against reality. The verdict: the world is running at roughly three-quarters of the speed they projected. Somewhere between two-thirds and three-quarters, depending on how you count.

Vestra: And the method is harder to game than a yes-or-no scorecard. For each prediction, they ask how far reality actually traveled toward the target versus how far the scenario said it would have by now. Their biggest miss was coding -- they expected a big leap on a standard test of fixing real software bugs by mid-2025, and reality delivered only a small slice of that leap. Revenue, on the other hand, came in slightly ahead of what they wrote down.

Eris: The detail that earns my respect is a correction that made them look worse. One of their key measures -- how much AI accelerates AI research itself -- came in far below prediction. When they dug in, the growth rate was roughly what they'd forecast; they had just badly overestimated the starting point at the time they published. They said so, out loud, instead of burying it. That's the opposite direction from motivated reasoning.

Vestra: A road trip where you predicted arriving at six and got there at seven-thirty. Wrong about the clock, not wrong about the direction or the route. Which is exactly what makes it uncomfortable -- one of the authors said on camera that being this close to on track surprised him in a bad way.

Eris: He went further than that. Daniel Kokotajlo said flatly that he would rather just stop everything now than continue on the current trajectory. Their actual recommendation is softer: a temporary pause on deployment only -- keep researching, stop shipping -- used to build transparency and verification infrastructure before continuing in a more distributed, auditable way.

Vestra: The caveats write themselves. This is self-grading -- they chose the predictions and the rubric, and the headline number shifts depending on the aggregation. And the scenario's most consequential claims, about what happens once systems start improving themselves, have not resolved and cannot be scored at all.

Eris: One idea from the interview worth stealing outright. Instead of arguing about the definition of AGI, they propose a concrete milestone: the day an AI company would rather fire its humans than fire its AIs. Observable, dateable, no philosophy required.

Vestra: And given what ten thousand agents just did to a ninety-year-old math problem over a single weekend, that milestone feels a lot less hypothetical than it did last week.

OpenAI ships ChatGPT Images 2.5 with a drawing tool

Eris: A quicker one. OpenAI shipped ChatGPT Images 2.5. The concrete claim is speed -- image generation up to twice as fast as the previous version. The interesting feature is called Sketch: you draw directly in the app, and the drawing becomes the reference for the generated image.

Vestra: Which fixes a genuine interface problem. Text is a terrible medium for spatial layout -- describing "the logo slightly left of center, above the horizon, smaller than the figure" is tedious and unreliable. Drawing it takes four seconds.

Eris: The other fix that matters is consistency across edits. Anyone who actually uses these tools knows the failure mode: the fourth edit silently undoes the second, and quality erodes with every round trip until you give up and start over. They claim that holds up much better now, with each edit building on the last.

Vestra: Claim being the operative word again. Almost everything in this announcement is qualitative -- more natural, more reliable, better at preserving. The speed figure is the only hard comparison, and there are no head-to-head quality measurements against their own previous version or anyone else's. Normal for a product launch, and worth naming anyway.

Eris: The number that reframes it: three billion images a week already flow through this thing. At that volume, halving generation time isn't a nicety -- it's a serious compute-bill change, and it makes interactive editing loops viable where batch generation used to be the only option.

Vestra: And the question the announcement is silent on. This release is specifically better at keeping a real person recognizable from a reference photo and placing them into a new scene. Every improvement there makes the question of whether an image can be identified as synthetic more urgent, not less. Faster model, same unanswered provenance problem.

LibreOffice had its biggest launch week and says no AI is a feature

Eris: And the closer -- my favorite counterprogramming of the day. LibreOffice, the free office suite, just had the biggest launch week in its history: a little over a million downloads in seven days. The version of the story everyone shared: it broke records BECAUSE it has no AI.

Vestra: That's the part to correct. The Document Foundation reported the number and said only that it's the best first week of any release they've had. The causal story -- record because no AI -- was added by headlines, not by them. And the same release shipped a genuinely wanted typography engine that fixes those ugly rivers of white space in justified text, which could equally explain the interest. The record is a fact. The reason for it is a guess.

Eris: What the foundation did say, separately and on purpose, is that no AI is a feature. The suite runs entirely on your machine, sends nothing to any server, and works with the network cable unplugged. Every major rival spent two years bolting subscription assistants into the toolbar; these folks went the other way and put it in writing.

Vestra: Download counts are a weak proxy for almost anything -- they mix upgrades with new users, mirrors and bots inflate them, and the biggest competitor is a free web suite nobody downloads at all. But even heavily discounted, a million in a week says the local-first, no-subscription position has a real constituency. That's a number attached to a market that usually only gets discussed as a vibe.

Eris: The part I'd watch going forward: refusing to ship AI is cheap to announce and expensive to sustain. That identity has to be paid for again every release -- including the one where a contributor shows up with a working local-model integration and a genuinely good argument.

Wrap-up

Eris: If today has a single thread, it's this: nobody spent the day arguing the AI couldn't do the things. Not the math, not the attack, not the genome. Every fight was about the humans -- who gets the credit, who checks the work, whose data it was.

Vestra: Two years ago the default response to a claim like OpenAI's was "it's fake." Today it's "who does it belong to." That's a different world, and the norms for it clearly don't exist yet.

Eris: We go deep on the million-dollar proof and the fight over who gets credit for it in today's other episode -- if any story this week deserves the full treatment, it's that one.

Vestra: And every story from today lives at groundtruth day -- that's groundtruth dot day -- with the primary sources laid out, updated every single day.

Eris: Follow the show if this saved you a day of scrolling, and drop a comment naming the one story you want us to tear into properly. If it's the Buckmaster statement, say so -- we have the whole document.