Ground Truth.
AI, checked against the source.

News · 2026-08-12

Claude raised the zeta critical-line bound to 67.2 percent, and Anthropic published the proof

An unreleased research version of Claude raised the proven lower bound on the fraction of Riemann zeta zeros lying on the critical line from 41.6 percent to 67.2 percent, and Anthropic published the full paper, a short note for experts, and a machine-checked Lean proof on August 10, 2026. The model did not solve the Riemann Hypothesis, which is what it was actually asked to attempt. The improved bound was a byproduct of failing at the larger problem.

Key facts

A correction first

On August 10 this publication ran a story headlined the viral Riemann result is not in the literature. It checked arXiv and the surrounding number-theory papers, found the 67.2 percent figure attached only to a different and conditional statement, and concluded the claim was unsupported. That was wrong. Anthropic published the paper on its own research page the same day, outside arXiv, with a formalization attached. The lesson is narrow and worth stating plainly: absence from a preprint server is not absence from the record, and a lab that publishes to its own domain will not show up in the searches you would normally trust.

What the number means

The Riemann Hypothesis, posed in 1859 and carrying a million-dollar Clay Institute bounty, says that every non-trivial zero of the zeta function sits on one particular vertical line in the complex plane. Those zeros control the distribution of prime numbers, which is why the conjecture matters far beyond itself. Nobody has proved it. What mathematicians do instead is prove partial statements of the form "at least this fraction of the zeros provably lie on the line," and grind that fraction upward. It had reached 41.6 percent, and getting it there took decades.

Moving it to two-thirds in one step is the kind of result that would headline a number theory conference. Anthropic's paper states the clean version as an unconditional two-thirds lower bound, with an optimized variant reaching the 67.2 percent headline figure.

How it got there

Claude's move was structural rather than computational. Prior approaches try to count the zeros on the line and separately bound the zeros off it, which means fighting the negative term directly. Claude instead built a single space of functions carrying a quadratic form, where zeros on the line contribute positive directions and zeros off it contribute paired positive-and-negative blocks, then wrote down an inequality on the rank of that form in terms of information you can actually compute from the primes.

The bookkeeping analogy is close enough to be useful. The old method audits income and expenses in two separate ledgers and has to bound the expense ledger tightly. Claude wrote one combined ledger in which the two partly cancel, then used a general rule about the ledger's rank to force the conclusion. Anthropic's post says the essential step was "the courage to treat the entire space, with positive- and negative-definiteness taken into account together," building on published work by Baluyot, Goldston, Suriajaya and Turnage-Butterbaugh, whose 2023 and 2025 papers removed an assumption that had blocked this line of attack, and on a 2000 paper by Bombieri.

The run

The human in the loop was Jarred Sumner, an Anthropic staff member who is not a mathematician. He asked Claude to take a real stab at the hypothesis and left the mathematical choices to the model. Claude generated and tried 650 ideas, all of which failed. Prompted to try again, it spent about a day and a half coordinating roughly 60 subagents: two developed the key ideas, 13 contributed to those two, 30 tried and failed to produce anything, 13 served as validators, and two wrote the initial paper. The subagents ran thousands of numerical checks against known zeta zeros, refereed each other, and downloaded 54 arXiv papers to check the finding was not already known.

Anthropic describes Sumner's contribution bluntly. His "input was mostly limited to sending Claude messages of encouragement (mostly variants of 'keep going' or 'believe in yourself')." Claude then volunteered to write the result up and recommended that a human number theorist validate it, which is roughly what a careful graduate student would do. The Lean formalization means the argument is machine-checkable rather than merely persuasive, which is the difference a proof assistant buys you.

The honest caveat

Anthropic wrote the post, ran the model, and employs the two mathematicians who validated the paper internally. Conrey and Goldston examined it "on short notice," and Goldston is an author of the prior work the result builds on, which makes him well placed to judge it and also not a disinterested party. The result has not yet been through journal peer review. Anthropic itself says it does not expect these techniques to lead to a proof of the Riemann Hypothesis. The most defensible reading is that a real, checkable improvement to a hard bound came out of a long-horizon agent run rather than a flash of insight, and that the interesting number is not 67.2 but 650: the count of ideas that had to fail first.


Primary source, verified: read the paper →

Key questions

Did Claude prove the Riemann Hypothesis?

No. Anthropic says the model was asked to take a real stab at the hypothesis and failed, and that the improved bound emerged as an unintended byproduct of the attempt. Anthropic also says it does not expect these techniques to lead to a proof of the hypothesis itself.

What does raising the bound from 41.6 to 67.2 percent actually mean?

It means the fraction of zeta zeros that mathematicians can prove sit on the critical line went from about two-fifths to about two-thirds. Nobody can prove all of them do, which is the hypothesis, so the field has spent a century pushing this partial figure upward, and the previous benchmark had stood at 41.6 percent.

How much compute did the run take?

Anthropic says the result came out of two sessions in Claude Code using 31 million output tokens in total. After 650 failed ideas, the model spent about a day and a half coordinating roughly 60 subagents that ran 2,400 shell commands and wrote hundreds of Python scripts.
Cite this

APA

Ground Truth. (2026, August 12). Claude raised the zeta critical-line bound to 67.2 percent, and Anthropic published the proof. Ground Truth. https://groundtruth.day/news/claude-raised-the-zeta-critical-line-bound-and-we-said-it-was-not-real.html

BibTeX

@misc{groundtruth:claude-raised-the-zeta-critical-line-bound-and-we-said-it-was-not-real,
  title  = {Claude raised the zeta critical-line bound to 67.2 percent, and Anthropic published the proof},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {aug},
  url    = {https://groundtruth.day/news/claude-raised-the-zeta-critical-line-bound-and-we-said-it-was-not-real.html}
}

Topics: mathematics · ai-for-science · proof-assistants · agents · anthropic · corrections

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.