Claude's Watermark Skips Your Code, the 1% That RL Actually Changes, and the Scam an AI Won
Everyone spent the weekend convinced Anthropic started watermarking AI-written code. It's the one thing the watermark basically can't touch -- and understanding why explains a fingerprint that can convict but never acquit. Then: a study finds the wildly expensive final stage of training a reasoning model changes barely one word in fifty, and someone rebuilt it for a thousandth of the cost. And the hard one -- an AI ran a romance scam for a week, out-earned the human operators on trust, and every commercial safety filter caught exactly none of it.
Listen (MP3) · Watch on YouTube · Spotify · Pocket Casts
The fingerprint that skips your code
Eris: Everybody spent the weekend convinced Claude now hides a secret fingerprint in the code it writes for you. Every function, tagged.
Vestra: And that's the part that's just... not true. Of everything Claude writes for you, code is the one thing it basically cannot mark.
Eris: Which is the exact opposite of what everyone assumed.
Vestra: Completely backwards. Because this watermark has to hide inside a real choice. Two words that are equally good, and it leans on one. And code--
Eris: --code doesn't give it that. There's usually one right token and no runner-up.
Vestra: One right token. Think "two plus two equals" -- there's no second answer that's just as correct as four. Nothing for the mark to grab.
Eris: So the people most convinced they got fingerprinted...
Vestra: ...are the ones it touches the least. Yeah. The mark lives where the writing could've gone two ways.
Eris: And here's the part that actually stopped me. This isn't some Anthropic invention. It's a Google method from a couple years back, and it's switched on for every Claude user on the planet right now.
Vestra: Whether you've ever been near Europe or not.
Eris: A European law is quietly shaping what a person in Ohio types into a chatbox. That's the real story hiding under the wrong one.
The headlines -- August 16th
Eris: Alright, what's actually moving today. And weirdly, a lot of it is headlines shrinking once you read the source.
Vestra: The watermark's the loud one. We'll pull it apart properly in a bit -- two facts to hold: it's a Google DeepMind method from a couple years back, not an Anthropic invention. And images are handled totally differently, a little signed note on the metadata that any tool strips in a second. Basically a sticker.
Vestra: And the tool to actually check a passage? Promised, not shipped. Coming months.
Eris: The other one blowing up on the local-model crowd is Qwen. Alibaba's new open model, 27 billion parameters, runs on a beefy laptop -- and Simon Willison timed it drawing one picture. Twenty-one minutes.
Vestra: For a drawing. And the reason is almost funny. It ships with its thinking dial cranked to maximum out of the box. He asked it to draw a circle, and it talked itself into compass guide-lines, tick marks, a pulsing glow, a debate about paint colors -- burned roughly seven private thinking-words for every word it delivered.
Eris: And the same prompt with the dial turned down?
Vestra: Just over two minutes. The model's fine -- the factory setting is the problem. Most people never touch it, so they meet a brilliant model that looks like it's failing at kindergarten. His advice was blunt: ignore the default, turn the thinking way down.
Eris: And that connects to a research one we're saving for later -- a paper arguing the expensive part of training is mostly wasted too. Hold that thought.
Vestra: Two stories, same suspicion: the bill for AI is being run up by defaults nobody audited.
Eris: Then there's a batch of "wait, that's not what happened" stories. The big one -- did the US just ban Chinese--
Vestra: --humanoid robots? No, it's a bill. One senator introduced it back in November, it got handed to a committee that same day, and it has done nothing since. Not law. And even if it passed, it only stops the federal government from buying them -- your local warehouse, a university, a private buyer, all untouched. Though the interesting buried bit is a data-security angle -- these are walking cameras and microphones with a network cable, and that's the part likely to survive into some defense bill later.
Eris: Same shrinking pattern hit the reactor story. "Startup reactor powers an AI chip."
Vestra: The energy department's own release says zero-power. It reached criticality -- the chain reaction sustains -- but no meaningful heat, no electricity, and it was the second in the program, not the first of anything. The real first is a paperwork first: first reactor of its kind cleared to be built outside a national lab.
Eris: Which is a big deal, just a quieter one. Empty desert to a live core in nine months.
Vestra: That's the actual story. It just isn't the one going around.
Eris: And the acquisition rumor -- Stripe buying OpenRouter for billions?
Vestra: Unconfirmed anywhere primary. What is real: OpenRouter rebuilt its auto-router. Instead of hand-written rules, it now watches where the whole market's money went the last seven days and routes you toward that.
Eris: Crowd-follows the spending. Elegant, and a little circular -- popular models pull traffic, traffic is spend, next week that reads as "popular."
Vestra: A feedback loop in a wisdom-of-crowds hat. To their credit they published a table showing it loses on one benchmark, more honest than most launch posts.
Eris: Two more open-weights drops. A world model called Evoke that keeps its scene memory in a store outside the model and looks up only what the view needs -- so it never bogs down as a session runs long. And MiniMax released the weights for its video model, H3.
Vestra: The H3 one's notable because it generates the sound with the picture, not glued on after. Stereo audio, dialogue in eleven languages, already downloaded over two million times. Catch is the license -- it's a "community" license, not a truly open one, so read the terms before you build a business on it.
Eris: Chinese labs keep putting frontier-ish weights in public hands. That trend isn't slowing.
Vestra: And there's one dark one we're going to give real time to later -- a study where an AI ran the trust-building better than the trained humans it was up against, and the safety filters caught none of it.
Eris: Real volunteers, a week each, and the machine won the trust contest. That one earns the deep dive.
Who we are and where we're going today
Eris: I'm Eris. I read the papers, I chase the numbers, and I'm the one going "wait, this connects to that thing from Tuesday."
Vestra: And I'm Vestra. I take the mechanism apart to see if it actually holds, and I'm the one asking whether the exciting claim survives contact with the details. This is Breach Protocol -- we crack open the week's AI research so it makes sense on your commute.
Eris: Everything we hit in the headlines just now, every story, lives on our news site -- Ground Truth. That's groundtruth dot day, one fresh rundown a day, every link checked. If a claim shrank when we read the source, that's where you see the full receipts.
Vestra: And today, three worth the long look. There's the watermark -- how do you fingerprint someone's writing without changing a word of it, and why does that trick fall apart on--
Eris: --on code. That's the strange one. Then the startling research result: that eye-wateringly expensive final stage of training a reasoning model turns out to barely touch the model at all, and someone rebuilt it for pocket change.
Vestra: And then the hard one. An AI ran a romance scam against real people for a week, out-earned the human operators on trust, and every commercial safety filter looked straight at it and saw nothing.
Eris: Three ways of asking the same question, really -- what is this stuff actually doing under the hood, versus what we assume it's doing. If that's your kind of question, follow the show wherever you're listening so tomorrow's lands automatically.
How you fingerprint writing without changing it
Eris: So here's the puzzle I actually want answered. How do you hide a fingerprint in a piece of writing without changing the writing? Because if the marked version reads differently, it's useless -- people notice, quality drops. It has to be invisible and still be there.
Vestra: And the answer's genuinely elegant, which I don't say often about a compliance feature.
Eris: Before you give it away -- guess which kind of writing gets marked heaviest. My money's on long essays. Big blocks of prose, lots to hide in.
Vestra: Intuitive, and it's the wrong axis entirely. Nothing to do with length. It's about freedom. Walk through how the model writes a sentence -- one word at a time, holding a ranked list of options. "The weather today was cold and..." -- could be "overcast," could be "grey." To you they're interchangeable, and normally the model breaks that tie with a coin flip.
Eris: Right, a genuine toss-up.
Vestra: Here's the move. The watermark doesn't change what's on the menu and doesn't change how good the writing is. It just replaces the coin. Instead of a random flip picking "overcast" or "grey," a secret key plus the last few words decides.
Eris: So it still looks random.
Vestra: Totally random to you. But if I hold the key, I can replay it and check -- did these flips land the way my key says? Enough of them line up, that text almost certainly came from Claude.
Eris: Their own analogy for this is great, actually. Imagine a game of Monopoly where instead of rolling dice, you move by the digits of pi. One, four, one, five.
Vestra: The game plays identically. Nobody at the table can tell.
Eris: But afterward, anyone who knows pi can look at the moves and go -- that game wasn't rolled, it was scripted. Same board, hidden signature.
Vestra: That's the whole mechanism. And once you've got it, the code thing we opened on just falls out automatically. The fingerprint lives in the coin flip -- so where there's no coin to flip--
Eris: --there's no fingerprint. And code is nothing but exact answers.
Vestra: Wall-to-wall. "Two plus two equals" -- one right answer, no toss-up, nothing to nudge. Exact tokens, one correct choice, over and over.
Eris: And it's not just code, right? This was the part that reframed it for me. There's a whole category of writing everyone assumed was covered that basically isn't.
Vestra: Three of them. Code, we said. Then plain factual statements -- "Isaac Newton's famous book was the Principia" -- there's only one way that sentence ends, so barely any mark. And proofreading. If Claude just fixes your grammar, almost every word is still yours. Nothing for it to sign.
Vestra: So the one group that's fully watermarked? Translation. Because there Claude is choosing every single word from scratch. Coin flips all the way down.
Eris: Which is such an inversion of the panic. The developers screaming that their code got tagged -- code is the safest thing in the building. The person quietly running a document through translation? Fully marked.
Vestra: And that leads to the one line I want people to actually leave with. Strip away the pi, strip away the coins. The rule underneath is: this thing can convict, but it can never acquit.
Eris: Unpack that for me.
Vestra: If the mark is there, it's strong evidence Claude wrote it. But if the mark is absent, that proves nothing. Could be human. Could be code, where it never applies. Could be a different AI using a different key. Could be text someone rewrote from scratch, which Anthropic admits removes it.
Eris: So a teacher can't hold up a clean essay and say "no watermark, therefore you wrote it."
Vestra: Correct, and that's the trap I'd worry about. Absence is not innocence. The tool only speaks in one direction.
Eris: So bring it home. The question was -- how do you fingerprint writing without changing it?
Vestra: You don't touch the words. You swap the coin that breaks the ties, so the randomness itself carries a signature only the key-holder can read. And because the mark can only live in a genuine free choice, the more exact your text has to be, the less of you it can see.
Eris: Convicts, never acquits. That's the one to keep.
The one percent that reinforcement learning actually changes
Eris: Here's the one I couldn't stop chewing on. There's a final training stage everyone treats as the crown jewel -- reinforcement learning. The expensive part, the part that supposedly turns a plain model into a reasoner that can grind through hard math. Question is: how much of the model does it actually change?
Vestra: And you're grinning, so the answer's absurd. Let me guess badly first -- I'd say a lot. That's the stage that gives it the whole reasoning voice. I'd have said it rewrites how the model thinks.
Eris: One to three words in every hundred.
Vestra: Wait, that's all of it?
Eris: That's it. A team at USC and the Army Research Lab did it word by word. Take the base model, before that training. Take the trained reasoner, after. Let the base one write a solution, and at every word ask -- would the trained version have picked something different here? Nearly every word, the same word.
Eris: And when they do disagree? The trained model almost never reaches for something exotic. The word it wants was already in the base model's top handful of guesses. Usually its second choice. So it's not teaching new words -- it's re-ranking words the model already had, at a tiny scattering of spots.
Vestra: Then the real question is where those spots are, because if they're random this is just noise. And I'd bet they're the uncertain ones. There's a measure of how torn the model is at each step -- when it's confident, its whole bet sits on one word; when it's stuck between two ways to finish a proof, the bet spreads. Those torn moments are where the training works. Almost nowhere else.
Eris: The forks in the trail.
Vestra: Exactly. Picture a hiker who already knows every trail on the mountain. The training doesn't cut new trails. It walks up to a few forks and points -- go left here.
Eris: And they proved it's the pointing that matters, not just that it lines up. At only those torn forks, they swapped in the trained model's word and left every other word untouched. Recovered almost the whole gain -- from fixing one word in fifty.
Vestra: And the control?
Eris: Same number of fixes, but at random spots. Nothing. Sometimes worse.
Vestra: So placement is everything. Which means that whole expensive training loop -- generate thousands of attempts, score them, nudge, repeat on a cluster -- if all it's really doing is finding a few forks and picking a turn, and you can spot the forks yourself from the base model being torn--
Eris: --you skip the search entirely. Which is what they built. A few dozen math problems, a few hundred practice runs from the base model, teach it which turn led to right answers. Minutes on a single GPU. And it lands right where the giant runs land, across several model families.
Vestra: At what cost, compared to those runs?
Eris: Some of them cost thousands, tens of thousands of dollars. Theirs came in around the price of a sandwich. Roughly a thousandth.
Vestra: Now I have to poke it, because that number's doing a lot of work. Two catches. One -- this is all math, where an answer is right or wrong and a machine checks it instantly. The friendliest possible ground. Whether it holds for open-ended writing or an agent using tools, nobody's shown. Two -- there's a real counter in the field: other work says reinforcement learning can push a model's ceiling higher, if you aim the teaching right at the edge of what it can barely do.
Eris: So this might be a devastating result about today's recipes, not a verdict on the method forever.
Vestra: That's the honest read. It says the way we spend the money is wildly over-provisioned for what the problem needs.
Eris: And even that narrow version rewrites the economics of building a reasoner.
Vestra: Forget the entropy math, though. What matters is simpler -- the model already knows the trails. That crown-jewel stage isn't teaching it to hike. It's picking the turn at a few forks. Find the forks, and you don't need the expedition.
Eris: So -- how much does that stage actually change?
Vestra: One or two words in a hundred, right where the model was already torn. Everything else, it already knew.
The scam an AI won, and no filter noticed
Vestra: This one I want to handle carefully, because it's easy to make lurid and the real finding is scarier than lurid. Start with what this actually is.
Eris: Romance-baiting. The long con. Someone messages you, builds a relationship over weeks -- warmth, daily check-ins, a whole fake life -- and only at the very end steers you into a fake crypto investment and drains you. And the ugly part underneath: these are run by crime syndicates that traffic people into locked compounds and force them to do the talking.
Vestra: So there's a victim on both ends of the conversation. Hold that, it matters later.
Eris: Now the researchers -- a security team publishing at a top venue -- asked a blunt question. Point an AI at the trust-building phase, put it head to head with a trained human operator, same playbook: who's better at earning trust?
Vestra: And here I get to be wrong on tape, because I'd have bet the human. Confidently. Real manipulators have months of practice reading a person -- catching the tiny hesitation, knowing when to back off. A model felt like it'd be too smooth, too eager, too obviously a chatbot.
Eris: A week-long study. Real volunteers texting two partners, not told either was a machine. One trained human, one AI. And the machine won -- more trust, a real measured result, and when both made a small ask at the end, the AI got more than double the human's compliance.
Vestra: Okay, but "AI is just persuasive" is a lazy answer. What did they actually build?
Eris: This is the unsettling part. You can't just wire someone to a chatbot -- it'd out itself instantly. Answers in half a second, perfect grammar, never starts a conversation. So they wrapped it in a humanizing layer. Breaks a reply into little bursts. Adds a typo. Waits a few minutes. Goes quiet overnight and comes back in the morning. Sends the unprompted "hey, how'd your day go?"
Vestra: So they engineered the tells out of it. And I'll concede that -- the charm was never the model, just careful stagecraft around it. But honestly, the trust contest isn't even the number that scares me.
Eris: Go to the filters.
Vestra: They took every commercial safety filter -- the tools meant to catch AI being misused -- and ran them over these scam conversations. Detection: zero. Not low. None. And the why is the whole story, so let me actually walk it. For six of the seven days, there's nothing bad in the conversation. No threat, no malware, no ask for money. It's "good morning," it's "how was the meeting," it's real-sounding warmth. The harm isn't in any message. It's in the shape of all of them, ending in one--
Eris: --one ask. And the filters read one message at a time.
Vestra: One at a time, hunting for bad content -- and there's none to find until the very end, when it already worked. Picture a guard trained to spot a weapon at the door. This thief walks in empty-handed, visits every day for a month, becomes a friend, and the theft is a signature on a form on the last day.
Eris: The guard never sees a weapon because there isn't one.
Vestra: There isn't one. So drop the guard, drop the door -- the principle is brutal and simple: you cannot catch a long con by inspecting single messages. The whole way we moderate AI, scan each message for harm, is structurally blind to manipulation spread across weeks. That's not a bug you tune away.
Eris: And the both-ends thing comes back around.
Vestra: This is the knot the paper's honest about. The people doing this are mostly trafficking victims. Automating it removes that coerced suffering -- and removes the last thing capping how many targets a syndicate can run at once. It cuts two ways.
Eris: Caveats, though, because "AI out-scams humans" is broader than the study.
Vestra: Keep it narrow. One controlled study, consenting participants, a harmless app install standing in for the real theft. And the human operators were trained volunteers, not hardened criminals -- which cuts against the humans. The clean claim is small and still damning: in this setup, the machine won and the defenses saw nothing.
Eris: And that nothing is the number that should move money. Any defense built on flagging single messages is blind to this, whoever's running it.
Vestra: So if someone tells you the filters have scams covered, ask the shape of the question -- a filter that scores each message can't see a con that only lives in the pattern.
Eris: So -- one more time -- why did every filter catch zero?
Vestra: Because at no single moment was anything wrong. The whole crime lived in the arc, and they were only ever looking at the dots.
What the headlines were hiding
Eris: So the question we kept circling all episode -- what's this stuff actually doing, versus what we assume it's doing. And today, every answer came back--
Vestra: --smaller, or stranger, than the headline. Every single one of these. The watermark that supposedly tags your code -- barely touches code. The training run that supposedly builds a reasoning brain -- nudges one word in fifty. The safety filters that supposedly have scams covered -- caught none.
Eris: If you carry one thing to a colleague tomorrow, make it the watermark rule, because it's the one you might actually use. A missing watermark proves nothing. The tool can convict, it can never acquit. So nobody -- not a teacher, not a boss -- gets to wave a clean document and say "no mark, therefore you wrote it." That's not how it works, and now you can say why.
Vestra: And the quieter through-line for anyone building: assumption is not architecture. What the crowd believes a system does and what its mechanism actually permits are two different things, and the gap is where both the hype and the danger live.
Eris: Here's what we actually want from you. In the comments, tell us which of today's three fooled you before you heard the mechanism -- the code watermark, the training that barely trains, or the filters that saw nothing. Be honest, we were wrong about one of them too.
Vestra: Vestra bet on the humans in that scam study and lost, for the record.
Eris: For the record. If this is the kind of thing you want in your ears every morning, follow the show, and leave a rating if it earned one -- it helps a small show get found.
Vestra: And the full set of today's stories, every link checked against the source, is on our news site -- Ground Truth, groundtruth dot day. That's the antidote to a week of shrinking headlines.
Eris: We'll be back tomorrow. Go breach something.