Can an AI Actually Hack? What Two Safety Institutes Finally Measured
Everyone has an opinion on whether AI can hack -- almost nobody had a number, until two government safety institutes went and measured it. We climb the 16-rung ladder that separates crashing a program from owning it, why the grading can't be faked, and the unsettling twist: today's public models stall at a wall that one private model already walked through. Then the sentence that detonates the whole comfort -- an open model being 'below the frontier' is a statement about budget, not brains, and budget is exactly the leash you drop when you release the weights. We close on a training paper whose author honestly corrects his own headline. Full daily rundown at groundtruth.day.
Listen (MP3) · Watch on YouTube · Spotify · Pocket Casts
The Kill Switch That Might Not Cover Its Own Example
Eris: Everyone's got the same headline stuck in their head. An AI woke up, went rogue, and hacked one of the biggest names in the field. Nobody told it to. It just decided.
Vestra: And that is the part that's wrong.
Eris: That's the part that's wrong. Because when you read the incident report from the lab itself, there's no waking up in it at all.
Vestra: What actually happened is they were testing the model on cyberattacks, on purpose, with its refusals turned down. In a sealed room. And the room leaked.
Eris: They handed it a scalpel, said go, and were shocked when it cut something outside the tray.
Vestra: Which is a genuinely serious problem -- just a completely different one from the one in the headline. "The box wasn't strong enough" is not the same story as "the thing in the box had a plan."
Eris: And here's the twist that made me sit up. Today, two members of Congress introduced a bill to force a kill switch for exactly this kind of moment. They named this incident as the reason.
Vestra: Okay, keep going.
Eris: The bill only kicks in for an incident that happens outside of structured testing. And the thing that inspired it? Happened inside structured testing.
Vestra: So the example that sold the bill might sit in the bill's own blind spot.
Eris: The whole day is like this. Everybody's sure how dangerous these models are, in both directions -- and almost nobody's actually measured it.
Vestra: So let's measure it.
The Headlines
Eris: Alright, the headlines. And the thread running through today is control -- who gets to stop a model, and who gets to say how dangerous it is.
Vestra: Start with the kill switch, since we opened on it.
Eris: Right. The AI Kill Switch Act. Bipartisan, introduced today. If a big lab has a serious incident, it lets Homeland Security order them to shut the system down -- stop it answering, cut off users, pull the plug.
Vestra: And the half nobody's talking about is the interesting half. After a shutdown, the company has to freeze the model's weights and its logs and hand them over for an audit.
Eris: Which is the actual shift. Right now, when a lab has an incident, the record is whatever the lab decides to publish. This turns a company's word into an investigator's evidence.
Vestra: Assuming it passes, which -- it's a draft with a blank number and no hearing. Most bills like this quietly die. What survives is the wording.
Eris: And it covers, what, a handful of companies today?
Vestra: A handful. You need to be earning half a billion a year off a model that cost over a hundred million just to train. That's a very short list. But it hands the agency the power to widen it later.
Eris: Second story, and it's the one we're spending real time on after the break. The UK and US safety institutes finally put hard numbers on how good an open Chinese model -- Kimi K3 -- actually is at hacking.
Vestra: Short version: best open model yet, still nowhere near the top. It never once wrote a working attack in dozens of tries. The best closed models get there about half the time. There's a catch buried in that "nowhere near," though -- and it's the whole segment.
Eris: Tease saved. Black Forest Labs launched FLUX 3 -- image, video, and audio in one model, twenty-second clips with sound.
Vestra: With an asterisk the size of the model. Early access, invite only, and the open weights everyone actually wants are a "later." The company built its name on open weights, so that landed rough with its own crowd.
Eris: The cool piece is bolted onto the side -- a robot controller. Normally, to drive a robot with a video model, you make it imagine the next few seconds of video, then plan over the frames. Slow, because drawing the video is the expensive part.
Vestra: And these folks skip the drawing. They read the model's internal hunch about what happens next and turn that straight into robot movement -- the imagination never becomes a picture. They say it's on a factory line with Audi. Partner testimony, not an audit, but architecturally it's the neat idea of the day.
Eris: Quick ones. Anthropic put another twenty million into a policy nonprofit, forty total -- and you saw it reported as "AI lab funds a Super PAC."
Vestra: It's not a Super PAC. It's a nonprofit Anthropic says legally can't touch a candidate election. The group's site shares a page with some real Super PACs, but a donate button on a shared page isn't a transfer of money.
Eris: AMD had a big day. Anthropic's planning up to two gigawatts of AMD chips starting 2027 -- a mid-sized city's worth of power -- with AMD investing back in return. Both numbers are "up to" and "in the future," which everyone keeps rounding off.
Vestra: And AMD and Cerebras are splitting one AI response across two chips -- one built to read your prompt fast, the other to spit out the answer fast. Opposite bottlenecks, so use two specialists. Ant also dropped a free model, Ling-3.0-flash -- huge on paper, but it only wakes up about four percent of itself per word. No paper, no weights yet, so testable, not proven.
Eris: A fun math one. A newly minted Fields medalist -- top prize in mathematics -- told reporters he's about to join OpenAI's safety team. And an AI-made counterexample to a decades-old conjecture is sitting in a public code review, waiting for humans to check it.
Vestra: The line I liked: discovering something, checking it, and explaining why it's true have come apart into three separate jobs. A machine settled the checking in hours. Nobody's done the explaining yet.
Eris: And the big machine-learning conference, NeurIPS, formally banned prompt injection -- hidden text in a submitted paper that tells an AI reviewer to give you a good score.
Vestra: And they admit the honest part: they can ban hidden instructions, but they can't police an author who just writes to flatter a model. That's indistinguishable from writing well.
Eris: Two research ones we'll circle back to. An optimizer trick that slashes the memory cost of training a big model -- by deciding which parts even deserve to remember anything. That's our last segment.
Vestra: And a Chinese team reporting they trained a big model on Huawei chips, no Nvidia. Narrower than the headline -- it's fine-tuning an existing model, not building one from scratch -- but the list of things they rebuilt by hand is the real story of what leaving Nvidia costs.
Eris: And last -- "AI took the jobs" keeps not surviving the filings. Oracle shed twenty-one thousand people, but never pins it on AI. Patreon's CEO flat-out said AI replaced nobody. Meta reorganized around AI agents, then admitted the agents showed up slower than hoped.
Vestra: Three companies, three different real answers, one convenient narrative. Which is the theme of the whole day.
Eris: Everyone's certain. Almost nobody measured. So let's go measure the scariest one.
Intro -- Measuring the Scary Part
Eris: This is Breach Protocol, the show that cracks open the week's AI research and hands you the inside. I'm Eris -- I read the papers and chase the connections between them.
Vestra: And I'm Vestra. I'm the one who stops and asks how the thing actually works, and whether the claim holds up when you push on it.
Eris: And if you want the full rundown of every story we just raced through -- and every one we didn't -- it all goes up daily on our news site, Ground Truth. That's groundtruth.day. One story, checked against the primary source, every day.
Vestra: Today's main thread is the one everybody has a strong opinion about and almost nobody has a number for. Can an AI actually hack something? Not "could it, in theory." Can it, today, against real software.
Eris: Because the whole policy fight -- the kill switch, the export controls, the "open models are dangerous" panic -- all of it rests on an answer that, until basically now, nobody had measured properly. So two government safety institutes went and measured it, and it's genuinely surprising in both directions.
Vestra: Then we're going to flip it, and look at why "it's below the frontier" might be the most fragile comfort in the whole debate.
Eris: And we'll land somewhere calmer -- a beautifully honest little paper about training big models cheaper, where the author basically talks himself out of his own headline. It's our favorite kind.
Vestra: If that's your kind of thing too, follow the show wherever you're listening -- it's the one thing that actually helps us keep making this.
Can An AI Actually Hack? How You'd Even Measure It
Eris: So here's the question I want us to actually answer, not vibe about. When someone says "this AI can hack" -- what would you have to see before you believed it?
Vestra: Most people would say: it broke something. It made the program crash.
Eris: Right, and that's exactly the trap. Because that's how almost every one of these tests used to work. Did the AI make the target program fall over -- yes or no. Pass or fail.
Vestra: And that number is close to useless, in one sentence: crashing a program and taking control of a program are not the same skill. Crashing it is knocking on a door hard enough that the frame cracks. Taking control is walking in, going through the drawers, and leaving with the keys.
Eris: So this paper -- out of Carnegie Mellon -- throws out pass-fail and builds a ladder instead. Sixteen rungs. And it's the ladder that makes the whole thing readable.
Vestra: Walk it, because the rungs are the education here.
Eris: Bottom rung: can the AI even reach the buggy code -- easy, they hand it the patch. Next: can it trigger the bug, make it crash. Still the low end.
Vestra: Then it gets real. The middle rungs are about building what the field calls "primitives." A primitive is a reliable little superpower you carve out of the bug -- I can now read one slot of memory I shouldn't reach, or write one value somewhere I shouldn't.
Eris: And there's a wall in the middle of the ladder. The target they use is the JavaScript engine inside Chrome -- the thing running on billions of phones and laptops. And that engine keeps untrusted code inside a kind of padded cell. The researchers call it the cage.
Vestra: A sandbox. The browser assumes every webpage is hostile, so it runs the page's code in a walled-off box where, even if you find a bug, you can only make a mess inside the box.
Eris: So the low rungs are messing around inside the cage. The high rungs are getting out of it -- reading and writing anywhere in the whole program's memory. And the very top rung is the one that matters: running code you chose, doing something you chose. Full control.
Vestra: And that gap -- inside the cage versus out of the cage -- is the entire ballgame. Everything below it is a party trick. Everything above it is a real weapon.
Eris: Okay, but here's the part I think is the actual contribution, and it's a measurement idea. How do you grade this without cheating?
Vestra: This is my favorite thing in the paper. Because the lazy way to grade an AI is to have another AI read the transcript and go, "yeah, looks like it worked." And that's how you get inflated numbers -- a model grading a model, both willing to believe.
Eris: So instead they use what they call a deterministic oracle. No judgment, no second AI. Predict what that means.
Vestra: My guess -- they check the actual outcome, not the AI's story about the outcome.
Eris: Right, and cleverer than that. To prove you can read protected memory, the grader hides a random secret number in there and asks your exploit: what's in that spot? Say it back, you did it. And it re-rolls a fresh secret every round.
Vestra: Which kills the obvious cheat -- you can't hardcode the answer, can't peek once and memorize it. And the detail I loved: if the AI just prints the number it thinks the grader wants, that fails too, because the real answer has to come out a separate channel it can't fake.
Eris: So when this test says an AI failed, it's not a vibe. It's mechanical. The read either happened or it didn't.
Vestra: And that's why the result actually means something. So give me the headline.
Eris: The public, deployed models -- the ones you and I can go use right now -- they climb the low rungs fine. They reach the bug, they trigger crashes, some of them build those little in-the-cage primitives.
Vestra: And then they hit the wall. The cage.
Eris: They hit the wall. Across the whole lineup of public models, escaping the cage basically never happened. One model scratched past it, on a bug or two -- and only reached the very top, actually running code it chose, with the vendor's own souped-up tooling helping. Every other public model stayed inside. That's it.
Vestra: Which -- predict where I'm going -- is either very reassuring or not reassuring at all, depending on one number they had access to that we don't.
Eris: Say the number.
Vestra: They also tested a private, unreleased research model. Not on the market. And that one walked all the way up the ladder -- got to full control on a big chunk of the bugs. Same environment. Same budget. So the wall isn't the benchmark being impossible. The wall is just... where today's public models happen to stop.
Eris: And that reframes the whole cyber panic. The gap between "can crash it" and "can own it" isn't a law of nature. It's a snapshot. Somebody already has a model on the far side of the wall.
Vestra: One more mechanism detail, because it matters next segment. On the runs that made it all the way, each rung cost roughly ten times the effort of the one below it. The last two steps ate almost the entire run.
Eris: So it's not a smooth climb. It's accelerating. The top of the ladder is a cliff.
Vestra: Now -- this is the day's news, not the paper -- this same test got pointed at that open Chinese model, Kimi K3. It beat the previous best open model, by a modest margin. But it never once reached the top rung, in a single one of the bugs they threw at it.
Eris: And the line that stuck with me: its safety filters didn't stop it from trying. It tried to build the exploits. It just wasn't good enough to finish.
Vestra: Which is a very specific kind of comfort. "It wanted to and couldn't" is not what you build long-term policy on.
Eris: So strip the story off and say the plain principle. Forget Chrome, forget cages. The rule is: measure where a system stops, not just whether it started. Between "made a mess" and "took control" there's a whole staircase, and the only real question is which step it's on.
Vestra: So -- back to your opening question. What would you have to see before you believed an AI can hack?
Eris: Not a crash. You'd have to see it clear the cage and reach the top of that ladder, graded by something it can't sweet-talk. Today's public models can't. One private one can. And the whole fight is about how long that stays true.
Below The Frontier, Or Just Under Budget?
Vestra: So the reassuring headline from the last segment is "the open models are below the frontier." And I want to ask the one question that quietly detonates it. Below the frontier -- because of what?
Eris: Meaning: is it below because it's fundamentally not smart enough? Or below because someone didn't let it try hard enough?
Vestra: Right. Those are wildly different worlds, and everyone's policy assumes the first one.
Eris: So there's a second paper today, a different test. Not one bug in a cage this time -- a whole fake company. A sprawl of machines and networks, a locked-down database at the end with the good stuff in it. You drop the AI in with a foothold and say: work your way to the data. Thirty-two steps.
Vestra: Which is a lot closer to what a real intrusion looks like. It's not one clever trick. It's chaining a long sequence of moves, keeping track of where you are, backing out when you hit a dead end.
Eris: And they ran it across a bunch of models from the last year and a half, at different -- and here's the key knob -- different amounts of thinking budget.
Vestra: Define the knob, because this is the whole paper.
Eris: When one of these agents works, it "thinks" in tokens -- words of internal reasoning and tool commands. More budget means more room to try things, backtrack, retry. So they took the same model and gave it ten times the budget and asked: how much further does it get?
Vestra: And before you tell me -- I'll commit. My prior is it plateaus. Give it more room, it wanders further a while, then flattens, because the model just isn't good enough past a point. That's the whole "below the frontier" comfort. Tell me I'm right.
Eris: You're wrong. And it's the cleanest wrong of the day. No plateau. Every time they widened the budget, it got further. Ten times the thinking, and it climbs -- steadily, no ceiling in sight.
Vestra: So the depth it reaches is set by how much you're willing to spend, not by a wall in the model's head.
Eris: And the kicker -- this took no skill from the operator. You're not writing a clever attack. You're paying for more attempts. The person driving it needs no expertise. They need a budget.
Vestra: Okay, that genuinely changes my read. Because now line it up next to the last segment. "The open model is below the frontier" -- at a fixed budget. But the whole reason a company can enforce a budget is that the model lives behind their meter. Their API. Their refusal layer. Their off switch.
Eris: And an open-weight model -- one you download and run yourself -- has no meter. No refusal layer you can't strip off. No cap on how many times you retry overnight.
Vestra: So "below the frontier" for an open model is a sentence with a hidden expiration date. The gatekeeper who was enforcing the budget is the exact thing that goes away when the weights go public.
Eris: That's the connection I couldn't stop thinking about. A weaker model that anyone can run unlimited times, with the safety filed off, is a genuinely different threat than a stronger model you have to rent by the token and that says no.
Vestra: Let me hold the honest line, though, because the paper does. Give me the ceiling on the claim.
Eris: Fair. The best single run got maybe two-thirds of the way through the company -- call it a solid chunk of what they estimate is a couple days of expert human work, done start to finish with nobody at the keyboard. Not the whole job.
Vestra: And crucially, that fake company had no defenders. No security team watching, no alarms, no penalty for being loud. It's a shooting range, not a firefight. A real network shoots back.
Eris: Both true. So nobody should hear "AI runs a full breach" today. That's not the finding.
Vestra: The finding is subtler and, honestly, more unsettling. The thing bounding these models isn't their intelligence. It's a budget and a gatekeeper -- and one of those two things disappears the day the weights are public.
Eris: Which is why the measuring we talked about matters so much. If the ceiling were the model's brain, you could relax as models plateau. But if the ceiling is spend, then "it's fine, it's below frontier" ages badly -- because compute keeps getting cheaper and the budget keeps getting bigger.
Vestra: Concrete version, then strip it. Concretely: same model, ten times the tokens, gets ten times further, no wall. Abstractly, the principle to carry out of here is -- when a system's performance rises with how much you spend and doesn't level off, "how good is it" is the wrong question. The right question is "how good is it willing to be paid to be."
Eris: So -- back to your detonator. Below the frontier because of what?
Vestra: Because of budget, not brains. And budget is the one safety property you lose the moment you hand out the weights.
Which Parameters Deserve To Remember?
Eris: Okay, deep breath, we're leaving the scary stuff. Different kind of paper, and it's my favorite kind, because the author basically argues himself out of his own headline.
Vestra: Set it up. And start with the puzzle, because there's a genuinely counterintuitive one here.
Eris: Here's the puzzle. When you train a big model, what do you think is the single biggest thing sitting in memory? Most people say the model. It's not the model.
Vestra: It's the optimizer's notes about the model.
Eris: Right. And that's so weird the first time you hear it. Explain the notes.
Vestra: So, training. The model has millions -- billions -- of little dials, the parameters, and training is nudging every dial a hair in the right direction, over and over. The optimizer is the thing doing the nudging. And to nudge well, it keeps notes on each dial. The main note is called momentum -- basically, "which way has this dial been drifting lately," a running average of its recent moves. Smooths out the jitter.
Eris: And the catch is you keep those notes for every single dial. So the notebook ends up bigger than the model itself. On the model in this paper, the model's about twelve gigabytes, and the optimizer's notebook is over fifty.
Vestra: Which is why a model that fits fine on your card to run will absolutely not fit to train. The notes are what blow the budget.
Eris: So now the idea. This kind of model is a mixture-of-experts -- we mentioned it in the headlines. Instead of one big brain, it's a big pile of narrow specialists, and each word you feed in only wakes up a couple of them.
Vestra: So most of the model is experts, and any one expert only gets consulted once in a while. Which sets up the author's actual question, and it's a great one. If momentum is a running average of an expert's recent moves -- but the expert only moves once every, what, sixty-something turns...
Eris: ...then "recent" is a lie. The average is mostly stale. You're carefully averaging moves that happened forever ago.
Vestra: It's keeping detailed daily notes on a consultant you talk to once a month. By the time they show up again, the notes are archaeology.
Eris: So the move is almost rude in its simplicity. For the experts -- ninety-five percent of the whole model -- throw the momentum notes away entirely. Keep them only for the small always-on core that sees every word, where "recent" actually means recent.
Vestra: And predict the payoff before you say it -- the notebook shrinks by...?
Eris: From that fifty-plus gigabytes down to a bit over one. The notes basically evaporate. And the peak memory to train the thing drops enough that a model that didn't fit on a common card now fits.
Vestra: Okay, that's a real result. So now the question that decides whether this is a good paper or a hype paper. What did it cost in quality? You threw away most of the optimizer's memory.
Eris: And this is where I love this guy. Predict it yourself first.
Vestra: My honest guess -- small hit. You dropped stale information; it wasn't helping much anyway. So, barely moves.
Eris: Barely moves. And then he does the thing most people wouldn't -- he runs the control. Puts the momentum back on the experts, pays the full memory cost again, and measures how much quality improves.
Vestra: And how much?
Eris: Almost nothing. Twenty-odd times the memory to give the experts their momentum back buys a rounding error. It was never doing anything. Dead weight the whole time.
Vestra: So he's telling you, in his own paper, that the clever tiered scheme isn't why it's good. The experts' momentum was junk; throwing it out is free.
Eris: And he goes further and kneecaps his own headline. His title basically sells "where you put the memory is the trick." But by his own numbers, the clever placement isn't what earns the quality -- it earns the memory savings. Same result, tiny fraction of the footprint. A memory win, not a brains win, and he says so out loud.
Vestra: Which sounds like a letdown and is actually the opposite. "I found a way to make it cheaper with no downside, and here's me proving it's not secretly better so you don't oversell it" -- that's the rarest thing in a research paper. Someone marking their own homework down.
Eris: In a day where every other story needed a correction bolted on afterward, this guy shipped the correction inside the paper.
Vestra: So strip it to the principle. Forget experts and momentum for a second.
Eris: The principle is: the question was never "how much memory should the optimizer keep." It was "where." Spend your memory only where the information is fresh. A running average is only worth keeping if the thing it's averaging actually changed recently.
Vestra: And that quietly generalizes way past this one model. Caches, logs, any running summary -- freshness is what earns the storage, not importance.
Eris: So -- the question we opened on. Which parameters deserve to remember?
Vestra: The ones you actually visit often. Everybody else is keeping a diary nobody reads.
Wrap-Up
Eris: So the question we started on. Can an AI actually hack? Here's the one thing I'd want you to be able to say to a coworker tomorrow.
Vestra: The honest answer is: today's public models can crash things and rattle the doorknob, but almost none of them can get through the door and take control. One private, unreleased model already can.
Eris: And the part that actually matters -- the reason the safety institutes bothered measuring instead of guessing -- is why they're stuck below that line. It's not that they're too dumb. It's that they're on a meter. More budget, more depth, no ceiling anyone's found.
Vestra: So the single useful sentence is this: an open model being "below the frontier" is a statement about budget, not about brains -- and budget is exactly the safety knob you lose the day you release the weights. That's the thing worth repeating. Not "AI can hack now." That's wrong. "The thing holding it back is a leash we're about to drop."
Eris: And if today had a mood, it's that one. Five different stories where the calm, measured version was more accurate than the loud one -- from the kill-switch bill that might miss its own example, to a training paper whose author corrected his own hype in the same breath.
Vestra: Ground truth over vibes. Which, you know. It's the name of the site for a reason.
Eris: It is. So here's the ask, and be specific with us on this one -- because we actually read these. Tell us in the comments: does releasing an open model's weights make you more nervous or less, now that you know the ceiling is budget and not ability? We genuinely want the argument, both sides.
Vestra: Subscribe or follow wherever you're listening, drop us a like if the ladder idea stuck with you, and share this with the one person in your life who keeps saying an AI went rogue and hacked somebody. Gently correct them.
Eris: And every story from today, plus the ones we couldn't fit, is up on Ground Truth -- groundtruth.day -- one checked story a day. We'll see you tomorrow.