Ground Truth.
AI, checked against the source.

News · 2026-10-03

Ataraxos beat Stratego's most decorated champion 15-1 after a training run that cost under $8,000

An AI system called Ataraxos beat Pim Niemeijer, the most decorated player in the history of the board game Stratego, by 15 wins to 1 with four draws in a 20-game series, according to a paper published in Nature on September 30. The authors estimate the whole training run cost under $8,000 at 2025 cloud prices, a small fraction of what the previous best Stratego AI is thought to have cost.

Key facts

Stratego looks like a simple war game, and it has been one of the harder tests in game-playing AI for a specific reason: you cannot see your opponent's pieces. Each player secretly arranges 40 pieces of different ranks on their side of the board. You can see where an enemy piece stands, but not whether it is a marshal, a scout or a bomb, until two pieces collide and both ranks are revealed. Chess and Go show everything on the board. Stratego, like poker, forces a player to reason about what the other side might be hiding and to bluff about what they themselves are hiding. The paper counts more than 10 to the power of 33 possible ways the hidden pieces could be arranged.

The previous landmark was DeepNash, a DeepMind system published in Science in 2022. It learned purely by playing against itself and made no attempt to search ahead during a game. The Ataraxos paper reports that DeepNash won 42 of 50 games on the Gravon online platform in 2022 but did not reach the site's top ranking, and that its opponents did not know they were playing a bot. The two systems never met: the Ataraxos authors say DeepMind told them DeepNash's old code was no longer functional.

What the team built

Ataraxos combines three ideas. First, like DeepNash, it learns by self-play, with one learning process for the opening setup of the army and a second for moving pieces during play. The authors deliberately shrink the size of each learning step as play improves, which they say keeps training stable when information is hidden.

Second, it learns a belief model: a transformer that looks at everything the player has seen so far and predicts the likely ranks of the enemy's unrevealed pieces.

Third, it searches at decision time. Before each move, Ataraxos samples several plausible hidden armies from its belief model, plays out candidate moves against each one, and updates its plan for that specific position. A useful comparison is a detective who, instead of trying to consider every possible suspect, writes down a handful of the most likely stories that fit the evidence and checks which action works best across all of them. In the match, that search ran on a single NVIDIA H100 GPU and took about 1.26 seconds per move on average.

Training took 16 H100 GPUs for one week for the game-playing networks, then four H100s for four days for the belief model. The authors estimate DeepNash cost between $3 million and $4.5 million to train, based on its reported use of 1,024 Google TPU nodes and a corresponding author's recollection of two to three months of training. That DeepNash figure is a retrospective estimate, not an invoice.

The match

The 20 games were played on the Strategus platform at a standard 15-minute clock with a three-second increment per move, spread across three weeks so Niemeijer could rest and prepare. He was told in advance that Ataraxos would not adapt to his play, which gave him the chance to hunt for a weakness and exploit it repeatedly. The paper quotes George Franka's assessment that Niemeijer is "the best Stratego player ever." Separately, at a 2025 Stratego World Championship demonstration, Ataraxos won 38 of 40 games against attendees and lost two.

The paper also reports results in Barrage Stratego, the cooperative card game Hanabi, and the Chinese card game dou dizhu, which supports the authors' claim that the method is not tied to one game. The code is public under an MIT license in the AtaraxosAI GitHub repository, and records of the 20 games are linked from the project site. Ars Technica's coverage and the Hacker News discussion carried the story to a wider audience.

Why it matters

The cost is the headline for anyone who builds systems. A problem that recently looked like it required a frontier lab's compute budget fell to an academic team with a few thousand dollars of rented GPU time. The recipe, which pairs a learned guess about hidden information with a short search over sampled possibilities, is the same pattern that matters for negotiation, security and other settings where an agent has to act without seeing the whole picture. Our lesson on imperfect-information games explains the underlying idea.

The caveat

Twenty games against one person is not proof that no human can beat Ataraxos. The authors themselves note that the games were not independent, because Niemeijer could learn from one game to the next, so the usual statistical test overstates the certainty. The GitHub release includes code and pretrained files, but it has not been confirmed that everything needed to reproduce the exact match-strength agent is there. And success in games with clear rules does not by itself show the method will work in messy real-world decisions.


Primary source, verified: read the paper → (arXiv 2511.07312)

Key questions

Did Ataraxos beat DeepMind's DeepNash?

No head-to-head match took place. The Ataraxos authors say DeepMind reported that DeepNash's old code no longer worked, so the two systems were never played against each other.

Why is Stratego harder for AI than chess or Go?

Stratego hides information: you can see where your opponent's pieces are but not what rank they are until they fight. An AI has to reason about many possible hidden armies at once, which the paper puts at more than 10 to the power of 33 possible piece configurations.

Can I run Ataraxos myself?

The code is public under an MIT license on GitHub with a pretrained directory, and it needs Linux and a CUDA GPU. Whether every file needed to reproduce the exact match-strength search agent is included has not been confirmed.
Cite this

APA

Ground Truth. (2026, October 3). Ataraxos beat Stratego's most decorated champion 15-1 after a training run that cost under $8,000. Ground Truth. https://groundtruth.day/news/ataraxos-beats-stratego-champion-15-1-on-a-budget.html

BibTeX

@misc{groundtruth:ataraxos-beats-stratego-champion-15-1-on-a-budget,
  title  = {Ataraxos beat Stratego's most decorated champion 15-1 after a training run that cost under $8,000},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {oct},
  url    = {https://groundtruth.day/news/ataraxos-beats-stratego-champion-15-1-on-a-budget.html}
}

Topics: research · games · reinforcement-learning · imperfect-information · search · self-play

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.