Learn · Intermediate
Imperfect-information games: how AI plays when it cannot see the whole board
An imperfect-information game is one where players hold private facts their opponents cannot see, such as the cards in a poker hand or the ranks of hidden pieces in Stratego. To play well, an AI must reason about what its opponent might be hiding, keep its own secrets unpredictable, and choose actions that work across every situation it cannot tell apart. These games are the testbed for AI that has to act on partial information, which describes most real decisions.
Perfect versus imperfect information
In chess and Go, everything is on the board. Both players see the same position, and in principle the best move depends only on that position. That is why game-tree search, as in Monte Carlo tree search, works so well there: you can evaluate a position on its own and back the values up the tree.
Poker breaks that. When you face a big bet, the right response depends on what cards your opponent holds, which you cannot see, and on how they would bet with every hand they might have. Game theorists group the situations a player cannot tell apart into an information set. You have to choose one policy for the whole set, because from your seat they all look identical.
Why ordinary search fails
Two problems appear at once. First, the value of a situation depends on beliefs. Holding a pair of kings is great against a random hand and terrible against someone who only bets big with aces. Second, the right strategy is usually mixed: a player who always bets with strong hands and always checks with weak ones is trivially read. Good play bluffs some fraction of the time, at frequencies that leave the opponent unable to profit by guessing.
A useful analogy is a penalty kick. If the striker always shoots left, the goalkeeper always dives left. The only unexploitable plan is to randomize in the right proportions. In two-player zero-sum games, a strategy profile where neither side can gain by changing alone is a Nash equilibrium, and finding or approximating one is the standard goal.
Counterfactual regret minimization
The breakthrough algorithm came from Martin Zinkevich, Michael Bowling and colleagues at the University of Alberta in 2007: counterfactual regret minimization (CFR). The AI plays the game against itself over and over. At every information set it tracks regret: how much better it would have done, on average, by always picking each alternative action. It then plays actions in proportion to their positive regret. In two-player zero-sum games, the average of the strategies it played converges toward a Nash equilibrium. It is a form of self-play, with a guarantee attached.
Belief states and search come back
Pure CFR needs the whole game written out, which is impossible for no-limit poker. The next step brought search back, carefully.
DeepStack (Matej Moravcik and colleagues, 2017) re-solved the game during play from the current situation, using a neural network to estimate values at a limited depth. Libratus (Noam Brown and Tuomas Sandholm at Carnegie Mellon) beat four top professionals at heads-up no-limit Texas hold'em in a 2017 match, combining a precomputed strategy with real-time solving of the parts of the game it reached. Pluribus (2019) extended superhuman play to six-player poker. ReBeL (2020) unified the ideas: it treats the probability distribution over everyone's private information, given the public actions so far, as a public belief state, and runs reinforcement learning and search over those beliefs the way AlphaZero runs them over board positions.
A belief state is the key concept. It is a probability distribution over the hidden facts, updated as evidence arrives, like a detective's ranked list of suspects that changes after each interview.
From cards to Stratego
Stratego is harder than poker in one sense: the hidden information is enormous. Each player privately arranges 40 pieces, and you learn an enemy piece's rank only when it fights. DeepMind's DeepNash (2022) reached expert human level with model-free self-play and no search at all, using a method called Regularised Nash Dynamics.
This week Ground Truth reported on Ataraxos, which beat the most decorated Stratego player in history 15 games to 1. Its recipe follows the pattern above: self-play to learn a strong policy, a learned belief network that predicts the ranks of hidden enemy pieces, and a short decision-time search that samples several plausible hidden armies from that belief model and checks which move does best across them. The training run cost under $8,000 by the authors' estimate.
Why it matters beyond games
Negotiation, auctions, cybersecurity, fraud detection and military planning all involve opponents with private information and an incentive to mislead. Card and board games give these problems clean rules and a clear score. The techniques that work there, especially explicit beliefs about hidden information plus search over sampled possibilities, are the ones researchers try first when an AI agent must act without seeing the whole picture. Related ideas appear in Markov decision processes, the fully observed case, and in Bayesian updating, the math of revising beliefs.
Limits to keep in mind
The equilibrium guarantees hold cleanly for two-player zero-sum games. With more players or cooperative goals, as in Pluribus's six-player poker or the card game Hanabi, an equilibrium strategy is no longer guaranteed to be the best choice, and systems rely more on empirical strength. Sampling a few hidden states, as Ataraxos does, is an approximation, not an exhaustive search. And a result against one champion over a limited number of games shows strength, not unbeatability.
Regret Minimization in Games with Incomplete Information (Zinkevich et al., 2007)
DeepStack: Expert-Level Artificial Intelligence in No-Limit Poker (Moravcik et al., 2017)
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals (Brown and Sandholm, 2018)
Superhuman AI for multiplayer poker (Brown and Sandholm, 2019)
Combining Deep Reinforcement Learning and Search for Imperfect-Information Games (ReBeL, Brown et al., 2020)
Mastering the Game of Stratego with Model-Free Multiagent Reinforcement Learning (DeepNash, Perolat et al., 2022)
Key questions
What makes a game one of imperfect information?
Why can't chess-style search just be applied to poker?
What is a belief state?
Cite this
APA
Ground Truth. (2026, October 3). Imperfect-information games: how AI plays when it cannot see the whole board. Ground Truth. https://groundtruth.day/learn/imperfect-information-games-and-belief-states.html
BibTeX
@misc{groundtruth:imperfect-information-games-and-belief-states,
title = {Imperfect-information games: how AI plays when it cannot see the whole board},
author = {{Ground Truth}},
year = {2026},
month = {oct},
url = {https://groundtruth.day/learn/imperfect-information-games-and-belief-states.html}
}