An AI system has finally conquered Stratego, the classic board game that stumped machines for decades — and it did it on a shoestring budget. A team of researchers from Carnegie Mellon University, MIT, New York University and Stanford University built an AI called Ataraxos that defeated Pim Niemeijer, widely regarded as the best Stratego player of all time, by a score of 15 games to 1 with four draws, Ars Technica reports.
The result is striking less for the scoreline than for the price tag. Ataraxos trained on just 16 GPUs for about a week, plus four additional GPUs for four days to train a companion model — a run the team says cost only a few thousand dollars. DeepMind's DeepNash, the 2022 system that first brought serious AI attention to the game, trained for two to three months on 1,024 of Google's specialized TPU chips, a run the Ataraxos team estimates would cost between $3 million and $4.5 million at 2025 prices. For more context on this story, see our ongoing AI industry coverage.
Why Stratego Stumped AI for Decades
Deep Blue beat Garry Kasparov at chess in 1997. AlphaGo defeated Lee Sedol at Go in 2016. Poker bots have beaten professionals for years. Stratego held out — even against DeepMind's considerable resources.
The game's difficulty comes from its unusual combination of hidden information, scale and length. Each player commands 40 pieces representing military ranks, from a marshal down to a spy, plus bombs and a flag. You win by capturing the opponent's flag. Your opponent can see where your pieces are, but not what they are. Identities are revealed only when two pieces collide, and the weaker piece is removed.
"In Stratego, there's 40 pieces on the board that could be in any order," Gabriele Farina, an MIT computer scientist and co-author of the study, told Ars Technica. That allows more than a decillion possible setups — dwarfing the 1,326 possible hands in Texas Hold'em, a game computers cracked years ago. Length compounds the problem: "In chess, usually the game lasts 40 moves, but in Stratego, a game can easily last 2,000 moves," Farina said.
Then there is bluffing. Moving a weak piece as if it were a marshal can scare an opponent off, but bluff too often and threats mean nothing; never bluff and you become predictable. That balancing act is what kept earlier systems like DeepNash from reaching the top.
"There's something super distinctive about Stratego, which is that it is a massive amount of hidden information that unfolds over a very long time scale," said Eugene Vinitsky, a researcher at NYU and another co-author.
A Second Neural Network That Guesses
Ataraxos learned the same basic way DeepNash did: by playing against itself, 163 million games in total, reinforcing moves that led to wins and damping those that did not. The team's first improvement was in how much the system adjusted after each game, because hidden information tends to send self-play algorithms in circles. Ataraxos made large strategy changes early in training and increasingly careful ones later.
The bigger breakthrough was something DeepNash never had: thinking ahead before each move. Search-based refinement before acting is what made AlphaGo powerful, but DeepMind could not make it work in Stratego because the search space was too vast.
The solution was a second neural network — a belief model, trained to guess the opponent's hidden pieces from how they had been moving. Instead of iterating through every possible arrangement, Ataraxos samples plausible ones, plays out candidate moves in each, and picks based on the outcomes. Lead author Samuel Sokota and Farina also wrote a custom simulator that runs millions of moves per second on graphics cards, which is how an academic lab could afford the training run. "At the scale that we are in academia, we don't really have access to an entire field of GPUs," Farina said.
The Human Champion Got One Back
Niemeijer, who holds four world championships and has spent more than 600 weeks as the world's top-ranked player, played 20 online games against Ataraxos over three weeks, earning $100 for each win. He knew the AI would not adapt to him, giving him time to hunt for weaknesses. He managed to win exactly once.
Researchers say that single loss was not a flaw. Good Stratego requires randomizing your setup, so luck always plays a role. "Even a perfect strategy, sometimes it will just lose," Farina said.
Human players at the 2025 Stratego World Championship fared far worse: attendees who challenged Ataraxos won only 2 of 40 games. The bot even changed how people play, skewing the metagame with unusual choices like tucking its flag into a corner behind just two bombs — a rarely played setup.
The system's name comes from the ancient Greek word for calm, and the researchers say that quality shows in its play. While a human might gamble wildly to recover from a deficit, Ataraxos works its way back slowly and methodically. "We would watch the bot 'bluff' its way back from like a two percent victory probability, very, very casually," Vinitsky said.
What It Means Beyond the Board
Beating humans at hidden-information games is more than a parlor trick. Techniques for reasoning about an opponent's private state underpin real applications where information is incomplete — negotiation systems, cybersecurity, economic modeling and multi-agent coordination, to name a few.
The efficiency angle may prove the most influential part of the story. A frontier lab spent months of compute on this problem in 2022; an academic team replicated and exceeded the result for roughly the price of a used car in 2026. As AI research becomes synonymous with massive compute budgets, Ataraxos is a reminder that algorithmic ideas — here, a belief model that tames a huge search space — can still matter more than raw scale.
The team's paper describes a machine that stays calm under uncertainty, ignores knowledge a human could not ignore, and bluffs better than the best in the world. Those are useful skills, on the board and off it.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →