Bright
← Living questionsAnalysis / Science

Ataraxos AI beat a Stratego champion by planning around hidden information

Ataraxos won 15 of 20 games against a Stratego champion. Its hidden-information planning and estimated training compute below $8,000, explained.

Maturity
Emerging, stage 2 of 4
Support
8 sources · paper, institution, dataset, repository
Evidence detail
How we know ↓
Conceptual Stratego-style board with visible blue ranks, flag and bomb, and concealed red piece identities. Three hypothesis panels show different possible identities at the same fixed locations. Text labels the image Ataraxos in Stratego and says it is not an actual match position.
explanatory · Original Bright illustration, AI-generated; conceptual explanation of hidden-information planning, not an actual match position. · Original Bright illustration approved for Bright publication. Not a recorded Ataraxos match, a game screenshot, measured output or a demonstrated real-world deployment. Source ↗ · View the full-size image ↗

You can look straight at a Stratego board and still have little idea what you are facing. The approaching piece might be powerful. It might be weak. Its owner might want you to mistake one for the other.

That makes the game a useful test of a familiar problem: how do you choose a sensible next step when another person knows something you do not? The research system Ataraxos has produced a striking answer within the game's rules. Its results point to progress in making computer decisions under uncertainty, with a training approach that could be easier for other researchers to explore. Carnegie Mellon's account explains the challenge.

The research appeared in Nature on September 30, 2026. An earlier version was posted in November 2025. This is a closer look at the work, rather than a claim that the games happened this week.

Knowing where a piece is isn't enough

Each player starts with 40 pieces, including a flag, bombs and pieces of different ranks. Players can see where their opponent's pieces stand, while many identities remain concealed as they move. Encounters reveal pieces. The main objective is to capture the opponent's flag.

A tempting move can therefore have very different consequences depending on what is hidden. There is a second complication: your moves give the other player clues about you. A bluff works partly because it is unexpected. Repeat it too predictably and its value changes. DeepMind's explanation of Stratego describes why this combination makes straightforward search difficult.

Practice first, then think about this position

Ataraxos combines extensive practice with planning at the moment a move is needed. During training, it plays against itself to learn a general strategy. During a game, it revisits the choice in front of it using a model of the likely identities of hidden pieces. CMU describes these complementary steps.

The planning step samples plausible hidden arrangements and considers how candidate moves could play out. That lets it concentrate computation on possibilities supported by the information available, instead of attempting to inspect every conceivable arrangement. MIT's account describes the learned strategy as the starting point for this extra calculation.

The useful distinction is between knowing and estimating. A plausible hidden board is a hypothesis. Keeping several possibilities in play gives the system a way to compare risks before committing to an action. That is an appealing design principle for decision tools: their reasoning should leave room for the parts of a situation they cannot observe.

A strong result with a clear opponent

The headline test was a 20-game series against Pim Niemeijer, a four-time world champion. Ataraxos won 15 games, lost one and drew four. Niemeijer knew the system would keep a fixed strategy, giving him an opportunity to look for weaknesses across the series. The researchers have made the game archive public.

That is persuasive evidence of strength against an exceptional player. A broader picture would include more repeated series against different leading players. An archive is also valuable in its own right: readers can inspect what the reported result refers to, rather than relying solely on a victory headline.

A smaller training bill matters

The researchers' detailed methods report one week on 16 NVIDIA H100 graphics processors for reinforcement learning, followed by four days on four H100s for the hidden-information model. They estimate that training run at less than $8,000 using 2025 rental prices. That is a compute estimate, rather than an accounting of salaries, development and every experiment behind the project.

Earlier work was already substantial. DeepMind's DeepNash reached an expert-level, all-time top-three ranking on the Gravon online platform in 2022. DeepMind documented that achievement. The Ataraxos authors' comparison is based on training resources and separate evaluations; the two systems did not play a direct match.

Lower compute requirements could let more university groups investigate these methods. There is something concrete to investigate: the public Stratego repository contains training instructions, an MIT software license and a pretrained-model directory. Its documented requirements include Linux and a CUDA-compatible GPU. Public code helps make a result inspectable; successful independent reproduction still has to be established. Bright has not run the software.

The question beyond winning

The paper also reports wins against leading humans in Barrage Stratego, and stronger benchmark results in Hanabi and dou dizhu. Those settings include cooperation and team play, broadening the evidence beyond one adversarial board game.

Real-world usefulness remains to be demonstrated. A decision tool would need a credible model of its setting, ways to recognize when its assumptions fail, and explanations people can inspect. The researchers themselves identify understandable, auditable recommendations as an important next step in MIT's account.

For now, the earned optimism is specific: a research team made planning around extensive hidden information work impressively in controlled games, and left tools other researchers can examine. That gives the next stage of work a stronger starting point.

For another perspective on access to AI, read Bright's guide to using AI on your own computer. Get more source-checked developments in Bright Weekly.

How we know8 sources · checked 2026-10-10 · no corrections

Original sources

  1. Scalable decision-making for games of imperfect information · Nature, September 30, 2026 ↗ · paper
  2. Scalable Decision Making for Games of Imperfect Information · author manuscript v2, October 4, 2026 ↗ · paper
  3. Author manuscript version history · first posted November 10, 2025 ↗ · paper
  4. CMU researchers develop AI that tackles hidden information in Stratego · October 8, 2026 ↗ · institution
  5. Game-playing AI brings a new champ to Stratego · MIT, September 30, 2026 ↗ · institution
  6. Ataraxos vs. Pim Niemeijer · public 20-game Stratego archive ↗ · dataset
  7. Ataraxos Stratego repository · training instructions, MIT license and pretrained files ↗ · repository
  8. Mastering Stratego, the classic game of imperfect information · DeepMind, December 1, 2022 ↗ · institution

Institutions: Carnegie Mellon University · Massachusetts Institute of Technology

Maturity
Emerging
Source published
2026-09-30
Captured
2026-10-10
Last source review
2026-10-10
Editorial method
AI-assisted source review
Place / relevance
University research in controlled games · unspecified

Bright compared this account with the linked original and supporting sources and kept reported, budgeted, projected, and observed claims distinct. Bright did not independently audit the underlying records.

Maturity describes the tested or operational setting. Confidence describes support for the particular claim; one does not determine the other.

Revision & correction history

2026-10-10T22:43:40.128Z · Published a substantial source-checked research explainer with controlled-game evidence, cost scope and comparison limits, rather than presenting the older research as breaking news.

No corrections recorded.

Keep exploring

Explore the shared question in another setting. These connections do not imply replication.

An editorial connection recorded in Bright’s evidence catalog.Your AI doesn’t have to live on someone else’s computer ↗Related development · not a replicationYour AI doesn’t have to live on someone else’s computer ↗

Keep looking closer.

See what changed at Bright ↗

Add Bright to your Google Preferred Sources ↗

Suggest a correction · Bright on TikTok