# Ataraxos AI beat a Stratego champion by planning around hidden information

Agent contract: 1.2.0

Ataraxos won 15 of 20 games against a Stratego champion. Its hidden-information planning and estimated training compute below $8,000, explained.

Canonical: https://brightaifuture.com/discoveries/ataraxos-stratego-hidden-information
Format: discovery
Source publication: 2026-09-30
Bright publication: 2026-10-10
Substantive update: None recorded
Evidence and review: Emerging; confidence: unassessed; approved; ai-assisted. Owner explicitly approved Bright publication and branded distribution after the source-checked explainer was reviewed. Final Nature publication metadata and indexed text were checked; direct final publisher PDF was not read. Detailed numerical claims use the accessible author manuscript v2 and primary match archive. September 30 Nature publication, October 8 CMU account and November 2025 original preprint are kept distinct. Under-$8,000 is an estimate for a specific training run at 2025 rental prices, excluding broader research and development. No direct DeepNash match, independent reproduction by Bright or demonstrated real-world deployment is claimed. Original AI-generated conceptual hero reviewed for rank symbols, concealed identities and non-documentary labeling.

## Editorial image

Conceptual Stratego-style board with visible blue ranks, flag and bomb, and concealed red piece identities. Three hypothesis panels show different possible identities at the same fixed locations. Text labels the image Ataraxos in Stratego and says it is not an actual match position.: https://brightaifuture.com/media/content/82f3a5fcf6a66a84ad0a0c729f1db544b6706183da3c553de618282e1240be60.png

Credit: Original Bright illustration, AI-generated; conceptual explanation of hidden-information planning, not an actual match position.. License: Original Bright illustration approved for Bright publication.

Source: https://brightaifuture.com/media/content/82f3a5fcf6a66a84ad0a0c729f1db544b6706183da3c553de618282e1240be60.png

explanatory; not a recorded Ataraxos match, a game screenshot, measured output or a demonstrated real-world deployment.

## The story

You can look straight at a Stratego board and still have little idea what you are facing. The approaching piece might be powerful. It might be weak. Its owner might want you to mistake one for the other.

That makes the game a useful test of a familiar problem: how do you choose a sensible next step when another person knows something you do not? The research system Ataraxos has produced a striking answer within the game's rules. Its results point to progress in making computer decisions under uncertainty, with a training approach that could be easier for other researchers to explore. [Carnegie Mellon's account](https://www.cs.cmu.edu/news/2026/ai-tackles-stratego) explains the challenge.

The research appeared in [Nature on September 30, 2026](https://www.nature.com/articles/s41586-026-11036-y). An earlier version was [posted in November 2025](https://arxiv.org/abs/2511.07312). This is a closer look at the work, rather than a claim that the games happened this week.

## Knowing where a piece is isn't enough

Each player starts with 40 pieces, including a flag, bombs and pieces of different ranks. Players can see where their opponent's pieces stand, while many identities remain concealed as they move. Encounters reveal pieces. The main objective is to capture the opponent's flag.

A tempting move can therefore have very different consequences depending on what is hidden. There is a second complication: your moves give the other player clues about you. A bluff works partly because it is unexpected. Repeat it too predictably and its value changes. [DeepMind's explanation of Stratego](https://deepmind.google/blog/mastering-stratego-the-classic-game-of-imperfect-information/) describes why this combination makes straightforward search difficult.

## Practice first, then think about this position

Ataraxos combines extensive practice with planning at the moment a move is needed. During training, it plays against itself to learn a general strategy. During a game, it revisits the choice in front of it using a model of the likely identities of hidden pieces. [CMU describes these complementary steps](https://www.cs.cmu.edu/news/2026/ai-tackles-stratego).

The planning step samples plausible hidden arrangements and considers how candidate moves could play out. That lets it concentrate computation on possibilities supported by the information available, instead of attempting to inspect every conceivable arrangement. [MIT's account](https://news.mit.edu/2026/game-playing-ai-stratego-new-champ-0930) describes the learned strategy as the starting point for this extra calculation.

The useful distinction is between knowing and estimating. A plausible hidden board is a hypothesis. Keeping several possibilities in play gives the system a way to compare risks before committing to an action. That is an appealing design principle for decision tools: their reasoning should leave room for the parts of a situation they cannot observe.

## A strong result with a clear opponent

The headline test was a 20-game series against Pim Niemeijer, a four-time world champion. Ataraxos won 15 games, lost one and drew four. Niemeijer knew the system would keep a fixed strategy, giving him an opportunity to look for weaknesses across the series. The researchers have made the [game archive public](https://ataraxosai.github.io/).

That is persuasive evidence of strength against an exceptional player. A broader picture would include more repeated series against different leading players. An archive is also valuable in its own right: readers can inspect what the reported result refers to, rather than relying solely on a victory headline.

## A smaller training bill matters

The [researchers' detailed methods](https://arxiv.org/html/2511.07312v2) report one week on 16 NVIDIA H100 graphics processors for reinforcement learning, followed by four days on four H100s for the hidden-information model. They estimate that training run at less than $8,000 using 2025 rental prices. That is a compute estimate, rather than an accounting of salaries, development and every experiment behind the project.

Earlier work was already substantial. DeepMind's DeepNash reached an expert-level, all-time top-three ranking on the Gravon online platform in 2022. [DeepMind documented that achievement](https://deepmind.google/blog/mastering-stratego-the-classic-game-of-imperfect-information/). The [Ataraxos authors' comparison](https://arxiv.org/html/2511.07312v2) is based on training resources and separate evaluations; the two systems did not play a direct match.

Lower compute requirements could let more university groups investigate these methods. There is something concrete to investigate: the [public Stratego repository](https://github.com/AtaraxosAI/stratego) contains training instructions, an MIT software license and a pretrained-model directory. Its documented requirements include Linux and a CUDA-compatible GPU. Public code helps make a result inspectable; successful independent reproduction still has to be established. Bright has not run the software.

## The question beyond winning

The [paper](https://www.nature.com/articles/s41586-026-11036-y) also reports wins against leading humans in Barrage Stratego, and stronger benchmark results in Hanabi and dou dizhu. Those settings include cooperation and team play, broadening the evidence beyond one adversarial board game.

Real-world usefulness remains to be demonstrated. A decision tool would need a credible model of its setting, ways to recognize when its assumptions fail, and explanations people can inspect. The researchers themselves identify understandable, auditable recommendations as an important next step in [MIT's account](https://news.mit.edu/2026/game-playing-ai-stratego-new-champ-0930).

For now, the earned optimism is specific: a research team made planning around extensive hidden information work impressively in controlled games, and left tools other researchers can examine. That gives the next stage of work a stronger starting point.

For another perspective on access to AI, read Bright's [guide to using AI on your own computer](https://brightaifuture.com/discoveries/your-ai-your-computer). Get more source-checked developments in [Bright Weekly](https://brightaifuture.com/newsletter).

## Provenance and history

{
  "dates": {
    "captureDate": "2026-10-10",
    "eventDate": null,
    "lastReviewedDate": "2026-10-10",
    "publicationDate": "2026-09-30"
  },
  "provenance": {
    "origin": "editorial",
    "externalId": "https://www.nature.com/articles/s41586-026-11036-y"
  },
  "revisions": [
    {
      "id": "revision:ataraxos-first-publication-20261010",
      "recordedAt": "2026-10-10",
      "sourceIds": [
        "source:ataraxos-nature",
        "source:ataraxos-author-v2",
        "source:ataraxos-history",
        "source:ataraxos-cmu",
        "source:ataraxos-mit",
        "source:ataraxos-games",
        "source:ataraxos-stratego-code",
        "source:ataraxos-deepnash-prior"
      ],
      "summary": "Published a substantial source-checked research explainer with controlled-game evidence, cost scope and comparison limits, rather than presenting the older research as breaking news."
    }
  ],
  "corrections": []
}

## Original sources

- [Scalable decision-making for games of imperfect information · Nature, September 30, 2026](https://www.nature.com/articles/s41586-026-11036-y)
- [Scalable Decision Making for Games of Imperfect Information · author manuscript v2, October 4, 2026](https://arxiv.org/html/2511.07312v2)
- [Author manuscript version history · first posted November 10, 2025](https://arxiv.org/abs/2511.07312)
- [CMU researchers develop AI that tackles hidden information in Stratego · October 8, 2026](https://www.cs.cmu.edu/news/2026/ai-tackles-stratego)
- [Game-playing AI brings a new champ to Stratego · MIT, September 30, 2026](https://news.mit.edu/2026/game-playing-ai-stratego-new-champ-0930)
- [Ataraxos vs. Pim Niemeijer · public 20-game Stratego archive](https://ataraxosai.github.io/)
- [Ataraxos Stratego repository · training instructions, MIT license and pretrained files](https://github.com/AtaraxosAI/stratego)
- [Mastering Stratego, the classic game of imperfect information · DeepMind, December 1, 2022](https://deepmind.google/blog/mastering-stratego-the-classic-game-of-imperfect-information/)

## Continue exploring

- [Your AI doesn’t have to live on someone else’s computer](https://brightaifuture.com/discoveries/your-ai-your-computer)
