AI Becomes New Stratego Champion

Category :

AI

Posted On :

Share This :

One day, human decision-makers may be assisted in choosing the best tactics to outwit adversaries in complex scenarios, such as military maneuvers, by a new AI system that excels at difficult games with hidden knowledge.

 

Researchers from MIT, Carnegie Mellon University, New York University, and Stanford University used improvements in machine learning to create an AI that significantly outperformed top-ranked human players in the board wargame Stratego—something no AI system had previously been able to accomplish.

 

Stratego, a two-player imperfect information game where the identity of the opponent’s pieces is concealed, is frequently used as a standard to evaluate how well strong AI models can think strategically.

 

The researchers coupled innovative methods designed for deliberate decision-making in hidden-information circumstances with effective training algorithms to create their model.

 

The AI system was significantly less expensive and required less processing power to train, and it outperformed the next best models in Stratego. Additionally, the system demonstrated its generalizability for a wide range of use cases by outperforming the best human players in additional strategic games with varied rules and designs.

 

The AI system might be modified to assist people in solving a variety of real-world issues involving hidden knowledge, like cybersecurity or business negotiations.

 

You frequently don’t have the luxury of listing every option in the kinds of incomplete information assignments you would encounter in real life. Simply put, there are too many. According to Gabriele Farina, an assistant professor in the Department of Electrical Engineering and Computer Science (EECS), principal investigator at the Laboratory for Information and Decision Systems (LIDS), and senior author of a paper on this AI system, “having general-purpose AI algorithms that can provably perform this challenging task so well is a big step forward.”

 

Lead author Samuel Sokota, a graduate student at Carnegie Mellon; assistant professor Eugene Vinitsky of NYU; professor Zico Kolter of Carnegie Mellon; Stanford graduate student Hengyuan Hu; and MIT graduate student Zhiyuan Fan of EECS accompany him on the paper. The study was published in Nature today.

 

Information That Is Hidden

Imperfect information problems abound in the world.

 

Some parties in these exchanges have information that others do not. For example, military personnel probably don’t fully understand enemy positions, and traders in financial markets might not understand the reasoning behind other people’s trades.

 

The decisions parties make, and the decisions they decide not to make, are so entwined when there is hidden information that it is very challenging to decide what should be done next.

 

The more you bluff, the less each bluff is worth and the more your opponent anticipates it. Sokota says, “It’s not clear how to reason about that.” “It differs greatly from a situation like chess, where the best move remains the best move regardless of how many times you’ve played.”

 

Stratego is frequently used to simulate scenarios with incomplete information. Players position 40 pieces on their side of a board and then move pieces across the board to capture their opponent’s flag in this board wargame that is similar to military chess.

 

However, each piece’s identity is kept a secret until they collide, at which point the lower-ranking piece is eliminated.

 

Stratego is very challenging for an AI system to play successfully since there are more than 10 to the 66th power conceivable piece configurations, which is considerably more than in chess.

 

Previous attempts, like Google’s DeepMind, depended on complex processes that were expensive and computationally intensive. However, despite spending millions of dollars on training, these models were still unable to outperform the best human Stratego players.

 

“You may have to deal with an explosion of potential universes when using Stratego. According to Farina, “AI techniques that were developed for games like poker could definitely not scale in this setting.”

 

The goal of the MIT researchers was to create a complete AI system that could outperform humans at a lower cost. They named it Ataraxos, which is a Greek word for “one who is unbothered or free from anxiety.”

 

Two-Pronged Strategy

The researchers used a method known as self-play reinforcement learning to train the model in order to create Ataraxos. To develop a solid “blueprint strategy” for succeeding at Stratego, the model repeatedly competes with itself.

 

They created particularly effective algorithms that allowed Ataraxos to learn far more quickly than previous techniques without becoming bogged down attempting to anticipate every potential move. This improves performance and lowers training expenses.

 

According to Farina, “our system achieves strictly higher playing strength than DeepNash (DeepMind’s system) while using less than one hundredth of the training examples and less than one thirtieth of the self-play games, indicating a massive improvement in efficiency.”

 

Ataraxos starts each round of play by setting up the board and considering its next move using the blueprint strategy.

 

However, it uses a process known as “decision-time planning” to fine-tune its decisions on the spot before taking action. Before deciding on the next move, the system makes use of a generative model that uses probabilities to estimate the likely identities of the opponent’s hidden pieces.

 

We employ decision-time planning to determine the most likely state of the board rather than only speculating. We can really focus on the particular board and opponent we are facing by using this generative model, according to Farina.

 

The final component that allowed Ataraxos to reach superhuman performance was the creative application of this generative model for decision-time planning.

 

Ataraxos scored a 39-2 record versus the best human players at the Stratego world championship and defeated the world’s strongest Stratego player by a record 15-1-4. “Unlike humans, ataraxos are adept at estimating risk. If their most precious piece is revealed, a person could get extremely alarmed, yet a bot can remain quite calm. According to Farina, it doesn’t overcorrect and reveal its secrets.

 

Additionally, the researchers modified Ataraxos for additional imperfect information games, such as Hanabi, a multiplayer cooperative card game, Barrage Stratego, a faster-paced variation with fewer pieces, and Dou dizhu, a game in which two players work together against a third.

 

The system demonstrated the generality of this approach by achieving superhuman performance in each case.

 

The researchers hope to incorporate interpretability features into Ataraxos in the future so that the system can provide a human-understandable explanation for its decisions.

 

“We need a way to audit the model’s decisions before adoption can occur because humans must have the last say in whether a recommendation is followed.” “I hope these algorithms can be the foundation for a lot more work to come, but we still have a long way to go,” Farina says.

 

The Office of Naval Research, the Department of Civil and Urban Engineering at New York University, the C2SMART Center, the National Science Foundation, and a Schmidt Sciences AI2050 Early Career Fellowship are some of the funding sources for this study.