> The algorithm also learned far faster—it played about 34 times fewer games than DeepNash, and still ended up much stronger.
Imo, this is the critical piece and what makes the AI work at all.
With hidden information games, the best move depends on information you don’t have. So a move could be good or bad, it just depends on something that’s impossible to know.
You’d like to search ahead, meaning “if I do this they will do that” but that’s impossible since you don’t even know what the opponent can do because you don’t know their hidden state.
If the possible hidden states are randomly distributed, you are screwed. It’s just like rock paper scissors: there’s no best move if your opponent is unpredictable.
However if you can quickly learn to predict their moves, it becomes possible to make informed decisions about what to do.
dmurray
Oh no! Stratego had been on my mind as something we just hadn't tried hard enough to make a winning bot for, including the DeepMind effort from 2022. I was planning to make the first one.
I thought this was slightly less crank-coded than trying to prove the Riemann Hypothesis, but maybe these days you just ask Claude to do that and it tells you there's a counterexample at 1 + πi that no one ever noticed before.
show comments
SwellJoe
I loved Stratego so much as a kid. But, I eventually couldn't find anyone to play with me because I crushed everyone, including my dad who was much better than me at chess. But, I never would have thought it'd be a game that models would have a hard time with. It feels relatively simple. And, inexplicably, I never even thought that there might be serious players...I've kept the board game, one of the very few things I have from young childhood, but it's been twenty years since I played. I guess it's time to find an online Stratego. Surely someone in the whole world can beat me.
Too bad I never played against an AI before they cracked it.
show comments
rovr138
> Now, a team of researchers from Carnegie Mellon, MIT, New York University, and Stanford University has done it. Their AI, called Ataraxos, beat Pim Niemeijer, arguably the best Stratego player of all time, 15 games to one, with four draws. And it took just 16 GPUs and a few thousand dollars to train it.
Just 16 GPUs, and a few thousand dollars?
What about “researchers from Carnegie Mellon, MIT, New York University, and Stanford University” this wasn’t just anyone.
show comments
smokel
This puts the earlier "Mastering the Game of Stratego with Model-Free Multiagent Reinforcement Learning", 2022 [1] in some perspective. Apparently the "mastering" in 2022 wasn't quite there yet. Four years later, the new approach seems to actually be better than humans.
This approach also works for Hanabi, which is a very interesting game. You can't see your own cards, but the other players can. I bought the game because someone on a reinforcement learning podcast [2] mentioned it, and actually played it multiple times.
I think that what makes these games beatable repeatedly is that they're static. Not saying an algorithm properly trained won't play better than the average player a game like MtG, or my own https://aethersummon.com (specially now while it has under 90 possible scrolls only) but if you have a regular release cadence (say weekly or bi-weekly) of relevant new "cards", then I think the playing field is much more even for humans.
Those new additions can invalidate the whole training data by a single new "card" that changes completely the dynamics and would be easy for a player to understand and incorporate but not for an algorithm (perhaps with enough compute to re-train it regularly it could) - that along with the decision trees being orders of magnitude deeper, wider and with more conditionalities than go, chess or stratego - even through the same turn with the same cards available and same table state - would probably pose much harder problems for a compute bound algo.
I recall playing this game as a preschooler. It was mostly psychology and bluff.
Very interesting.
show comments
gavinlilly
I wonder how capable current AIs are with the "silent defense" variant of Stratego [1]? The article states the high level of uncertainty presents a challenge. With silent defense the uncertainty is even higher.
[1] https://www.hasbro.com/common/instruct/Stratego.PDF
"When an attack is made, the attacker is the only player who has to declare the number of his or her piece. The defender does not reveal the number of his or her piece, but resolves the attack by removing
whatever piece has a lower number from the gameboard. Players keep their own captured pieces. Exception: when a Scout attacks, the defender must reveal the number of his or her piece.
show comments
osti
Wait, wasn't there that strong stratego bot that came out from deepmind in 2022?
While interesting from a technology stand point I get disheartened each time I see this sort of news.
I am particularly disappointed that it has influenced how people play the game.
The joy comes from the journey and the experience.
Look at competitive chess and Go and how they have fundamentally been transformed. It's not better and now the box is opened, it can't be closed.
show comments
askjdfksdbfhk
I would really love to see a serious research effort take a crack at contract bridge. Bridge, like Stratego, is an imperfect information game with a big hidden information space. Bridge also adds another wrinkle of explainability which is, I think, very interesting.
Bridge is played as a pair vs pair game, with North/South and East/West being the two pairs and seated around the table in these compass directions. A bridge hand consists of two phases: there is first an auction phase, where players go around the table bidding on contracts (agreeing to take a certain number of tricks with a certain trump suit) until a final contract is decided. Then there is the cardplay phase, where the player who won the auction is the declarer, their partner is the dummy, and the other pair are defenders. The dummy's hand is placed face up on the the table and the declarer controls which cards are played from dummy, so the cardplay phase is effectively played by only three players now, with each of the three knowing one common hand (dummy) and one private hand (their own) and not knowing the other two hands.
In both the auction and (for the defense) the cardplay phases, it is important for players to exchange some information about their hand to their partner. However, any information you exchange about your own hands also helps your opponents. You might naturally conclude that you want to come up with some secret scheme to exchange information which your opponents don't know (and it is even possible to exchange encrypted information which your opponents can't know--if the defense is known to hold a certain card, but declarer doesn't know in which hand it is, the defense could say that a signal means one thing if the card is in one defender's hand, but means a different thing if it's in the other defender's hand).
But it turns out that this ends up being very uninteresting to play, so instead, when playing bridge, there is an important rule: all of your partnership agreements must be public. If a certain bid that I make promises that I have at least 5 spades in my hand, it is the opponents' right to know that this is our agreement. You must be able to explain the information which your action provides, and you must be able to use the information that the opponents give you themselves.
This poses several problems for self-play reinforcement learning. First, a naive self-play approach will produce agreements that cannot be explained to a human. What really needs to happen is that your partner, when determining what hands you might have as part of search, must not do so simply by sampling its own system (ie by asking what it itself would have done with hand X or hand Y). The information and possibilities really need to be mediated by some kind of intermediate, rules-based description, which can be provided to the opponents as well.
You also need to be able to encode and ingest the opponents' agreements, and to use this information to inform your own decisions. And you need, in particular, to be able to handle a wide variety of agreements from your opponents; it's not enough to force them to play the same system as you.
You must also account for deceit. If, for example, I have a bid which promises that I have at least 2 cards in every suit, it's perfectly legal for me to lie and make this bid when I only have 1 card in some suit--as long as my partner is in the dark about this just as much as the opponents. So if you make this bid, and your machine opponents assume there is a 0% probability of you having lied about your hand, it is possible that they will make gross errors by not accounting for this possibility (for example, they may be in a position where all of their actions are equivalent if you told the truth, but where one action is clearly better if you didn't--a human player will naturally take this action, but a robot may just select an action randomly).
It's an interesting game and a very interesting AI challenge.
> The algorithm also learned far faster—it played about 34 times fewer games than DeepNash, and still ended up much stronger.
Imo, this is the critical piece and what makes the AI work at all.
With hidden information games, the best move depends on information you don’t have. So a move could be good or bad, it just depends on something that’s impossible to know.
You’d like to search ahead, meaning “if I do this they will do that” but that’s impossible since you don’t even know what the opponent can do because you don’t know their hidden state.
If the possible hidden states are randomly distributed, you are screwed. It’s just like rock paper scissors: there’s no best move if your opponent is unpredictable.
However if you can quickly learn to predict their moves, it becomes possible to make informed decisions about what to do.
Oh no! Stratego had been on my mind as something we just hadn't tried hard enough to make a winning bot for, including the DeepMind effort from 2022. I was planning to make the first one.
I thought this was slightly less crank-coded than trying to prove the Riemann Hypothesis, but maybe these days you just ask Claude to do that and it tells you there's a counterexample at 1 + πi that no one ever noticed before.
I loved Stratego so much as a kid. But, I eventually couldn't find anyone to play with me because I crushed everyone, including my dad who was much better than me at chess. But, I never would have thought it'd be a game that models would have a hard time with. It feels relatively simple. And, inexplicably, I never even thought that there might be serious players...I've kept the board game, one of the very few things I have from young childhood, but it's been twenty years since I played. I guess it's time to find an online Stratego. Surely someone in the whole world can beat me.
Too bad I never played against an AI before they cracked it.
> Now, a team of researchers from Carnegie Mellon, MIT, New York University, and Stanford University has done it. Their AI, called Ataraxos, beat Pim Niemeijer, arguably the best Stratego player of all time, 15 games to one, with four draws. And it took just 16 GPUs and a few thousand dollars to train it.
Just 16 GPUs, and a few thousand dollars?
What about “researchers from Carnegie Mellon, MIT, New York University, and Stanford University” this wasn’t just anyone.
This puts the earlier "Mastering the Game of Stratego with Model-Free Multiagent Reinforcement Learning", 2022 [1] in some perspective. Apparently the "mastering" in 2022 wasn't quite there yet. Four years later, the new approach seems to actually be better than humans.
[1] https://arxiv.org/abs/2206.15378
This approach also works for Hanabi, which is a very interesting game. You can't see your own cards, but the other players can. I bought the game because someone on a reinforcement learning podcast [2] mentioned it, and actually played it multiple times.
[1] https://en.wikipedia.org/wiki/Hanabi_(card_game)
[2] https://www.talkrl.com/episodes/jakob-foerster
I think that what makes these games beatable repeatedly is that they're static. Not saying an algorithm properly trained won't play better than the average player a game like MtG, or my own https://aethersummon.com (specially now while it has under 90 possible scrolls only) but if you have a regular release cadence (say weekly or bi-weekly) of relevant new "cards", then I think the playing field is much more even for humans.
Those new additions can invalidate the whole training data by a single new "card" that changes completely the dynamics and would be easy for a player to understand and incorporate but not for an algorithm (perhaps with enough compute to re-train it regularly it could) - that along with the decision trees being orders of magnitude deeper, wider and with more conditionalities than go, chess or stratego - even through the same turn with the same cards available and same table state - would probably pose much harder problems for a compute bound algo.
Related:
https://news.ycombinator.com/item?id=49918148
I recall playing this game as a preschooler. It was mostly psychology and bluff. Very interesting.
I wonder how capable current AIs are with the "silent defense" variant of Stratego [1]? The article states the high level of uncertainty presents a challenge. With silent defense the uncertainty is even higher.
[1] https://www.hasbro.com/common/instruct/Stratego.PDF "When an attack is made, the attacker is the only player who has to declare the number of his or her piece. The defender does not reveal the number of his or her piece, but resolves the attack by removing whatever piece has a lower number from the gameboard. Players keep their own captured pieces. Exception: when a Scout attacks, the defender must reveal the number of his or her piece.
Wait, wasn't there that strong stratego bot that came out from deepmind in 2022?
Looks like a game archive exists here: https://ataraxosai.github.io/
While interesting from a technology stand point I get disheartened each time I see this sort of news.
I am particularly disappointed that it has influenced how people play the game.
The joy comes from the journey and the experience.
Look at competitive chess and Go and how they have fundamentally been transformed. It's not better and now the box is opened, it can't be closed.
I would really love to see a serious research effort take a crack at contract bridge. Bridge, like Stratego, is an imperfect information game with a big hidden information space. Bridge also adds another wrinkle of explainability which is, I think, very interesting.
Bridge is played as a pair vs pair game, with North/South and East/West being the two pairs and seated around the table in these compass directions. A bridge hand consists of two phases: there is first an auction phase, where players go around the table bidding on contracts (agreeing to take a certain number of tricks with a certain trump suit) until a final contract is decided. Then there is the cardplay phase, where the player who won the auction is the declarer, their partner is the dummy, and the other pair are defenders. The dummy's hand is placed face up on the the table and the declarer controls which cards are played from dummy, so the cardplay phase is effectively played by only three players now, with each of the three knowing one common hand (dummy) and one private hand (their own) and not knowing the other two hands.
In both the auction and (for the defense) the cardplay phases, it is important for players to exchange some information about their hand to their partner. However, any information you exchange about your own hands also helps your opponents. You might naturally conclude that you want to come up with some secret scheme to exchange information which your opponents don't know (and it is even possible to exchange encrypted information which your opponents can't know--if the defense is known to hold a certain card, but declarer doesn't know in which hand it is, the defense could say that a signal means one thing if the card is in one defender's hand, but means a different thing if it's in the other defender's hand).
But it turns out that this ends up being very uninteresting to play, so instead, when playing bridge, there is an important rule: all of your partnership agreements must be public. If a certain bid that I make promises that I have at least 5 spades in my hand, it is the opponents' right to know that this is our agreement. You must be able to explain the information which your action provides, and you must be able to use the information that the opponents give you themselves.
This poses several problems for self-play reinforcement learning. First, a naive self-play approach will produce agreements that cannot be explained to a human. What really needs to happen is that your partner, when determining what hands you might have as part of search, must not do so simply by sampling its own system (ie by asking what it itself would have done with hand X or hand Y). The information and possibilities really need to be mediated by some kind of intermediate, rules-based description, which can be provided to the opponents as well.
You also need to be able to encode and ingest the opponents' agreements, and to use this information to inform your own decisions. And you need, in particular, to be able to handle a wide variety of agreements from your opponents; it's not enough to force them to play the same system as you.
You must also account for deceit. If, for example, I have a bid which promises that I have at least 2 cards in every suit, it's perfectly legal for me to lie and make this bid when I only have 1 card in some suit--as long as my partner is in the dark about this just as much as the opponents. So if you make this bid, and your machine opponents assume there is a 0% probability of you having lied about your hand, it is possible that they will make gross errors by not accounting for this possibility (for example, they may be in a position where all of their actions are equivalent if you told the truth, but where one action is clearly better if you didn't--a human player will naturally take this action, but a robot may just select an action randomly).
It's an interesting game and a very interesting AI challenge.
Man I loved Hasbro’s 1998 Windows Stratego
https://archive.org/details/STRATEGO