With most information hidden, the game Stratego had stumped AI until now

240 points by PaulHoule · 23 hours ago · 116 comments · arstechnica.com ↗
Loading article...

Comments (116)

I loved Stratego so much as a kid. But, I eventually couldn't find anyone to play with me because I crushed everyone, including my dad who was much better than me at chess. But, I never would have thought it'd be a game that models would have a hard time with. It feels relatively simple. And, inexplicably, I never even thought that there might be serious players...I've kept the board game, one of the very few things I have from young childhood, but it's been twenty years since I played. I guess it's time to find an online Stratego. Surely someone in the whole world can beat me.

Too bad I never played against an AI before they cracked it.

> But, I eventually couldn't find anyone to play with me because I crushed everyone

I think this is a problem with many deeper games. I remember going over to other people's houses and they would pull out obscure-board-game-xyz. The rest of the night for me was trying to figure out how to play, while others were 30 steps ahead of me.

I think you need something like chess club - stratego club - where people have all gotten to the nuance level and you can play with people of appropriate level.

I feel like the depth of a lot of board games gets undersold simply because of how inconvenient it would be to play a large number of games physically. And/or because it's easy at the start to try something complicated, fail to get it right and end up concluding wrongly that only the simple things actually work.

At one point in my life I got in literally thousands of games of Dominion, and I was definitely still not as good as some of the people I discussed the game with online, but in person I could barely find anyone to play with, ever. (In fact, of the people I knew who were familiar with it, few even seemed to think it was any good as a game.)

I've played a lot of Dominion with people who have played thousands of games. I've even beat them with some frequency (it's not that hard!).

I don't think Dominion is a good game. It's good if you maybe think cubing is good. And it can be a lot of fun!

But as a game? It's not good. If you have no idea what you are doing and are totally unfamiliar, you can even win pretty easily with some luck. But you won't understand why you won, or rather why others lost.

There are many, many games where everyone can read the rules together and be, more or less, on even footing. Dominion can sometimes be that way, but it can pretty easily be extremely opposite of that.

>If you have no idea what you are doing and are totally unfamiliar, you can even win pretty easily with some luck. But you won't understand why you won, or rather why others lost.

Interesting, this is a _feature_ of poker games like Texas Hold'em. If the losing players didn't go on convincing winning streaks, they'd stop playing. And those winning streaks can convince them they don't need to learn anything else about the game in order to be a winning player.

I’ve played a lot of Hold’em and don’t really see that. Sure, any single hand could be won with luck by a clueless player, but even over a short amount of time, the lesser player will lose.
The key difference is gambling.

Have you ever gambled over a game of Dominion? I wouldn't.

Have you ever played Hold'em without gambling? I wouldn't either.

Or another way of saying it, when the default no-thinking baseline strategy can win frequently enough by chance that it becomes difficult to learn from a single game. In this case, just 'buy silver, then gold, then provinces) can win too often just by chance.
Digital Board Games solve this imho very elegantly, since the rules are tracked by the engine for the players (and is always impartial). It also can avoid the onboarding friction.

Just a pity it does not come close to the haptic pleasure of playing with a real cardboard, polished resin figures, and sturdy cards with dice (or whatever accessories are supplied) :-)

Indeed. Online games, Board Game Arena especially, are great for the ease of the rules enforcement and getting together with remote friends. But if given an option, I would always choose the in-person physical version. I'm spending enough hours in front of a monitor already.
Board games really get so much better once you get a vaguely-regular group. If nothing else, you can walk in one night and everyone just knows how to play a lot of really great games right off the bat.
Does anyone else remember playing Electronic Stratego? The version where the identity of each piece was encoded on the bottom with a series of bumps? It was an even more incomplete information game: attacking a piece would only let you know if the piece was higher, lower, or even, instead of revealing the piece's identity. I played it so much as a kid that I can't even think of Stratego without hearing that electronic bomb sound in my head. I feel it is actually a better game than the original.

https://en.wikipedia.org/wiki/Stratego#Electronic_Stratego

That is the rules we used for playing stratego. Attacker reveales his piece number and the defender declared who won, or obviously if both died they were even.
Electronic Stratego didn’t even reveal the number.
I love Stratego. I used to play with my girlfriend and my brothers. We started playing with different rules too to switch it up. Like moving bombs or the flag. Or playing with the pieces reversed. Reversed pieces were interesting because the hidden information was swapped, made for some funny games.
Question: I need to play this again but there are many different versions. It's impossible to decide which one to get. Which is the best one to purchase?
> including my dad who was much better than me at chess.

My experience as well unfortunately. No matter how many books on chess I read, no matter how many games I played against Battle Chess to try to get better, he always won.

Chess...I'm in the top 1 percent among regular players on chess.com/lichess, yet impossibly far away from Master/GrandMaster level as to be impossible; their minds are completely alien to me.
Impressive. Keep at it. I hope you'll find a Eureka moment. I also hope they're not just cheating using AI.
Becoming such is not my goal; I have too many other interests. It was more commentary about how one might be good, but no where even near competitive, and the journey, even if mentally capable, of improving further would take years and year of full time dedication...because chess is that deep, and because "good" is always context dependent.
I can relate. I'm in the top percentile on stack overflow (now killed by AI). My only goal was to be "good" in the way you described.
I remember the day I beat my dad the first time at chess and also my grandfather. Those were big days for me after losing so many times.
Makes me wanna play against Swell Joe's Dad \s

But about having nobody to play because you crush everyone - as long as you want to play to enjoy and not just always go full-on try-hard mode - just stop crushing everyone and you'll have people to play with.

Let people win every third game. If you're good enough (at it), they won't notice. They'll have a good time and you will. In Tekken it works to have them win some rounds (extra points if you let them have two out of three), but still win the match, tho maybe only if you're not just 2 ppl playing.

That's a rule established from rat behavior observation study, you can hear about it in every third jbp lecture.

Ever since implementing it, I never end up with nobody to play with. Just curb your wanting to always win and focus on maximizing the amount everyone enjoys playing in the long run.

Even a game where there's always a winner and a loser doesn't have to be a zero sum game.

"But about having nobody to play because you crush everyone - as long as you want to play to enjoy and not just always go full-on try-hard mode - just stop crushing everyone and you'll have people to play with."

Sure, I'll go back and tell my 12 year old self to do that.

But, actually, I did do that, to some degree. As I mentioned, I would drop hints and talk about my strategies in games, usually after the game, so he could be more competitive and understand how I was thinking about things. That also has a cost on the fun for both parties, though. Smart people can recognize when you're going easy on them. Nobody normal likes being pandered to.

I don’t know stratego, but could you play to put yourself in a tough position and then dig yourself out?

This is what I do when training jiujitsu with less experienced people.

You could give yourself a terrible setup. A significant portion of the game is how you place the pieces and there are definitely challenging/risky options. Flag in front with no bombs protecting it is bound to lose, but the game won't last very long either.

And, I did that kind of thing sometimes, though I was a kid and still liked to win, I doubt I used any truly suicidal setups.

I'd also heard about that rat experiment (probably on HN) and also tried applying it to fighting games since my friend I usually played with seemed to get really mad most of the times we played. I think it worked somewhat, but then I told him what I did and he didn't take it well. I kinda get why he'd be mad about losing, but then he never plays without me, doesn't spend time in training mode, doesn't watch videos on the games, etc. Not sure what he expects then, I don't think he's gonna improve much jumping right into a real match while rusty and possibly not even understanding all the fundamentals. Not sure there's a solution here since you can't really force people to care about or like things. Ideally we'd both keep striving to improve and competing with each other, but instead I hear "I'll never be better than you anyway". I'm not even god-tier at FGs, I'll lose 90% of the time if I do online matchmaking. I'm just good enough to have trouble getting new people to play with me.

The rat strat may be worth revisiting but it's not very fun letting people win, and it feels dishonest and disrespectful as well. The closest thing I've found is to play super defensive. If I try to take as few hits as possible and drag out the battle it still serves as practice, and I kinda let them set the pace.

I can’t believe I’m giving social advice on HN, but maybe it will help.

> I told him what I did and he didn't take it well

This was a mistake. Ask yourself why you did it. If you don’t have a good answer, one possibility is that it was your ego. Another is your very strong adherence to a certain code of behavior that others do not relate to.

> we'd both keep striving to improve and competing with each other

I’m not sure competing with someone who is not on your level is a normal approach.

I also doubt your friend sees the situation the same way as you. They are more of a recreational player. It’s a level below where you are, but for some reason you want them to care more.

> I'll lose 90% of the time if I do online matchmaking

And yet playing in person is important to you. It’s fun?

> it's not very fun letting people win, and it feels dishonest and disrespectful as well

At work, you want people on your team to pull their weight even though you could do the job better. How do you keep them from being discouraged? A strategy like this works.

Stratego has a special place in my heart. I have fond memories of playing it with an old family friend. Later in life I sought out an old set off eBay so I could play it with my kids. I really love that game!
I also enjoyed it. We even used to play live Stratego at camp with index cards in our socks with our ranks.

What was the secret of your success?

I don't actually know. I had a few strong setups that I rotated through, but even after I started explaining my strategy after every game, and even giving hints during the game, my best childhood friend would still lose so much it wasn't fun for either of us so we stopped playing. He's smart, and was fine at most games, but no match in Stratego. I guess I'm just good enough at the various aspects of the game, memory, bluffing, planning, that it added up to being a pretty strong player even without much in the way of study or practice. I'm sure I'd crumble against an actual serious player, and would have back then, too, it was just that I wasn't around anyone else who liked the game enough to become good at it.
I'm curious. I could kind of see that being fun, but I could also see 1 person moving with 79 standing still, and two players who get to make all the decisions. Did you guys modify the rules?
Very little standing still. Everyone is running around. It's a bit like capture the flag, except when you tag someone you compare index cards to see who is higher rank.
It has such a strong story despite being so simple.
I, too, remember figuring out a strategy when I was around 11, and never lost a game after that.

It's been a loooong time, and I don't recall all the details. But it revolved around doing probing attacks to determine where the ranks were in the enemy formation, and then having "channels" in my side to move up a soldier that outranked by 1 a targeted attack.

Color me surprised that it would be difficult to write a program to play it.

Most of my setups, though not all, revolved around bombs and strong soldiers held in reserve behind the bombs, sort of a rope-a-dope...sometimes, I'd kill all of their miners before the bombs surrounding the flag could be disarmed which guarantees no quick win for them. There are risks to that, in that the strong pieces can be stuck while the opponent picks off everything else, but even once people knew I regularly did that, I could still win pretty much always, so it's more than a strong opening layout.
Sounds good. My "channels" were a way to move the right pieces to where they were needed. I did this repeatedly, and my opponents never learned.

One thing Stratego did was implement the "fog of war". My Empire game took inspiration from Stratego and Risk.

In hindsight, maybe Stratego in childhood is part of why I love Civ games so much, which, I think, tickles a lot of the same brain parts. The setting up the board part is heavily why I liked Stratego so much, and in Civ games, you get to plan your cities and empire. Or, maybe it's just the kind of game I find satisfying, and Stratego was the first instance of it I played.

And, I'm also working on a game that kind of pulls on that same kind of upfront planning, where how you place your pieces to start is as important as how you move them; a settling and strategy game that's heavy on clever defense and using terrain and the interplay of pieces.

> The algorithm also learned far faster—it played about 34 times fewer games than DeepNash, and still ended up much stronger.

Imo, this is the critical piece and what makes the AI work at all.

With hidden information games, the best move depends on information you don’t have. So a move could be good or bad, it just depends on something that’s impossible to know.

You’d like to search ahead, meaning “if I do this they will do that” but that’s impossible since you don’t even know what the opponent can do because you don’t know their hidden state.

If the possible hidden states are randomly distributed, you are screwed. It’s just like rock paper scissors: there’s no best move if your opponent is unpredictable.

However if you can quickly learn to predict their moves, it becomes possible to make informed decisions about what to do.

In most good full information games, the best move also depends on knowledge you don't have - a full tree of all possible game states from a position. AI playing Go or Chess can't make the best move because they don't know what it is, they're technically just guessing. There isn't any reason to think AI have more or less trouble with hidden information games.

The practical difference is how many rules a game has and how easy it is to implement the engine. Implementing a chess bot is relatively easy because the amount of state tracking required to set up a simulation is basically nothing (I think just whether the king has made a move yet or not). That makes it easier to implement than something with a lot of signals that need to be recorded. Something like DoTA or Starcraft takes serious engineering effort.

You are confusing the inability to compute a full game tree with not knowing anything at all. In fact there are many positions in chess where we can compute the full game tree. Forced mates, and tablebases of positions with 7 pieces or less. And even if we can’t compute the full tree, errors get smaller with depth.

> There isn't any reason to think AI have more or less trouble with hidden information games.

How about the fact that a child can beat the best rock paper scissors player in the world in a game, but no human can beat the best chess engine? Same thing with poker, a novice could get lucky and win a hand against the best poker player.

> How about the fact that a child can beat the best rock paper scissors player in the world in a game, but no human can beat the best chess engine? Same thing with poker, a novice could get lucky and win a hand against the best poker player.

How about that indeed. What are you trying to say here? The probability of a novice beating an expert doesn't tell us anything except how decisive skill is in a game. We can come up with skilful games where novices sometimes beat experts - in fact, that includes basically all games.

The point of something like an Elo rating is that on very rare occasions a novice will beat a grandmaster at chess. A human child may well beat the top chess AI on a long enough timescale, we really don't know how that will go. No reason it can't happen. They just have to get really lucky enough times in a row.

This has nothing to do with perfect or imperfect information. We can come up with a hidden information game where the odds of a novice beating an expert are approximately zero. And we can come up with a perfect information game where experts lose easily (for example, Go has a handicap system to equalise any skill gap).

It would be easy enough for a player to be purely random if that was all it took. I think the tension is that piece rank makes some layouts and move strategies more equal than others and calculating the best ones for what has been uncovered so far makes the best layouts not the best layouts.
Oh no! Stratego had been on my mind as something we just hadn't tried hard enough to make a winning bot for, including the DeepMind effort from 2022. I was planning to make the first one.

I thought this was slightly less crank-coded than trying to prove the Riemann Hypothesis, but maybe these days you just ask Claude to do that and it tells you there's a counterexample at 1 + πi that no one ever noticed before.

I wonder why people act as if Riemann hypothesis is somehow already solved, when it is still highly probable that it is impossible for humans and slightly less impossible for AI overminds.

(Of course, tomorrow Google might announce that it has solved it.)

RH would require completely novel insight and a brilliant original thought. Unlike Navier-Stokes it's not a gradual effort and you cannot piggyback off of other mathematician's works, so "AI" will not solve it anytime soon.
There has been a lot of intermediate progress on the Riemann hypothesis, for example ruling out particular kinds of zeros, and some of those partial results even involve LLMs https://www.anthropic.com/research/riemann-zeta so I don't see why I can't possibly be a gradual effort.
On what basis do you make these claims? Honest question.
Which claims? The general feeling in the mathematical community was that significant progress towards Navier-Stokes was being made, and that the last "piece" to settle the question was imminent for some years now. Riemann Hypothesis seems virtually as unreachable today as it was 50 years ago.
I think now it's a good time for people to work on the Riemann hypothesis. There are actually a lot of ideas, no one knows if any of them goes anywhere--but I don't think it's completely impossible to make progress especially with new tools haha.
recently we got Starcraft 1 RL bot and Advance Wars bot, both on upper human level
The frontier is moving, but there's still no human-level starcraft bot that plays entirely through vision, like a human.

IMO it also needs to use a real mouse before I think it's a true comparison, even a casual player would have a massive advantage if they could issue selections and unit commands via query.

They already cap the actions per second a bot can take, sometimes to very low numbers. I don't remember if it was AlphaStar that limited how often you can move your camera or not, but some work definitely did.

I personally don't see why vision is important, if anything I'd frame vision as useful to humans, rather than being the baseline.

The capped actions cheat in EPM afaik. It does crazy micro
Your statement is severely underselling Plutos APM.

Pluto uses Ghosts to counter Dragoons. Any SC:BW player should know how stupidly impossible this is for humans.

For the non-players out there. Ghosts have an ability called Lockdown that can stop mechanical units (like Dragoons) from attacking or moving for long periods of time.

The problem: it requires you to click on a Ghost, then click on a Dragoon.

No human in the whole history of SC:BW would seriously do this. At best, Ghosts in practice are used vs capital ships (like Carriers or Battlecruisers), because in smaller numbers humans are fast enough to click and perform these feats. (Carriers and Battlecruisers are very expensive, very very powerful, units. So they are weak to specialized attacks like this as you cannot build many of them due to their incredible cost in resources).

But dragoons? This is the mass production unit of Protoss. This is so far outside the scope of human-level APM that the games Pluto plays are useless. Lockdown simply isn't feasible to be used vs what amounts to the standard / stock infantry unit of Protoss. There's simply too many opponent dragoons to even click on.

-------

I dunno what kind of "APM rules" we need to give SC:BW bots to allow them to have a fair fight vs humans.

But the rules we have in place today do NOT represent a fair fight at all. Computers have stupid amounts of speed and precision far in excess of a human player.

We got an Advance Wars bot? Say more...
Seems to be "JuggerMaster_BOT" on AWBW.
I only saw this vid https://www.youtube.com/watch?v=hdxAmJsspcA

clickbait title and the bot isn't at "top players struggle to beat it" level - but it's no longer a casual walk in the park like the default ai is, so the progress is massive

On my way to buy stratego off ebay
> Now, a team of researchers from Carnegie Mellon, MIT, New York University, and Stanford University has done it. Their AI, called Ataraxos, beat Pim Niemeijer, arguably the best Stratego player of all time, 15 games to one, with four draws. And it took just 16 GPUs and a few thousand dollars to train it.

Just 16 GPUs, and a few thousand dollars?

What about “researchers from Carnegie Mellon, MIT, New York University, and Stanford University” this wasn’t just anyone.

That's a bit out of context, no? The preceding sentence is > Even DeepMind, with its exceptional budget, couldn’t build a machine that reliably beat the best human players.

so the contrast here is the budget available, not the quality of the talent, if we accept the premise that DeepMind and the universities have approximately similar level of talent.

It’d be less budget and resources than DeepMind.

So clearly it’s due to talent as well as advances, of course.

This puts the earlier "Mastering the Game of Stratego with Model-Free Multiagent Reinforcement Learning", 2022 [1] in some perspective. Apparently the "mastering" in 2022 wasn't quite there yet. Four years later, the new approach seems to actually be better than humans.

[1] https://arxiv.org/abs/2206.15378

I remember playing Stratego as a kid at a friends house, and eventually found out some of his pieces were subtly marked (chip in the pieces). Unhiding information was very unfair :)
In real life this leads into the wonderful arena of deception: I create an information side channel, which you detect and begin to rely on, thinking it is secret, but I can then lie about.
This approach also works for Hanabi, which is a very interesting game. You can't see your own cards, but the other players can. I bought the game because someone on a reinforcement learning podcast [2] mentioned it, and actually played it multiple times.

[1] https://en.wikipedia.org/wiki/Hanabi_(card_game)

[2] https://www.talkrl.com/episodes/jakob-foerster

I think that what makes these games beatable repeatedly is that they're static. Not saying an algorithm properly trained won't play better than the average player a game like MtG, or my own https://aethersummon.com (specially now while it has under 90 possible scrolls only) but if you have a regular release cadence (say weekly or bi-weekly) of relevant new "cards", then I think the playing field is much more even for humans.

Those new additions can invalidate the whole training data by a single new "card" that changes completely the dynamics and would be easy for a player to understand and incorporate but not for an algorithm (perhaps with enough compute to re-train it regularly it could) - that along with the decision trees being orders of magnitude deeper, wider and with more conditionalities than go, chess or stratego - even through the same turn with the same cards available and same table state - would probably pose much harder problems for a compute bound algo.

> Those new additions can invalidate the whole training data by a single new "card" that changes completely the dynamics

This doesn't follow. You're basically proposing that new combo decks be added all the time, and it's far simpler for an agent to scan the new cards for potential interactions with the thousands of other cards in circulation than for a human to remember all of them.

Your analogy is akin to saying that all you have to do is keep landing new code all the time, and since the agents weren't trained on the code they won't be able to identify and respond to security vulnerabilities in it as fast as humans, which hasn't turned out to be correct

No, well, in MtG you could interpret it as meaning such but what I mean is that if in the training set sequence A-B-B-A when state is C-A-X-Y is the play 80% of the time, then you have a new card (that doesn't need to be combo) that by sheer mechanics thwarts that then that strategy won't stick by the addition of that single card to the opposing deck (that you can't know if your opponent is playing or not) and having one or 2 or 3 or 10 different cards renders every calculation very problematic as a play can be the best or the worst depending on such simple things diluting further the best play as the pool grows. Then you need to take into account in MtG shuffling and drawing. I think it's fair to say it's much more difficult to model... And while an agent can learn new combos, you just need to read the card once, the agent needs to be retrained.
Doesn't this entirely depend on the latent embeddings of strategies and game space in the AI model, which may not be so concrete and explicit as you've described? That's kind of the magic of LLMs with coding, they can generalize because the abstract patterns are encoded in latent space, not the specifics.
I might be wrong but what I was thinking was that in chess (or even imperfect information games with a much smaller "range" such as Stratego,) a model can calculate all possibilities for all moves and following moves, by itself and opponent up to a depth that the human cannot. So it can see everything that can happen if it does move X-Y, then Y-Z, then A-C and figure out one that is unbeatable no matter what (or at worse leads to a draw).

But on MtG in particular that never really applies in full due to drawing new cards. You can play perfectly and still lose due to sheer randomness of draws.

The latent space I'm not sure how it translates to a game playing bot, but I would imagine that it would open it up to fail in the same ways a human fails.

On the game I'm designing it could do that (calculate all possibilities up to X depth, for all possible scrolls and table states) but it would be extremely expensive to do so (not a very good argument if compute power keeps increasing), but more than that, in contrast to something like chess, there can be many more paths and decision points where a bad decision turns into a loss, so if it assumes that the best play is X at some point, a sequence that it discarded due to not being the most probable can exist and the bot can never be sure, so if it makes a decision that plays into a "trap" he can't undo to a favourable position. While in Chess it's much clearer what is possible from a given state, it's unambiguous and the rules are fairly limited.

In stratego you have a 10x10 board game, a very clear objective and at most 40 pieces (with repeated pieces and simple mechanics amongst them), while in MtG and similar games a single piece (card) can have probably hundreds of different interactions depending on everything else going (and everything else hidden), at many points of decision. In stratego it also seems that for humans at least, most moves are "inconsequential", as it probably plays more at the psychological/bluff level. Maybe a human player that was given the same budget for training could spend a month training against bots might fare better as the strategies might be then better understood (by the article it's mentioned that the agent recovered from bad positions, so it seems that it was mostly human error, as the human was playing better up to that point).

While on MtG or Asummon, although there can be inconsequential moves (they don't matter given the context/stage of the game), every move carries with it a possibility of being consequential in unpredictable ways. Anyway, there should be ways of training models with just a rule abiding client for these games, without codifying all rules, that they can just keep playing to figure out the interactions, so if that theory is true then it should be possible to create an unbeatable bot - I'm just not sure it is without infinite time/compute and less so if the "meta" keeps changing rendering possible training inconsequential regularly.

You can definitely try to regularize against ruleset changes by generating a bunch of cards and making the agent play in randomized subsets of those cards.

I didn't look for prior work on this, but my estimate is that it's probably within 2-3 orders of magnitude of additional training compared to a static game. (Still a lot!)

But wouldn't (couldn't) the model then hallucinate play patterns and get itself into problems when playing against a real opponent?
Well, if your training includes regularization against ruleset changes, the model should simply handle it. (that would be the expensive option, and require vastly more training)

When the Dota 2 bot was made, they retrained the bot only partially when new patches came in, so it was definitely cheaper to adapt.

There are very few missing pieces for a game like MTG. The main reasons we don't have a Stockfish for MTG is that it's a PITA to implement the rules and that nobody cares (or at least not enough to make it happen.)

There is nothing that, in principle, makes MTG different from poker or bridge, and we have superhuman engines for both.

>There is nothing that, in principle, makes MTG different from poker or bridge, and we have superhuman engines for both.

We don't have superhuman play for bridge.

Poker and bridge are quite different from each other in terms of solving them. Among other things, the hidden information space in poker (at least, in hold'em) is far smaller than in bridge (or Stratego, for that matter, as discussed in the linked paper). This makes hold'em solvable using CFR, an algorithm which essentially optimizes play by considering all the possible holdings than the opponent might have and their best strategy with each one. Even going from two to four hidden cards per player (Omaha) requires a slightly different approach although you can still use CFR as the basis for the search algorithm.

Bridge has 13 hidden cards per player which makes CFR basically impossible to apply, at least in any obvious way--just way too many states. Similarly you see it's not used at all in this Stratego paper.

Sure, but you can do Monte Carlo with a double-dummy solver.

The point is that, especially for games perceived as being lower-status like MTG and other board games, I'm more inclined to believe the answer is closer to "nobody is willing to pour in the resources to seriously try" as opposed to "we definitively cannot with current science and technology."

I think computer bridge will take off soon. LLMs have allowed double-dummy solvers to move from ~0.1sec to microseconds (if you allow 99.9% accuracy) and parsing human bidding system descriptions must either be possible or v close.
> There is nothing that, in principle, makes MTG different from poker or bridge

There is - metagame. There is no universal optimal strategy in a trading card game, because what is optimal depends on what decks and strategies other people are playing.

I'm sure you could train a neural network to play a specific deck within a specific metagame of a specific card game, but you would probably have to keep re-training it when there are new decks/combos/releases/rotations/banlists/metagame shifts.

MTG is also severely constrained (small hand, mana -> possible moves) although I don't think it's anywhere near the same. In my opinion the rules are effectively what change the whole dynamics. You can't plan as efficiently without knowing what your opponent holds and having to take into account all possibilities (with infinite energy/compute time perhaps)... I don't doubt you can train a model to play well, I just think it should be much more level to the human player. In MtG you also have the randomness which is not easy to model nor account for - the perfect play by an LLM can be the worse once the opponet draws next.

In my own game you don't have shuffle/draw randomness but the pool of options is statistically tending to infinite (if I would have 500 or 1000 scrolls designed and MtG depending on the format has that depth) when compared to something like chess, or this game. On the other hand in my own game you have to account for much more depth on the possible options your opponent has.

There's only so many card interactions that strong players actually think about.

Ex: you don't really care if the opponent plays Giant Growth or Chastise. The effect is that the opponent is playing a combat trick, and combat has moved from attackers favor into defenders favor.

To defeat an instant speed combat trick requires a combat trick of your own, or a generic counter spell of some kind. Some have interactions (ex: Doom Blade beats Giant Growth but not Chastise), but the overall gist is that opponents can do things after combat is declared. You only need to keep track of how many combat tricks you think the opponent has.

---------

Other situations are card advantage (ex: 2 for 1. If the opponent spends 1 cards to defeat only 2 cards of yours). The traditional card for this is Mindrot, but well placed counterspell can turn a combat trick into. 2-for-1 reversal.

You don't necessarily keep track of how your opponent makes 2-for-1 opportunities. You just have vague gists of them.

---------

Good spells have huge applicability. Doom blade or Murder is high because killing opponent creatures at instant speed handles the vast majority of creature buffed combat tricks, and also serves as a way to stop enemy combos and other such tricks.

In contrast, chastise is very niche. If the opponent were playing like Swords to Plowshares (powerful white instant speed removal), it's pretty much always better than chastise.

If the opponent plays chastise instead, you take that as a win because you know they could have had a deck of better cards. But for whatever reason decided to play with weaker cards...

I agree in a way, but at the same time, and I think it's a bit more applicable to MtG due to the limit of cards you can have as possible plays at any given time (outside of combos), and I believe too that you can train a bot to be good, better than average - I doubt arena doesn't have bots - but I still think that without unbound compute/time it's a game where human players have much better odds to outsmart an AI if they're good players. MtG has for the past 10 or more years been re-hashing the same play patterns, while introducing some new mechanics on most cycles, but pretty much you have staples throughout most editions that are just variations on that - card advantage, denial, combat tricks, removal, curve and then the rarity enabled bombs/combos

But even then (not saying I'm right) I think the depth of choices, effects and so on, on a format like modern, or legacy, would be very difficult for an AI to top against pros. If you add draft into the mix it gets worse for the AI in my view too.

Because a good play in most situations can easily be a bad play under others. That doesn't happen in chess for instance, given enough decision depth to the algos to see the future game. In my own game I think those situations can occur much easier due to you always having your full deck available. Also, in MtG it's easy to get into table states that are either ahead/behind and then you kinda just have to protect your position (like with denial decks). Then you have the effects that you might remove a creature threat (graveyard) but then that enabling a combo you weren't expecting that needs a creature on the grave, or enabling delve cards or whatever have you. It's much less clear cut for a probabilistic model to make the optimal play at every single interaction. So the more you train the model on all the variations and possible follow ups, the more you dilute its certainty isn't it? In chess, or this game, or RTS such as starcraft, that doesn't really happen in my view.

I'm also a Poker player and the way Poker AIs solved this problem was by making the best estimate of the Nash Equalibrium and playing around it.

No human can possibly keep up with all the possibilities or combinations that are accounted for.

Games of incomplete information have been IMO soft-solved as of.... Maybe 5 years ago? As in, stronger than any human can possibly reach (ie: massive GB-sized matricies accounting for all information iterated over millions of iterations of "he thinks that I think that he thinks that I think that....")

It's not a true Nash Equalibrium, which remains outside of the realm of even computers to compute. But a computer can always reach a closer / better estimate of any Nash Equalibrium, which covers all games of incomplete information.

--------

For Poker, it turns out that a few types of bet sizes (3x pot, 1.5x pot, pot, half pot, quarter pot) covered enough betting patterns to reach superhuman.

And frankly, MtG is simpler than the bluffing game in Poker. Like MtG has bluffs but it's no where close to Pokers level.

There's no crazy deep game for Red Deck Wins vs Control. The game basically plays itself out (Red tries to win before Control comes online. Control tries to stall before Red Deck Wins). There are some games with complex board states but they're largely a game of bluffing + card counting (opponent holds 4 cards, two of which were since the start of game and 2 were top decked in the last two turns. He at best has only planned for 2 responses or got lucky with the other two newest cards. Do I have a play that beats two cards yet?)

The card counting is harder since we can hide which cards were top decked or held since the start by just rearranging/shuffling our hand. We could have 0-4 counters or setups waiting for the gating decision.

The possible game state is also much larger than poker. Deck construction alone is 60 cards out of about 30,000 unique cards. Sure, not all cards are viable in all decks, but we can have 1-4 copies of a card in our deck.

So we might not even know if we’re playing against mono-red or multicolored since the decklist is unknown. You can think you’re playing mono-red and then they suddenly play a Plains. Poker at least always has the same 52 cards to reason about.

I do think AI could be great at coming up with decklists though.

Yeah I am of the same opinion. And of those unique cards (although formats will limit the total number) they all can play differently in different contexts/states of the game - even a "bear" (2/2 vanilla), can be just a bear, or part of a strategy (if other cards pump those specific cards being played, or enhance them), while poker they're always evaluated in aggregate from 52 cards that are split between players, so you can always remove the ones you're holding, the ones on the table. So 48 cards to calculate possibilities after initial deal + whatever is on the table. The fact that they're shared also means you can exclude immediately what is revealed and what you hold on your hand from those calculations.
Hm, I don't really agree. I think that bluffing feels more important in poker because of the usual monetary value attached, and is more of a strategic hindrance when playing against bots, while it's very important between humans because (usually) of the money involved - even in this article there's a mentioning that the bots don't care about bluffing and that for humans that is part of the normal gameplay and mastery.

There's also the decision points, in poker it's way less. You're dealt cards, the table reveals cards, you bet/ante, move to the next, bet/ante. It's a very finite sequence of moves until disclosure. Plus it's 52 cards divided by the table players while on MtG it's 60 per players that you can't know (even lands can interact beyond being a resource and in competitive lists usually they do, specially in older formats) - although with enough/infinite time/memory you can probably generate a table for all possibilities, you still have to contend with the interactions other than the card types. A 9 Diamond is a 9 diamond. A Skull Clamp is a Skull Clamp but the way it interacts, its value/threat is highly dependent on context that can/might be hidden.

While I think that MtG is indeed "poker" like underneath the keywords, it's many more levels and I think bluffing can be way more "complex" but simply isn't because even at the pro-tour level prize pools are insignificant when compared to serious poker tables. Some MtG players are known to also dabble/play poker regularly.

I also think that the structure of poker game-play is more prone to be exploited by a competent bot - if you have a "budget" and you assign a bot to a table where the antes are "in-line" with the "budget" it has, it can mathematically (within a very high degree of probability) always turn a profit - ultimately humans fail in part because they enter "bluff" kingdom against a bot as the bots can just rely on mathematical probabilities. Made up numbers but the idea being, you have $200 to play. Choose a table where this allows you to play X games at least, say antes of cents, it should be able to make money most of the time at some point.

Yes, but as the other reply mentioned, the thing is you don't know if it's red deck wins, or a RDW with a tweak for the metagame and building the "he thinks that I think that he thinks" tables would probably require for practical terms what could amount to infinite storage and any of these chains, if followed through, can land the bot in a losing position hard to come back from. Now, to be honest, most MtG players aren't that good either, they play it more like a hobby/fun game, rather than approach it as poker/probabilities.

> There is nothing that, in principle, makes MTG different from poker or bridge

Only in the most general form they are games with cards and hidden information with a state space that some form of tree search can theoretically play out.

The difference is the size of the search space. In MTG the search space is unimaginably huge. It would make Go's search space look like a spec of hydrogen in the middle of the universe.

It would require completely different techniques to produce a computer good at MtG than one that is good at bridge.

Hearthstone is absolutely swarming with bots that beat humans regularly.
Which humans? Many people are simply not that good at Hearthstone. Are the bots regularly achieving high legend ranks?
I wonder how capable current AIs are with the "silent defense" variant of Stratego [1]? The article states the high level of uncertainty presents a challenge. With silent defense the uncertainty is even higher.

[1] https://www.hasbro.com/common/instruct/Stratego.PDF "When an attack is made, the attacker is the only player who has to declare the number of his or her piece. The defender does not reveal the number of his or her piece, but resolves the attack by removing whatever piece has a lower number from the gameboard. Players keep their own captured pieces. Exception: when a Scout attacks, the defender must reveal the number of his or her piece.

As a kid, a friend of mine had "Electronic Stratego"[1], the biggest gameplay change was that you could carry out fights without revealing the strength of either piece to the other side. I found this made for a much more interesting game and we had quite a bit of fun playing it.

1. https://boardgamegeek.com/boardgame/3513/electronic-stratego (We generally banned the use of the 'probing' feature)

I recall playing this game as a preschooler. It was mostly psychology and bluff. Very interesting.
yes back then the game was hard partly because you couldn't remember all of your opponents pieces that you had seen. An AI would never forget though.
Depending on the rule set you play, lesser pieces could be removed without them being declared.
eventually though, resimulations will never to revolve around trying to inject bad context into the the stream and try to distract them.
Oh I loved Stratego. Got it as a present from my (now dead) granddad. Some great memories of playing it with him and my brother.
> Players were surprised, for example, by how often it tucked its flag into a corner behind just two bombs, a rarely played setup.

I like Stratego a lot. Next time, I will try this and put three bombs in the other corner to keep the opponent guessing and wasting an expected 4 extra moves.

I speculate a big part of the game is moving your power pieces (1, 2, 3) to places where they can actually attack the opponent safely. Any other placement of the flag and bombs creates a bottleneck for left-right movement on your side.

Looks like a game archive exists here: https://ataraxosai.github.io/
Play against the computer, made in 5 mins by z.ai

https://chat.z.ai/space/p1dcc82cr101-art

The rule deciding the outcome of a battle between two pieces is backwards.
No, it is correct. Different versions of the game label the pieces differently, either by rank (1 is highest) or strength (10 is highest).
Wait, wasn't there that strong stratego bot that came out from deepmind in 2022?
The article talks about that. The new bot required two orders of magnitude less training data and plays better.
Awesome, I will read this carefully later today. I'm always excited by AI research applied to games.
Philippines has a very similar game called Game of the Generals.
There is no information in this article on how modern LLMs like GPT-6 Astra or Claude Opus 5.5 perform at this game. I have a hard time believing they are bad at it.
I would really love to see a serious research effort take a crack at contract bridge. Bridge, like Stratego, is an imperfect information game with a big hidden information space. Bridge also adds another wrinkle of explainability which is, I think, very interesting.

Bridge is played as a pair vs pair game, with North/South and East/West being the two pairs and seated around the table in these compass directions. A bridge hand consists of two phases: there is first an auction phase, where players go around the table bidding on contracts (agreeing to take a certain number of tricks with a certain trump suit) until a final contract is decided. Then there is the cardplay phase, where the player who won the auction is the declarer, their partner is the dummy, and the other pair are defenders. The dummy's hand is placed face up on the the table and the declarer controls which cards are played from dummy, so the cardplay phase is effectively played by only three players now, with each of the three knowing one common hand (dummy) and one private hand (their own) and not knowing the other two hands.

In both the auction and (for the defense) the cardplay phases, it is important for players to exchange some information about their hand to their partner. However, any information you exchange about your own hands also helps your opponents. You might naturally conclude that you want to come up with some secret scheme to exchange information which your opponents don't know (and it is even possible to exchange encrypted information which your opponents can't know--if the defense is known to hold a certain card, but declarer doesn't know in which hand it is, the defense could say that a signal means one thing if the card is in one defender's hand, but means a different thing if it's in the other defender's hand).

But it turns out that this ends up being very uninteresting to play, so instead, when playing bridge, there is an important rule: all of your partnership agreements must be public. If a certain bid that I make promises that I have at least 5 spades in my hand, it is the opponents' right to know that this is our agreement. You must be able to explain the information which your action provides, and you must be able to use the information that the opponents give you themselves.

This poses several problems for self-play reinforcement learning. First, a naive self-play approach will produce agreements that cannot be explained to a human. What really needs to happen is that your partner, when determining what hands you might have as part of search, must not do so simply by sampling its own system (ie by asking what it itself would have done with hand X or hand Y). The information and possibilities really need to be mediated by some kind of intermediate, rules-based description, which can be provided to the opponents as well.

You also need to be able to encode and ingest the opponents' agreements, and to use this information to inform your own decisions. And you need, in particular, to be able to handle a wide variety of agreements from your opponents; it's not enough to force them to play the same system as you.

You must also account for deceit. If, for example, I have a bid which promises that I have at least 2 cards in every suit, it's perfectly legal for me to lie and make this bid when I only have 1 card in some suit--as long as my partner is in the dark about this just as much as the opponents. So if you make this bid, and your machine opponents assume there is a 0% probability of you having lied about your hand, it is possible that they will make gross errors by not accounting for this possibility (for example, they may be in a position where all of their actions are equivalent if you told the truth, but where one action is clearly better if you didn't--a human player will naturally take this action, but a robot may just select an action randomly).

It's an interesting game and a very interesting AI challenge.

Very detailed, accurate, and insightful summary of the interesting dynamics in bridge. So well written I would have clocked it as AI except in my experience AI is terrible at discussing game dynamics.

How does AI, in a game as complex as bridge, manage to deal with a human partner, or even an AI partner? Seems like an answer we could find out.

> The information and possibilities really need to be mediated by some kind of intermediate, rules-based description, which can be provided to the opponents as well.

This is standardized at tournaments (I'm sure you know that.)

>So well written I would have clocked it as AI

Still a few humans writing on the internet... at least for now ;) I'm sure the green username doesn't help either, I just switched to a new account a few days ago.

>This is standardized at tournaments (I'm sure you know that.)

Right, in human play we have convention cards (pieces of paper that are essentially big forms for specifying common agreements), although these don't cover every situation.

What I was getting at was more along the lines of some kind of general schema that would allow fully describing any individual bid. But then there's also a problem where such a schema inherently limits the creativity available in constructing a bidding system, if all the bids must fit into what is expressible by this schema.

My personal approach would be to try to decompose it into two subproblems: learning a set of agreements, and optimizing results given a fixed set of agreements. Then you could try to solve the former problem using the results from the latter one. But even playing well under a fixed set of agreements isn't so easy to solve.

Man I loved Hasbro’s 1998 Windows Stratego

https://archive.org/details/STRATEGO

While interesting from a technology stand point I get disheartened each time I see this sort of news.

I am particularly disappointed that it has influenced how people play the game.

The joy comes from the journey and the experience.

Look at competitive chess and Go and how they have fundamentally been transformed. It's not better and now the box is opened, it can't be closed.

> Look at competitive chess and Go and how they have fundamentally been transformed.

For people like me, competitive chess killed chess long before Deep Blue. It only feels fair that competitive chess players now feel like I did :-)

I had a math professor who played competitive chess in his youth. He told me he realized the demands for competitive chess were such that he couldn't really dedicate himself to math (or any other discipline) at the same time. So one day, he gave up chess - and refused to play it for the rest of his life.

> I am particularly disappointed that it has influenced how people play the game.

> The joy comes from the journey and the experience.

> Look at competitive chess and Go and how they have fundamentally been transformed.

Go is better since AlphaGo. Tools are better, it's easier to learn from your games, we're better at it. The AI makes sick fucking moves and we get to see.

The journey is still there, the experience is still there.

Chess I doubt is worse off either, but I don't know chess that well.