In 2019, an AI named Pluribus, developed by Carnegie Mellon University and Facebook AI, defeated top human professionals in six-player No-Limit Texas Hold'em. This was a monumental milestone. Unlike Go, poker is a game of hidden information. And unlike heads-up (two-player) poker, six-player poker lacks a computable Nash equilibrium that guarantees a non-losing strategy.

To win, Pluribus had to learn how to bluff, and more importantly, how to remain unpredictable in a chaotic, multi-agent environment.

Pluribus's Move 37 was its aggressive use of "donk betting"—a move considered so amateurish by human pros that it's named after a donkey. Pluribus proved the humans were wrong.

The Alien Strategy: Donk Betting

In human poker theory, if you call a bet on the flop, you typically check to the aggressor on the turn. Leading out with a bet (a "donk bet") is considered mathematically unsound because it gives away information and leaves you vulnerable to a raise.

Pluribus discarded this heuristic. It utilized donk betting significantly more than human pros. By calculating the expected value across thousands of simulated futures, it found that strategically employing this "amateur" move completely shattered the established human meta-game, throwing pros off balance.

Strategy Human Professional Consensus Pluribus Execution
Donk Betting Rarely used (considered a leak) Frequent, mathematically balanced
Bet Sizing A few standard sizes (e.g., 1/2 pot) Massive overbets (e.g., 2x pot)
Bluff Frequency Based on intuition and "tells" Perfectly randomized Game Theory Optimal

Self-Play in the Dark

Like AlphaStar, Pluribus trained via self-play. However, because it could not solve for a perfect Nash equilibrium in a 6-player game, it used a technique called Monte Carlo Counterfactual Regret Minimization (MCCFR). It played against copies of itself, constantly analyzing: "If I had chosen a different action in this past situation, would my payout have been higher?"

FAQ: Pluribus

Can Pluribus read human "tells"?
No. Pluribus did not have cameras to see physical tells, nor did it track the specific historical habits of individual players. It won purely by playing a strategy so mathematically robust that it exploited the inherent flaws in human play.
Did Pluribus adapt to human opponents during the game?
Surprisingly, no. It computed a "blueprint" strategy offline and only performed real-time search for the current hand, assuming the opponents were also playing optimally. Its base strategy was simply too strong.