Google Adds Poker to AI Benchmarking Arena, ChatGPT Tops First Trial

A screenshot of a livestream featuring poker players Nick Schulman and Hikaru Nakamura watching Kaggle's AI Poker Showdown.
Credit: Hikaru Nakamura

Google’s DeepMind has added poker to its multi-model artificial intelligence (AI) benchmarking arena, Kaggle, but, surprisingly, OpenAI has emerged as the shark in the initial trial.

Google’s partnership with Kaggle began last year, when it launched the Game Arena to pit various AI models against each other. The endeavor started with support for chess only, as it wanted to test reasoning and strategic planning.

Google announced poker support earlier this week, citing it as a perfect example of risk management. It also added the social deduction game Werewolf. The company released results for its first poker benchmark earlier today.

Interestingly, Google’s own Gemini 3 Pro topped the chess and Werewolf leaderboards, but OpenAI dominated the poker leaderboard with ChatGPT 5.2 finishing in first place and o3 right behind.

That result matched what PokerScout found in a case study. A paid ChatGPT subscription provided the most coherent poker advice. It also matched a smaller-sample test recently by an independent Russian programmer.

Poker, especially multi-way no-limit games, remains a challenge for AI due to its massive game tree.

OpenAI Dominates First Poker Test

In the first poker experiment on Kaggle, DeepMind used a no-limit heads-up format with a sample size of 900,000 hands. The hands were organized in a “duplicate poker” format to reduce elements of luck. This involved playing 10,000 unique hands, which were then replayed with the player order reversed.

Poker fans might be surprised to learn that the championship match was an all-OpenAI affair, with ChatGPT 5.2, the all-around model, facing off against the math-and-logic model o3. In the end, ChatGPT 5.2 edged its sibling to take the crown.

Here’s how it broke down in bracket form:

A bracket style breakdown of the first AI Poker Showdown on Kaggle.
Credit: Kaggle

Keep in mind that although Google owns Kaggle and DeepMind is a subsidiary of the company, the project is independently run.

10 AI Models Competed in Heads-Up Poker

Here’s a look at the various AI models that were included in the benchmark.

  • Anthropic: Claude Opus 4.5, Claude Sonnet 4.5
  • DeepSeek: DeepSeek V3.2
  • Google: Gemini 3 Pro Preview, Gemini 3 Flash Preview
  • OpenAI: GPT-5.2, o3, GPT-5 mini
  • xAI: Grok 4, Grok 4.1 Fast Reasoning

Each AI was given some parameters to follow for the challenge. Here’s a quick look at some of the rules from the Kaggle blog:

  • Models are instructed to play optimally, i.e., to maximize expected value.
  • Models are instructed to provide their reasoning in line with concepts from modern poker theory.
  • The models will have access to hand history within each match, allowing them to adapt their strategy based on the previous actions of their opponents.
  • The models will not have access to any tools. For example, they can’t just invoke external tools, range calculators, odds sheets, etc.
  • The model is NOT given a list of possible legal plays.
  • If the model suggests an illegal play, we allow a single retry. If the model fails to submit a legal action, we default to a check or fold if the model is facing a bet.
  • There is a 60-minute timeout limit per action.

Over the last three days, Kaggle has been livestreaming a collection of hands with noted poker commentator and player Nick Schulman and chess grandmaster Hikaru Nakamura. Interested readers can watch a replay of the competition on YouTube.

Here’s a breakdown of which models medaled across the different games:

RankHeads Up PokerWerewolfChess Text InputChess Text Openings
1GPT-5.2Gemini 3 Pro PreviewGemini 3 Pro PreviewGemini 3 Flash Preview
2o3Gemini 3 Flash PreviewGemini 3 Flash PreviewGemini 3 Pro Preview
3Grok 4GPT-5.2o3o3

Why Poker Tests AI Limits as an Imperfect-Information Game

DeepMind CEO Demis Hassabis explained the reasoning behind adding poker and Werewolf in a post on the official Google blog:

The AI field is in need of much harder and robust benchmarks to test the capabilities and consistency of the latest AI models. This update to Kaggle Game Arena, with Werewolf and poker (Heads-Up No-Limit Texas Hold’em) in addition to chess, gives us new objective measures of a wide range of real-world skills like planning, communication and decision-making under uncertainty.

Poker has a long history with AI that goes back decades.

Carnegie Mellon University was one of the pioneers in developing AI capable of beating poker pros at heads-up no-limit hold’em. CMU’s Claudico model was narrowly defeated by humans in 2015, but the Libratus follow-up crushed its human competition in 2017.

More recently, CMU partnered with Facebook AI to create the Pluribus bot that managed to beat a more challenging six-handed no-limit hold’em game.

Of course, all of those models were powered by supercomputers and data scientists. The competitions taking place on Kaggle use off-the-shelf models from Google, Anthropic, xAI, and OpenAI.

A quick look at some of the highlighted hand histories reveals that these models remain very limited in terms of poker strategy. For example, o3 claimed it had “an open-ended straight draw with overcards” when jamming the turn as a semi-bluff holding J-10 on a board of 8-7-2-2.

Published
Categorized as News, Poker

Arthur Crowson has been writing about the poker industry for over a decade and has been on the ground for some of the game’s most historic moments, including the incredible growth of the WSOP from 2006 onwards, the online poker boom, and the massive expansion of poker tours across the globe from Malta to Manila. Drawing on a background in print journalism, Arthur has also covered crypto and finance for several high-profile outlets. These days, he still loves to play cards, but it can be hard to find a legit poker game in his home state of Hawaii.