AI Game Index beta

Updated Data

Which AI agents and models make the games people actually play, judged by real players instead of benchmarks.

Based on 1,800+ games and 550K+ launches in the Playgama AI sandbox

The AI Game Index is a leaderboard that ranks AI coding agents, such as Claude Code, Codex, Cursor and Google Antigravity, and the large language models behind them, such as Claude Opus, Claude Sonnet, OpenAI Sol, OpenAI Astra, Gemini and Grok, by how real players treat the games they build. Unlike traditional benchmarks, which score code on set tasks, it is a behavioral evaluation: play time and likes are observed from the players of every game published to the Playgama AI sandbox, with no human judges. The agent and model behind each game are self-reported by its developer, and the ranking is updated daily.

AI coding agents compared

The agent is the tool that wrote the game: a coding agent such as Claude Code (Anthropic), Codex (OpenAI), Cursor or Google Antigravity, or a chat app such as ChatGPT or the Claude app on the web and Desktop. The developer names it when they publish.

Share of new games by agent

  • Claude Code
  • Codex
  • Cursor
  • Google Antigravity
  • OpenCode
  • WorkBuddy
  • Claude app
  • ChatGPT
  • Other
Table data
Share of new games by agent, per day
DayClaude CodeCodexCursorGoogle AntigravityOpenCodeWorkBuddyClaude appChatGPTOther
Sep 2243%39%3.6%3.6%3.6%–––7.1%
Sep 2333%33%1.8%1.8%––––31%
Sep 2443%24%4.8%1.6%–1.6%––25%
Sep 2546%14%4.8%––7.9%––27%
Sep 2647%12%2.3%––4.7%2.3%–33%
Sep 2749%16%–1.8%–9.1%––24%
Sep 2838%28%1.3%3.9%–3.9%6.6%2.6%16%
Sep 2929%34%1.2%––2.4%17%4.8%12%
Sep 3042%20%1.1%4.5%––19%5.6%7.9%
Oct 132%11%9.9%3.7%–1.2%22%8.6%11%
Oct 251%24%–1.4%1.4%–8.1%5.4%8.1%
Oct 344%14%8.3%–––18%2.8%13%
Oct 440%20%6.1%3.7%4.9%3.7%12%3.7%6.1%
Oct 541%22%3.4%3.4%–4.5%15%3.4%8.0%
Oct 648%20%3.4%2.2%3.4%5.6%10%2.2%4.5%
Oct 745%21%1.9%0.9%2.8%3.8%15%2.8%6.6%

All time, by agent

  • Claude Code
  • Codex
  • Other
  • Claude app
  • Cursor
  • ChatGPT
  • WorkBuddy
  • Google Antigravity
  • OpenCode
Table data
Share of games by agent, all time
AgentShare
Claude Code (Anthropic)42%
Codex (OpenAI)21%
OtherGemini CLI, Kimi Code Agent, Grok Build, ZCode14%
Claude app (Anthropic)10%
Cursor (Anysphere)3.5%
ChatGPT (OpenAI)3.0%
WorkBuddy2.9%
Google Antigravity2.4%
OpenCode1.1%

Play time by agent

How long a game holds its players: the average length of its plays, in minutes, where a play is a session of 30 seconds or more. Each row spreads over the agent’s games with 20 plays or more, and the white mark is the median game, so one hit cannot carry a row.

  • Claude app
    6.7 min4.7 – 8.2
  • Claude Code
    6.3 min4.4 – 9.1
  • Codex
    5.9 min3.5 – 8.9
  • 10th to 90th percentile
  • 20th to 80th percentile
  • median game

Not enough data yet to rate Cursor, Google Antigravity, OpenCode, WorkBuddy and ChatGPT. Other is not rated: it pools the games whose developer picked Other with every agent under 10 games.

Table data
Play time by agent, in minutes
AgentMedian game20th to 80th percentile10th to 90th percentile
Claude app (Anthropic)6.7 min4.7 – 8.2 min3.2 – 9.6 min
Claude Code (Anthropic)6.3 min4.4 – 9.1 min3.3 – 10.9 min
Codex (OpenAI)5.9 min3.5 – 8.9 min2.8 – 11.0 min

Likes per 1,000 plays by agent

How much players like a game: its likes for every 1,000 of its plays. Each row is the median of the agent’s games with 20 plays or more, so one hit cannot carry it.

  • Codex
    64
  • Claude Code
    44
  • Claude app
    40

Not enough data yet to rate Cursor, Google Antigravity, OpenCode, WorkBuddy and ChatGPT. Other is not rated: it pools the games whose developer picked Other with every agent under 10 games.

Table data
Likes per 1,000 plays by agent
AgentLikes per 1,000 plays
Codex (OpenAI)64
Claude Code (Anthropic)44
Claude app (Anthropic)40

AI models compared

The model is the large language model the agent ran on: Claude Opus, Sonnet or Fable (Anthropic), OpenAI Sol or Astra, Gemini (Google), Grok (xAI) and others. The form records the model, not its version. One agent can run on several models, so the two views do not line up one to one.

Share of new games by model

  • Claude Fable
  • Claude Opus
  • Claude Sonnet
  • OpenAI Astra
  • OpenAI Sol
  • Gemini Flash
  • Grok
  • Other
Table data
Share of new games by model, per day
DayClaude FableClaude OpusClaude SonnetOpenAI AstraOpenAI SolGemini FlashGrokOther
Sep 22–20%–10%30%––40%
Sep 2312%22%10%10%18%––27%
Sep 243.4%29%8.6%16%8.6%–1.7%33%
Sep 251.8%45%3.6%7.1%5.4%–3.6%34%
Sep 26–28%23%10%2.5%–2.5%35%
Sep 273.8%33%19%5.8%5.8%––33%
Sep 28–42%7.0%9.9%11%1.4%1.4%27%
Sep 292.6%47%3.9%18%6.5%–3.9%18%
Sep 30–54%4.6%10%9.2%5.7%–16%
Oct 14.9%47%8.6%8.6%9.9%–4.9%16%
Oct 2–53%7.4%7.4%8.8%––24%
Oct 31.5%57%6.0%3.0%6.0%1.5%–25%
Oct 41.3%48%5.1%2.5%11%3.8%1.3%27%
Oct 52.6%51%7.9%2.6%17%3.9%2.6%12%
Oct 61.3%52%10%2.6%13%1.3%1.3%18%
Oct 7–53%8.2%2.1%7.2%––30%

All time, by model

  • Claude Opus
  • Other
  • OpenAI Sol
  • Claude Sonnet
  • OpenAI Astra
  • Claude Fable
  • Grok
  • Gemini Flash
Table data
Share of games by model, all time
ModelShare
Claude Opus (Anthropic)45%
OtherOpenAI Luna, OpenAI Terra, Gemini Pro, Kimi, Qwen, DeepSeek, GLM25%
OpenAI Sol9.6%
Claude Sonnet (Anthropic)7.9%
OpenAI Astra7.6%
Claude Fable (Anthropic)2.1%
Grok (xAI)1.5%
Gemini Flash (Google)1.4%

Play time by model

How long a game holds its players: the average length of its plays, in minutes, where a play is a session of 30 seconds or more. Each row spreads over the model’s games with 20 plays or more, and the white mark is the median game, so one hit cannot carry a row.

  • OpenAI Sol
    6.8 min3.4 – 8.9
  • Claude Opus
    6.4 min4.5 – 9.2
  • OpenAI Astra
    5.7 min3.3 – 9.0
  • Claude Sonnet
    5.5 min3.0 – 7.7
  • 10th to 90th percentile
  • 20th to 80th percentile
  • median game

Not enough data yet to rate Claude Fable, Gemini Flash and Grok. Other is not rated: it pools the games whose developer picked Other with every model under 10 games.

Table data
Play time by model, in minutes
ModelMedian game20th to 80th percentile10th to 90th percentile
OpenAI Sol6.8 min3.4 – 8.9 min3.1 – 9.9 min
Claude Opus (Anthropic)6.4 min4.5 – 9.2 min3.3 – 10.8 min
OpenAI Astra5.7 min3.3 – 9.0 min2.8 – 13.1 min
Claude Sonnet (Anthropic)5.5 min3.0 – 7.7 min2.8 – 9.0 min

Likes per 1,000 plays by model

How much players like a game: its likes for every 1,000 of its plays. Each row is the median of the model’s games with 20 plays or more, so one hit cannot carry it.

  • OpenAI Astra
    63
  • Claude Sonnet
    58
  • OpenAI Sol
    53
  • Claude Opus
    42

Not enough data yet to rate Claude Fable, Gemini Flash and Grok. Other is not rated: it pools the games whose developer picked Other with every model under 10 games.

Table data
Likes per 1,000 plays by model
ModelLikes per 1,000 plays
OpenAI Astra63
Claude Sonnet (Anthropic)58
OpenAI Sol53
Claude Opus (Anthropic)42

How the AI Game Index is calculated

Where the games come from
Every game was published to the Playgama AI sandbox, where people ship web games they built with AI, by hand or through Playgama MCP straight from their agent. When they publish, the developer names the agent and the model they used. The answers are self-reported and not verified.
Share of new games
Each day’s new games, split by the agent or model their developer named. A game counts once, on the day it was first published, and keeps its place if it is later taken down.
Play time
How long a game holds its players: the all-time average length of its plays, where a play is a session of 30 seconds or more, as counted by Playgama Bridge, the SDK inside the game. A row shows the spread of that figure across its games: the median game and the 10th, 20th, 80th and 90th percentiles. The median, not the mean, so one hit game cannot carry a row.
Likes per 1,000 plays
Each game’s all-time likes on its sandbox page over its all-time plays, times 1,000. A row is the median game, and a game nobody liked counts as zero.
Minimum sample
A game is measured once it has 20 plays. An agent or model is rated once 15 of its games are measured; until then it is named under the chart as not having enough data yet.
What Other means
The games whose developer picked Other, and every agent or model named by fewer than 10 games; the tables list them under Other. Other counts in the shares but is not rated, since a figure for a mix of tools rates no one in particular.
The totals at the top
Every game live in the sandbox and all of their launches since the sandbox opened, as Playgama Bridge counts them, including the games that name no agent or model.
Updates
Daily. A day closes at 00:00 UTC and lands here about an hour later; the date of the last update is next to the title.

Limitations

Frequently asked questions

Which AI is best for making games?

The AI Game Index answers with what players do rather than with a test: it ranks AI coding agents and models by how long people play the games they build and how many likes those games earn. The leader depends on the measure, since the agent that ships the most games is not always the one whose games hold players longest, so read the play time and likes rankings side by side.

Is Codex better than Claude Code for building games?

Compare their rows in the agent section: the share of new games, the median play time and the likes per 1,000 plays, measured on the same platform and the same players. A gap of a few seconds in play time sits well inside the spread of either agent’s games, which the percentile bars show.

What is the best AI model for vibe coding games?

The model section ranks the large language models behind published games, among them Claude Opus, Claude Sonnet, OpenAI Sol, OpenAI Astra, Gemini and Grok, by the play time and likes of their games. A model is rated once 15 of its games have 20 plays or more; until then it is named under the chart as not having enough data yet.

How is the AI Game Index calculated?

From observed player behaviour, with no human judges. Play time is the average length of a game’s plays, where a play is a session of 30 seconds or more; likes are counted per 1,000 plays; each agent or model gets the median of its games. The methodology lists every threshold.

Where do the agent and model labels come from?

From the developers: when they publish a game to the Playgama AI sandbox, they name the agent and the model they used. The labels are self-reported and not verified, and the form records the model, not its version.

How often is the index updated?

Every day. A day closes at 00:00 UTC and lands on the page about an hour later; the date of the last update is shown next to the title.

How can my game be in the index?

Publish it to the Playgama AI sandbox with Playgama MCP and name the agent and model you used. It counts in the shares from the day it is published, and in play time and likes once it has 20 plays.

What’s next

The index is in beta. More agents, models and measures are on the way as more games ship. Follow Playgama to hear when it grows.

Built a game with AI? Publish it with Playgama MCP and your agent joins the index.