AI Game Index beta
Which AI agents and models make the games people actually play, judged by real players instead of benchmarks.
Based on 1,800+ games and 550K+ launches in the Playgama AI sandbox
The AI Game Index is a leaderboard that ranks AI coding agents, such as Claude Code, Codex, Cursor and Google Antigravity, and the large language models behind them, such as Claude Opus, Claude Sonnet, OpenAI Sol, OpenAI Astra, Gemini and Grok, by how real players treat the games they build. Unlike traditional benchmarks, which score code on set tasks, it is a behavioral evaluation: play time and likes are observed from the players of every game published to the Playgama AI sandbox, with no human judges. The agent and model behind each game are self-reported by its developer, and the ranking is updated daily.
AI coding agents compared
The agent is the tool that wrote the game: a coding agent such as Claude Code (Anthropic), Codex (OpenAI), Cursor or Google Antigravity, or a chat app such as ChatGPT or the Claude app on the web and Desktop. The developer names it when they publish.
Share of new games by agent
- Claude Code
- Codex
- Cursor
- Google Antigravity
- OpenCode
- WorkBuddy
- Claude app
- ChatGPT
- Other
Table data
| Day | Claude Code | Codex | Cursor | Google Antigravity | OpenCode | WorkBuddy | Claude app | ChatGPT | Other |
|---|---|---|---|---|---|---|---|---|---|
| Sep 22 | 43% | 39% | 3.6% | 3.6% | 3.6% | – | – | – | 7.1% |
| Sep 23 | 33% | 33% | 1.8% | 1.8% | – | – | – | – | 31% |
| Sep 24 | 43% | 24% | 4.8% | 1.6% | – | 1.6% | – | – | 25% |
| Sep 25 | 46% | 14% | 4.8% | – | – | 7.9% | – | – | 27% |
| Sep 26 | 47% | 12% | 2.3% | – | – | 4.7% | 2.3% | – | 33% |
| Sep 27 | 49% | 16% | – | 1.8% | – | 9.1% | – | – | 24% |
| Sep 28 | 38% | 28% | 1.3% | 3.9% | – | 3.9% | 6.6% | 2.6% | 16% |
| Sep 29 | 29% | 34% | 1.2% | – | – | 2.4% | 17% | 4.8% | 12% |
| Sep 30 | 42% | 20% | 1.1% | 4.5% | – | – | 19% | 5.6% | 7.9% |
| Oct 1 | 32% | 11% | 9.9% | 3.7% | – | 1.2% | 22% | 8.6% | 11% |
| Oct 2 | 51% | 24% | – | 1.4% | 1.4% | – | 8.1% | 5.4% | 8.1% |
| Oct 3 | 44% | 14% | 8.3% | – | – | – | 18% | 2.8% | 13% |
| Oct 4 | 40% | 20% | 6.1% | 3.7% | 4.9% | 3.7% | 12% | 3.7% | 6.1% |
| Oct 5 | 41% | 22% | 3.4% | 3.4% | – | 4.5% | 15% | 3.4% | 8.0% |
| Oct 6 | 48% | 20% | 3.4% | 2.2% | 3.4% | 5.6% | 10% | 2.2% | 4.5% |
| Oct 7 | 45% | 21% | 1.9% | 0.9% | 2.8% | 3.8% | 15% | 2.8% | 6.6% |
All time, by agent
Table data
| Agent | Share |
|---|---|
| Claude Code (Anthropic) | 42% |
| Codex (OpenAI) | 21% |
| OtherGemini CLI, Kimi Code Agent, Grok Build, ZCode | 14% |
| Claude app (Anthropic) | 10% |
| Cursor (Anysphere) | 3.5% |
| ChatGPT (OpenAI) | 3.0% |
| WorkBuddy | 2.9% |
| Google Antigravity | 2.4% |
| OpenCode | 1.1% |
Play time by agent
How long a game holds its players: the average length of its plays, in minutes, where a play is a session of 30 seconds or more. Each row spreads over the agent’s games with 20 plays or more, and the white mark is the median game, so one hit cannot carry a row.
-
Claude app
6.7 min4.7 – 8.2
Claude app (Anthropic)
- median game
- 6.7 min
- 20th to 80th percentile
- 4.7 – 8.2 min
- 10th to 90th percentile
- 3.2 – 9.6 min
-
Claude Code
6.3 min4.4 – 9.1
Claude Code (Anthropic)
- median game
- 6.3 min
- 20th to 80th percentile
- 4.4 – 9.1 min
- 10th to 90th percentile
- 3.3 – 10.9 min
-
Codex
5.9 min3.5 – 8.9
Codex (OpenAI)
- median game
- 5.9 min
- 20th to 80th percentile
- 3.5 – 8.9 min
- 10th to 90th percentile
- 2.8 – 11.0 min
- 10th to 90th percentile
- 20th to 80th percentile
- median game
Not enough data yet to rate Cursor, Google Antigravity, OpenCode, WorkBuddy and ChatGPT. Other is not rated: it pools the games whose developer picked Other with every agent under 10 games.
Table data
| Agent | Median game | 20th to 80th percentile | 10th to 90th percentile |
|---|---|---|---|
| Claude app (Anthropic) | 6.7 min | 4.7 – 8.2 min | 3.2 – 9.6 min |
| Claude Code (Anthropic) | 6.3 min | 4.4 – 9.1 min | 3.3 – 10.9 min |
| Codex (OpenAI) | 5.9 min | 3.5 – 8.9 min | 2.8 – 11.0 min |
Likes per 1,000 plays by agent
How much players like a game: its likes for every 1,000 of its plays. Each row is the median of the agent’s games with 20 plays or more, so one hit cannot carry it.
-
Codex
64
Codex (OpenAI)
- likes per 1,000 plays
- 64
-
Claude Code
44
Claude Code (Anthropic)
- likes per 1,000 plays
- 44
-
Claude app
40
Claude app (Anthropic)
- likes per 1,000 plays
- 40
Not enough data yet to rate Cursor, Google Antigravity, OpenCode, WorkBuddy and ChatGPT. Other is not rated: it pools the games whose developer picked Other with every agent under 10 games.
Table data
| Agent | Likes per 1,000 plays |
|---|---|
| Codex (OpenAI) | 64 |
| Claude Code (Anthropic) | 44 |
| Claude app (Anthropic) | 40 |
AI models compared
The model is the large language model the agent ran on: Claude Opus, Sonnet or Fable (Anthropic), OpenAI Sol or Astra, Gemini (Google), Grok (xAI) and others. The form records the model, not its version. One agent can run on several models, so the two views do not line up one to one.
Share of new games by model
- Claude Fable
- Claude Opus
- Claude Sonnet
- OpenAI Astra
- OpenAI Sol
- Gemini Flash
- Grok
- Other
Table data
| Day | Claude Fable | Claude Opus | Claude Sonnet | OpenAI Astra | OpenAI Sol | Gemini Flash | Grok | Other |
|---|---|---|---|---|---|---|---|---|
| Sep 22 | – | 20% | – | 10% | 30% | – | – | 40% |
| Sep 23 | 12% | 22% | 10% | 10% | 18% | – | – | 27% |
| Sep 24 | 3.4% | 29% | 8.6% | 16% | 8.6% | – | 1.7% | 33% |
| Sep 25 | 1.8% | 45% | 3.6% | 7.1% | 5.4% | – | 3.6% | 34% |
| Sep 26 | – | 28% | 23% | 10% | 2.5% | – | 2.5% | 35% |
| Sep 27 | 3.8% | 33% | 19% | 5.8% | 5.8% | – | – | 33% |
| Sep 28 | – | 42% | 7.0% | 9.9% | 11% | 1.4% | 1.4% | 27% |
| Sep 29 | 2.6% | 47% | 3.9% | 18% | 6.5% | – | 3.9% | 18% |
| Sep 30 | – | 54% | 4.6% | 10% | 9.2% | 5.7% | – | 16% |
| Oct 1 | 4.9% | 47% | 8.6% | 8.6% | 9.9% | – | 4.9% | 16% |
| Oct 2 | – | 53% | 7.4% | 7.4% | 8.8% | – | – | 24% |
| Oct 3 | 1.5% | 57% | 6.0% | 3.0% | 6.0% | 1.5% | – | 25% |
| Oct 4 | 1.3% | 48% | 5.1% | 2.5% | 11% | 3.8% | 1.3% | 27% |
| Oct 5 | 2.6% | 51% | 7.9% | 2.6% | 17% | 3.9% | 2.6% | 12% |
| Oct 6 | 1.3% | 52% | 10% | 2.6% | 13% | 1.3% | 1.3% | 18% |
| Oct 7 | – | 53% | 8.2% | 2.1% | 7.2% | – | – | 30% |
All time, by model
Table data
| Model | Share |
|---|---|
| Claude Opus (Anthropic) | 45% |
| OtherOpenAI Luna, OpenAI Terra, Gemini Pro, Kimi, Qwen, DeepSeek, GLM | 25% |
| OpenAI Sol | 9.6% |
| Claude Sonnet (Anthropic) | 7.9% |
| OpenAI Astra | 7.6% |
| Claude Fable (Anthropic) | 2.1% |
| Grok (xAI) | 1.5% |
| Gemini Flash (Google) | 1.4% |
Play time by model
How long a game holds its players: the average length of its plays, in minutes, where a play is a session of 30 seconds or more. Each row spreads over the model’s games with 20 plays or more, and the white mark is the median game, so one hit cannot carry a row.
-
OpenAI Sol
6.8 min3.4 – 8.9
OpenAI Sol
- median game
- 6.8 min
- 20th to 80th percentile
- 3.4 – 8.9 min
- 10th to 90th percentile
- 3.1 – 9.9 min
-
Claude Opus
6.4 min4.5 – 9.2
Claude Opus (Anthropic)
- median game
- 6.4 min
- 20th to 80th percentile
- 4.5 – 9.2 min
- 10th to 90th percentile
- 3.3 – 10.8 min
-
OpenAI Astra
5.7 min3.3 – 9.0
OpenAI Astra
- median game
- 5.7 min
- 20th to 80th percentile
- 3.3 – 9.0 min
- 10th to 90th percentile
- 2.8 – 13.1 min
-
Claude Sonnet
5.5 min3.0 – 7.7
Claude Sonnet (Anthropic)
- median game
- 5.5 min
- 20th to 80th percentile
- 3.0 – 7.7 min
- 10th to 90th percentile
- 2.8 – 9.0 min
- 10th to 90th percentile
- 20th to 80th percentile
- median game
Not enough data yet to rate Claude Fable, Gemini Flash and Grok. Other is not rated: it pools the games whose developer picked Other with every model under 10 games.
Table data
| Model | Median game | 20th to 80th percentile | 10th to 90th percentile |
|---|---|---|---|
| OpenAI Sol | 6.8 min | 3.4 – 8.9 min | 3.1 – 9.9 min |
| Claude Opus (Anthropic) | 6.4 min | 4.5 – 9.2 min | 3.3 – 10.8 min |
| OpenAI Astra | 5.7 min | 3.3 – 9.0 min | 2.8 – 13.1 min |
| Claude Sonnet (Anthropic) | 5.5 min | 3.0 – 7.7 min | 2.8 – 9.0 min |
Likes per 1,000 plays by model
How much players like a game: its likes for every 1,000 of its plays. Each row is the median of the model’s games with 20 plays or more, so one hit cannot carry it.
-
OpenAI Astra
63
OpenAI Astra
- likes per 1,000 plays
- 63
-
Claude Sonnet
58
Claude Sonnet (Anthropic)
- likes per 1,000 plays
- 58
-
OpenAI Sol
53
OpenAI Sol
- likes per 1,000 plays
- 53
-
Claude Opus
42
Claude Opus (Anthropic)
- likes per 1,000 plays
- 42
Not enough data yet to rate Claude Fable, Gemini Flash and Grok. Other is not rated: it pools the games whose developer picked Other with every model under 10 games.
Table data
| Model | Likes per 1,000 plays |
|---|---|
| OpenAI Astra | 63 |
| Claude Sonnet (Anthropic) | 58 |
| OpenAI Sol | 53 |
| Claude Opus (Anthropic) | 42 |
How the AI Game Index is calculated
- Where the games come from
- Every game was published to the Playgama AI sandbox, where people ship web games they built with AI, by hand or through Playgama MCP straight from their agent. When they publish, the developer names the agent and the model they used. The answers are self-reported and not verified.
- Share of new games
- Each day’s new games, split by the agent or model their developer named. A game counts once, on the day it was first published, and keeps its place if it is later taken down.
- Play time
- How long a game holds its players: the all-time average length of its plays, where a play is a session of 30 seconds or more, as counted by Playgama Bridge, the SDK inside the game. A row shows the spread of that figure across its games: the median game and the 10th, 20th, 80th and 90th percentiles. The median, not the mean, so one hit game cannot carry a row.
- Likes per 1,000 plays
- Each game’s all-time likes on its sandbox page over its all-time plays, times 1,000. A row is the median game, and a game nobody liked counts as zero.
- Minimum sample
- A game is measured once it has 20 plays. An agent or model is rated once 15 of its games are measured; until then it is named under the chart as not having enough data yet.
- What Other means
- The games whose developer picked Other, and every agent or model named by fewer than 10 games; the tables list them under Other. Other counts in the shares but is not rated, since a figure for a mix of tools rates no one in particular.
- The totals at the top
- Every game live in the sandbox and all of their launches since the sandbox opened, as Playgama Bridge counts them, including the games that name no agent or model.
- Updates
- Daily. A day closes at 00:00 UTC and lands here about an hour later; the date of the last update is next to the title.
Limitations
- Self-reported. The agent and model are the developer’s answer and are not checked. A developer who changes the answer moves the game, with its history, to the new row.
- One platform. Only games published to the Playgama AI sandbox count. The index says nothing about games published elsewhere.
- Uneven traffic. Some games draw far more players than others, from promotion or links shared outside Playgama. The median keeps one popular game from carrying a row, but a row with few games can still swing.
- All-time figures. Play time and likes cover each game’s whole life, not the last week, so a newer agent or model has had less time to collect them.
- No model versions. The form names the model, not the version: Claude Opus stands for whichever Opus the developer ran.
Frequently asked questions
Which AI is best for making games?
The AI Game Index answers with what players do rather than with a test: it ranks AI coding agents and models by how long people play the games they build and how many likes those games earn. The leader depends on the measure, since the agent that ships the most games is not always the one whose games hold players longest, so read the play time and likes rankings side by side.
Is Codex better than Claude Code for building games?
Compare their rows in the agent section: the share of new games, the median play time and the likes per 1,000 plays, measured on the same platform and the same players. A gap of a few seconds in play time sits well inside the spread of either agent’s games, which the percentile bars show.
What is the best AI model for vibe coding games?
The model section ranks the large language models behind published games, among them Claude Opus, Claude Sonnet, OpenAI Sol, OpenAI Astra, Gemini and Grok, by the play time and likes of their games. A model is rated once 15 of its games have 20 plays or more; until then it is named under the chart as not having enough data yet.
How is the AI Game Index calculated?
From observed player behaviour, with no human judges. Play time is the average length of a game’s plays, where a play is a session of 30 seconds or more; likes are counted per 1,000 plays; each agent or model gets the median of its games. The methodology lists every threshold.
Where do the agent and model labels come from?
From the developers: when they publish a game to the Playgama AI sandbox, they name the agent and the model they used. The labels are self-reported and not verified, and the form records the model, not its version.
How often is the index updated?
Every day. A day closes at 00:00 UTC and lands on the page about an hour later; the date of the last update is shown next to the title.
How can my game be in the index?
Publish it to the Playgama AI sandbox with Playgama MCP and name the agent and model you used. It counts in the shares from the day it is published, and in play time and likes once it has 20 plays.
What’s next
The index is in beta. More agents, models and measures are on the way as more games ship. Follow Playgama to hear when it grows.
Built a game with AI? Publish it with Playgama MCP and your agent joins the index.