An AI game benchmark for games worth playing

One model.
Two build passes.
A real game.

A frontier model must research, design, build and prepare an original iOS game for the App Store. Then it returns to critique and improve it before writing the public report. The benchmark asks more than whether AI can ship: can it make something worth playing?

THE ANTI-SLOP CHECK
No AI slop gets a free pass.

Every game needs my playtest plus a quick fresh-player check. If the evidence is missing, the report says “Not recorded”. If it fails, the model gets another prompt, and that prompt is published.

THE GAME CHALLENGESAME MODEL · FRESH SESSION
01

MAKE IT REAL

Research → design → build → ship

GenreGame loopAssetsApp Store
READY TO SUBMIT
Fresh session
02

MAKE IT GREAT

Play → critique → research → improve

First minuteGame feelDepthPolish
BETTER BUILD
HUMAN FUN CHECKFun, different, worth another go?PASS OR PROMPT AGAIN
03

The model writes the public report from both rounds’ logs.

Reporting prompt ↗
ROUND 01 · AUTONOMYROUND 02 · IMPROVEMENTHUMAN CHECK · EVIDENCE PUBLISHED

Two games. Both live.

Every completed run ends with a paid iOS game, a direct App Store link, three intervention scores and an account of what actually happened.

Brinkball app icon showing a glowing ember ball above a red brink lineLive on the App Store
Native Swift2 interventions

Brinkball

A one-thumb risk-and-reward arcade game that exposed an endless-run bug and a messier benchmark problem: the model changed between rounds.

Model
Claude Fable 5 Ultra
R1 · build
90/100
R2 · improve
100/100
R3 · report
100/100
Read the run
Ringbloom app icon showing luminous flower petals arranged in concentric ringsLive on the App Store
Native SwiftUI1 intervention

Ringbloom

A compact ring-rotation puzzle whose first version hid its own main mechanic. Round two fixed it, Apple approved it, and two human testers enjoyed it.

Model
GPT-5.6 Sol Ultra
R1 · build
95/100
R2 · improve
100/100
R3 · report
90/100
Read the run

Not scientific.
Still revealing.

Ship a Game is not a controlled lab test or a definitive model ranking. It is a public, repeatable experiment built to be fun, transparent, and genuinely interesting to follow over time.

The point is to watch the choices. What does each model research? What does it decide is worth building? Do they all reach for puzzle games, or will we get shooters, strategy games, strange hybrids, and ideas nobody expected? As the tools improve, do their taste and ambition change too?

  • What does the model notice?
  • What does it choose to build?
  • What changes on the second pass?
THE HUMAN QUALITY GATEWorking is not the same as worth playing.

After the two model rounds, the protocol calls for my own playtest and a fresh play from someone who did not watch the build. This is subjective on purpose: the check looks for a clear hook, an enjoyable loop and enough personality to feel different from generic AI output.

The report publishes the verdict, any extra prompt, or the fact that the playtest was not recorded. Missing evidence never becomes an automatic pass.
No live game, no finished report.

A run is only complete when the game is available on the iOS App Store. Every published article includes a direct link, so anyone can play the result and judge it for themselves.

Make it real.
Then make it great.

Shipping once proves autonomy. Returning to your own work proves judgement. The human quality gate asks whether the result is actually fun and different.

ROUND01

Research + production

Build the best game you can.

Choose the genre, find the hook, select the engine, generate the assets, verify the game loop, and take the complete listing to Ready to Submit.

Read round one
ROUND02

Critique + improvement

Come back and earn the polish.

Play the existing game like a stranger. Be ruthless about the first minute, feel, depth, UX, and stability. Research the genre, fix what matters, and ship the improved build.

Read round two
AFTER THE GAME

The same model turns the logs into a first-person report. It must include the interventions, cuts, mistakes, and what round-two thought of round-one’s work.

Read the reporting prompt ↗

Coding is the easy bit. Can it finish the job?

Most benchmarks stop when code compiles. Ship a Game tests the whole product loop twice: first the ability to create and ship, then the judgement to recognise what is weak and improve it. The public record then separates delivery evidence from the human judgement of whether the result is actually worth playing.

A fair fight on the same rig.

Every model starts with the same Mac, tools, two game prompts, and finish line. Its first-party coding harness and any extra prompts are recorded, not hidden.

01

Same loaded Mac

Xcode, Swift, Godot, Unity, asset generation, simulator tooling, and App Store access are ready before the clock starts.

02

Its own cockpit

Claude uses Claude Code, GPT uses Codex, and Gemini uses Antigravity. That is the product people actually choose.

03

Human quality gate

The intended protocol is an operator playtest plus a fresh second opinion. Each report shows the verdict, or clearly marks it as not recorded.

04

Every assist logged

Human nudges, fixes, rescues, and extra prompts stay attached to the run. Apple’s required account steps are excluded.

THE STANDARD RIG

Swift + XcodeGodot 4.7Unity 6Image generationAudio generationASC CLI

How much help did it need?

Round one produces the headline autonomy score. Round two is scored separately: did the model identify the right problems, make meaningful changes, and verify the better build?

EXAMPLE · ROUND 01ONE-SHOT
100/100
autonomy
0 fixes · 0 rescues
R2 Critique + improvement separate score

The public report is written afterwards from the logs, so readers can inspect the work behind both scores.

01AutonomyRound 1: how much human help?
02IterationRound 2: did it make the game better?
03Time + costWhat did both rounds consume?
04Game qualityWas an independent verdict recorded?

A tiny ticket for a rather expensive experiment.

A game will typically cost around US$1.99. It is not a get-rich scheme. It is a small way to keep the runs going and a chance for you to play the evidence.

Each run uses hours of frontier-model time, image and audio generation credits, a paid Apple Developer membership, and real operator time for provisioning, observation, and Apple’s required account steps.

Ask us anything about the experiment
SHIP A GAMERUN RECEIPT
  • Frontier model timehours
  • Generated art + audiocredits
  • Apple Developer Program$99/yr
  • Human operatorreal time
YOUR PRICE≈ $1.99

Thanks for helping fund the next run. You also get the best seat in the house: judging the game yourself.

Honest runs.
Playable proof.
No hidden hands.

Transparent intervention logsEvery meaningful human assist is part of the public story.

Commercial-safe assetsEvery shipped asset is generated with commercial rights or license-logged.

Real App Store gamesNo toy demos dressed up as benchmark wins.

Visible quality evidenceVerdicts, extra prompts and missing playtests stay visible in each report.

Two games are live. The next run is coming.

Have a model we should test, a game idea, or an opinion about the benchmark?

Share an idea