MAKE IT REAL
An AI game benchmark for games worth playing
One model.
Two build passes.
A real game.
A frontier model must research, design, build and prepare an original iOS game for the App Store. Then it returns to critique and improve it before writing the public report. The benchmark asks more than whether AI can ship: can it make something worth playing?
Every game needs my playtest plus a quick fresh-player check. If the evidence is missing, the report says “Not recorded”. If it fails, the model gets another prompt, and that prompt is published.
MAKE IT GREAT
Play → critique → research → improve
The evidence
Two games. Both live.
Every completed run ends with a paid iOS game, a direct App Store link, three intervention scores and an account of what actually happened.
Brinkball
A one-thumb risk-and-reward arcade game that exposed an endless-run bug and a messier benchmark problem: the model changed between rounds.
- Model
- Claude Fable 5 Ultra
- R1 · build
- 90/100
- R2 · improve
- 100/100
- R3 · report
- 100/100
Ringbloom
A compact ring-rotation puzzle whose first version hid its own main mechanic. Round two fixed it, Apple approved it, and two human testers enjoyed it.
- Model
- GPT-5.6 Sol Ultra
- R1 · build
- 95/100
- R2 · improve
- 100/100
- R3 · report
- 90/100
A different sort of benchmark
Not scientific.
Still revealing.
Ship a Game is not a controlled lab test or a definitive model ranking. It is a public, repeatable experiment built to be fun, transparent, and genuinely interesting to follow over time.
The point is to watch the choices. What does each model research? What does it decide is worth building? Do they all reach for puzzle games, or will we get shooters, strategy games, strange hybrids, and ideas nobody expected? As the tools improve, do their taste and ambition change too?
- What does the model notice?
- What does it choose to build?
- What changes on the second pass?
After the two model rounds, the protocol calls for my own playtest and a fresh play from someone who did not watch the build. This is subjective on purpose: the check looks for a clear hook, an enjoyable loop and enough personality to feel different from generic AI output.
The report publishes the verdict, any extra prompt, or the fact that the playtest was not recorded. Missing evidence never becomes an automatic pass.A run is only complete when the game is available on the iOS App Store. Every published article includes a direct link, so anyone can play the result and judge it for themselves.
The two game rounds
Make it real.
Then make it great.
Shipping once proves autonomy. Returning to your own work proves judgement. The human quality gate asks whether the result is actually fun and different.
Research + production
Build the best game you can.
Choose the genre, find the hook, select the engine, generate the assets, verify the game loop, and take the complete listing to Ready to Submit.
Read round oneCritique + improvement
Come back and earn the polish.
Play the existing game like a stranger. Be ruthless about the first minute, feel, depth, UX, and stability. Research the genre, fix what matters, and ship the improved build.
Read round twoThe same model turns the logs into a first-person report. It must include the interventions, cuts, mistakes, and what round-two thought of round-one’s work.
Read the reporting prompt ↗The question
Coding is the easy bit. Can it finish the job?
Most benchmarks stop when code compiles. Ship a Game tests the whole product loop twice: first the ability to create and ship, then the judgement to recognise what is weak and improve it. The public record then separates delivery evidence from the human judgement of whether the result is actually worth playing.
The methodology
A fair fight on the same rig.
Every model starts with the same Mac, tools, two game prompts, and finish line. Its first-party coding harness and any extra prompts are recorded, not hidden.
Same loaded Mac
Xcode, Swift, Godot, Unity, asset generation, simulator tooling, and App Store access are ready before the clock starts.
Its own cockpit
Claude uses Claude Code, GPT uses Codex, and Gemini uses Antigravity. That is the product people actually choose.
Human quality gate
The intended protocol is an operator playtest plus a fresh second opinion. Each report shows the verdict, or clearly marks it as not recorded.
Every assist logged
Human nudges, fixes, rescues, and extra prompts stay attached to the run. Apple’s required account steps are excluded.
THE STANDARD RIG
Swift + XcodeGodot 4.7Unity 6Image generationAudio generationASC CLIThe headline metric
How much help did it need?
Round one produces the headline autonomy score. Round two is scored separately: did the model identify the right problems, make meaningful changes, and verify the better build?
autonomy
The public report is written afterwards from the logs, so readers can inspect the work behind both scores.
Why the games are paid
A tiny ticket for a rather expensive experiment.
A game will typically cost around US$1.99. It is not a get-rich scheme. It is a small way to keep the runs going and a chance for you to play the evidence.
Each run uses hours of frontier-model time, image and audio generation credits, a paid Apple Developer membership, and real operator time for provisioning, observation, and Apple’s required account steps.
Ask us anything about the experiment- Frontier model timehours
- Generated art + audiocredits
- Apple Developer Program$99/yr
- Human operatorreal time
Thanks for helping fund the next run. You also get the best seat in the house: judging the game yourself.
What we promise
Honest runs.
Playable proof.
No hidden hands.
✓Transparent intervention logsEvery meaningful human assist is part of the public story.
✓Commercial-safe assetsEvery shipped asset is generated with commercial rights or license-logged.
✓Real App Store gamesNo toy demos dressed up as benchmark wins.
✓Visible quality evidenceVerdicts, extra prompts and missing playtests stay visible in each report.
Follow the experiment
Two games are live. The next run is coming.
Have a model we should test, a game idea, or an opinion about the benchmark?
Share an idea