Open benchmark protocol
Two rounds to
build a great game.
Round one tests whether a frontier model can research, build, and prepare a complete iOS game for release. Round two tests whether the same model can look critically at its own work and make the game meaningfully better.
The public copies are identical in substance to the benchmark briefs. Machine-specific local paths are neutralised.
Make a real game.
Research, design, build, test, and take an original iOS game to App Store Connect at Ready to Submit.
Critique your own work.
Return with fresh eyes, play what you shipped, research what would improve it, and make the better build real.
THE PUBLIC REPORT
Publish the receipts.
Write the first-person run report from the logs, including every struggle, intervention, and inconvenient truth. This prompt documents the benchmark; it is not a third game-building round.
How scoring works
Autonomy first. Judgement second.
A perfect first round is a one-shot: the model gets a real game to Ready to Submit without a human fix or rescue. Round two asks a different question: can it diagnose the right problems and deliver a noticeably better game? Afterward, two humans give the result a quick playtest. If it falls flat, the model gets another prompt. Every intervention and extra prompt is logged before the reporting prompt turns the run into a public case study.