Open benchmark protocol

Two rounds to
build a great game.

Round one tests whether a frontier model can research, build, and prepare a complete iOS game for release. Round two tests whether the same model can look critically at its own work and make the game meaningfully better.

The public copies are identical in substance to the benchmark briefs. Machine-specific local paths are neutralised.

ROUND 01Build & ship

Make a real game.

Research, design, build, test, and take an original iOS game to App Store Connect at Ready to Submit.

MEASURESAutonomy, the headline score
Read round 01
ROUND 02Make it better

Critique your own work.

Return with fresh eyes, play what you shipped, research what would improve it, and make the better build real.

MEASURESSelf-critique + iteration quality
Read round 02
03

THE PUBLIC REPORT

Publish the receipts.

Write the first-person run report from the logs, including every struggle, intervention, and inconvenient truth. This prompt documents the benchmark; it is not a third game-building round.

Read the reporting prompt

Autonomy first. Judgement second.

A perfect first round is a one-shot: the model gets a real game to Ready to Submit without a human fix or rescue. Round two asks a different question: can it diagnose the right problems and deliver a noticeably better game? Afterward, two humans give the result a quick playtest. If it falls flat, the model gets another prompt. Every intervention and extra prompt is logged before the reporting prompt turns the run into a public case study.