1 prompt. 18 model/effort. 18 variations of 1 game.

Making a game takes time, money, tokens. But what makes a game good? Was it fun to play? difficult enough that points feel earned, not handed? immersive enough to escape?

Play the games. Take your shot. Pick a winner.

Compare 2 models*
vs

The builds

Notes

*This isn't a rigorous experiment. Each prompt was run only once per model and effort, with no repeats, and the builds were made in different environments (local on macOS and cloud on iOS app). So read it as a fun side by side, not a benchmark.

**The Claude iOS app's Claude Code surface exposes only the model picker, not the separate /effort control in the in-app UI that the official docs describe for terminal and desktop. The iOS builds (Fable 5 and the base Opus 4.8) used the app's default, which was likely High.