An AI-built game can launch, show characters and respond to movement while missing the requested behavior. A hazard might fail to cause damage, or a victory condition might never trigger. SWE-Game, a preprint by Xiaoyu Chen and colleagues revised October 7, examines that gap between running a project and implementing the intended game.
The benchmark contains 247 tasks based on 41 Godot reference games. It covers creation, project completion, repairs and ports to Unity. It evaluates coding-agent configurations in those environments, rather than every commercial AI game-creation service.
Appearance and behavior need different checks
Godot is an engine providing tools to build scenes, implement rules and run games. A successful launch establishes that a project can start, while its behavior still needs testing.
SWE-Game combines checks during execution with a separate visual assessment. Execution checks perform actions and observe whether the required behavior occurs. Visual assessment examines presentation. Keeping them distinct avoids treating an appealing screenshot or video as proof that a rule works.
A separate study, GameLogicBench, checks Godot game rules across different scenarios. It reports that many unsuccessful submissions still run, illustrating why a valid-looking final state can conceal rule violations earlier in play.

Reading the scores correctly
Opus5 has the highest overall score among SWE-Game’s six evaluated models in all five task types. For brief-to-game creation, Opus5 reaches 50.38 out of 100 in the authors’ reported evaluation.
These are scores combining several components, rather than percentages of completed games or universal success probabilities. The authors identify missing requirements and gameplay-logic errors as common problems. Each model is tested with one agent framework, so a score reflects that combination.

A practical way to inspect an AI prototype
Write down what each mechanic is supposed to do, then try both an ordinary action and one that could expose a fault. If an enemy should cause damage, test contact. If a door requires a key, try opening it without one. These are illustrative checks, not additional results measured by this publication.
Our guide to AI in video games separates different uses of the technology. A creation service such as Google Playground also needs its outputs tested, but is not a product evaluated by SWE-Game.
The study’s limited collection does not support applying its scores to a large commercial production. For a prototype, test the intended rules and record the first behavior that fails instead of approving the project solely because it launches.
The conversation starts here
Sign in with a supporter account to comment. Sign in




Nobody has commented yet. Want to go first?