GPT-6 vs. Claude Opus 5.5: Who Builds a Better Game From One Prompt?

Which model gives you the better playable game after a single prompt? As of September 2026, the frontier comparison is OpenAI GPT-6 Astra versus Anthropic Claude Opus 5.5. GPT-6 Sol and Luna can build games too, but OpenAI still identifies Astra as its best model across the board, making it the appropriate head-to-head rival for Opus 5.5.
The practical answer is split. Opus 5.5 produced the richer and more game-like first build in the most direct public same-prompt comparison available. Astra is unusually strong at the full build-run-observe-fix loop, especially when its agent can operate a browser or game engine. Cost does not have a fixed winner: the model that decides to do more work can produce the better showcase and the larger bill at the same time.
What “one-prompt game creation” actually means
This is not merely asking a chatbot to print one HTML file. A capable coding agent has to interpret the brief, design game state and controls, create a project, run it, inspect the real screen, recover from errors and leave a playable build. The quality of its terminal, browser and engine access matters alongside the model itself.

A fair comparison therefore includes first-run success, controls, game feel, system depth, visual presentation, multi-file consistency, browser testing, regression repair, mobile behavior and asset licensing. A conventional coding benchmark can illuminate part of that pipeline, but it cannot tell us whether a jump feels satisfying or whether a level remains fun after the third revision.
The most direct test: two prompts, four models
The strongest direct evidence is Moe Lueker’s September 25 same-prompt test. Each model received an empty folder and an isolated Git repository. The evaluator pasted the same prompt into separate sessions and gave no follow-up fixes. The tasks were a deliberately simple runner with a flying level and a cozy island builder with more interacting systems.
In the runner round, Opus 5.5 added jumping, diving, braking, sprinting, checkpoints, a mushroom boost, a hold-to-launch glider, a swing action and original background music. Astra produced real art, a parallax background and a climbable glide section, but fewer mechanics. The evaluator counted 2,405 lines of game code from Opus and 358 from Astra.
In the island-builder round, Opus added placement bonuses, a night mode with lit houses, walking residents, sailing boats and an offshore whale. Astra offered a cleaner but much simpler building game. The evaluator preferred the Opus result in both rounds.
Better first builds can consume more time and money
These figures are not subscription charges. The evaluator took the local build-turn logs and priced them at API list rates; follow-up prompts were excluded.

On the runner, Opus generated roughly 256,000 output tokens versus Astra’s 36,000—more than seven times as much. Opus has lower per-token API rates, but it did so much more work that the build cost $15.25 versus Astra’s $7.09 and took 62 minutes instead of 24. On the island builder, the output gap narrowed to roughly three times and the bill flipped: $5.40 for Opus versus $6.88 for Astra.
The useful metric is cost per completed outcome, not price per token. Opus can spend its budget filling in mechanics, animation and audio that the prompt did not explicitly request. Astra may finish a narrower interpretation faster, then need more follow-up prompts to reach the same level of richness. Either behavior can be preferable depending on whether scope discipline or showcase impact matters more.
Why this is evidence, not a universal model ranking
The test controlled the prompt and isolated the folders, but it was not a laboratory benchmark. Opus ran in Claude Code while the GPT-6 models ran in Codex. Luna used max effort; the other models used xhigh for the runner and high for the island. The harnesses have different system instructions, tool behavior, context management and default testing habits.
The sample is also just two games. A model that excels at a compact 2D browser project may not keep the same advantage in Unity 3D physics, a Godot scene tree, multiplayer synchronization or mobile memory optimization. Treat this as a concrete, reproducible-style case study—not proof that one model wins every game-development task.
What adjacent coding benchmarks add
Anthropic’s official Opus 5.5 page reports 66.4% for Opus 5.5 and 57.9% for GPT-6 Astra on Terminal-Bench 4.0. That supports Opus’s strength in long, terminal-heavy coding work. The same vendor table prevents a universal conclusion: Astra reaches 41.4% on AutomationBench versus Opus’s 40.0%, and Astra leads Terminal-Bench-Science 64.6% to 58.7%.
Artificial Analysis’s independent release comparison places the models close under matched evaluation infrastructure. Opus 5.5 high scores 54 on its Intelligence Index and Astra max scores 53. Its Terminal-Bench 4.0 measurement is 59.6% for both. These results support the claim that both are frontier coding agents, but none directly measures game feel, art direction or the fun of a one-shot build.
Astra’s edge appears after the code is written
OpenAI’s GPT-6 Astra announcement emphasizes computer use: creating a website, running frontend QA, installing and testing software, and troubleshooting what the model sees on screen. Lovable’s early testing says higher effort bought more iterations on fresh builds and more verification through browser testing.
The OpenAI Developers game-building case shows why that matters. It used named test scenes, Playwright screenshots, real control paths and recorded state changes. Numerical terrain checks did not replace browser playtesting for shader compilation, GPU uploads, motion or transition feel. The model investigated, implemented and reran measurements while a person judged what looked and felt right.
Playco’s Astra customer story describes a similar engine-native loop in Unity and Godot. Playco says Astra produced three themed prototypes from one gray-box foundation, most worked on the first take, and manual fixes fell 50% compared with the previous model. It is a vendor-published customer case, not an independent controlled trial, but it demonstrates the type of integrated testing workflow Astra targets.
Why Opus 5.5 can make a richer first impression
The Claude Opus 5.5 model documentation positions it for long-running agentic coding and knowledge work. It has a 1-million-token context window, up to 128,000 output tokens, always-on adaptive thinking and medium default effort.
In the public game test, Opus did more than satisfy the minimum brief. It inferred the kinds of controls, rewards, ambience, moving objects and music that make a prototype feel like a designed game. That initiative is valuable for a one-shot demo. It can also create scope creep: an agent that adds seven systems without being asked leaves more code to test and maintain. Visual richness and production simplicity are not the same objective.
Which model fits which U.S. workflow?

- For the richest single-pass prototype: The current public case favors Opus 5.5. It is more likely to fill in mechanics, presentation and audio without waiting for detailed follow-ups.
- For repeated browser or engine testing: Astra is a strong choice. Its advantage depends on an agent environment that can actually run the game, inspect the screen and feed test results back into the model.
- For generating many ideas before choosing one: GPT-6 Sol is the practical cost tier. It did not make the best games in the comparison, but the two build turns totaled $2.89 and 37 minutes.
- For very small, high-volume experiments: Luna can generate simple candidates cheaply, but it is a weaker choice as the sole owner of a complex game.
ChatGPT or Codex subscription allowances, Claude plan limits and API billing are separate systems. In a subscription workflow, five-hour or weekly limits, interrupted runs and retries may matter more than the dollar figures above. Compare models under the same prompt, effort level, tool access and time limit on your own account.
A mixed-model pipeline is often better than a team sport
You do not have to choose one vendor for every stage. Opus can produce the design-heavy first build, while Astra performs browser QA and difficult repair. A lower-cost version is to have Opus write a detailed design and acceptance-test plan, let Sol produce several prototypes, and use Astra or Opus only on the finalists.
Human playtesting still closes the loop. Mobile input, frame pacing, save integrity, accessibility, package security and the licenses for music, fonts and images are not guaranteed because the game launches. Teams should keep internal records of generation tools, prompts, selected assets and human edits, and screen generated visuals for close resemblance to protected characters or interfaces.
Bottom line
If the question is “Which model is most likely to create the most impressive game from one prompt?”, the best direct public comparison currently favors Claude Opus 5.5. If the question is “Which agent should run the build, inspect it and keep repairing it?”, GPT-6 Astra has a compelling end-to-end case. If the goal is “How do I turn many ideas into playable candidates without losing cost control?”, GPT-6 Sol is the more realistic production tier.
The decisive test is your own brief. Track first-run success, follow-up prompts to acceptance, regressions, playtest time, token use and human repair time. That turns a viral one-shot demo into a model-selection decision you can defend.
Sources and use notice
- Moe Lueker, Claude Opus 5.5 vs GPT-6 Astra same-prompt game test
- OpenAI, GPT-6 Astra announcement
- OpenAI Developers, Building games with Astra
- OpenAI, Playco game-prototyping customer story
- Anthropic, Claude Opus 5.5 official page
- Claude Platform, Opus 5.5 model documentation
- Artificial Analysis, Opus 5.5 and GPT-6 Astra release comparison
- OpenAI, GPT-6 Sol and Luna launch and API pricing
OpenAI, GPT, Codex, Anthropic, Claude and related product names are marks of their respective owners. This independent article is not sponsored or endorsed by either company. It distinguishes official specifications, vendor customer stories, independent measurements and one creator’s same-prompt test. It does not reproduce third-party game screens, official product interfaces, X screenshots or external photography.



