The article describes the “Gauntlet Loop,” a prompting method for getting AI agents to improve complex outputs through repeated, independent evaluation. Its motivating example is a Call of Duty-style game that the author says Claude Code with Claude Opus 5 produced from one prompt: the agent worked for many hours, spawned subagents, wrote roughly 55,000 lines of code, and generated the game’s textures, meshes, animations, and sounds in code. After skeptics questioned the result, the prompt and code were published, and other users reportedly reproduced working games with adapted versions of the approach. The method begins with an actual agentic harness, such as Claude Code or Codex, that can inspect files, run code, render results, use tools, and spawn subagents. The user supplies a goal rather than a detailed implementation plan, allowing the lead agent to choose the architecture and work breakdown. Crucially, the agent also receives a concrete quality bar, such as real Call of Duty screenshots, a test suite, a latency target, or strong reference writing. The lead agent divides the task into independently judgeable pieces, then assigns each piece a builder and a separate critic with fresh context. The critic compares the real output against the reference, identifies the largest gap, and sends the work back for another round instead of allowing the builder to grade itself. The loop continues until the user decides the result is ready, rather than stopping after a preset number of rounds. The article also recommends a live progress page for long runs and describes an optional final agent that smooths inconsistencies among separately improved components. It presents the method as applicable to games, websites, product design, marketing, writing, research, and backend engineering, and includes a meta-prompt that can help users select a quality bar and generate a short prompt for Claude Code or Codex.
