ANALYSIS / AI-assisted production · Practice report

Beyond the AI game demo: combat, levels and testing become the real work

Antti Kananen's latest production log follows an AI-assisted game project beyond asset generation and into combat, traversable levels, UI, audio and regression testing. The practical lesson is that agents can expand execution, but playable quality still depends on explicit acceptance criteria, repeatable evidence and human judgment.

Published Event
An aerial view of a low-poly playable world with a walled city, surrounding settlements, roads, hills and a river.
A playable world assembled and reviewed through the agent workflow.
Who should pay attention

For independent developers and technical teams moving AI-assisted coding or content generation into a playable prototype, with practical lessons on task breakdown, regression tests, model routing and cost limits.

On September 14, 2026, Games Alchemy author Antti Kananen published a new production report about moving an agent-driven game workflow beyond assets and demonstrations into combat, levels, UI, audio and testing. Once those parts had to cooperate inside one playable system, the difficult production problems became much easier to see.

The measurements and results come from Kananen's own project and have not been independently reproduced, so the report is most useful as a production case. It shows how much implementation agents can already take on, while also exposing the remaining distance between a build that runs and a game that is worth playing.

Combat turns “feature complete” into many acceptance checks

Putting a character and weapon on screen is relatively easy. A working attack also depends on input, collision, animation, camera behavior, effects, audio and damage rules. Kananen encountered sword animations that looked correct while missing the enemy, a bow that reached its ready pose but ignored real mouse input, and enemies that faced the player while pointing their weapons sideways.

His response was to stop treating combat as one task. The workflow instead checks contact, direction, timing windows, input paths, damage, feedback and rejection cases separately. One test uses real mouse input rather than calling an internal fire function; fixed cases also exercise melee directions and archer alignment repeatedly.

Two low-poly characters face each other in an isolated combat test scene while a fire state is visible on the sword wielder.
A combat QA scene for hit and state feedback.

Kananen reports that one accepted build eventually passed 16 packaged regression suites. Those checks can establish whether a blade connected, a wall blocked an attack or an enemy entered the right state. They cannot decide whether an attack has weight, a bow release feels satisfying or the twentieth encounter with the same enemy becomes boring. Mechanical paths can be automated; game feel still needs a person to play and direct it.

Levels need to be checked along player routes

The same gap appeared in level work. Agents connected roads, walls, terrain, districts, bridges, interiors and underground areas into a traversable scene. One version covered a playable area more than a kilometre across, according to the report, with route audits running through thousands of samples. Controller tests then walked entrances, stairs and longer paths.

Those checks did not remove visual problems. Reversed walls, awkward pivots, grass inside floors, compressed ceilings and backwards furniture still required human review. Some routes passed structural tests while failing from the player's viewpoint; later walks uncovered stair corners and a missing bypass.

A third-person character stands inside an underground hall containing stone walls, a chandelier and a large model table.
An underground level slice after route and camera corrections.

The resulting process is a loop of planning, generation, traversal, evidence, human rejection and correction. Coordinate connectivity can answer whether a route exists. It cannot decide whether a space reads clearly, the composition works or local details support the setting.

UI and audio fail in use, not in isolation

An agent can assemble the information structure, controls and basic responsiveness of a settings screen quickly, but Kananen found that early interfaces tended to feel like complete development tools rather than finished game screens. The model still needed clear direction about what players should see first, what could remain hidden, how the screen should adapt and where it needed visual emphasis.

A dark game settings screen shows display, resolution, frame rate, brightness, UI scale and audio controls.
An early game settings interface study.

Audio followed a similar pattern. Generated music supplied a useful emotional direction quickly. A sound effect that worked in isolation could be too long during repeated attacks, tiring in a loop, or wrong once loudness, spatial position and voice limits applied. One audio pass produced 63 compact mono variants and used automated checks for file structure, peaks, loops and routing. Headphones inside the running game remained the acceptance gate.

Multi-model work adds a production-management layer

Kananen routes bounded coding, repair, data and test work to lower-cost models, while reserving more capable models for visual judgment, large spatial problems and final review. Every handoff needs a goal, exact files, accepted references, constraints, tests, known defects and a definition of done. Without that packet, multiple models repeatedly spend context learning the same project.

The arrangement creates a new budget problem as well. A difficult task may return several uneven results, and repeated rolls can consume context and allowance without improving the outcome. Kananen's controls include setting a budget before a task, limiting weak retries, recording usage by task type and keeping a cheaper or deterministic fallback path.

For teams testing AI-assisted game development, the report suggests four concrete starting points: split large features into observable acceptance checks; make automated tests exercise real input and player routes; preserve fixed test scenes and accepted references; and set stop conditions for retries and model upgrades. AI can move a project to a playable build faster, but quality from that point depends increasingly on tests, decisions and production rules that accumulate over time.

Sources (1)
  1. Games Alchemy · AI Game Development Got Real

No sponsorship or affiliate links in this article.