ANALYSIS / Audio-driven scenes · Research case

EchoForge turns spatial audio into explorable 3D scenes in Unity

EchoForge combines sound-event recognition, directional cues, an explicit scene graph and procedural rules to turn a spatial recording into an explorable Unity environment. It generates a plausible interpretation of a soundscape rather than reconstructing exact geometry.

Published Event
A bright procedural forest scene in Unity contains a pond, trees and an on-screen marker for a detected bird vocalization.
An audio-driven scene generated by EchoForge.
Who should pay attention

For technical artists, sound designers and level designers exploring spatial audio as an input for procedural worldbuilding, XR visualization or accessible interfaces.

On September 11, 2026, 80 Level published a detailed account of EchoForge from its two researchers. The research prototype starts with a spatial recording and produces an explorable Unity environment without asking the user to write a text prompt.

The EchoForge paper, by ESGI students Anamaria Buliga and Kamil Tebani, appeared at SIGGRAPH 2026 in July and won first place in the undergraduate category of the SIGGRAPH Student Research Competition. The recent development is the public explanation of its method and limits, rather than a newly launched commercial product.

Four spatial-audio waveforms lead to scene-graph data and then to birds placed in a right-side source zone in Unity.
Audio, a scene graph and a bird zone to the right.

Hear the recording before deciding what belongs in the scene

The system first analyzes when sounds are active and their approximate directions. In parallel, YAMNet recognizes events such as birds, ducks, insects, wind, rustling vegetation and water. EchoForge does not turn every classifier label directly into an object; it groups related detections into broader semantic families such as forest, water, wind, bird vocalization and waterfowl.

It then combines semantic labels with directional evidence. A sound that repeatedly arrives from a stable direction can become a localized source zone. Diffuse sounds such as wind are more useful as environmental context. Weak evidence stays ambiguous instead of being forced into a specific object.

The scene graph separates model evidence from generation rules

EchoForge's intermediate representation is an explicit audio scene graph. It records the inferred environment, candidate sources, confidence, spatial hints, temporal evidence and consistency checks. A visual decision can therefore be traced either to something detected in the recording or to a stated procedural rule, instead of leaving only an opaque generated result.

The graph is converted into Unity-readable layout instructions for terrain, vegetation clusters, water regions, source zones, ambience and representative objects. In the public farm example, bird, farm-bird and quacking labels combine with a stable right-side cue to produce a bird zone on the right of the scene.

The EchoForge pipeline runs spatial audio through probing, semantic tagging and semantic-spatial fusion, then builds an audio scene graph, procedural layout and Unity demo.
The pipeline from spatial audio to a Unity scene.

A plausible interpretation, not a reconstruction

Audio alone cannot reveal the exact shape of a tree, the size of a pond or a measured distance between a bird and the recorder. EchoForge uses acoustic evidence for coarse direction, while procedural rules decide scale and visual representation. The same recording may support more than one plausible scene, so the result should not be treated as a metric reconstruction of the place that was recorded.

The researchers also expose the limits of the current vocabulary. The system can recognize an insect or cricket-like sound in the dusk example, but when the generation stage has no matching asset and rule, the scene graph leaves it unresolved and places no concrete object. Recognizing a sound and knowing how to represent it in a world are separate problems.

The public material shows three qualitative cases: farm birds lead to a clear placement, a windy forest produces a useful but non-unique interpretation, and the cricket case reveals a missing rule. These examples explain the pipeline, but they do not establish generation speed, coverage across environments or production stability.

For game teams, it is a prototyping input

This approach can turn an ambience recording into a starting point for early worldbuilding. A sound designer could provide a recording with clear directional changes, while technical artists and level designers inspect how the scene graph interprets it, which detections remain unstable and how procedural rules turn those decisions into a rough walkable layout.

The same representation could help XR teams see where spatial sounds are being located, or translate sound type, direction and localized-versus-ambient status into visual cues for deaf and hard-of-hearing users. A production game would still need designers to decide terrain, paths, scale, mechanics and final assets. EchoForge supplies an inspectable scene hypothesis, not a finished level.

Sources (3)
  1. 80 Level · EchoForge Uses Spatial Sound to Build 3D Worlds in Unity
  2. ACM · EchoForge: Interpretable Audio-Conditioned Procedural Scene Generation from Spatial Sound
  3. SIGGRAPH 2026 · Student Research Competition results

No sponsorship or affiliate links in this article.