Inworld AI

Connect speech, an LLM and game actions to prototype one bounded realtime NPC.

Evidence statusPublic-evidence reviewMedium confidenceLast checked Sep 1, 2026By MakeGameWithAI

PRODUCT OVERVIEW

What Inworld AI is—and how it works.

Inworld AI is a runtime stack for realtime interactive agents. It can receive a player's microphone audio, turn it into text, send the conversation and character context to a chosen language model, then stream a generated voice back. Memory, knowledge, turn-taking, back-channel behavior and tool calls can sit in the same observable pipeline.

Unity and Unreal templates provide a shorter path to a talking character. You describe the role, voice, motivations and constraints, connect the Runtime, and start from a text or push-to-talk demo. The game still owns the character model, animation, navigation, quest facts, permissions, rewards and saved state. Inworld is not a button that generates a complete NPC or game system.

This makes a first voice NPC approachable even if you cannot build the entire STT–LLM–TTS stack yourself. Start with one guide, witness, companion or optional hint character in a small scene. Programming and engine judgment become more important when dialogue can trigger actions or serve many players, but they are not prerequisites for exploring the first bounded interaction.

HOW YOU USE IT

The indexed Unity route starts at Unity 2022.3 LTS and recommends 6000.0.41 or newer; the Unreal route lists 5.4–5.7 and a C++ project. Both begin with an API key and sample character. Current engine documentation redirected to login when checked, so confirm package versions and release notes inside the account before changing a production project.

Run speech in and speech out

Stream microphone audio through speech recognition, a selected LLM and TTS in one session. Turn detection and streaming reduce plumbing, but real game latency still depends on model, network, context and client behavior.

O05O06

Shape memory, context and voice

Character context, knowledge, conversation history, long-term memory and voice direction can shape an answer. Longer context can also increase token use, and generated statements are not authoritative game facts.

O05O06O08

Connect tools to game logic

The Runtime can request registered tools without breaking the audio session. Treat each request as a proposal: your game validates the action, arguments, permissions and resulting state change.

O02O05O06U04

Start through Unity or Unreal

The indexed quickstarts describe a Unity package with text/push-to-talk demos and an Unreal plugin with Character/Chat templates. Shipping still needs version compatibility, a credential backend, logging, reconnect behavior and target-device checks.

O03O04O15O16
REVIEW SCOPE

A first bounded voice NPC, a small Unity or Unreal prototype and a transparent public-rate budget. This page does not establish current Chinese quality, end-to-end latency, long-session stability, production concurrency, a measured bill or an accepted-conversation success rate.

Beginner friendlinessSome basics needed
Examples reduce voice-pipeline setup, but a first engine NPC still involves plugins, scene configuration and keys; web auditioning is simpler.
Agent integrationEasy to integrate
Official Realtime API documentation covers authentication and streaming sessions for agent-orchestrated voice workflows, not complete NPC generation.
Scope, integration requirements and sourcesChecked

Beginner scope: One constrained voice NPC using an engine example or WebSocket template.

Agent scope: Create a voice/text session and receive output events.

Confidence: learning Medium · Agent Medium

Access
Configure a Runtime API key and protect it behind a backend instead of shipping a long-lived key in the client.
Billing
Check voice, model, session usage and capacity; offline audio pricing is not the cost of a live conversation.
Getting results
Open a WebSocket, send input, receive text/audio events and close the session.
Limitations
The project still implements actions and game rules; engine SDK and individual service maturity must be checked separately.

These are independent editorial judgments, not an overall score or a claim of hands-on agent integration testing.

KEY FINDINGS

What the evidence supports—and what it does not.

01
Official factHigh confidence

One runtime reduces plumbing, not game design

Inworld combines speech recognition, a selected LLM, streaming voice, memory and tools in one observable path. Engine templates can get a character speaking sooner. Character assets, navigation, quests, permission checks, state transitions and failure UX remain your game's work.

O02O03O04O05O15
02
User reportsMedium confidence

Open conversation can feel immersive and still fail the task

Historical developers, players, journalists and a 48-participant study found novelty, emotional engagement or freedom in bounded scenes. The same evidence includes recognition errors, off-topic or incorrect answers, interruptions and connection failures. Positive overall experience does not mean reliable quest completion.

U01U05U06U08U09U10
03
Editorial inferenceLow confidence

The current Agent Runtime lacks enough independent game evidence

Most detailed game reports describe the 2023–2024 Character Engine. Recent records include one unresolved Runtime debugging question and two non-game TTS uses. That is enough to justify a small prototype, not a current success rate, Chinese-quality claim or production-maturity label.

O02O14U10U11U12U13
04
Editorial inferenceMedium confidence

Generated intent should request an action, not become authority

The Cygnus team added its own intent layer for concrete game behavior, while media and research cases still saw misunderstandings or invented information. Register narrow tools, validate their arguments and keep consequential state changes deterministic.

O05O06U04U08U10
05
Official factHigh confidence

Cost is speech in, speech out and model tokens together

STT is billed by player-audio time, TTS by generated characters and the chosen LLM by input/output tokens. Plans also change rates and request capacity. A low nominal session cost does not include subscription cash, retries, backend hosting, growing context or the share of conversations players accept.

O07O08

EDITORIAL VERDICT

Our take

Medium confidence

Inworld is worth considering when voice conversation is part of the game experience and you want one runtime instead of wiring speech recognition, a model and voice output from scratch. It gives beginners a real place to start: one character with one bounded job in a small scene. Keep rewards, inventory, combat, quest completion and saved progress behind deterministic game logic. The current Agent Runtime is newer than most independent game evidence, so prototype before making it a core production dependency.

Conclusion scopeOwner-approved, medium-confidence public-evidence review. Fourteen independent records across ten platforms support bounded workflow and failure-mode conclusions. Current Runtime performance remains low-confidence and no first-party product test was performed.

BETTER FIT

Worth auditioning when

  • Creators, including beginners, who want to prototype one voice-first character without building the whole speech-and-model stack.
  • Guides, witnesses, companions, optional quest hints and VR interactions with a clear, bounded role.
  • Games that let dialogue stay flexible while game code validates actions, facts, rewards and saved state.
  • Small audiences whose voice usage, model tokens, peak sessions and backend requirements can be measured before expansion.

POORER FIT

Do not depend on it yet when

  • Projects requiring a completely offline character or no continuing service and backend dependency.
  • Critical dialogue and story branches that must always be identical without a scripted fallback.
  • Letting generated replies directly award items, alter saves, execute purchases or decide combat outcomes.
  • Assuming vendor latency claims, old demos or nominal per-session math prove current production readiness.

WORKFLOW FIT

Where it fits in a production workflow.

01

Define one bounded character job

Choose a guide, witness, companion or hint character. Write what it knows, what it must not claim and which one low-risk tool it may request. Prepare role questions, out-of-scope questions and interruption cases before integrating a larger scene.

GuardrailUse fictional, non-sensitive material. Do not let generated dialogue directly change rewards, purchases or saved progress.
02

Connect the smallest engine sample

Use the current Unity text/push-to-talk demo or Unreal Chat template in a disposable project copy. Add subtitles, one no-side-effect action and per-stage logs before custom animation, memory or multiple characters.

GuardrailKeep the Base64 credential on a backend and issue short-lived client tokens. Confirm package versions inside the current documentation account.
03

Measure a limited playtest before launch

Invite a small audience and record recognition, time to first audio, full response time, interruptions, tool acceptance, disconnects and credits. Use those results to estimate accepted-conversation cost and peak capacity.

GuardrailProvide a scripted fallback and an exit path when dialogue is slow or unavailable. These are recommended checks, not tests completed by this site.

RECOMMENDED WORKFLOW

Use this tool inside a complete production workflow.

Start from a testable brief, then move through prototyping, assets, audio, testing, and release with an explicit handoff and human check at every step.

Open the complete method

SHORTEST RESPONSIBLE PATH

Start with one NPC, one safe action and twenty test prompts

This is a low-risk starting workflow synthesized from the evidence.

  1. 01

    Pick one bounded character role and write a short source-of-truth sheet: identity, known facts, forbidden claims, tone and one allowed tool request.

  2. 02

    Create twenty prompts covering ordinary dialogue, role boundaries, repeated questions, interruptions, silence, Chinese or accent needs and an unauthorized action request.

  3. 03

    Connect the current Unity or Unreal sample in a disposable project copy. Keep credentials on a backend, add subtitles and expose only one action with no irreversible side effect.

  4. 04

    Run the prompts on the target network and device. Record transcription, first-audio time, full response, rule compliance, tool validation, reconnect behavior and actual credits.

  5. 05

    Let a few people try the same bounded scene. Only widen access after defining an accepted conversation and calculating real cost, peak sessions, privacy notice and a scripted fallback.

PRICING & RIGHTS

What you pay for—and what one conversation can consume

Pricing and terms last checked: Sep 1, 2026

  • Self-serve monthly plans are On-Demand $0, Creator $25, Builder $100, Developer $300 and Growth $1,500. Paid plans include credits equal to the monthly fee; Enterprise is custom.
  • TTS-2 per million characters is $25 / $20 / $17.50 / $15 / $12.50 across those tiers. TTS-2 Flash is $15 / $10 / $9 / $8 / $7. The page uses the vendor's budgeting assumption of about 1,000 characters per audio minute.
  • STT is $0.15 per audio hour on On-Demand and $0.10 on the paid self-serve tiers. A complete voice NPC also consumes the selected LLM's input and output tokens.
  • Request limits are 5 / 10 / 50 / 150 / 500, with vendor-estimated concurrent sessions of 20 / 40 / 200 / 600 / 2,000. These are capacity labels, not monthly output or a latency SLA.
  • On-Demand advertises up to 70 TTS minutes or 400 STT minutes free. The public copy did not make reset cadence, simultaneous eligibility or full-pipeline deduction clear enough to subtract it from the durable budget below.
  • Subscription credits can roll over for up to three months while the same or a higher plan stays active. Additional purchased credits expire after twelve months; remaining credits are lost when a downgrade or cancellation takes effect.
  • The model directory showed different counting language from the pricing page. This review uses one listed model and exact rates rather than claiming a fixed number of available models.
COST TO OUTPUT

One explicit ten-minute voice-NPC scenario

Illustrative public-rate arithmetic, not measured billing: one ten-minute session contains five minutes of player speech and five minutes of NPC speech (5,000 characters), ten turns, 12,000 input tokens and 1,500 output tokens using Gemini 2.5 Flash at $0.30 / $2.50 per million tokens. Exclude free allowances, retries, memory/RAG growth, moderation, tools, backend, tax and labor.

O07O08
The actual Creator subscription$25 / month

You pay $25 and receive $25 in monthly credits at Creator rates, with ten concurrent requests and an estimated 40 user sessions. Unused value follows the rollover rules; it is not a pay-only-for-what-you-used plan.

One dollar with TTS-2 Flash$1 → 15.22 sessions nominal

One session consumes about $0.06568: $0.00833 STT + $0.05000 TTS + $0.00735 LLM. This is a credit division under the stated assumptions; Inworld does not sell a separate one-dollar session pack.

One dollar with TTS-2$1 → 8.64 sessions nominal

One session consumes about $0.11568: the same STT and LLM assumptions, with $0.10000 of TTS. Voice choice changes cost before quality, retry or acceptance is considered.

Use all $25 Creator credits380 Flash or 216 TTS-2 sessions

These are whole-session theoretical ceilings under the same ten-minute workload, with no other account use. They are not 380 or 216 accepted conversations and do not override request or session capacity.

Iterate, then invite ten people30 sessions → $1.97–$3.47 nominal

Assume 20 creator sessions plus ten visitor sessions. Creator cash paid remains $25 that month. The twenty-session iteration count is a planning assumption, not a first-pass success rate or a complete NPC development cost.

One hundred ten-minute visits$9.49–$14.49 on demand

Creator rates consume about $6.57–$11.57, but cash paid is still $25; On-Demand uses the displayed rates and excludes the unclear free allowance. This is 100 total visits, not 100 simultaneous sessions.

Nominal usage is not accepted output. The real denominator is conversations that meet your character, task, voice and player-experience criteria. Actual cost must include failed and repeated sessions, growing context, external services, backend hosting, engineering and operations. Peak capacity is separate: Creator lists an estimated 40 concurrent user sessions, while 100 simultaneous sessions would require a higher published tier or a confirmed custom arrangement and load test.

Commercial-use condition

The pricing page offers a commercial-license path, and the general terms assign Inworld's rights in compliant outputs to the customer. The same public materials also mention internal-business use, deletion of Services, Models and Outputs after termination, and an older SDK agreement that ties Project Files to the Inworld Platform. Confirm current Runtime redistribution, post-termination game content, third-party model/voice rights and any cached audio before release. This review is not legal advice.

Player speech and chat can enter a data path involving audio, possible voice biometrics, research/training and third-party models. Do not assume an ordinary self-serve workspace has zero retention. Confirm player notice or consent, region and age requirements, DPA/ZDR settings, deletion, model-provider handling and the checkout terms that apply to the actual account.

PRODUCTION RISKS

Resolve these before production use.

The product is between generations and status labels differ

The public product line has moved from Character Engine toward Agent Runtime and general realtime services. WebSocket may be GA while the overall Realtime API, Router or STT carry preview labels. Confirm the exact engine package, changelog and supported target instead of treating one GA label as maturity for the whole stack.

O02O03O04O14O16

Recognition, turn-taking and disconnects need visible fallbacks

Historical evidence repeatedly includes recognition problems, waits, speaking over characters and sessions that stop responding. Current incidence is unknown. Add subtitles, interruption handling, timeout copy, retry limits, scripted dialogue and an exit path before testing immersion.

U05U06U08U09U10U13

Conversation must not own consequential game state

A persuasive character can still be wrong. Keep quest truth, inventory, rewards, combat, payments and saves in deterministic systems. Tool calls need narrow permissions, schema validation, idempotency and failure rollback.

O05O06U04U08U10

Credentials and player data require a backend plan

Do not ship a reusable Base64 key in the client. Use a backend to keep credentials and issue short-lived tokens, then document what text, audio, metadata and logs leave the game, how long they remain and which model providers receive them.

O10O11O12O15O16U01U14

Distribution and termination terms need written confirmation

The commercial-license label, general service terms and older SDK agreement do not answer every current Runtime distribution question in the same words. Before release, confirm plugin redistribution, cached outputs, post-termination use, attribution and third-party voices or models for the exact plan.

O07O09O10O13

NOT VERIFIED

Claims this page does not make

RESEARCH METHOD

Research-reviewed from public evidence

Fourteen independent records across ten platforms and five evidence types: developer/community cases, player/review records, two independent media hands-ons, an academic study and a technical repository. Steam contributes 3/14 (21.4%) and Reddit 2/14 (14.3%); no platform reaches half the set. Searches also covered Epic Developer Community, GameDev.net, G2 and Hacker News, where results were absent, too weak, incentivized, derivative or inaccessible. Vendor stories and sponsored media were excluded from independent counts, while repeated demos, one game and the study's earlier thesis were clustered or deduplicated. The recurring historical themes are useful for workflow and failure-mode guidance, but recent exact-version game evidence is insufficient; current Runtime performance remains low-confidence.

Research windowPrimarily September 1, 2025–September 1, 2026. Earlier Character Engine projects and demos are retained only as historical design evidence; their failure rates do not describe the current Agent Runtime.
OFFICIAL SOURCES16
  1. O01
    Official
    Inworld AI

    Current product entry point for realtime speech, model routing and agent infrastructure. Vendor performance and ranking claims are not treated as independent results.

    Checked Sep 1, 2026
  2. O02
    Official
    Introducing the Inworld AI Runtime for Unreal

    Runtime graph, templates, observability and the Unreal release. Unity was described as early access at publication, not as a current status guarantee.

    Checked Sep 1, 2026
  3. O03
    Official
    Unity Agent Runtime quickstart

    Search-indexed Unity requirements, package installation, API key and text/push-to-talk demos. Direct access redirected to login when checked.

    Checked Sep 1, 2026
  4. O04
    Official
    Unreal Agent Runtime quickstart

    Search-indexed Unreal 5.4–5.7, C++ project, plugin, Character/Chat template and trace requirements. Direct access redirected to login when checked.

    Checked Sep 1, 2026
  5. O05
    Official
    Current Inworld documentation index

    Public index for current TTS, STT, Router, Realtime API, memory, back-channel and tool documentation.

    Checked Sep 1, 2026
  6. O06
    Official
    Inworld Realtime API

    Single-session speech pipeline, model choice, tool calls and transport status. Published latency numbers remain vendor claims, not game measurements by this site.

    Checked Sep 1, 2026
  7. O07
    Official
    Inworld pricing

    Current self-serve monthly prices, credits, TTS/STT rates, request limits, estimated sessions, rollover and commercial-license label.

    Checked Sep 1, 2026
  8. O08
    Official
    Inworld model catalog

    Current model list and input/output token rates. The page budget uses Gemini 2.5 Flash as a transparent example, not a quality recommendation.

    Checked Sep 1, 2026
  9. O09
    Official
    Inworld Terms of Service

    Service license, inputs and outputs, material use, training exceptions, restrictions, fees and termination requirements.

    Checked Sep 1, 2026
  10. O10
    Official
    Inworld service-specific terms

    Additional voice-input, voice-output, user voice model, content-rights and prohibited-health-data terms.

    Checked Sep 1, 2026
  11. O11
    Official
    Inworld Privacy Policy

    Chat, audio, possible voice biometrics, research/training, retention and cross-border processing. Enterprise processor data has a separate scope.

    Checked Sep 1, 2026
  12. O12
    Official
    Inworld security

    SOC 2 Type II, encryption and Trust Center entry points. Zero data retention is not presented as a default on every self-serve plan.

    Checked Sep 1, 2026
  13. O13
    Official
    Inworld SDK License Agreement

    Older public SDK agreement covering Project Files, distribution and platform restrictions. Its applicability to the current Agent Runtime needs confirmation.

    Checked Sep 1, 2026
  14. O14
    Official
    Inworld product availability status

    Official status labels distinguish TTS GA, Router/STT research preview, Realtime API research preview and WebSocket GA.

    Checked Sep 1, 2026
  15. O15
    Official
    Inworld authentication guidance

    Search-indexed guidance not to expose Base64 credentials in clients and to issue short-lived JWTs from a backend.

    Checked Sep 1, 2026
  16. O16
    Official
    Official multimodal companion sample

    Archived official sample documenting a session backend, WebSocket, retry, timeout and conservative concurrency handling. Archival does not prove the Runtime ended.

    Checked Sep 1, 2026
INDEPENDENT EVIDENCE RECORDS14
  1. U01
    Community case
    Adding Inworld AI characters to an existing game

    An independent JavaScript narrative-game integration needed a Node token service. The author found basic connection manageable and characters engaging, while warning that open chat can distract from the game.

    Checked Sep 1, 2026
  2. U02
    Community case
    Unity discussion — character-system limitations

    An open-world story developer liked the early character setup but questioned whether its limits could support a story that evolves with player interaction.

    Checked Sep 1, 2026
  3. U03
    Community case
    Unity3D discussion — connecting AI-driven NPCs

    A developer connected both Inworld and Convai but described the process as challenging and highlighted the STT–LLM–TTS stack and recurring connection costs.

    Checked Sep 1, 2026
  4. U04
    Community case
    Cygnus real-time AI NPC showcase

    The game team combined Inworld speech with its own intent layer for gathering, combat, skills and events. It explicitly separated generated conversation from executable NPC behavior.

    Checked Sep 1, 2026
  5. U05
    Review platform
    Origins player discussion — engaging until dialogue stopped

    One player found the official Origins demo compelling, then reported that all NPCs stopped responding and a restart lost progress. This is an older demo case.

    Checked Sep 1, 2026
  6. U06
    Review platform
    Origins player discussion — repeated response failure

    A second Origins player reported NPC responses stopping after roughly ten minutes and recovering only after restart. It is a reproduction clue, not a current Runtime test.

    Checked Sep 1, 2026
  7. U07
    Review platform
    Cygnus player review

    A 29.2-hour player disliked the game's AI NPC integration and noted that it could be disabled. The review also criticizes other systems, so not all dissatisfaction can be assigned to Inworld.

    Checked Sep 1, 2026
  8. U08
    Independent test
    Axios hands-on with Origins

    A journalist found one effective on-path exchange, while also encountering multi-second waits, characters speaking over each other and mixed off-topic responses. The demo was too rough to prove replacement of written dialogue.

    Checked Sep 1, 2026
  9. U09
    Independent test
    4Gamer GDC hands-on

    A Japanese reporter compared an earlier accent failure with a newer live demo that still looped or misunderstood topics but felt more continuous and emotionally engaging.

    Checked Sep 1, 2026
  10. U10
    Academic case
    Player experience with LLM-powered NPCs in VR

    A 48-participant VR study published in 2026 used a 2023 Inworld build. Overall experience was positive, while 77% reported recognition issues and notable minorities reported unhelpful or incorrect answers.

    Checked Sep 1, 2026
  11. U11
    Review platform
    Product Hunt — Lore Machine uses Inworld voice

    A product developer reported choosing Inworld voice for Korean narrative users. This is a current, low-weight TTS clue, not evidence for the game Runtime.

    Checked Sep 1, 2026
  12. U12
    Review platform
    Product Hunt — Toyo uses Inworld voice

    A developer used Inworld voice in a calling agent and described it as more conversational. The non-game case cannot establish NPC engine quality.

    Checked Sep 1, 2026
  13. U13
    Community case
    Inworld Community — debugging pipeline latency

    A current Runtime user could not identify which pipeline stage caused delay. Staff pointed to per-node debug logs, but the thread records no final measured resolution.

    Checked Sep 1, 2026
  14. U14
    Technical repository
    Decentraland Inworld SDK7 scene

    An independent open-source scene includes a remote server, proxy/WebSocket, custom chat UI and environment endpoints. It shows integration glue, not output quality, and uses an older API path.

    Checked Sep 1, 2026

DISCLOSURE

How this research was supported

Owner Review passed on September 1, 2026, accepting this bounded public-evidence review and its medium-confidence conclusion. This work introduced no affiliate link, sponsorship, vendor account, free credits, equipment, interview or technical support, and no Inworld SDK or model was tested. The owner's prior use and commercial relationships remain unstated; approval does not resolve the listed Runtime, pricing, distribution or data unknowns.

Research review and Owner Review are complete.

This public research review will be revisited when pricing, terms, product versions, or material new evidence changes.