Business FAQs

How to architect LLM powered NPC dialogue in a browser game?

0
(0)

Never call the LLM API directly from client code – proxy it through your own backend to hide the API key, cap and summarize conversation history per NPC to control token cost, and keep generation calls async so they don’t block Canvas/WebGL rendering.

At a glance

Fact Value Source
API key must never sit in client JS server proxy required
Conversation memory fades without full history resend or summarize context ar5iv.labs.arxiv.org
In-browser inference needs WebGPU support WebLLM + WebGPU browser llm.mlc.ai

Put the LLM call behind your own backend, never in the browser bundle. A browser game’s JS is fully visible to players, so any API key shipped client-side will be extracted and abused within hours. Your server holds the key, forwards the player’s message plus the NPC’s context to the model, and returns only the text (or audio) the client needs to render. This also lets you rate-limit, cache repeated prompts, and swap providers without a client update.

If you build or iterate on your browser game using an AI coding agent like Codex, Claude Code, Cursor, or VS Code, the Playgama MCP server connects your agent directly to the Playgama developer cabinet to upload builds and publish sandboxes.

What decides the architecture for a browser game specifically

  • Context management: NPC memory fades after enough turns unless you resend or summarize the full history to the LLM each call – decide per NPC whether to keep full history, a rolling window, or a vector-store summary.
  • Async, off the render loop: fire the request, keep Canvas/WebGL rendering, resolve the promise into a dialogue UI state – never await synchronously inside a frame update.
  • Client engine binding: Unity WebGL calls browser JS via a .jslib plugin; Godot 4 uses JavaScriptBridge; Phaser/PixiJS call fetch() directly since they already run as JS.
  • In-browser inference (WebLLM, WebGPU) avoids server calls entirely but needs a WebGPU-capable browser and adds load time and GPU memory cost – usually not worth it for dialogue-only NPCs.
  • Content safety: constrain NPC outputs to the game’s themes; an LLM playing an NPC role carries reputational risk if it drifts into toxic language.

Test the full loop – key hidden, context capped, UI non-blocking – before wiring more than one NPC.

Sources

Can I run the LLM entirely in the browser without a server?

Yes, with WebLLM: it runs inference in-browser via WebGPU with an OpenAI-compatible API, but it needs a WebGPU-capable browser and adds significant load time and memory use.

How do I stop NPC dialogue costs from scaling with player count?

Cap or summarize conversation history per NPC instead of resending full transcripts every turn, and cache or reuse responses for common dialogue branches.

Does the same dialogue system work if my game is iframed into a portal?

Yes, as long as calls go to your own backend and not a hardcoded client key; the iframe origin doesn’t change how the API proxy works.

Last updated: 30 September 2026


How useful was this post?

Click on a star to rate it!

Average rating 0 / 5. Vote count: 0

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

Your email address will not be published. Required fields are marked *

Games categories