Move the model load and inference calls into a Web Worker so the main thread keeps running requestAnimationFrame uninterrupted; send prompts and results with postMessage, and only apply the AI output on the next frame once the worker replies.
At a glance
| Fact | Value | Source |
|---|---|---|
| Workers run scripts off the main thread | background thread | developer.mozilla.org |
| Worker communication method | postMessage + onmessage | developer.mozilla.org |
| Worker message data handling | copied, not shared | developer.mozilla.org |
Put the model and every inference call inside a Web Worker, and keep your game loop (update/render, usually driven by requestAnimationFrame) on the main thread untouched. The main thread posts a message with the prompt or game state; the worker runs inference and posts back a result; your loop reads that result on whatever frame it arrives and applies it, instead of waiting for it. This is exactly what Web Workers are for: “laborious processing can be performed in a separate thread, allowing the main (usually the UI) thread to run without being blocked/slowed down”.
If you build with AI coding agents, Playgama MCP lets the agent upload builds and publish sandboxes straight from the developer cabinet.
What actually decides this in a browser
- Data passed with postMessage is copied, not shared, so don’t post large buffers every frame – post compact JSON (a short prompt, a small state object) and let the worker hold the model weights.
- In-browser LLM runtimes such as WebLLM are built with this pattern in mind: they run on WebGPU and, per the MLC docs, provide “built-in support for web workers to separate heavy computation from the UI flow”. A WebGPU-capable browser is required.
- Never call `await` on inference inside your render/update tick – queue the request, keep rendering, and consume the worker’s reply asynchronously on a later frame.
- On weaker mobile browsers, a second worker thread still competes for CPU/GPU with rendering, so budget frame time and consider throttling how often you request inference.
Sources
- Web Workers API – MDN
- WebLLM Javascript SDK
- BrowserAI – GitHub
- MDN: Game development
- Playgama Bridge SDK – getting started
Related questions
Can I share the model weights between the main thread and a Web Worker without copying?
Structured data sent via postMessage is copied by default; check MDN’s Web Workers API docs for transferable objects if you need to move large buffers without duplicating them.
Does WebLLM need WebGPU to run inside a worker?
Yes – per the WebLLM docs, a WebGPU-compatible browser is needed to run WebLLM-powered applications, whether the engine runs on the main thread or inside a worker.
Last updated: 30 September 2026