Use a library that targets WebGPU – Transformers.js, WebLLM, or ONNX Runtime Web’s WebGPU execution provider – which downloads a model, runs it client-side via the GPU, running via WASM on the CPU where WebGPU isn’t available.
At a glance
| Fact | Value | Source |
|---|---|---|
| WebGPU needs Chromium v113+ on Windows | Chromium 113+ | onnxruntime.ai |
| Firefox/Safari WebGPU support status | behind flag / preview | onnxruntime.ai |
| Default execution without WebGPU | runs on CPU via WASM | huggingface.co |
Load a JavaScript library that has a WebGPU backend, let it detect GPU support, and have it fetch the model weights at runtime instead of shipping them in your build. The practical choices are Transformers.js (uses ONNX Runtime under the hood, functionally similar to the Python transformers library), WebLLM (built on the MLC inference engine, purpose-built for LLMs), and ONNX Runtime Web’s WebGPU execution provider for ONNX models like Whisper or Segment Anything. All three run entirely client-side with no server call after the model is cached.
If you build your web game with an AI coding agent, Playgama MCP lets agents like Claude Code, Cursor, or VS Code manage developer cabinet tasks such as uploading builds and publishing a test sandbox.
What decides whether it works in the browser
- Check WebGPU support first. It works out of the box in current Chrome/Edge on Windows, macOS, Android and ChromeOS; on Firefox it’s behind a flag and on Safari it’s in Technology Preview. Transformers.js and ONNX Runtime default to WASM on CPU when WebGPU is unavailable.
- Run inference off the main thread. WebLLM defaults to a Web Worker so model loading and inference don’t freeze the UI – important for a game loop.
- Weights are large and cached, not bundled. Models are fetched from a hub at runtime and cached by the browser; first load is the slow one, so trigger it during a loading screen, not mid-gameplay.
- Test in the actual target browser, especially on non-Chromium browsers, since WebGPU is still experimental there and a model that runs fine in WASM can still fail on WebGPU.
If the game runs inside an iframe on a portal, confirm the portal’s frame allows WebGPU and doesn’t restrict worker threads before you ship.
Sources
- Running models on WebGPU – Transformers.js
- Local Inference | WebLLM.io
- Using the WebGPU Execution Provider | ONNX Runtime
- Get started with ONNX Runtime Web
- sauravpanda/BrowserAI
- Playgama Bridge SDK docs
Related questions
What happens if WebGPU is not supported?
Transformers.js and similar libraries run on CPU via WASM when WebGPU is unavailable, so the model still runs, just slower.
Can I run Whisper or image models in-browser too, not just LLMs?
Yes – ONNX Runtime Web’s WebGPU provider has demos for Whisper tiny.en (speech) and Segment Anything (image segmentation), both running client-side.
Where do the model weights come from if there’s no server?
Libraries fetch weights from a model hub at runtime and the browser caches them, so only the first load is slow.
Last updated: 30 September 2026