Send less to the model: truncate or summarize old conversation turns, cache stable context (system prompt, NPC lore), use cheaper models for routine lines, and batch non-urgent requests. Most cost in chat-based NPCs comes from re-sending full history on every call, not from the reply itself.
At a glance
| Fact | Value | Source |
|---|---|---|
| Main cost driver in chat loops | resent history tokens | ai.google.dev |
| Caching cuts repeated-prompt cost | cache hits cost less than standard input | docs.anthropic.com |
| Async non-urgent requests discount | Batch API discounts non-urgent requests | ai.google.dev |
The biggest cost in an LLM-powered NPC or dialogue system is not the single reply – it’s resending the whole conversation history as input tokens on every call. Cut that by: capping how many past turns you send (keep the last N exchanges, drop or summarize the rest), caching the parts that don’t change (system prompt, character bio, world lore), and routing routine lines to a cheaper/smaller model while reserving a stronger model for key moments.
If you build with an AI coding agent, Playgama MCP lets the agent create the game, upload builds, and publish a sandbox test link directly.
What actually reduces token spend
- Trim history server-side. Keep a rolling window of recent turns and summarize older ones into a short state note instead of replaying full text.
- Use prompt caching for content that repeats across calls (system instructions, character sheets); Anthropic reports a cache hit costs a fraction of the standard input price, so caching pays off once the same prefix is reused. See Claude pricing docs.
- Batch what isn’t real-time – background NPC chatter, lore generation – through async batch APIs, which run at a discount versus synchronous calls (Gemini API optimization).
- Count tokens before sending if your model provider offers a token-counting endpoint, so you can cap context length before it grows.
- Keep the API key off the client. Route LLM calls through your own backend; never ship a raw key in a WebGL or browser build.
Check your model provider’s current pricing and caching rules directly – they change.
Sources
- Gemini API optimization and inference
- Pricing – Claude Platform Docs
- Context caching | Gemini API
- API overview – Claude Platform Docs
- Billing | Gemini API
- Playgama wiki: MCP Server
Related questions
Should I store full chat history client-side in the browser?
No – keep it server-side so you control what gets resent to the model and can summarize or trim it before the next call, reducing input tokens.
Does caching help if every player has a different conversation?
Yes for the shared parts: system prompt and character setup repeat across players and sessions, so caching those portions still cuts cost even if the dialogue itself varies.
Last updated: 01 October 2026