02/06—LLMOps · cost
- Role
- Design & implementation
- Context
- Metarune Labs · multi-tenant platform
- When
- In production
01Problem
After the first LLM feature shipped, API spend spiked. Some tenants were heavy users, most barely touched it, and identical requests kept hitting the provider. There was no visibility into who was spending what.
02What I built
Per-tenant cost tracking first, then Redis response caching, per-tenant monthly token budgets with graceful degradation, and routing of simple tasks to smaller, cheaper models.
03Architecture
- 01Requesttenant-scoped
- 02Cache lookupnormalised-input hash
- 03Budget checktenant token budget
- 04Model routertask complexity
- 05Usage ledgercost per tenant
04Technical decisions
- 01
Measure before optimising
Cost-per-tenant tracking went in before any optimisation. Only then did it become visible that a large share of traffic was cacheable duplicates.
Rejected
Hard per-user rate limits: they frustrate power users and ignore the real cause, duplicate requests.
- 02
Budgets as a product surface
When a tenant exhausts its budget, it receives cached or fallback responses and a notification, not a surprise bill or a hard failure. Visible usage makes tenants self-regulate.
- 03
Exact-match hashing now, embeddings later
Normalised-input hashing captured most of the win. True embedding-based semantic similarity adds complexity that was not yet earned, so it stays on the roadmap.
Rejected
Self-hosting a smaller open model: GPU cost and quality trade-offs outweighed managed APIs at this scale.
05Impact
~20%
of requests turned out to be cacheable duplicates once cost was measured per tenant.
Predictable spend per tenant, with no hard failures at the limit.
Stack
Node.js / TypeScript / Redis / PostgreSQL