Skip to content
Pasindu Lanka
All case studies

02/06—LLMOps · cost

Role
Design & implementation
Context
Metarune Labs · multi-tenant platform
When
In production

01Problem

After the first LLM feature shipped, API spend spiked. Some tenants were heavy users, most barely touched it, and identical requests kept hitting the provider. There was no visibility into who was spending what.

02What I built

Per-tenant cost tracking first, then Redis response caching, per-tenant monthly token budgets with graceful degradation, and routing of simple tasks to smaller, cheaper models.

03Architecture

  1. 01Requesttenant-scoped
  2. 02Cache lookupnormalised-input hash
  3. 03Budget checktenant token budget
  4. 04Model routertask complexity
  5. 05Usage ledgercost per tenant

04Technical decisions

  1. 01

    Measure before optimising

    Cost-per-tenant tracking went in before any optimisation. Only then did it become visible that a large share of traffic was cacheable duplicates.

    Rejected

    Hard per-user rate limits: they frustrate power users and ignore the real cause, duplicate requests.

  2. 02

    Budgets as a product surface

    When a tenant exhausts its budget, it receives cached or fallback responses and a notification, not a surprise bill or a hard failure. Visible usage makes tenants self-regulate.

  3. 03

    Exact-match hashing now, embeddings later

    Normalised-input hashing captured most of the win. True embedding-based semantic similarity adds complexity that was not yet earned, so it stays on the roadmap.

    Rejected

    Self-hosting a smaller open model: GPU cost and quality trade-offs outweighed managed APIs at this scale.

05Impact

  • ~20%

    of requests turned out to be cacheable duplicates once cost was measured per tenant.

  • Predictable spend per tenant, with no hard failures at the limit.

Stack

Node.js / TypeScript / Redis / PostgreSQL