Skip to content
Pasindu Lanka
All case studies

01/06—LLM applications

Role
Design & implementation
Context
Metarune Labs · multi-tenant platform
When
In production

01Problem

Product wanted AI summarisation. Engineering's question was how to ship it without making a slow, non-deterministic, swappable third-party API a load-bearing dependency of every request.

02What I built

A single LLM service interface that feature code calls, with provider adapters behind it, hard timeouts with deterministic fallbacks, output validation before anything reaches a user, and prompts stored as versioned config.

03Architecture

  1. 01Feature codecalls one interface
  2. 02LLM servicetimeout · retries
  3. 03Provider adapterOpenAI / Anthropic
  4. 04Validatorlength · format · rules
  5. 05Fallbacktruncated source text

04Technical decisions

  1. 01

    Timeouts are a product feature

    If the model has not answered in 3 seconds, the request returns a deterministic fallback instead of blocking. Users take a slightly worse result over an unbounded spinner.

    Rejected

    Calling the model inline: 2–5 seconds added to every request and no behaviour when the provider is slow or down.

  2. 02

    Validate before you surface

    Model output is treated as untrusted input. Length, format and content rules run before anything is shown, and failures route to the same fallback path.

  3. 03

    Prompts are versioned config, not code

    Prompts live in versioned files so they can iterate without a deploy, and a regression can be traced to a prompt version and rolled back.

    Rejected

    Fine-tuning our own model: operational cost and complexity that a well-prompted general model did not justify.

05Impact

  • Switching provider for cost reasons meant changing one implementation, not dozens of call sites.

  • No request path blocks on the model for more than 3 seconds.

Stack

TypeScript / Node.js / Redis / OpenAI / Anthropic