01/06—LLM applications
- Role
- Design & implementation
- Context
- Metarune Labs · multi-tenant platform
- When
- In production
01Problem
Product wanted AI summarisation. Engineering's question was how to ship it without making a slow, non-deterministic, swappable third-party API a load-bearing dependency of every request.
02What I built
A single LLM service interface that feature code calls, with provider adapters behind it, hard timeouts with deterministic fallbacks, output validation before anything reaches a user, and prompts stored as versioned config.
03Architecture
- 01Feature codecalls one interface
- 02LLM servicetimeout · retries
- 03Provider adapterOpenAI / Anthropic
- 04Validatorlength · format · rules
- 05Fallbacktruncated source text
04Technical decisions
- 01
Timeouts are a product feature
If the model has not answered in 3 seconds, the request returns a deterministic fallback instead of blocking. Users take a slightly worse result over an unbounded spinner.
Rejected
Calling the model inline: 2–5 seconds added to every request and no behaviour when the provider is slow or down.
- 02
Validate before you surface
Model output is treated as untrusted input. Length, format and content rules run before anything is shown, and failures route to the same fallback path.
- 03
Prompts are versioned config, not code
Prompts live in versioned files so they can iterate without a deploy, and a regression can be traced to a prompt version and rolled back.
Rejected
Fine-tuning our own model: operational cost and complexity that a well-prompted general model did not justify.
05Impact
Switching provider for cost reasons meant changing one implementation, not dozens of call sites.
No request path blocks on the model for more than 3 seconds.
Stack
TypeScript / Node.js / Redis / OpenAI / Anthropic