05/06—Distributed systems
- Role
- Software Engineer · frontend team lead
- Context
- Metarune Labs
- When
- 2023 — now
01Problem
A real-time chat platform serving 100K+ users had to stay reliable under bursty load, while one small team onboarded multiple client products onto the same codebase.
02What I built
Full-stack features for the chat platform, event-driven serverless workflows on AWS Lambda and SQS, and shared-schema multi-tenancy on PostgreSQL — with the frontend team (7+ engineers) held to consistent review and testing standards.
03Architecture
- 01Clientweb · React Native
- 02API layertenant context injected
- 03SQSbuffering under load
- 04Lambdaevent-driven workers
- 05PostgreSQL + Redistenant_id scoped
04Technical decisions
- 01
Row-level tenancy over schema-per-tenant
A tenant_id on every tenant-owned row, enforced by middleware from the authenticated session. Simple to reason about, with a central tenants table for flags, limits and branding.
Rejected
Schema- or database-per-tenant: N migrations for N tenants and operational overhead the team could not afford yet.
- 02
Queues absorb the spikes
Work that does not need to complete in the request path moves through SQS to Lambda workers, so load spikes degrade latency gracefully instead of failing outright.
- 03
Cache only what you can invalidate
A small set of entities are cached in Redis with tenant-prefixed keys and invalidate-on-write. Live data is never cached.
05Impact
100K+
users served.
$10M+
revenue impact.
7+
engineers led on the frontend.
Stack
AWS Lambda / SQS / TypeScript / Node.js / PostgreSQL / Redis / React Native / Playwright