03/06—Agentic workflow
- Role
- Design & implementation
- Context
- Metarune Labs · multi-tenant platform
- When
- In production
01Problem
AI-generated drafts were not accurate enough to ship directly, and one bad output in front of a client would damage trust. Reviewing everything by hand, though, would create a permanent backlog.
02What I built
A queue-backed pipeline where every generated item carries a confidence score. High-confidence items auto-approve against a per-tenant threshold; the rest enter a review dashboard, and every human edit is captured as an (original, edited) pair.
03Architecture
- 01Generatedraft + confidence
- 02QueuePostgreSQL-backed
- 03Thresholdper-tenant, e.g. ≥ 0.9
- 04Review UIapprove · edit · reject
- 05Feedback storeoriginal ↔ edit pairs
04Technical decisions
- 01
Confidence thresholds are a product decision
Each tenant tunes its own auto-approval threshold to its risk tolerance. Stricter and looser tenants run on the same pipeline with no code changes.
Rejected
Fully automated or fully manual. The first risks trust; the second creates a backlog for outputs that were mostly fine.
- 02
Capture feedback before you know how you'll use it
Every edit stores the delta between model output and human correction. These pairs became the most useful input for prompt iteration.
- 03
Invest in the reviewer's tool
A slow review UI makes reviewers avoid the queue. The dashboard was built to be fast and clear because the human step is the bottleneck.
05Impact
Review effort is spent only on the edge cases, not on every generated item.
A growing corpus of corrections that feeds prompt improvements.
Stack
TypeScript / Node.js / PostgreSQL