AI Customer Service
An AI transformation programme inside a bank-owned B2B platform moving over $200M a month — where every automated decision has an auditor attached to it.
THE PROBLEM
Transaction volume grew faster than the team reviewing it. Every growth plan on the roadmap ended at the same staffing ceiling, and nobody had separated the reviews that genuinely needed judgement from the ones that were pattern-matching.
CONTEXT
The platform ran a high volume of B2B transactions on manual review queues. Operations cost scaled linearly with volume, and every growth plan hit the same ceiling.
APPROACH
We mapped the queue into decision classes, then automated only the classes where a model could be audited after the fact. LLM agents draft the decision, a rules layer holds the boundary, and a reviewer confirms anything outside confidence.
MY ROLE
I owned the roadmap, the confidence thresholds and the rollout sequence — including the decision to leave two queues manual because the error cost was not worth the saving.
HOW I WORKED
DECISION LOG
The calls I owned, the alternative I rejected, and the cost I accepted for each.
Automate by decision class, not by queue
Rules layer holds the boundary; the model drafts inside it
Two queues stay manual, permanently
Ship the audit trail before the first automated decision
WHAT MOVED
OUTCOME
WHAT I PRODUCED
WHAT I'D DO DIFFERENTLY
I ran the shadow period on aggregate agreement rate, which hid that one class disagreed on a narrow but expensive slice of inputs. I would segment the shadow analysis by input type from week one — we caught it in week six, and it cost us a rollout slot.