Skip to content
ENTERPRISE AI

AI Customer Service

An AI transformation programme inside a bank-owned B2B platform moving over $200M a month — where every automated decision has an auditor attached to it.

ROLEAI Product Manager
PERIOD2025 — present
ORGANISATIONTejarat Shayan (Bank Tejarat)

THE PROBLEM

Transaction volume grew faster than the team reviewing it. Every growth plan on the roadmap ended at the same staffing ceiling, and nobody had separated the reviews that genuinely needed judgement from the ones that were pattern-matching.

~40%of operations headcount sat inside manual review queues

CONTEXT

The platform ran a high volume of B2B transactions on manual review queues. Operations cost scaled linearly with volume, and every growth plan hit the same ceiling.

APPROACH

We mapped the queue into decision classes, then automated only the classes where a model could be audited after the fact. LLM agents draft the decision, a rules layer holds the boundary, and a reviewer confirms anything outside confidence.

MY ROLE

I owned the roadmap, the confidence thresholds and the rollout sequence — including the decision to leave two queues manual because the error cost was not worth the saving.

HOW I WORKED

01Queue auditTwo weeks shadowing reviewers; 1,200 sampled items coded by decision type and error cost.
02Class modelSix decision classes, each with an explicit error-cost estimate signed off by risk.
03Threshold designConfidence bands set per class with the risk team, not by the model's own metric.
04Shadow runEight weeks of the model deciding silently beside humans; disagreements reviewed weekly.
05Staged rolloutOne class at a time, each with a rollback trigger defined before launch.

DECISION LOG

The calls I owned, the alternative I rejected, and the cost I accepted for each.

01
DECISION

Automate by decision class, not by queue

ALTERNATIVE REJECTEDAutomate the highest-volume queue end to end
WHYVolume and error cost are not correlated. Classing the work first let us automate 60% of items and leave the expensive 40% untouched.
COST ACCEPTEDA slower headline number in the first quarter.
02
DECISION

Rules layer holds the boundary; the model drafts inside it

ALTERNATIVE REJECTEDLet the model decide on a confidence threshold alone
WHYA threshold is a statistical promise. Auditors need a deterministic one — the rules layer is what we can defend in a review.
COST ACCEPTEDExtra engineering, and some correct model calls get blocked.
03
DECISION

Two queues stay manual, permanently

ALTERNATIVE REJECTEDAutomate everything and correct after the fact
WHYError cost in those classes exceeded the salary saved by roughly 4×. Automating them would have been a net loss dressed as progress.
COST ACCEPTEDThe programme's savings headline is smaller.
04
DECISION

Ship the audit trail before the first automated decision

ALTERNATIVE REJECTEDLog retroactively once volume justified it
WHYReconstruction after the fact is not reconstruction. Every automated decision has carried its full input snapshot since day one.
COST ACCEPTEDSix weeks before any automation went live.

WHAT MOVED

BEFOREAFTER
Items reviewed manually per day1,400560
Median review latency6.2 h11 min
Operations cost per 1,000 itemsindex 100index 90

OUTCOME

−10%operational cost across automated queues
$200M+monthly GMV running through the platform
3 queuesfully automated with audit trails intact

WHAT I PRODUCED

Decision-class taxonomyError-cost model (risk-signed)Confidence threshold specAudit trail schemaRollout & rollback planWeekly disagreement review

WHAT I'D DO DIFFERENTLY

I ran the shadow period on aggregate agreement rate, which hid that one class disagreed on a narrow but expensive slice of inputs. I would segment the shadow analysis by input type from week one — we caught it in week six, and it cost us a rollout slot.