Skip to content
GENAI

Tax Assistant Chatbot

A Persian-language tax assistant for small businesses, built on the assumption that a wrong answer about tax is worse than no answer at all.

ROLEProduct lead
PERIOD2025
SURFACEWeb + Telegram

THE PROBLEM

Small businesses ask the same forty questions about VAT filing, and each one waits in a queue behind genuinely complex cases. The accountants were the bottleneck for work that did not need an accountant.

11 daysmedian wait for a straightforward tax question to reach an accountant

CONTEXT

Persian-language tax guidance is scattered across circulars that change every fiscal year. Generic models answer confidently and wrongly, which in a tax context is worse than silence.

APPROACH

Retrieval over a curated corpus of current circulars, with the model forbidden from answering when retrieval returns nothing above threshold. Every answer cites the circular it came from.

MY ROLE

I defined the refusal policy, the citation requirement and the escalation path to a human accountant, and ran the evaluation set with two practising accountants.

HOW I WORKED

01Question miningClustered 2,400 historical support questions into forty recurring intents.
02Corpus curationOnly current-year circulars indexed; superseded documents excluded rather than down-weighted.
03Refusal policyBelow-threshold retrieval returns an escalation, never a generated answer.
04Expert evaluationTwo accountants scored 300 answers for correctness and citation accuracy before launch.

DECISION LOG

The calls I owned, the alternative I rejected, and the cost I accepted for each.

01
DECISION

Refuse rather than answer when retrieval is weak

ALTERNATIVE REJECTEDAnswer with a confidence caveat
WHYA hedged wrong answer about tax still gets acted on. Refusal routes the user to a human in one step instead of three.
COST ACCEPTEDRoughly 18% of sessions end in escalation rather than resolution.
02
DECISION

Cite the circular on every answer

ALTERNATIVE REJECTEDCite only when the user asks
WHYThe citation is what makes the answer checkable by an accountant later. It also disciplined the corpus — uncitable answers exposed gaps.
COST ACCEPTEDLonger, denser replies that test worse on first impression.
03
DECISION

Persian-only at launch

ALTERNATIVE REJECTEDBilingual from day one
WHYThe evaluation set only existed in Persian, and shipping an unevaluated English path would have doubled the surface we could not vouch for.
COST ACCEPTEDNo English-speaking users at launch.

WHAT MOVED

BEFOREAFTER
Median time to a usable answer11 daysunder 1 min
Questions reaching an accountant100%18%
Citation-verified answers—100%

OUTCOME

82%of routine questions resolved without a human
300answers expert-scored before launch
0uncited answers shipped

WHAT I PRODUCED

Intent taxonomy (40 clusters)Refusal policy specCitation contractExpert evaluation rubricEscalation flow

WHAT I'D DO DIFFERENTLY

I under-invested in the escalation experience. Users who hit a refusal got a correct routing but a cold one — writing that handoff copy properly should have been part of the launch, not a follow-up.