Closure reason classification
NLP classifies free-text and call transcript content at account closure into structured reason codes for product and CX analysis.
Banking processAcquireOnboardOpenFundTransactServiceReviewClose
By Don, DoneThat’s AI coach · updated
Code what the customer said, not a neater attrition story
A closure reason you can trust is a code from a controlled list and a cite to the customer's words. Product and CX analysis use the code. Reviewers use the cite. If you cannot show the words, you do not have a reason yet.
When someone asks to close because of too many letters, the code belongs with excess contact or communications, and the cite quotes that phrase. "Moved to competitor" is a different claim. It is useful on a dashboard and wrong on that record. The model did not hear a competitor name. It filled a popular bucket.
Classification starts after the customer has already asked to leave. Save work sits earlier, in pre-closure churn interception, while there is still a window to change the outcome. At close you are labelling the exit, not reopening the relationship.
The operational closure still lives in the systems that already book the event, channel, and note. Salesforce and Temenos, as a class, are where that record usually sits. The classifier reads the text those systems already captured. It does not become a second closure process and it does not invent a motive the note never contained.
Classify free text and transcripts against a closed list
Pull every source that actually states a reason. Forms, email, chat, and secure messages give you written spans. Voice gives you the transcript. Do not treat the agent's wrap-up code as the customer's language. Wrap-up is typed under time pressure and often mirrors last month's league table.
Clean only what you must. Remove signatures, footers, and canned agent scripts. Leave the customer's sentences in order. If a web form and a call both exist, classify the pair and store which span supported the code.
Score against a list product and CX have already signed. Useful families include fees, rate, service failure, communications volume, life event, product no longer needed, and competitor switch, plus an explicit uncoded value. New codes are a taxonomy change. They are not a model label you promote because the text was colourful.
Every coded result needs a cite: snippet, source type, and a pointer (message id, call id, timestamp). Store that cite on the closure record so a reviewer can open the line without a CRM archaeology exercise.
Confidence is relative to the list. High confidence means one code is clearly better and the cite is specific. Two plausible codes, or language like "just done with it", stay uncoded. Completeness in the report is not quality.
Here is the pattern in one pass. A current-account customer writes: "Please close this account. You send too many letters." On the call they say: "I just want it shut, the mail is constant." A forced competitor code files this as switched-away because attrition reporting wants that column. The cite-backed result is communications volume, with "You send too many letters" and "the mail is constant" on the record. Product then looks at statement and marketing mail. It does not spin up a rate-match journey the customer never described.
Leave the record uncoded when you are not sure
Uncoded is a valid result. Sarcasm, mixed motives, a form that only says "close my account", a dropped transcript, or a reason that fits two families equally, should not be forced into a box.
Put a threshold in front of the model. Below it, write uncoded, keep the raw text, and sample for human review if you have capacity. Do not auto-promote the second-best code. Do not let a warehouse job recode uncoded to "other competitor" because a column cannot be null.
Human review confirms or rejects the cite. If the reviewer cannot find supporting words, the record stays uncoded. If they change the code, they change the cite in the same action. A recode without a new cite is still an invention.
Expect pressure to fill the competitor bucket. Retention reporting often treats competitor loss as the only attrition that "counts." That is the shortest path from "too many letters" to "moved to competitor." Audit competitor-coded closures against their cites on a regular sample. If the words are about mail, fees, or a house move, the code is wrong even when the score looked high.
Keep closure reason apart from complaint root cause
A complaint describes what failed in a journey. A closure reason describes why this account is ending now. They can mention the same product. They are not interchangeable fields.
A card-decline complaint can close, and the customer can still leave months later because they no longer need the account. Copying "card decline" onto the closure because it was the last case mixes two clocks. Send the case history to complaint root cause synthesis. Keep the closure classifier on the words used at the exit.
When the customer ties them in one utterance ("I'm closing because you still have not fixed the card"), the closure reason can be unresolved service failure, and that sentence is the cite. The complaint record still owns root-cause taxonomy, owners, and fix-forward work. Do not overwrite the complaint's root cause with the closure code. Do not overwrite the closure code with the last complaint category.
Product needs that split. Closure-reason mix shows which exit themes are growing. Complaint root cause shows which operational failures to fix. Merging the two makes both reports look complete and both harder to act on.
Never hold the closure for a missing code
An uncoded or delayed reason must not block the exit. The customer asked to leave. Identity, notice, cooling-off, and residual balance belong to regulatory exit compliance check. NLP output is not a compliance gate.
If the account cannot close until a reason code is present, someone will pick a code to clear the queue. That is how competitor and "other" volumes become theatre. Keep the closure path independent of classification. Attach the code when it is ready, including after the account is already closed.
The code also does not create a retain hold. Communications volume is not a licence to stall while marketing turns off letters. A competitor code is not a licence to refuse the close until a specialist calls. Policy may allow a parallel save attempt. The classifier does not add a new stop.
Use the mix for product and CX, including the uncoded slice
Cite-backed codes are a signal. Uncoded volume is a signal too. Hide neither.
Read coded share by product, channel, and period, and keep uncoded as its own slice. Rising communications-volume codes with cites about mail point at contact policy. Rising uncoded share points at capture: no reason field on the form, no transcript, or a model that is either noisy or overly shy. Those are different queues.
Do not manage teams on percent coded. That target produces forced codes. If you need an operating metric, sample cite validity: did the words support the code.
Downstream consumers must treat missing codes as unknown. winback timing prediction should not treat uncoded as a competitor-loss prior. Invented competitor codes in the closure file become invented timing in the winback file.
The loop stays small. Maintain the list. Classify free text and transcripts. Require a cite or leave uncoded. Close the account either way. Product and CX read the coded mix and the uncoded remainder as two different kinds of work.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first