AI Adoption GuideBankingService
Complaint root cause synthesis
NLP clusters complaint text and call transcripts to surface systemic product and process failures, producing a prioritized issue log for product owners.
Banking processAcquireOnboardOpenFundTransactServiceReviewClose
By Don, DoneThat’s AI coach · updated
A cluster is not a closed complaint
The usable output of complaint root-cause synthesis is a named theme with citations to real complaints and call transcripts, assigned to a named product owner. It is not a finding that the customer's case is finished. Clustering tells you where the same failure is repeating. It does not replace the case owner's duty to the person who complained, including any Financial Ombudsman Service (FOS) clock that is already running.
Treat the cluster as a quality artefact. A product owner can act on it: change a fee disclosure, fix a journey, or open a process investigation. The individual complaint stays open until that customer's outcome is resolved, remediated if needed, and recorded. If a cluster exists, you still answer the customer in front of you.
Case systems such as Salesforce and core platforms such as Temenos already hold the complaint, the product, and often the transcript pointer. Synthesis sits on top of those records. It does not invent a new case type, and it does not close the source cases when the cluster is published.
Cluster first, then refuse unlike harms in the same bucket
Start from the text customers actually used: complaint letters, web forms, chat, and call transcripts. Group records that share the same described failure, not the same product code and not the same agent wrap-up. Product and process codes are useful filters. They are a poor substitute for what the customer said.
Keep fraud, scams, and unauthorised-use language in their own clusters even when the customer also mentions an app screen or a branch visit. A sentence that includes "I did not recognise this payment" is not evidence that onboarding copy is unclear. Merging those records into a UX theme hides a different control failure and can delay the path those cases need, including dispute pre-triage automation when the text points to a transaction dispute rather than a service niggle.
Separate vulnerable customer detection signals from theme formation. Vulnerability may change how you handle the case. It is not itself a product root cause unless the customer describes a specific barrier the product created.
Run clustering as a proposal, not a verdict. A reviewer who knows the product should reject a cluster that mixes unlike harms, split a blob that actually contains two failures, and refuse a label that the cited lines do not support. If the text only says the fee appeared without warning, do not name the root cause as "incorrect APR calculation." You do not have that evidence yet.
Illustrative example: a week of retail complaints mentions "the fee." Some transcripts say the customer was not told a foreign-transaction fee would apply on a contactless tap abroad. Others say a third-party payment appeared on the account and the customer wants it reversed. A volume-led model will happily call the whole set "fees and charges UX." That label is wrong for the second group. Split the set. One theme covers overseas-fee disclosure in the card journey, with cites only from the disclosure complaints. The other set must follow dispute and fraud handling, not a product-copy backlog.
Name the theme with cites a product owner can open
A theme name is a claim. Write it as a failure a product owner can verify against the cited lines, not as a slogan. Prefer "overseas contactless fee not shown before the tap is confirmed" over "fees are confusing." The first can be checked against the journey and the quotes. The second invites a workshop and no change.
Every cluster that leaves operations should carry a short theme title that matches the cited text; a handful of complaint and transcript excerpts, with case IDs the owner can open; the products, channels, and process steps those cites actually mention; what the text does not establish, so nobody fills the gap with a guessed root cause; and a named product owner, not a queue name.
Do not auto-generate a "root cause" field that goes beyond the evidence. "Customers describe an unexpected overseas fee at the point of tap" is a synthesis. "The authorisation message omitted the FX markup because of a configuration error" is an investigation hypothesis. Put hypotheses in the owner's follow-up, not in the cluster title, until someone has checked the system.
Volume is a weak ranking on its own. A small cluster that describes unauthorised payments, or that sits on a FOS-bound complaint, outranks a large cluster about a cosmetic app label. Rank with harm, regulatory clock, and whether the same failure will keep generating new cases. Use volume as a tie-breaker among comparable quality issues, not as the only sort.
Assign a product owner without completing the customer file
Ownership of the cluster is product or process ownership of the repeating failure. Ownership of each complaint remains with the case handler until that customer is done. Publishing the issue log must not trigger a bulk close, a bulk "root cause recorded," or a bulk FOS-complete flag.
When you assign the cluster, tell the owner what "done" means for the theme: a change, a documented decision that no change is warranted, or a handover to a control owner if the cites support fraud or operational failure rather than a journey bug. None of those outcomes closes the complaint from Tuesday. That case still needs an answer, any redress, and a closure reason that reflects that customer's outcome, not the existence of a workstream. Closure coding is a separate discipline from synthesis; see closure reason classification.
Do not treat the cluster as a closed FOS case. FOS cares about the individual complaint: what happened to that customer, what you offered, and whether you met the timescales. A well-named theme in a product backlog is useful context for a subject-matter expert. It is not a determination, and it is not a file you can send instead of the case papers.
If frontline teams use tier-1 service deflection to answer repeat "how do I see my fees?" questions, keep those contacts out of the systemic-failure log unless the customer is actually complaining. Deflection reduces known how-to volume. Root-cause synthesis is for failures the bank should fix, not for questions the bank should answer once.
Failure modes that look like a clean issue log
Merging fraud into UX is the fastest way to publish a tidy theme that is wrong. If any cited complaint alleges unauthorised use, social engineering, or "I did not make this payment," do not park it under a design theme. Split it, route it, and keep the product-owner cluster limited to cites that support a product or process failure.
Ranking by volume only rewards the noisiest, easiest-to-describe theme. A quiet cluster with high harm or an ombudsman clock should sit at the top of the owner's list even when it is small.
Treating the cluster as a closed FOS case confuses a quality artefact with a customer outcome. Creating a ticket in Salesforce, an epic, or a product brief does not complete the complaint. Auto-closing source cases because they were included in a numbered RCA pack is a conduct failure dressed as efficiency.
Inventing a cause the text does not support is what over-confident labels do. If the excerpts do not say it, the cluster does not get that name. The honest output is a theme plus an explicit unknown.
Letting the issue log become a dumping ground for every dissatisfied contact dilutes ownership. Synthesis is for repeating product and process failures. One-off service misses, once classified, belong in handling quality, not in a product owner's systemic log.
The quality bar is simple to inspect. Open the cluster. You should see real customer words, a theme those words actually support, a named human owner, and every cited complaint still alive in the case system until that customer is resolved.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first