Churn-risk prediction model
Multi-signal machine learning scores accounts on usage, sentiment, and ticket trends 90 days before renewal, using tools like ChurnZero, Gainsight, or Velaris.
Sales processProspectQualifyDiscoverProposeNegotiateCloseHandoffRenew
By Don, DoneThat’s AI coach · updated
Treat the score as a hypothesis, not a fact
A churn-risk score earns its place when it is a ranked hypothesis with receipts. It fails when the team treats the number as a verdict. Ninety days before renewal, the model's job is to put a score on the account and attach the signals that produced it so a customer success manager can accept, reject, or rewrite the story.
Quality is not a red, yellow, or green badge. Quality is a scored account with cited signals. The CSM still decides whether the account is at risk, whether the conversation should change, and whether anyone should offer a commercial concession. The score is an input to that judgment, not a substitute for it.
Platforms in this class (ChurnZero, Gainsight, Velaris, and health fields you already keep in Salesforce) can compute and display that score. None of them should close the loop without a human. If a threshold alone fires a discount, a save sequence, or an executive email, you have turned a quality model into an automation that trains customers to wait for a red score.
Four signal families on one account
Four families of signal should land on the same record. One family alone produces confident, wrong scores.
Usage is the easiest to over-weight. Product events, seat activation, feature adoption, and login cadence show whether the contract is being consumed, not whether the buyer still cares. An account that looks healthy on usage can still be a flight risk if the champion is gone or if every ticket is a workaround. Empty usage is a different problem. If the product has not been implemented, or if tracking is broken, a usage-based score looks like churn when the real issue is onboarding. Pair usage with an implementation risk flagger before you trust a cold product graph.
Tickets are directional only if you keep the text, not just the count. Volume up can mean adoption. Volume down can mean the customer stopped asking. What matters is theme and trajectory: repeated severity, unresolved defects, and time-to-resolution stretching as renewal approaches. Feed ticket trends as structured cites (category, last similar tickets, open versus closed) rather than a single support-intensity number.
Sentiment is not a star rating. It is language from QBRs, emails, survey comments, and call notes that a model can classify as friction, expansion interest, or indifference. Treat it as noisy. A frustrated champion still fighting for you is not the same as a quiet economic buyer who has already switched. If you already generate a QBR pack, reuse that narrative as a cite rather than inventing a second story. See auto-generated QBR.
Champion health is the signal most scores omit until it is too late. Role changes, stalled replies, a new evaluation thread from procurement, or a named champion who no longer appears in usage logs should move the score even when seats look full. Run this alongside champion-departure monitoring so the model is not discovering a departure from a calendar invite two weeks before the date.
Build a score the CSM can open
The output should be one number plus a stack of cites the CSM can expand in order. A cite is a pointer, not a slogan. "Usage down" is not a cite. "Weekly active seats on the core workflow fell across the last two billing cycles; last product event from the named champion was on this date; three high-severity tickets in this category remain open" is a cite.
Keep the score additive and inspectable. Show a contribution from each family (usage, tickets, sentiment, champion) so the CSM can see which lever moved. Do not hide the mix behind a black-box percentile. If usage is empty, mark the usage contribution unknown, not zero. Unknown is a data-quality flag. Zero claims nobody is using the product.
Ninety days is the right window for most annual contracts because it leaves time for a human conversation before legal and procurement lock the path. Align the score run with whatever you already do for an auto-renew and anomaly detector so a sudden usage drop and a rising churn score are not two separate alerts for the same account.
Store the score on the Salesforce account (or the CSM system of record) with a timestamp, model version, and a link to the cite bundle. Health platforms in this class can compute it. The CRM should own the field seen in standup, so nobody hunts a second console to learn why an account went red.
Illustrative path: a regional operations team is ninety days from renewal. Seat usage is flat. The model still raises risk because ticket themes shifted from how-to questions to a blocked payroll workflow, the last QBR notes mention a competing evaluation, and the original champion accepted a role outside the buying group. None of those facts is a churn event. Together they are a reason to book a working session, not a reason to pre-approve a discount. The CSM opens the cites, confirms the champion change with the account team, and decides the next conversation.
Put every material score in a human queue
Scoring without a queue is theater. Every account that crosses a risk threshold, and every account whose mix of signals is internally inconsistent (healthy usage, toxic tickets), should land in a human review queue with an owner and a due date.
The queue is a review, not a save campaign. The CSM confirms or kills each cite, adds context the model cannot see (a merger, a budget freeze, an escalation already run), and chooses a next action: watch, executive alignment, product intervention, or a commercial conversation. If a save motion is warranted, hand off to a save-play recommender after the human has named the actual risk. Do not let the recommender fire because the score was red.
Set a rule that no commercial offer is generated from the score alone. A red score is permission to look, not permission to discount. Discounting on a model output trains customers to perform distress and pollutes labels, because saved accounts will include people who were never leaving.
Review timing should match remaining runway. At ninety days, a few business days is enough for a CSM to open cites. Inside thirty days, the same queue should page a manager, because the cost of a missed cite is a surprised renewal call.
Failure modes that look like progress
Score without cites. A dashboard of red accounts feels operational. CSMs will ignore it, or they will invent a story that fits the color. If you cannot click from the number to the usage window, tickets, sentiment snippets, and champion status, you do not have a quality model. You have a traffic light.
Scoring on empty usage. New logos, stalled implementations, sandbox-only tenants, and broken instrumentation all look like non-adoption. If you score those accounts as churn risks, you flood the queue and miss the accounts that are actually leaving. Gate usage on observed events and skip accounts still in not-started implementation. When usage is empty, raise a data or onboarding flag, not a churn flag.
Treating a red score as permission to discount. Finance, customers, and the next cycle's labels will all notice when risk correlates with the last offer you made. Keep commercial authority on a separate path. The score can recommend a conversation. Only a human looking at cites can recommend a price.
What the CSM does after the score lands
Open the cites in order of contribution, confirm what is still true, and write a one-line thesis a peer could challenge: usage is stable, but the buyer is new and payroll tickets are unresolved. Choose the smallest next step that tests it. Shopping means a documented product conversation. No value means implementation, not a save email. A departed champion means a new map of influence. The model does not pick among those. Combine usage, tickets, sentiment, and champion health, cite every contribution, queue a human. ChurnZero, Gainsight, Velaris, and Salesforce compute, store, and display. The CSM still decides.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first