AI Adoption GuideInsuranceRenew
Churn propensity scoring at renewal
ML scores the renewal book for flight risk and flags high-value at-risk accounts for retention outreach.
Insurance processQuoteUnderwriteBindIssueBillServiceRenewClaim
By Don, DoneThat’s AI coach · updated
Load the renewal book with vintage attached
Churn scoring at renewal starts with a dated book, not a model run. Load every account that will reach a renewal decision in the window you are staffing, and attach vintage to payment history, claims history, and the terms that will be offered. If vintage is missing, do not emit a score.
The book is the cohort whose renewal effective dates fall in that window, including accounts already in a notice period. A warehouse dump of in-force lives is the wrong grain. You need identifiers, proposed terms for this cycle, and the history that belongs to those accounts with timestamps intact.
Vintage is the calendar behind the facts. A late payment last month is not the same signal as a late payment years ago under a different billing plan. A recent claim is not the same signal as one closed before the last two renewals. Terms that will change at this renewal (deductible, limit, endorsement, billing frequency) belong on the same row as the history, because you are scoring flight risk under those terms, not under last year's contract.
Policy administration, rating, and customer systems (Guidewire, Duck Creek, Earnix, Salesforce) usually already store fragments of this record. Treat them as a class of source systems, not as a ranked set of scoring products. If one extract has claim counts and another has billed premium, join on policy and account keys and keep the dates. Drop rows you cannot date rather than filling dates from neighboring accounts.
A score with no vintage is the first failure mode. It looks operational because every account has a number. It cannot tell a producer why this account is high risk now, and it cannot be audited when a large account asks what changed. Refuse to publish that run. Score only the subset where dates exist.
Cite payment history, claims vintage, terms, and grouping
A usable score is a cited score. For each account that receives a number, the output must name the payment and claims history vintage that went into it, the renewal terms assumed, and the grouping rule that set the comparison set. If any of those three cannot be stated, leave the cell empty.
Payment history vintage is the window of invoices, returns, and delinquencies you actually scored, with start and end dates. Claims history vintage is the occurrence or report window, not a lifetime count with no period. Renewal terms are the package you are asking the customer to accept: premium direction, deductible and limit changes, coverage grants or restrictions, and billing changes. The grouping rule is how you decided which accounts are alike enough to share a peer set: line, class or occupancy, tenure band, size band, or geography. State that rule next to the score.
Citing is what the retention lead reads before a call. A usable brief names the facts: payments slipped in the last two cycles, a claim sits inside the current vintage, and proposed terms change the deductible relative to similar accounts. A decimal with no cites is not.
Do not invent a churn percent. Propensity is an ordered risk within the group you defined, given the vintage and terms you loaded. It is not a forecast of how much of the book will leave. If leadership wants a book-level view, report coverage (how many accounts scored versus blank), the distribution of bands inside each grouping, and the premium in the high-risk band. Leave observed non-renewal and rewrite rates to the measurement that waits for actual outcomes.
Related work on lapse propensity scoring answers a different question: whether the policy is likely to lapse for non-pay, versus whether the customer will shop or refuse the renewal. Keep those pipelines separate so payment-lapse cites do not get copied onto a shopping-risk score.
Grouping exists so you do not compare a small personal auto account to a layered commercial package. Publish the rule. If you change it mid-cycle, rescore and mark the version. A score that cannot name its group is as empty as a score that cannot name its vintage.
Leave thin history blank
Empty stays empty when history is thin. New business on a first renewal, accounts with incomplete billing feeds, boarded books whose claims never received occurrence dates, and midterm acquisitions with stub payment files should not receive a blended score from the rest of the group.
Thin means you do not have enough dated payment or claims observations inside the vintage you declared, or you cannot attach proposed terms, or the account does not fit any grouping rule you will stand behind. The correct output is a blank propensity and a reason code: insufficient vintage, missing terms, or out of group. Do not impute a mid-pack score. Do not copy the group's average. Do not let a platform default fill the cell.
Blank is operationally useful. It tells the desk this account was not scored, so the play is judgment, not a false sense that risk is average. Retention still calls. A blank high-value account is a call-list item, not a skip. The score orders scarce outreach where the cites are strong enough to talk about. It does not remove unscored names from human work.
If a later extract backfills vintage, score those accounts in a dated rerun. Do not silently overwrite blanks in the original file without a run id. People will already have called, or already have skipped, from the first list.
Retention still calls
High-value at-risk accounts go to outreach. The score flags. A person decides whether to call, write, visit, or wait, and what to say. Treating the score as if the policy were already non-renewed is the second failure mode. It shows up as suppressed notices, blocked rewrite paths, or producers told not to spend time because the model already non-renewed them.
Keep non-renewal, cancellation, and underwriting decline on their own authorities. Propensity can inform a conversation about personalized renewal communication or a referral into a pricing discussion alongside dynamic renewal pricing. It does not flip those switches. If underwriting has a separate non-renewal rule, that rule must cite its own facts. A high propensity band is not that rule.
Outreach should use the cites, not the number. Talk about the payment pattern in the vintage, the open or recent claim as the customer experienced it, and the terms on the renewal. If there is a genuine coverage hole the customer might fill rather than leave, hand that to the coverage gap cross-sell at renewal motion as a separate offer with its own eligibility. Do not bundle a flight-risk label into a cross-sell pitch.
Staff the list by value at risk and by whether a human can actually reach the account in time. A mid-pack score on a small account can wait. A blank or high band on a large account does not wait for a prettier model. Do not auto-non-renew from the propensity file.
One account on the renewal desk
A mid-market manufacturers package sits 45 days from renewal. The book load shows two returned payments inside the current two-cycle vintage, a property claim that closed inside that same vintage, and proposed terms that raise the wind deductible while holding limit. The grouping rule places it with other light-manufacturing packages in the same tenure and size band. The score is high relative to that group. The cites print those three facts. Retention calls. The conversation is about the deductible change, the claim experience, and whether billing can be stabilized. Nobody quotes a model percentage. Nobody treats the row as already non-renewed.
If that same account had arrived on an extract with claim counts and no occurrence dates, you would not score it. A number without vintage is not a retention brief. If payment and claims dates were missing but the account were still large, the cell would stay blank and retention would still call. Blank is not "safe."
After the cycle, compare who left, who rewrote, and who paid against the bands and the blanks, with the grouping rule held constant. Until those outcomes exist, do not invent a book-level churn percent from the scores.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first