Skip to main content
DoneThat

AI Adoption GuideInsuranceBill

Lapse propensity scoring

ML scores policyholders for non-payment or lapse risk at billing cycle start so retention teams can intervene early.

Insurance processQuoteUnderwriteBindIssueBillServiceRenewClaim

By Don, DoneThat’s AI coach · updated

Run the score when the cycle opens

Lapse propensity scoring ranks billed lives at cycle start so retention can intervene before the installment fails or the policy walks. Load payment history and current billing terms first. Score second. Call third. Do not wait for a returned draft to decide who is at risk.

The work sits on the billing calendar. Carriers already run policy admin, billing, CRM, and pricing stacks in this class: Guidewire, Duck Creek, Salesforce, Earnix. Extract payment history and billing terms from the systems that hold them. Score in a separate, cited job. Do not treat any of those products as the lapse decision.

Freeze the extract at a stated cycle-open timestamp. Do not rescore mid-cycle because a card was updated, unless you also rewrite vintage. A moving score with a stale cite is the same failure as a score with no vintage.

A returned payment is a different queue. Hand that work to failed payment recovery sequencing. Cycle-start propensity is the earlier cut. It asks who is likely to miss or lapse on this draft, given history and terms, while the policy is still in force.

If you only score after the NSF, retention inherits a collection problem. The useful window was the days before the bank was hit.

Cite vintage, billing terms, and the grouping rule

Every score that ships must carry three cites on the same record: payment-history vintage, billing terms, and the grouping rule.

Vintage is the window of installment events the model actually used, with an as-of date and a lookback. It is not a label such as "full history." Terms are the live billing contract: mode, method, remaining installments, and the grace or lapse clock billing already computes. The grouping rule is the cohort you compared this life against: product, jurisdiction, payment method, term length, or the slice you truly used. Write the rule in words a supervisor can read.

Load history and terms in one extract. Join on policy and billing account. Score only rows where all three cites can be written back. If a cite cannot be produced, do not emit a score.

A score with no vintage is the first failure mode. Retention sees a high rank and treats it as current. Months later nobody can say whether the model used last year's book, a converted book with missing drafts, or a test file. Suppress that number. An unscored row is honest. An unexplained rank is not.

Do not invent a lapse percent to make the score look like a forecast. Propensity is a rank inside a defined group at a defined time. It is not a published lapse rate for the product, the state, or the company. If finance wants a lapse percent, point them at the actuarial study. Do not back a rate out of a quantile.

Empty stays empty when history is thin

New business, endorsements that reset the billing account, converted books with broken installment archives, and payment-method changes with no drafts on the new method all produce thin cells. Leave those cells blank.

Blank means the extract ran and the model declined to score. Retention still works the account the way they work any billed life they cannot rank: they call, they send the notice billing already allows, or they wait for the next successful draft.

Here is one path through the rule, not a measured result. A six-pay personal auto policy bound six weeks ago has one successful EFT and no prior drafts on this billing account. Grouping would drop it into a tiny, unstable cell. Vintage would be a handful of days. Do not build a score from a credit file, a household rank on another product, or a book average. That row stays empty. A second policy in the same household, on the same method, with eighteen months of drafts and two returns last spring, can be scored. Put the cites on that row: vintage through last spring, monthly EFT with four remaining installments, grouped with other monthly EFT auto. Retention gets the second policy on the callback list. The first policy does not inherit the sibling's rank.

Do not fill blanks with a default band such as "medium." A default band is a fake cite. Someone will treat it as a decision.

A score is not a lapse

The score does not cancel, lapse, or non-renew. It orders a human queue. Retention still calls.

Name a retention owner on each scored row. If the owner is a pool, name the pool. An ownerless high score sits in a report and never becomes a call.

Send the output to a work object a person owns: a callback, a billing-review task, or a supervisor list. The caller confirms the method still works, offers a mode or due-date change billing already supports, or records that the household intends to keep the policy. If nobody reaches the household, the existing lapse clock in billing continues. The model does not speed it up.

Treating the score as a lapse is the second failure mode. Operations maps "high" to a batch lapse or a withheld renewal. That turns a ranking error into a coverage error. High propensity means talk to them before this draft, not that they are already gone. If you need a lapse, use the contractual rules on the billing term after missed payment and required notice. Those rules are the contract. A quantile is not.

Calls that start as "why did you draft me" or "can I skip this month" belong with billing inquiry voice agent handling, then a person when the answer is about keeping the policy. Do not let that path auto-lapse from a propensity flag.

Keep recovery, inquiry, and renewal work on their own calendars

Cycle-start propensity is one billing-stage rank. Do not fold it into renewal churn work or into recovery after a fail.

Renewal ranking asks whether the household will rewrite, not whether this installment will pay. Use churn propensity scoring at renewal on that calendar. Language and timing for people who stay belong with personalized renewal communication. One shared risk flag that billing, retention, and product all read differently will be argued in the same spreadsheet and trusted in none.

Keep the loop short. Extract history and terms when the cycle opens. Score only rows that can cite vintage, terms, and grouping. Write blanks where history is thin. Push scored rows to retention as calls, not as lapses. After a miss, use recovery sequencing. At term end, use the renewal models and letters.

If a vendor screen shows a dashboard lapse percent next to the propensity chart, leave that percent out of the operating report. You ranked a group. You did not measure a rate. Report queue size, contact attempts, and whether billing later recorded a miss or a lapse under existing rules. That is what happened. A model-implied lapse rate is the third failure mode: a number invented to make the score feel finished.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first