Customer cohort LTV segmentation
Embeddings cluster customers by behavior to feed differentiated retention and revenue assumptions.
Finance processPlanBudgetInvoiceCollectPayCloseReportAudit
By Don, DoneThat’s AI coach · updated
What the segment is allowed to change
The output that belongs in a plan is a labeled set of billed-customer segments, each citing the behavior fields that defined it, plus draft retention and revenue assumptions that FP&A can accept, edit, or reject. A cluster is not a lifetime value, a price, or a forecast. It is a grouping of paying customers who look similar on fields you named in advance.
Similarity does not produce a rate. The plan still needs human-owned churn, expansion, contraction, and mix assumptions. Those assumptions may differ by segment when the segment is dense enough and the field cite is something a planning partner would put in a driver tree. They must not include a model-invented lifetime value. Lifetime value folds tenure, margin, discount rate, and a horizon. The embedding did not choose a horizon.
Keep systems in their roles. Salesforce holds commercial history and can store a segment label on the account. Workday holds cost and headcount context if you later need contribution views; it is not a clustering source. Anaplan and Pigment hold the versioned driver tree. Move a segment identifier and a field cite into that tree. Leave the assumption cells for FP&A. Asking any of those systems to "compute LTV from an embedding" is how a research grouping becomes a fake valuation.
If a cohort is too thin, the matching assumption cell stays empty. Empty is a quality outcome. Filling it with a company-wide blend, a neighbor cluster's rate, or a generated lifetime value is how four noisy logos become next year's revenue story.
Freeze billed customers before you cluster
Cluster only customers who have already been billed. Trials, prospects, unpaid seats, partner comps, and internal tenants change membership from week to week. A plan version needs a membership that can be reconstructed.
Freeze a billed set at a stated date: customer identifier, first paid invoice date, current booked revenue, product family, and every behavior field you will cluster on. Persist the freeze date next to the segment catalog. When someone asks why a logo is missing, the answer is the freeze, not a model defect.
Do not add unbilled rows to "get more data." Extra rows are not extra signal for a plan that describes paying customers. A four-account "expansion" cluster that includes two unpaid pilots will look decisive and will be wrong.
Walk through one working pattern, not a scored case study. A revenue-ops lead freezes every account with at least one paid invoice in the last twelve closed months. They drop trials, comps, and internal tenants. They cluster on four named fields: months since first paid invoice, count of expansion invoices, trailing support-ticket volume, and seat-utilization band. They keep list price, discount percentage, and any precomputed LTV column out of the feature set. After clustering, one group is long-tenured with repeated expansions, low tickets, and high utilization. Another is recently billed, with no expansions, high tickets, and low utilization. A third group has four logos that sit between those poles. That third group is not a plan cut. Its churn and expansion cells stay empty. The two denser groups receive draft rates labeled as proposed from the cluster and citing those four fields. The FP&A partner decides whether the drafts enter the driver tree.
Cluster on named fields and cite them
Embeddings can group customers. The quality bar is that every published segment cites the behavior fields used. If you cannot name the columns, you cannot defend a differentiated assumption.
Write the cite in the segment catalog: exact field names, freeze date, and whether each field is raw, clipped, or binned. "Behavior embedding" is not a cite. "Expansion-invoice count, trailing ticket volume, utilization band, tenure in months since first paid invoice" is a cite.
Choose fields FP&A already treats as drivers of retention or revenue: tenure, expansion events, usage intensity, support load, product mix, remaining contract term. Do not cluster on fields that are really credit, collections, or pricing decisions. Current discount, open dispute flags, and dunning stage belong with collections and dispute-likelihood scoring, not with plan segmentation. If two clusters differ only on a field you would never put in a driver tree, do not create separate plan rates.
Re-run clustering only on a new freeze. Do not let daily embedding drift relabel accounts inside a locked plan version. Label drift inside a version looks like performance and is actually a membership change.
Thin cells stay empty
A cluster with a handful of customers is a watchlist, not a planning dimension. Planning from n=4 is the failure mode this page exists to prevent. Four logos can all churn, all expand, or all get acquired in the same quarter. None of those paths is a rate you can take into the annual build.
Set a minimum cell size before the model is opened. Below that floor, the segment may live in an appendix as an account list. It does not appear as a plan cut. Do not dump thin cells into an "other" bucket and then apply a special rate to "other." Either the remainder is large enough to be a real segment with its own field cite, or it inherits the company-wide draft and is labeled inherited, not clustered.
Empty stays empty. That rule bans interpolated lifetime values, nearest-neighbor rates from a denser cluster, and "directional" plus-or-minus notes that someone will later paste into an assumption cell. If leadership wants to see the four logos, send the list. Do not send a cohort rate.
Drafts enter the plan; FP&A still owns the numbers
The handoff is a draft. Each dense segment can carry proposed retention, expansion, and contraction rates, with the field cite attached. FP&A maps those drafts onto the driver tree the same way they treat driver-tree assumption suggestions: as suggestions, versioned, with an owner who can reject them.
Do not output a lifetime-value number. Inventing LTV from cluster membership turns a similarity group into a valuation. There is no honest discount rate or horizon in the embedding. If someone needs a payback or a CAC recovery view, they build it in the planning model from rates they own, not from a score attached to a cluster id.
Do not treat a cluster as a pricing decision even when the cell is dense. A large group with low utilization and high tickets is a watchlist for CS and product. It is not a license to cut list price in Salesforce or to write a new price into Anaplan or Pigment. Segments can change who you watch. They cannot set price.
Connect the cut to the forecast instead of a slide. Differentiated retention by billed cohort is an input to an ML rolling revenue forecast, not a replacement for it. If payment behavior diverges on the same accounts, keep that work in adaptive dunning sequences. Do not bake dunning tactics or dispute risk into an LTV cell.
The planning system receives segment IDs, field cites, and draft rates. Salesforce can show the label to GTM. None of those writes should silently restate booked revenue because a cluster assignment moved after the freeze.
Failure modes that kill the plan cut
Three errors show up in review, and they are operational rather than statistical.
Planning from n=4. Suppress the cell. An account list is allowed. A rate is not.
Inventing a lifetime value. Publish rates and cites. Do not publish a DCF, a blended LTV, or a cluster value that finance will paste into a board pack.
Treating a cluster as a pricing decision. Send pricing a watchlist if you must. Do not write a price, a discount floor, or a packaging change into the segment file.
If those three stay closed, the rest is procedure: freeze billed customers, cluster on named fields, cite those fields on every segment, keep thin cells empty, and let FP&A own every number that enters the plan.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first