ML rolling revenue forecast
Multivariate ML projects revenue from CRM pipeline, historicals, and external signals, using tools like Anaplan PlanIQ, Pigment, or Runway.
Finance processPlanBudgetInvoiceCollectPayCloseReportAudit
By Don, DoneThat’s AI coach · updated
What quality looks like on this forecast
A quality rolling revenue forecast is a cited projection, not a scoreboard. Every line that has a number should point at frozen CRM pipeline, closed actuals, and any external series you have explicitly allowed. If a required driver is missing, that cell stays empty. The model does not invent a win rate, a stage conversion, or a seasonal bump to keep the grid full.
FP&A still owns the number that goes to the board. Planning platforms in this class (Anaplan, Pigment, Runway, Workday, Salesforce) can score a horizon from multivariate inputs. They do not replace the judgment that a forecast is presentable, conservative enough, or honest about gaps.
The useful output is a rolling view: last closed actuals locked, open pipeline as of a freeze timestamp, and a forward curve that names what moved it. Accuracy as a published rate is not the quality bar here. Traceability is. If you cannot say which pipeline snapshot and which actuals period produced a cell, the cell is not ready.
Freeze actuals and pipeline before you score
Start by locking two books, not by hitting run.
Close actuals at a named period. Revenue that has already recognized, bookings that have already closed, and any recognized deferred you treat as a run-rate input should sit in that freeze. If last month is still being adjusted in the subledger, wait or forecast from the prior locked close. Scoring against a moving actuals file is how you get a model change that is really a late credit memo.
Freeze CRM pipeline at the same instant. Snapshot opportunity amount, stage, close date, owner, and the fields your model is allowed to use. Salesforce is typically the system of record for that snapshot. A stale stage is a silent driver error. An opportunity that moved to commit in the live CRM but still sits in discovery in last week's extract will misstate near-term revenue. Do not score against a live, ticking pipeline if the rest of the pack is frozen.
Name the allowed external series up front. Those are the only non-company inputs the model may use: a published price index, a contracted capacity series, a calendar of known billing events. If a series is not on the list, it is not a driver. Do not pull an unnamed macro print at run time to fill a weak pipeline month.
This freeze is also where driver-tree assumption suggestions belong. The ML score should sit on the same driver tree FP&A already uses for price, volume, and mix, not on a parallel set of unnamed features. If the tree says win rate is an assumption you maintain, the model may consume the last agreed rate. It may not mint a new one because pipeline coverage looks thin.
Score the horizon and require a cite on every line
Run the model on the frozen pack only. Multivariate scoring in PlanIQ-style and adjacent planning tools is a class of workflow: historical actuals, open pipeline, and allowed externals in, a rolling curve out. Treat vendor labels as interchangeable for this purpose. You are evaluating whether the run produced cites, not which brand scored it.
Every forecast cell should carry a cite pack: actuals period, pipeline freeze identifier, stages and amounts that contributed, external series identifiers if any, and the model version. If the tool cannot attach that pack, export it beside the grid. A number without a cite is a draft, not a forecast line.
Walk the near term first. Next month should be mostly pipeline that is already in late stages plus in-period actuals you already know. If the model fills next month primarily from a learned typical conversion rather than named opportunities, you are looking at an invented win rate wearing a forecast format. Stop and empty those cells until the opportunities or the agreed rate is on the freeze.
Further out, cites get thinner. That is expected. Thin is not a license to interpolate a smooth curve. If coverage drops below what your process requires for a cited line, leave the months empty or mark them as not scored. Empty is honest. A filled grid that cannot name its drivers is not.
Accept, override, or leave the cell empty
FP&A reviews the cited curve the way you review a junior analyst pack. Accept a cell when the cites match the freeze and the story is one you would sign. Override a cell when the model used the right inputs but the judgment is wrong: a known slip that is not yet in CRM, a one-time deal that should not persist, a segment you will not take to the board at the scored level. Record the override as a human line with a reason, not as a silent edit to the model output.
Leave the cell empty when a driver is missing. Missing win rate for a new motion, missing stage history after a CRM redesign, missing actuals for a newly acquired book: those stay empty, not guessed. Inventing a win rate to complete the roll-forward is a quality failure even if the chart looks complete.
Never promote the raw model output to the board pack. The scored curve is an input to the forecast you own. If leadership asks what the model says, answer with the cited score and the accepted number, and show where they differ. Treating the model as the board number skips the only control that makes this a quality outcome.
One illustrative path: an FP&A lead freezes Salesforce open pipeline at close, locks last closed actuals, then scores a rolling view in the planning layer. One segment has no agreed stage-to-close rate after a mid-quarter process change. The scored cells for that segment stay blank. The board pack shows named pipeline for the rest of the book and an explicit gap for the changed motion, rather than a blended conversion the model inferred from the old stage names.
That same freeze would have been wrong if the extract still used last month's stage labels after reps had already recategorized deals. The model would have looked calibrated and still been citing a pipeline that no longer exists.
How the rolling score feeds variance and scenarios
A cited rolling forecast is the plan-stage sibling of later variance work. When actuals land, you should explain miss or beat with the same driver language: pipeline that did not close, mix that did, an external series that moved. That is the handoff into variance report with driver attribution, not a separate forensic rebuild.
Scenarios should start from the accepted forecast, not from a second unconstrained model run. If you need a downside, change named drivers (slip a set of opportunities, cut an allowed external, hold win rate at the last agreed value) and regenerate. Natural-language scenario generation is useful when it restates those driver edits in words leadership can challenge. It is not useful when it asks the model to invent another full curve with no freeze.
If retention and expansion are material to the book, keep new-logo pipeline separate from existing-customer motion. Cohort structure from customer cohort LTV segmentation belongs in the driver tree so the ML score does not average a land-and-expand book into a single conversion. Mixing those populations is another way a complete-looking forecast hides a missing driver.
Re-freeze on a calendar you can defend: close plus pipeline snapshot, score, accept or override, publish the owned number. Live CRM restaging between freezes is analysis, not a new official forecast, until you lock again.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first