Skip to main content
DoneThat

AI Adoption GuideSalesClose

Probabilistic forecast roll-up

AI replaces gut-feel forecasts with data-driven deal probability and ARR-weighted commit, using tools like Clari, Aviso, or BoostUp.

Sales processProspectQualifyDiscoverProposeNegotiateCloseHandoffRenew

By Don, DoneThat’s AI coach · updated

A probability-weighted roll-up is a range you inspect, not a number you submit

Gut-feel forecast calls break down when the book is large, late in the period, or mixed across reps who price risk differently. A probabilistic roll-up does a narrower job: multiply each open deal's remaining ARR by a deal-level close probability, sum those products, and show a range around that sum. That output is a quality view of what the current book implies. It is not the commit the VP of Sales will defend.

The useful question is not "what will we close." It is: given the probabilities we will stand behind on each deal, what band of ARR does this pipeline currently support, and how far is that from the number I am willing to own. Forecast software can show expected ARR and a spread. The manager still names the number they will be held to.

Weight every open deal by its own probability

Start from open pipeline only. Closed-won is already in the book. Closed-lost is out.

For each remaining opportunity, take the ARR that would land if the deal closed in this period, and a probability that belongs to that deal, not to its stage label. Stage-based weights treat every late-stage opportunity as interchangeable, which hides the gap between a paper-complete expansion and a first-time logo with an unsigned order form.

Where that probability comes from can vary: CRM fields, rep judgment, historical close patterns, or a model sitting on top of Salesforce. MEDDIC auto-extraction makes the qualification evidence behind a probability inspectable. Deal risk scoring can flag slippage, missing stakeholders, or a stalled next step. Those are inputs to a probability, or triggers to inspect. They are not a second forecast stacked on the first. If you apply a risk haircut on top of a probability that already encodes the same risk, you double-count and the roll-up drifts low without new information.

Sum remaining period ARR times deal probability across the open book. That sum is the probability-weighted point: expected ARR implied by current probabilities, not commit.

Then show a range, not a single line. A workable range uses more than one cut of the same book: a conservative band from high-probability deals, the weighted point, and an upside band that includes lower-probability coverage. Those three layers are for inspection. Only a person chooses which deals inform the commit they will defend.

Suppose a VP is looking at four named deals still open late in the period: a $400k expansion with verbal confirmation, a $250k new logo in legal, a $180k competitive renewal, and a $90k add-on added last week with a thin mutual plan. Stage names put three in late stages. Deal-level probabilities do not: likely expansion, unfinished logo, uncertain renewal, coverage add-on. Weighting remaining ARR by those probabilities lands well below the stage-weighted total. The conservative band is mostly the expansion. The weighted point includes a partial contribution from the logo and the renewal. The upside band only looks full if the add-on is treated as real. The VP still says which deals they will defend, and whether to haircut for concentration, slip timing, or a product gap the probabilities never saw.

That walk-through is a method check, not a measured result. Weight by deal, show the range, keep commit as a separate call.

Build the range so commit has somewhere else to live

The range is a quality-control surface. It shows whether the book, under stated probabilities, can support the number leadership wants, and where the mass of expected ARR sits. If the conservative band is thin and two deals carry the weighted point, the conversation is concentration and inspection, not a request to believe the model more.

Commit is the number the VP will stand behind if those deals slip. It can sit near the conservative band, near the weighted point, or rarely above it when the VP has evidence the probabilities are too low, such as a signed order still in processing. Auto-rewriting commit from the weighted roll-up trains the org to treat expected ARR as a promise.

Upside is leftover coverage. Treating upside as commit is how this process usually fails in a forecast call: the slide shows a full funnel, the weighted math is ignored, and the spoken number is the sum of everything that could close. The range exists so "could" has a place to live that is not the commit cell.

Last-mile deal coaching is how you move a probability by changing the deal. A disqualification recommender is how you pull deals that should never have been in the weighted book. Those actions update inputs. They do not substitute for the VP's commit call.

Forecast platforms and CRM stages are inputs, not the call

Clari, Aviso, and BoostUp sit in the same class: they take CRM pipeline, often Salesforce, and produce deal-level probabilities and rolled-up views a sales leader can inspect. Salesforce can hold stage, amount, close date, and custom probability fields. Use this class when stage-weighted CRM views are not enough and the org wants a consistent, inspectable roll-up.

Treat them as a class, not a ranked shortlist. The operating contract is the same whichever tool produces the probabilities: deal-level weights, an ARR-weighted sum, a visible range, and a human-owned commit the tool must not silently overwrite.

Demand inspectability. A rep and a manager should see why a deal's probability moved, whether the roll-up uses remaining ARR for this period, and whether commit is a separate field. If the commit tile is the weighted sum with a new label, that separation is already gone.

Three ways this roll-up gets corrupted

Replacing commit with the model. The weighted point looks precise, so it gets pasted into the forecast submission. Precision is not ownership. Expected ARR includes deals the VP is not willing to defend. Keep commit as an explicit call, even in weeks when it matches the weighted point.

Double-counting risk scores. A deal is already at a low probability because the champion left and legal is unsigned. A separate risk flag then applies another haircut. The book looks safer than the evidence warrants, and reps stop trusting either number. Use risk scores to explain a probability, to trigger inspection, or to set the probability. Do not stack them unless they encode different information.

Treating upside as commit. Coverage deals and "if everything hits" scenarios belong in the high band of the range. Speaking that band in the commit slot recreates gut-feel forecasting with better charts. If the org needs a stretch number, label it stretch.

Pulling closed-won ARR back into the weighted open book inflates the point. So does weighting a likely slip into this quarter with no slip path. Win/loss reason classifier work after the period helps recalibrate probabilities next time. It is not a reason to rewrite this period's commit after the fact.

How the VP still sets commit

After the weighted point and the range are on the table, the VP still names a commit: deal-level yes or no decisions, plus judgment about concentration, timing, and anything the probabilities cannot see.

Inspect the deals that carry most of the weighted mass. Confirm remaining ARR and close-period eligibility. Set commit from deals that are personally defensible, not from the upside band. Leave the range visible so finance and RevOps can see the gap. If the gap is large, inspect, coach, or remove junk. Do not slide commit up to meet the range.

The quality outcome is the range and the probability-weighted roll-up next to a commit the manager still owns. Anything that auto-fills commit from the model has missed the point.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first