AI Adoption GuideConsultingAnalyze
Multi-Scenario Outcome Modeler
ML models outcome distributions under differing strategic scenarios using client operational and financial inputs.
Consulting processSellScopeStaffKickoffAnalyzeRecommendDeliverClose
By Don, DoneThat’s AI coach · updated
Replace the single-line forecast with a range the client can argue with
The job is to show how much the answer depends on assumptions, not to crown a winner with a number that looks finished.
A partner who arrives with one NPV, one payback, and one recommended path has already hidden the argument. The client will challenge the drivers anyway. Put the drivers on the table first, then the range of outcomes under each strategic choice, using the client's own operational and financial inputs. If those inputs are missing, the honest output is that you cannot score the option yet.
This step is quantitative analysis, not the write-up of a case. A strategic option comparison synthesizer can array qualitative trade-offs. This modeler sits underneath: same options, explicit drivers, outcome distributions. Later, an implementation business case generator can package a preferred path. If that pack collapses the range into a single point, you threw away the quality of the work.
The model usually lives where the client already plans. Excel, Anaplan, Pigment, and Palantir Foundry are that class of place: workbooks, connected plans, operational objects, financials. Treat them as the system that holds drivers, not as ranked scenario products. Do not assume any of them ships a particular distribution engine. The work on top is the same: named scenarios, named assumptions, a range per scenario, and sensitivity before the recommendation.
Machine learning does not create driver data. A history of the same drivers can support a fitted distribution around volume or cost. Sampling from an interval the team made up only restates the fiction. If the client cannot supply the driver, say so.
One assumption sheet per scenario, still editable in the room
Do not run three scenarios off a shared base case with a few cells toggled on a hidden tab. Each scenario gets its own assumption sheet. Keep it client-facing and editable while people are in the room.
For every driver, record:
- Name and unit (overtime premium per hour, not "ops")
- Which scenario it belongs to
- The value or range the client supplied, or blank
- Source (GL, plant manager, contract, not supplied)
- Who in the room will defend it
- What a blank does to the run: drop the option, hold it out of the ranking, or show an unscored branch
A MECE hypothesis tree generator helps decide which drivers belong on the sheet. It does not fill the cells. Figures from data rooms and board packs can come through an unstructured document extraction pipeline. Someone still has to label last year's overtime rate as history, not a forecast.
Do not treat an upside / base / downside fan on one plan as three strategic scenarios. Those are three versions of the same choice. Different choices get different sheets.
Invented inputs are how this work goes wrong first. If the client cannot supply fully loaded plant cost, regional demand, or contract unit economics, write "not supplied" and keep that option out of the ranking. Filling cells so the model can run means the ranking is about your fill-ins.
When the CFO changes a policy in the session, change the cell and let the range move in front of them. A locked file only the analyst can touch trains the room to argue with the slide.
Output a range, not a fake-precise point
For each scenario, publish a range on the few measures the decision uses: cash over a horizon the client can actually see, capacity risk, contribution. Add a distribution only when you have history that supports one. Do not publish a single NPV to two decimals.
If two scenarios overlap on the measures that matter, you do not have a numerical ranking. You have a decision the numbers have not settled.
Match the horizon to the life of the assumption. Mix, freight, and labor you can observe over the next 18 to 36 months can support a range. A ten-year projection presented as precise is precision theater. Error compounds in the far years. Those years are not evidence. If the client wants a long DCF, show the near-term range as the finding and label the tail as a sensitivity they own.
If you simulate, publish the inputs to the simulation, not only the fan chart. A smooth distribution around invented drivers is still invented.
Run sensitivity before you write the conclusion
Show which inputs move the ranking before anyone names a preferred scenario.
Hold the other supplied drivers fixed. Move one driver across the range the client will defend. Record whether rank order changes. Repeat for the handful of drivers that could matter. If a plausible swing in freight or demand mix flips B past C, that flip is the finding. The recommendation is then "resolve this driver" or "choose on grounds the model does not own," not a quiet "we recommend B."
When you can, rebuild the sheet for a decision the client already made, using what they knew then. Check whether the model would have ranked the option they took. Agreement does not prove the next decision. Disagreement means a missing driver or a broken structure.
Once a preferred path exists, a recommendation adversarial stress-tester attacks the prose. This modeler should already have listed the drivers that change the answer and the ones that do not.
Example: overtime, a second plant, or contract overflow
Illustrative only. No results, no case study.
A manufacturer needs more finished goods in two regions over the next two planning years. Three choices sit on the table:
- Second shift and overtime at the existing plant
- A second plant closer to one region
- Contract overflow to a third party
The sheets differ. Overtime needs utilization, premiums, yield on extra hours, and freight from the current site. The second plant needs capex timing, ramp, local labor, and a freight map that exists only if regional shipments are real. Overflow needs a real unit cost, minimums, and who owns quality escapes.
If West volume is a story with no shipment history, you cannot rank the second plant against overflow on freight. Do not invent a West share to make the model run.
What you owe the room is overlapping ranges on cash and service risk over those two years, then sensitivity. The useful sentence is not "the model prefers the second plant." It is "freight and West volume flip the rank; capex timing flips overtime versus the second plant; contract unit cost is not supplied, so overflow is not scored." The partner can ask for the missing cost or decide on operational grounds with the gap visible.
Sophistication that wins the room can still be wrong
A polished model can win the meeting while being wrong. That is worse than a simple model people argue with.
- Invented inputs. Blanks filled so every option has a number. The ranking is a story about the fill-ins.
- Ten-year precision theater. Annual decimals on a decade nobody can defend. The decision rides on years that are not evidence.
- Complexity as credibility. Linked modules, simulation, a learned demand model, a Foundry pipeline, an Anaplan or Pigment app. None of that substitutes for a driver the plant controller will not sign. The room stops arguing because the artifact looks expensive.
The work is doing its job when each scenario has a sheet the client edited, the output is a range, sensitivity sits on the page before the conclusion, and missing drivers stay missing. It is failing when the slide has one number, the workbook hides estimates, or the cleverest model in the room cannot say what would change its mind.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first