Skip to main content
DoneThat

AI Adoption GuideFinancePlan

Natural-language scenario generation

LLM agent rebuilds scenario branches from prompts like "model 15% EMEA cut", using tools like Runway or Pigment.

Finance processPlanBudgetInvoiceCollectPayCloseReportAudit

By Don, DoneThat’s AI coach · updated

A scenario is usable only when it cites drivers and a baseline

A generated scenario is ready for review when it names the baseline it started from, lists every driver it changed, and leaves every other line equal to that baseline. If the prompt names a driver the model does not contain, that request stays empty. Completeness is not a reason to invent a P&L.

Quality is citation, not fluency. A paragraph that says "EMEA down 15 percent" without a driver ID, a version, and units is a comment, not a scenario. FP&A still owns which branch is official. The language model drafts a candidate. It does not promote that candidate to the live plan.

Planning suites such as Runway, Pigment, Anaplan, and Workday already store versions, driver trees, and scenario branches. Those objects are the source of truth. The prompt is an instruction to mutate a copy, not a substitute for the model.

Lock the version you will measure against

Lock a named baseline before anyone types a prompt. Do not aim at "current" unless current is a frozen version with a name. Use the latest board-approved plan, the latest forecast snapshot, or an explicitly labeled working copy. Record the version name, the close date it represents, and whether revenue on that version is a judgment call or an ML rolling revenue forecast.

Locking does two jobs. It gives the rebuild a starting cube. It also gives the delta a denominator. Without a locked baseline, "15 percent" has nothing to be 15 percent of, and two reviewers can disagree about which file they are looking at.

If your process already produces driver-level assumption notes, use those notes as the dictionary the parser will resolve against. The companion page on driver-tree assumption suggestions is the place to keep names, units, and default directions consistent so a prompt can match a real node.

After the lock, refuse to mix versions inside one branch. A scenario that applies an EMEA cut on Q2 actuals and a hiring freeze on the FY plan is two stories glued together. Split the work or re-lock.

Map the prompt onto drivers that already exist

Take a request such as "model 15% EMEA cut." That sentence is a sketch, not a measured result. Before any cell changes, parse it into fields the model already knows:

  • Geography or entity: EMEA, only if EMEA is a dimension member.
  • Driver: which named node moves. Revenue, bookings, opex, or a more specific leaf such as regional marketing spend.
  • Direction and magnitude: minus 15 percent, taken off the locked baseline for that driver, not off a made-up total.
  • Scope: which periods, products, or legal entities inherit the change.
  • Hold-constant list: everything the prompt did not name.

If EMEA is not in the dimension, or "cut" could mean volume, price, or headcount and the model has no rule for that ambiguity, stop. Return empty for the unmatched part. Do not guess a P&L so the conversation can continue.

Headcount is a frequent trap. "Cut EMEA 15 percent" is not permission to invent an FTE figure. If headcount is a driver in the model, resolve it through the same named node you would use in the capacity and headcount optimizer. If it is not a driver, leave it blank and say so. Inventing a round number to make payroll land is how unofficial scenarios become unofficial headcount plans.

Agree units before you write. Minus 15 percent of revenue is not the same as minus 15 percent of contribution margin, and a signed integer in thousands will not match a percent driver. If the driver only accepts a rate, write a rate. If it only accepts a currency amount, do not send a percent and hope the engine converts it.

Write the parse back to the requester in driver language before you rebuild. "I will apply -15 percent to driver Regional_Revenue on member EMEA, version locked to FY26_Plan, periods Jan-Dec, no other drivers." That sentence is the contract. If they meant opex, they can correct it while the cube is still untouched.

Rebuild the branch, then show the delta

Copy the locked baseline into a new scenario branch. Apply only the parsed driver writes. Recalculate dependent nodes with the same engine the official plan uses. Do not re-forecast lines the prompt did not touch.

Then show the delta in the same grain as the model: driver, member, period, baseline value, scenario value, variance, and the rule that produced the variance. A single enterprise total without that trail is not enough to review. The reviewer should be able to point at one changed driver and see every line that moved because of it.

Spot-check a few lines the prompt did not name and confirm they still equal the baseline. If they moved, the write hit the wrong driver or the engine has a side path you did not intend.

Keep the narrative off this step. A later board-pack narrative draft can describe the branch once FP&A has accepted the numbers. Generating commentary in the same pass as the rebuild is how a draft gets treated as the pack.

If a dependent node cannot calculate because an input is missing, surface the break. Filling the gap with a plausible figure is the same failure as inventing a P&L.

Empty is the correct answer when the model cannot match

An empty return is a quality outcome when the named driver, member, or unit is not in the model. It is not a product failure. The alternative is a complete-looking statement that no one can audit.

Common unmatched cases: a region the chart of accounts does not use, a "cut" that has no mapped driver, a percentage with no locked baseline, or a request to change a calculated node as if it were an input. In each case, say what did not resolve and leave the branch unbuilt.

Do not widen the prompt to nearby drivers to be helpful. Nearby is how an EMEA revenue cut becomes a global opex story.

FP&A still chooses the official branch

Only a person on the FP&A team marks a branch official. The agent can propose as many named branches as the suite allows. Until that mark, the output is a draft sitting next to the locked baseline.

Treat the draft as a board pack and you skip the control that matters: a human confirming that the cited drivers were the ones leadership asked for, and that empty stayed empty where the model could not follow. After that confirmation, the narrative draft and any downstream forecast refresh can run against the signed branch, not against the chat transcript.

The prompt never becomes the plan. The signed scenario does.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first