Skip to main content
DoneThat

AI Adoption GuideFinanceBudget

Zero-based budget challenger

LLM probes line owners with structured "why this number" questions and tests justifications.

Finance processPlanBudgetInvoiceCollectPayCloseReportAudit

By Don, DoneThat’s AI coach · updated

What a quality challenge produces

A zero-based budget challenger produces questions, not a new number. Each question must cite three things that already exist in the file: the line (account, cost center, or other grain you budget at), the prior actuals your close produced for a named period, and the justification text the owner submitted. If that text is missing, the challenge stays empty. The model does not invent a story so it has something to interrogate, and it does not invent a cut so the review can "finish."

FP&A still owns the challenge. You choose which lines enter the queue, which questions actually go to owners, which replies are good enough, and whether the requested amount changes. Treat the model as a consistent drafter of probes, not as a substitute reviewer.

This is a quality workflow. It tests whether a submitted justification can survive contact with the line and the actuals. It is not a substitute for budget anomaly flagging, which is the right place when a number looks off before anyone has written a story. It is not a cut engine. When you need to move dollars after leadership has set a target, use reallocation agent during cuts as its own process.

Planning workbooks and models in systems such as Anaplan, Datarails, Vena, or Workday already store the line, the last closed actual, and the comment field. Read those fields. Do not scrape a slide or a remembered figure from an earlier cycle.

Load the line, the actuals, and the justification text

Do not generate questions against a P&L screenshot. Assemble a payload for the one line you intend to challenge.

Include the line identifiers your owners recognize (account, cost center, product, project: whatever they will see when they reply). Include the requested amount and currency. Include prior-period actuals that the close actually posted, with the period named. Include the owner's justification verbatim. Include any driver the owner already attached (volume, rate, headcount). Stress-testing those drivers is a different job; send that work to driver-tree assumption suggestions instead of burying driver math inside a "why this number" question.

Then apply a hard gate. If the justification is empty, or is only a placeholder ("TBD", "same as last year" with no reference to an amount or a period), do not emit questions. Empty stays empty. FP&A can require a first submission. The model does not write that submission, and it does not treat a blank comment as evidence that the line is padded.

When the justification exists, every question must do three citations: name the line, name the actuals period and the posted figure, and quote or tightly paraphrase a specific clause from the owner's text. A question that cannot point at all three is not a zero-based challenge. It is an open-ended interview, and owners will answer it with another slogan.

Ask, capture the reply, then let FP&A decide

Keep the cycle sequential. Do not collapse it.

Load. Confirm the payload and the empty-justification gate. If the gate failed, stop. The output for that line is blank.

Ask. Draft a short set of cited questions. Each question should be answerable from operational facts the owner already has (scope, timing, mix, who does the work). Do not ask the owner to reverse-engineer a cut you have not decided to make.

Owner replies. Capture the reply as new text on the line. Leave the original justification in place. You need the trail of what was claimed before the probe versus what was claimed after it. If the owner still cannot answer because the first text was too thin, that is an FP&A process call, not a cue for the model to fill gaps.

FP&A decides. Accept the number, send one tighter follow-up that still cites the file, or change the amount. The model does not choose among those three. It does not score the reply into a recommended reduction.

Stop after the questions until a human owner has answered. A challenger that jumps from probes to a suggested cut is no longer a quality tool. Vendor rates belong in the questions only when the justification itself talks about vendors or rates. An external comparison is vendor spend benchmark. Do not paste a bench into a challenge unless that bench is already a field in the file you loaded.

A professional-services line that already has a story

This is an illustration of the cycle, not a measured result and not a template for a cut.

An FP&A lead opens a cost-center line for external professional services. The requested amount is above the prior-year actual the close produced for that same account and cost center. The owner has already written a justification: a named systems implementation, a mix of existing retainers and incremental specialist days, and a claim that the work cannot slip without delaying go-live.

The challenger may ask why this line's requested amount exceeds the named prior-year actual; which clause about "incremental specialist days" maps to that delta; and whether the go-live date in the text is still the date on the project plan. Each question names the line, cites the actuals period and posted figure, and quotes the clause it is testing.

The challenger may not say the line should come down to a round number, invent a "more typical" prior-year actual, or ask why the owner failed to justify the line if the comment field is blank. On a blank comment, the correct output is no questions. FP&A requests a first write-up through the normal submission rule.

The owner replies in the planning system. FP&A reads that reply against the original text and the posted actuals, then clears the line, asks one follow-up, or changes the number. The disposition has a human name on it.

Failure modes that look like diligence

Challenging a blank as if it were padded. This is the fastest way to kill the process. An empty comment is a missing submission, not proof of fat. Generating "why is this padded versus last year" on a blank line teaches owners that silence gets punished and that the model will supply a motive. Gate on empty. Chase the write-up separately.

Auto-cutting from the questions. Once probes exist, it is tempting to score answers and emit a recommended reduction. Do not. The questions are the product. A labeled "suggestion" will be treated as the answer in a compressed review meeting. If leadership wants a lower total, run a cut process in the open.

Inventing last year's actual. Models will round, annualize, or "remember" a figure that is not in the payload, especially when the justification mentions run-rate or a partial year. If the close figure is missing, do not write the question that needs it. Do not substitute a calculated proxy unless that proxy is already a named field in the planning system and you cite it as that field.

Related slips: quoting a driver the owner never attached; citing another cost center's actual; treating a one-line "inflation" comment as a full zero-based narrative and burying the owner in questions the file cannot support.

What FP&A still decides

You still pick the queue: materiality, variance to posted actuals, first-year zero-based lines, contested scope. You still decide when a reply is good enough. You still own the number that goes forward.

Use the challenger where justifications already exist and you need faster, more consistent probes than one reviewer can type by hand. Do not use it to pretend the organization is doing zero-based review if owners are not writing justifications yet. Fix the submission rule first.

Keep the artifact narrow on purpose: line identifier, cited actuals, quoted justification clause, question, owner reply, FP&A disposition. A recommended cut, a filled-in story, or a remembered actual is out of scope.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first