AI Adoption GuideGovernmentPlan
Policy impact simulation
LLM-orchestrated simulation models downstream economic, social, and demographic effects of policy options before political commitment.
Government processPlanFundAuthorizeDeliverInspectEnforceReportClose
By Don, DoneThat’s AI coach · updated
The quality bar: cited brief, cited assumptions, cited vintage
A policy impact simulation is usable in a briefing only when every filled effect traces to three locked artefacts: the option brief, the assumption table, and a named data vintage. If an assumption is missing, the cell stays empty. The run does not invent an employment percentage, a fiscal delta, or a demographic shift to look complete. Ministers still choose.
That is the quality outcome, not a prettier dashboard. Spreadsheet models, consultant workbooks, and LLM-orchestrated pipelines fail in the same place when they skip it: a number appears, nobody can reconstruct which input produced it, and the pack starts to read as if the politics were already settled.
Treat platforms as hosts, not as authors of missing figures. Palantir, Microsoft, OpenGov, and Excel-class models can store parameters, scenarios, and outputs. None of them licenses a jobs line that has no vintage, no assumption row, and no option identifier. If the system of record cannot show those three citations next to a cell, the cell is not ready for a minister.
A simulation answers what might happen under stated conditions. It does not rank options, pick a winner, or convert a blank into a zero. A zero is a claim. A blank is an honest gap.
Lock the option set and the assumption table first
Do not start a run until the option list is closed. Each option needs a short brief: what would change in statute, funding, or administrative practice; who is in scope; when it would take effect; and what is explicitly out of scope. If two options differ only in a talking point, collapse them. If an option is still being rewritten in a minister's office, it is not in this run.
Then lock the assumption table. Typical rows include labour-supply response, take-up, displacement, lag between enactment and observed effect, geographic grain, and which populations are held constant. Each row needs an owner, a source, and a status: cited, pending, or excluded. Pending and excluded rows do not generate numbers. They generate blanks and a named gap.
This lock is where briefing risk is usually born. An analyst who "just runs it" with a default elasticity is inventing a jobs number with extra steps. An analyst who treats the simulation as the decision skips the lock because the politics feel urgent. An analyst who runs without a data vintage produces a chart that cannot be audited when the statistics agency revises the series.
Once options and assumptions are locked, freeze the input pack and give the run an identifier. Later political edits create a new run. They do not patch yesterday's workbook in place. If consultation evidence would change the briefs themselves, stop and update the briefs before you spend another cycle. A public consultation synthesizer can show which publics contest which assumptions. A regulatory gap scanner can show whether an option is lawful as drafted. Neither replaces the impact run. Either can retire an option before you model it.
Run only cited inputs; empty stays empty
The engine may use only inputs that appear in the locked pack with a citation. A citation is a source, a vintage (publication date or extract date), and the series or table name. Administrative extracts need the same discipline: system, extract date, coverage, and known gaps.
When an assumption is missing, leave the cell empty. Do not interpolate from a neighbouring jurisdiction. Do not borrow last year's elasticity because it looks close. Do not let a language model complete a blank because the prompt asked for a full table. The output should show the option, the affected channel (economic, social, or demographic), and either a modelled range with its cited drivers or an explicit blank that names the missing assumption.
Keep channels in separate columns. A housing-supply option can move construction activity, rents with a lag, household formation, and school-age population. Mixing those into one impact index hides which assumption failed. A minister can reject one channel without discarding the rest only if the channels were never blended.
If the run is LLM-orchestrated, the model is a clerk. It maps the option brief onto the assumption table, pulls the cited series, and formats the gaps. It does not estimate a missing elasticity. Human review checks that every filled cell has a vintage and that every blank names the missing row. If the clerk filled a jobs line from "typical multipliers," delete the line and restore the blank.
Time pressure is not a reason to fill gaps. The cheaper recovery is a shorter pack: fewer options, fewer channels, every cell cited or empty. That pack can go the same afternoon. A complete-looking table with an invented employment figure cannot be walked back once it is on a slide.
Illustrative run: three density options, no invented jobs line
A planning unit compares three options for allowing more homes within a stated distance of rapid transit: a modest height lift, a larger height lift with inclusionary rules, and no change. The option briefs are locked. The assumption table includes construction-cost pass-through, expected take-up by landowners, a lag from permission to completion, and a school-place yield per additional home. The demographic series is an official population projection with a stated vintage. The rent series has a vintage. The labour-market series that would support a "jobs created in construction" line does not have an agreed elasticity for this geography.
The correct simulation fills channels that have cited drivers: additional dwellings under each take-up assumption, a range for completion timing, and a school-place implication that cites the yield row and the projection vintage. The jobs line stays empty. The output names the missing elasticity and does not substitute a national multiplier. A minister can still prefer the larger height lift on housing grounds without being handed a fabricated employment percentage.
That walk-through is a pattern, not a result. Change the domain (a skills visa, a congestion charge, a benefit taper) and the same rule holds. If the assumption is not in the table, the cell is blank. Do not dress the blank as "indicative" or "illustrative impact." Those labels are how invented jobs numbers travel.
The human presents; the simulation does not decide
The analyst, or the senior official who owns the brief, presents the pack. The pack includes the option briefs, the assumption table with statuses, the vintage list, the filled channels, the blanks, and the sensitivity notes. The presenter says what would change if a pending assumption arrived, and what would not.
Do not send model output as a decision memo. A decision memo names a preferred option. A simulation pack names consequences under stated conditions. Palantir dashboards, Microsoft workspaces, OpenGov records, and Excel-class workbooks can all be the system of record for the run. The attachment does not convert a blank into a recommendation, and it does not convert a filled channel into a ministerial choice.
If two options look close on the filled channels, send them through a scenario comparison matrix so differences in assumptions and blanks stay visible. Do not break a tie by inventing the missing jobs line. After a preferred option exists politically, a resource allocation optimizer can test delivery envelopes. That is a later question. Impact simulation answers what an option might do. It does not answer how to staff it.
Refuse a run with no data vintage on any series that would fill a cell. Refuse to populate employment, fiscal, or demographic effects from model common sense. Refuse to collapse blanks into a dash or a zero. Refuse language that says the model recommends or selects an option. Quality here is deliberate: cited brief, cited assumptions, cited vintage, blanks where the table is incomplete, and a human who still owns the choice.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first