Skip to main content
DoneThat

AI Adoption GuideGovernmentDeliver

Service demand forecaster

Time-series ML predicts volume spikes by service type and region, enabling proactive staffing and resource positioning before queues form.

Government processPlanFundAuthorizeDeliverInspectEnforceReportClose

By Don, DoneThat’s AI coach · updated

Lock the series before anyone trains a model

A service demand forecast is only as trustworthy as the series you freeze. Before a model runs, name the service types, the regions, the clock (day, week, or shift window), and the cutoff of the last observation allowed into training. That cutoff is the vintage. If two teams pull "the same" counter from different extracts, you do not have one series. You have two arguments wearing the same label.

Lock it in an order a duty manager can audit. List every service type that gets its own forecast, including types you are tempted to roll up because they share a lobby. List every region that has its own counter, and mark regions that do not. Daily arrivals are a different series than booked appointments. Freeze the last observation datetime on the extract. Training, scoring, and publication all point at that freeze. A later warehouse load is a new vintage, not a quiet correction.

Licensing walk-ins, booked appointments, mail-in packets, and call-center tickets are different demand even when they share a building. Mixing them because they hit the same lobby produces a number no roster can use. A region that reports in a different system, or that started logging last month, is not a peer of a region with years of counters. When Tuesday's forecast disagrees with Wednesday's, you should be able to point at the vintage, not at the algorithm.

Microsoft, Palantir, Salesforce Government Cloud, and Excel-class models can all sit on a locked series. None of them create the series for you. A spreadsheet that cites extract date and service code is more usable than a platform chart that redraws against a warehouse view that moved overnight. Treat those products as a class of place to compute. Do not treat a vendor name as proof the freeze happened.

If an anomaly-surfacing dashboard later flags a spike, the first check is whether the spike lives in the locked series or in an unversioned extract. Forecasting and anomaly review share counters. They fail together when those counters drift.

A forecast cell that can be audited

Quality here is a cell, not a dashboard score. Each predicted volume must carry three citations in the same place a manager reads the number: which locked series, which vintage (last observation date, and the model run date if they differ), and which service type. A cell that shows a volume with no vintage is not a forecast. It is a rumor with a decimal.

Forecast only after the lock. Write predicted volume into the cell together with the three citations. Do not publish a grid a manager has to reconstruct from a filename and another tab. If the platform cannot hold vintage next to the number, stamp it in the header and keep an Excel-class copy that does.

Do not add a wait-time percent to dress the cell up. Wait time depends on arrivals, staffing, channel mix, and same-day shocks. The model predicts volume for a named service type in a named region. If public language needs wait-time rules, that belongs in a separate product, closer to a citizen-facing policy assistant that explains published policy, not a percentage invented from the volume column.

Illustrative path, not a measured case: a motor vehicle operation forecasts next week's in-person renewals separately from title work, by district. The renewal cell for District North cites series renewals-in-person, vintage observations through Friday close and a Monday morning model run, and service type renewal. The title cell cites titles-in-person with the same vintage and its own service type. A manager can disagree with either number and still know what was predicted. They cannot do that if the cell is a blended "counter traffic" figure with no vintage.

Train and score against the locked series only. If a later extract revises last month's counts, that revision is a new vintage, not a silent overwrite. Republish the cell. Keep the old vintage so a roster written on Tuesday can be compared with what Tuesday actually used.

Empty stays empty when a region has no history

A new office, a newly split district, or a service type never logged in that region has no series. The correct forecast is blank. Do not backfill from a neighbor, a statewide average, or the first week of walk-ins. Inventing volume for a new office is the failure mode that looks helpful in a steering meeting and then puts staff in the wrong building.

Blank is a quality signal. It tells the roster process to use another method: a labeled manager estimate, a short pilot window, or a delay until counters exist. Those are staffing judgments. They are not model outputs. If the grid cannot render empty, the grid is wrong. Force the empty through even when a vendor interface prefers a zero. Zero is a prediction that demand will be absent, which is a different claim than "we have no history."

The same rule applies to sparse types. A region with history for renewals but none for commercial plates should forecast renewals and leave commercial plates empty. Filling the empty cell from renewal seasonality is a different claim than the model was hired to make.

Inspection and field programs hit the same honesty problem. A risk-based inspection scheduler that scores sites with no inspection history should not invent a demand-like volume for those sites either. Empty is cheaper than a confident wrong number.

The roster is still a staffing decision

The forecast is an input to the roster. It is not the roster. After cells are published, a service operations manager still decides who is on the counter, who is on phones, who sits in a float, and which sites get overtime or borrowed staff. Leave, training, union rules, skill mix, and building hours are not in the time series. Treating the forecast as the roster is the second failure mode: the model printed a volume, the schedule copied it into hours, and a specialist call-out emptied the title window anyway.

Use the cited cells to position people before queues form. That is the reason to forecast at all. If District North's renewal cell is high relative to its own history, move float toward that district before Monday opens. If title cells are quiet, do not strip title-capable staff just because the number is lower than last week. Service type matters. A body on a renewal window does not clear a title queue.

A resource allocation optimizer can propose moves across sites once the forecast cells exist. It still consumes the same citations and the same blanks. It does not get to fill empty regions so the optimizer has a complete matrix. Completeness is not quality.

Write the handoff in one direction. The model run produces cited cells. The manager accepts, overrides, or holds a blank. Roster software, or a spreadsheet, records the human decision. If you skip the accept step, you will not know whether a bad Monday was a bad forecast or a roster that never looked at the vintage.

What breaks when the number looks finished

A number with no vintage cannot be compared to actuals. When Wednesday's volume arrives, you will not know whether you are scoring Monday's model, Tuesday's refresh, or a dashboard that rolled forward at noon. Refuse to publish that cell.

Treating the forecast as the roster collapses volume into headcount. Volume is arrivals. Headcount is a policy about service level, skill, and risk. Two districts with the same renewal forecast can need different staff if one runs appointments and the other runs a walk-in queue. The manager sets that policy.

Inventing volume for a new office poisons the new site and the old ones. Borrowed history becomes actuals in the next vintage if you are not careful, and then the model learns the fiction. Keep the office out of the locked series until it has its own counters. Staff it as a stand-up, not as a forecast.

The remaining work is to keep service types split, keep regions honest, republish on a known vintage, and let the manager staff before the line forms.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first