Demand and capacity forecast
ML model predicts compute, storage, and license demand 12-18 months out using historical utilization and headcount signals.
IT processPlanSelectDeployProvisionSupportUpgradeReplaceRetire
By Don, DoneThat’s AI coach · updated
What the forecast produces and what it does not
A demand and capacity forecast gives FinOps and capacity planning leads a 12-18 month view of compute, storage, and license need derived from historical utilization and headcount signals. The output is a planning artifact: projected units by resource class, each line tied to the utilization vintage and headcount source that drove it. It is not a purchase order, not an approval to spend, and not a substitute for vendor negotiation or architectural review.
The model extrapolates from what you already consumed and who you already employ. It does not invent a growth rate when headcount or utilization history is thin. Empty fields stay empty. That honesty is the point. A forecast you can defend in a budget meeting cites its inputs; one you cannot usually hid assumptions inside a single blended percentage.
Use this page when you need numbers for annual planning, refresh cycles, or budget scenario modeling, not when you need real-time alerting or rightsizing of live workloads.
Signals you need before you run it
Two signal families drive the forecast: utilization and headcount. Both must carry provenance, because downstream readers will ask when the utilization snapshot was taken and which headcount feed was used.
Utilization covers compute (CPU, memory, instance hours), storage (allocated vs. used capacity), and license consumption where metering exists. Typical sources include observability platforms such as Datadog, cloud and data platforms such as Snowflake, and virtualization stacks such as VMware. The critical metadata is utilization vintage: the date range and aggregation window of the underlying metrics. A forecast built on Q2 averages differs materially from one built on a 30-day spike in Q4.
Headcount provides the people-side constraint: engineering headcount by cost center, department growth plans from HR, or role counts tied to license bundles. ServiceNow and similar ITSM or HRIS integrations often supply org structure; finance may supply approved headcount plans separately. The forecast must record which feed was used (actuals vs. planned, as-of date, granularity).
If either signal is missing for a resource class, the forecast leaves that slice blank rather than imputing a default. Planning still proceeds with partial coverage; it just does not pretend completeness.
Loading utilization and headcount into the forecast
Start by defining the forecast horizon (12 or 18 months) and the resource classes that matter for your environment: virtual machines, container clusters, object storage tiers, database capacity, SaaS seat counts, and any metered platform licenses.
Step 1: Export or query utilization history. Pull consistent intervals (weekly or monthly) for at least four quarters where possible. Tag each export with utilization vintage: start date, end date, aggregation (peak, average, p95), and scope (region, business unit, tag set). Observability tools like Datadog supply time-series rollups; VMware and hypervisor management APIs supply host and VM allocation; Snowflake and warehouse platforms expose credit or storage consumption over time. Normalize units before merge (gigabytes vs. terabytes, vCPU vs. physical cores).
Step 2: Attach headcount at matching granularity. Align headcount rows to the same periods as utilization. Use stable identifiers (cost center, department ID) so license-per-seat models can join cleanly. When HR provides only annual plans, load the plan as a separate series and label it explicitly; do not silently replace actual headcount with plan figures without marking the source change.
Step 3: Run the forecast with cites enabled. The model projects forward from utilization trends and headcount-linked demand drivers (for example, seats per engineer, storage per analytics role). Each projected line should echo utilization vintage and headcount source in the output metadata or footnotes, not only in a cover sheet.
Step 4: Review blanks before sharing. Any resource class without sufficient utilization history or headcount mapping remains empty. Document why (no meter, new service, org change mid-period) in planning notes, not inside the model as a guessed growth rate.
Reading the output: cites, gaps, and the planning handoff
Treat the forecast as input to human decisions. FinOps validates unit costs and currency; capacity planning validates physical and logical limits (rack space, region quota, license true-ups); procurement validates lead times and contract windows.
An illustrative walkthrough, without fabricated metrics: suppose utilization vintage spans January through December of the prior year, averaged monthly, sourced from Datadog for compute and VMware for on-prem VMs. Headcount comes from an HR export as of March 1, engineering departments only. The forecast shows rising compute need tied to headcount in two departments, flat storage with a blank line for a new analytics warehouse because Snowflake history starts mid-year, and license seats tracking headcount with a cite to the HR feed. Planning sees where confidence is high (compute with full vintage) vs. where someone must still estimate (storage for the new warehouse). They might pair the compute curve with a benchmark-driven standard config generator to translate abstract vCPU growth into SKUs, and send license rows to a license assignment optimizer to stress-test seat models before buy.
The forecast never auto-procures. It does not open tickets, trigger POs, or resize clusters. Planning owns the buy: they choose timing, vendor, and quantity after cross-checking replacement timing optimizers for hardware and contract calendars for software.
Failure modes that break trust
Forecast with no utilization vintage. If readers cannot see the observation window, they will assume the worst case or the best case to suit their agenda. Always publish vintage alongside numbers. Refuse to sign off on a deck that strips cites for simplicity.
Treating the forecast as a PO. Procurement needs quotes, approvals, and often legal review. A projected seat count is a planning hypothesis until finance releases budget and a buyer selects a SKU. Escalations happen when automation or over-eager stakeholders skip that chain.
Inventing a growth rate. When headcount is flat and utilization is noisy, a tempting shortcut is "apply 15% YoY." That fabricates demand and poisons the next cycle when actuals diverge. Leave the cell empty, flag the gap, and let planning supply a scenario in the budget scenario modeler with explicit assumptions rather than embedding fiction in the base forecast.
Mixing planned and actual headcount without labels. A sudden jump in projected license need often traces to swapping HR actuals for hiring plan mid-run. Keep sources separate or version the forecast when inputs change.
Single-vendor blind spots. Datadog sees what is instrumented; VMware sees what is virtualized; Snowflake sees warehouse usage; ServiceNow sees CMDB and request patterns. None alone is the whole estate. Gaps in coverage produce blanks, not guesses.
Operating rhythm for capacity and FinOps leads
Run the forecast on a fixed cadence aligned to budget gates (quarterly refresh for operational tuning, annual deep pass for 18-month capital plans). Each run should be diffable: utilization vintage moved, headcount source updated, new blanks appeared because a service landed without meters.
In review meetings, lead with what changed and what is still unknown. Stakeholders care less about precision to the third decimal than about whether storage for a new platform was assumed or left open. Pair forecast output with scenario work when blanks are material; one base forecast plus explicit scenarios beats one over-smoothed curve.
Capacity planning retains authority over buys. FinOps retains authority over cost assumptions. The forecast connects them with cited, time-bounded signals so neither team argues from incompatible spreadsheets. When signals improve (full Snowflake vintage, tagged FinOps chargeback), re-run and shrink the blank regions. Until then, empty stays empty, and planning still decides what to purchase and when.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first