AI Adoption GuideOperationsSchedule
Historical duration calibration
ML model refines task duration estimates from completed execution data, reducing schedule overruns over time.
Operations processIntakePrioritizeScheduleExecuteVerifyDeliverConfirmClose
By Don, DoneThat’s AI coach · updated
Overview
Planned task durations rarely stay accurate for long. Work changes, crews rotate, tooling differs by site, and the estimates baked into the calendar slowly diverge from what actually happens on the floor. Historical duration calibration uses completed execution records to propose better duration values for recurring task types, so the next schedule is built on evidence rather than stale defaults.
Schedulers remain in control. The model suggests duration updates; people decide what enters the live calendar.
Why planned durations drift from reality
Most schedules start with standard times: a work package template, a planner’s judgment, or a one-time study that was never refreshed. Those numbers are useful until conditions change. A new fixture, a different skill mix, longer travel between bays, or tighter quality checks can push actual finish times past the plan without anyone updating the estimate library.
Overruns compound. When one task runs long, dependent work slips, buffers get consumed, and the rest of the day looks worse than it is. Underestimates also hide capacity: if durations are padded everywhere, the schedule looks full while crews finish early and idle. Either failure mode is a schedule quality problem, not only a timekeeping problem.
Calibration addresses the gap between what the plan assumes and what completion data shows for comparable work. It does not invent new tasks or resequence the day on its own. It focuses on making the duration inputs more honest so constraint-based planning and matching work have better raw material.
What historical duration calibration does
The workflow is straightforward. Completed executions are grouped by a stable task identity (type, product family, site, or another agreed key). For each group with enough history, a model estimates a refined duration distribution or a single recommended planning duration, often with a confidence band. Those recommendations surface as proposed updates against the current estimate catalog.
Staff review the proposals. Accepted values become the new planning defaults for that task identity. Rejected or deferred proposals leave the calendar unchanged. Over successive cycles, accepted calibrations reduce systematic bias in planned durations for work that recurs often enough to learn from.
The outcome of interest is schedule quality: fewer surprises relative to the plan, tighter alignment between booked time and realized time, and less firefighting caused by estimates that never matched the work. Throughput gains may follow, but the primary job is to stop feeding the scheduler bad numbers.
Calibration is most useful when the same kinds of tasks repeat under comparable conditions. One-off projects with unique scopes benefit less until enough similar completions exist to form a trustworthy cohort.
Inputs, outputs, and when the model stays quiet
Useful inputs typically include planned duration, actual start and finish (or effort), task type or template ID, location or line, skill or crew attributes when available, and flags for rework, interruptions, or holds. Clean completion timestamps matter more than fancy features. If “done” is recorded loosely, the model will learn the wrong lesson.
Outputs should be proposal-shaped, not auto-applied: a recommended planning duration, optional percentile or risk-aware values for buffers, the sample size behind the estimate, and a short rationale (for example, which cohort and time window were used). Pairing each proposal with the prior default makes review fast.
Empty output is the correct behavior when history is too thin. Sparse cohorts, brand-new task types, heavy missing timestamps, or completions dominated by exceptional events should produce no recommendation rather than a confident guess. A blank result tells the scheduler to keep the current estimate or run a manual review. That is safer than a low-sample “update” that looks precise and is not.
Define thinness in operational terms before go-live: minimum completed instances per cohort, a maximum age for included jobs, and rules for excluding aborted or massively interrupted runs. Publish those thresholds so planners know why a task type has no proposal this cycle.
How schedulers review proposed updates
Treat proposals as advisory. A typical review queue shows task identity, current planning duration, proposed duration, sample size, recent actuals summary, and any caveats. Reviewers accept, reject, or accept with an edit (for example, round to a planning increment the shop floor understands).
Human acceptance protects against blind spots the model cannot see: upcoming process changes, temporary staffing shortages, seasonal volume spikes, or a known data quality issue in one plant. The calendar stays authoritative until someone intentionally updates the estimate library or the next planning run’s duration table.
Governance helps. Limit who can accept changes that affect many future jobs. Log accept/reject decisions so you can audit whether the model is being trusted for the right cohorts. Prefer batch review windows (daily or weekly) over interrupting live dispatch with constant micro-updates.
After acceptance, regenerate or refresh schedules that still use the old defaults for open work where policy allows. Do not silently rewrite in-progress assignments without a planner check; mid-job duration edits confuse crews and supervisors.
Measuring whether calibration is working
Judge success with schedule quality metrics, not model loss alone. Track planned-versus-actual duration error by cohort, share of jobs finishing within a defined tolerance of plan, residual overrun after known holds are excluded, and how often thin-history cohorts correctly return no proposal. Also watch acceptance rate: very high acceptance with no improvement may mean reviewers are rubber-stamping; very low acceptance may mean proposals are unusable or poorly explained.
Compare error before and after accepted updates for the same task identities over comparable volumes. Segment by site and shift so a global average does not hide a plant that is getting worse. When a process change lands, expect temporary empty or low-confidence output until new completions refill the cohort.
Keep the feedback loop closed. Rejected proposals and post-hoc notes about why actuals diverged (material delay, missing tooling, scope creep) improve the next cycle’s filters. The goal is a living estimate library that stays close to how work really runs, with people still owning what the calendar promises.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first