AI Adoption GuideConsultingScope
Effort Estimation from Historical Actuals
ML predicts task-level hours from scope inputs trained on the firm's historical project actuals and delivery patterns.
Consulting processSellScopeStaffKickoffAnalyzeRecommendDeliverClose
By Don, DoneThat’s AI coach · updated
Train on hours at the grain you sell
You cannot predict workstream hours from a project total. If you sell phases, workstreams, or role-weeks, the training rows have to be booked time at that same grain. Partners already estimate from memory of past jobs. This model only earns a seat at pricing if it replaces that memory with closed actuals, not with a smoother version of last year's quote.
Timesheet or delivery actuals have to exist at the grain that matches how you sell. Phase, workstream, and role are the usual commercial units. If people dump eight hours onto a single project code every Friday, you do not have workstream history. You have a project-level average, which is what the partner already carries in their head.
CRM "estimated hours" are not training data. That field is what someone typed into the opportunity before anyone had done the work. Train on it and the model learns to reproduce optimism. Use closed engagements whose hours were booked against the WBS you actually delivered, then map that WBS onto the packages you still sell.
A comparable past scope retriever is the sibling when a partner wants the jobs on the table, not a coefficient. Retrieval shows the evidence. Estimation turns that evidence into a range the pricing desk can argue with.
Pilot the offerings you already repeat
Do not start with the whole firm. Start with a handful of repeatable offerings that share a WBS, a staffing pattern, and enough closed jobs that a partner would recognize the shape. Packaged diagnostics, a standard data-migration workstream, a recurring regulatory filing support package: those are candidates. Bespoke strategy work that never looks the same twice is not.
A new service line has no history. Mapping a neighboring offering onto it is how the model anchors to the wrong work. If you have never delivered "AI operations standup" as a product, do not train on cloud-migration actuals and call the output an estimate. Return "no comparable actuals" and let the partner price it as a first of type. Pretending otherwise is worse than a blank cell.
You will also see grain drift inside an offering that looks stable. Last year's "data migration" workstream included reconciliation; this year's template moved reconciliation into a separate phase. If you train across that change without splitting the WBS, the model will split the difference and satisfy nobody.
Before you train, audit booking quality on those pilot offerings only:
- Hours land on the workstream and role codes you sell against, not a Friday dump-code.
- Overtime, contractor, and nearshore time exist in the extract if they exist on the job.
- The WBS on closed jobs still matches the package you sell today.
If the codes are fiction, stop. Change the booking practice on the next few jobs of that offering, then train. A model cannot recover structure that was never recorded.
Backtest the same way a partner would argue. Hold out closed jobs, pretend you are scoping them from the SOW that was signed, and compare the predicted range to what was actually booked. Consistent bias (always light on a phase, always heavy on a role) is something you can correct. Scatter with no pattern means the grain or the offering boundary is still wrong. Do not publish a midpoint until that check is boring.
A packaged workstream with two different histories
A pricing lead is scoping another ERP data-migration workstream. The firm has sold that package many times, always as a named workstream inside a larger transformation. The partner's memory is simple: one senior and two analysts for about six weeks, because that is how the last few proposals were written.
Closed timesheets for that package tell a more useful story, and they do it without a single headline number. The migration build itself is fairly stable when the source systems and the target are ones you have seen before. Hours move around when the SOW does not name a client data steward, or when data-quality remediation was left as an assumption instead of a deliverable. The steward-dependent slice is not "a bit more of the same work." It is a different risk, and it should not inherit a tight range from the build.
That is the job of the estimator: predict the slices you have actually delivered, and refuse a false point estimate on the slices you have not. Pair it with a scope risk classifier so client-dependency risk widens the range instead of disappearing into a midpoint. If the discovery call never named the steward, an assumption gap detector should surface that before hours are locked. A SOW drafter from discovery transcript can capture the workstream language; it should not invent the hours.
Keep effort figures out of generated prose. The hours come from actuals. The words come from the call.
The partner still owns the number in the proposal
A model that partners override every time has not failed at machine learning. It has failed at evidence. Pricing leads will ignore a midpoint with no comparables, no grain, and no "we have never done this." That is rational. You do not want a proposal whose only defense in a partner review is "the model said so."
Building trust is a data problem first. Show the closed jobs, the WBS lines, the date, the team mix, and whether the job overran. If timesheets are rounded to the nearest day and coded to the wrong phase, fix that before you train. A partner who has watched delivery eat a low bid will not come back because you shipped a nicer dashboard.
Watch for a model that anchors low. Happy-path packages, missing overtime, and junior-heavy bookings all pull predictions under what the next job will consume. If the training set quietly drops the ugly jobs, the output will look comforting and then lose money. Prefer a range, and prefer to say which slice is tight and which slice is a guess.
The partner still types the number that goes in the proposal. Treat the model as a second reader that arrived with the timesheets. If the partner moves the number, capture why. "Client has no data steward" and "this is a first-of-type offering" are different reasons, and only one of them should change the training set. Silent overrides teach you nothing and guarantee the next partner will override too.
Once hours exist at workstream and role grain, staffing can consume them. A utilization and bench risk predictor is only as honest as the effort you just put into the pipeline. Optimistic hours flow into optimistic utilization the same way CRM estimates flow into optimistic quotes.
Sit the model on the PSA you already run
The actuals already live in the professional-services stack: PSA and resource platforms such as Kantata, Planview, Forecast.app, and Mosaic. That is where time got booked, where assignments sat, and where the project structure usually lives. The estimator is a model on top of those records, not a replacement for them.
Do not expect any of those products to ship a magic estimator for your commercial WBS. They hold time, roles, and plans. Your offerings, your phase names, and your habit of booking (or not booking) to workstream codes are firm-specific. Export or warehouse the closed actuals at the grain you sell, join them to the package you sold, and train there.
If the PSA grain is project-only, stop. Change booking practice on the pilot offerings until a workstream or role row is trustworthy, then train. Pulling a firm-wide extract and hoping the model finds the structure is how you get a number nobody will defend in a partner meeting.
Keep the loop closed at the end of the job. Predicted versus booked, by workstream, on the same WBS you used in the proposal. That is the only accuracy conversation worth having. If partners keep overriding in the same direction, the training data or the grain is wrong. Fix that before you ask them to trust the next midpoint.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first