AI Adoption GuideConstructionAward
Contractor Delivery Risk Scoring
ML predicts delay and cost overrun probability per bidder from historical project performance data.
Construction processBidAwardPlanMobilizeBuildInspectHandoverClose
By Don, DoneThat’s AI coach · updated
Score delivery risk from closed jobs you actually ran
A delivery-risk score at award is a ranked note built from your first-party history, with cites the commercial lead can open. It is not a forecast of how many days this package will slip, and it is not a gate that drops a bidder from the list.
Start with completed work you ran with each bidder: programme outcome versus the baseline you held at award, cost outcome versus the award sum (and versus approved variations), and the reasons recorded at close-out. The model's job is to order those files by delay and cost-overrun risk for this package type, not to invent a delay probability you can quote as a measured result.
If the bidder has not closed comparable work with you, the score is incomplete. Label the history thin and keep them in the pack. A new specialist with no file is exactly the case where exclusion would punish lack of prior spend, not delivery risk.
The same retrieval step that feeds this score is vendor past performance retrieval. Do not skip it because a dashboard already shows a tidy rank. The rank is only as good as the jobs behind it.
Project records usually live in the systems already used for cost, drawings, and site reporting, including platforms such as Autodesk and Procore. Treat those systems as the source of closed-job facts. Do not treat a vendor export as a substitute for reading the cites the panel will be asked to defend.
Cite the projects; keep thin files labelled thin
Every rank line should name the jobs that moved it. For each cited project, the note should show: package type and location class, award date and practical completion, programme variance with the recorded cause, cost variance split between contractor performance and instructed change, and whether the team was the same entity you are scoring now (novation, name change, or a different regional company).
Thin history stays labelled thin. That label is a quality of the evidence, not a risk grade. Repeated closed jobs of similar scope make a file. One purchase-order for a day rate does not. A joint-venture partner's record is not automatically the bidding entity's record.
When a single delay sits in the file, read the cause before the model's pattern weight. Weather, a late client variation, or a utility that arrived after the date in the contract is a one-off unless the same cause repeats. Treating that delay as a pattern is how you mark a reliable contractor as risky and then award to someone whose file is simply quieter.
Pair this note with abnormally low bid detection when price and delivery risk pull in opposite directions. A cheap bid with a clean but thin file is a different problem from a mid-price bid with repeated overruns on similar packages.
Illustrative pack: three bidders on one award
A secondary-school extension is at award. Three bidders remain after compliance.
Bidder A has closed several education fit-out packages with you. The cites show completion close to the approved programme, cost movement mostly in instructed change, and the same regional company on the bid form. The note ranks A lower on delivery risk and lists those jobs.
Bidder B is a specialist envelope contractor you have not used. Their submission is technically strong. The model has almost nothing first-party to score. The note does not invent a delay probability and does not drop B. It labels the history thin and tells the panel that any risk call on B is a judgement about references and capability, not about your closed jobs.
Bidder C has one delayed hospital job in the file. Close-out records a late variation that moved the envelope sequence. The note cites that job, quotes the recorded cause, and does not treat the delay as a repeating pattern. If the panel still wants a programme buffer, that is a commercial decision, not an auto-exclusion.
The pack is a ranking with footnotes, not a shortlist. The PM lead can still award to B or C. The quality bar is that the panel can see which first-party jobs were used, which files were thin, and which delays were one-offs.
Failure modes that quietly wreck the award
Excluding a new specialist with no history. Procurement pressure likes a complete score. A blank file is not a complete score. If you convert "thin" into "out," you shrink the market to firms you already employ and you never build history with anyone else. Keep specialists with no file in the pack, and send the panel to references and method statements instead of to a fabricated rank.
Treating a one-off delay as a pattern. Models overweight rare events when the sample is small. One delayed hospital job with a documented client variation is not evidence that the contractor always overruns education packages. If the note cannot show repetition under similar scope and similar cause codes, write "single event, cause recorded" and leave the weight with the panel.
Skipping past-performance retrieval because the score looks clean. A tidy rank with no job list is the failure mode that looks like efficiency. If nobody retrieved the closed jobs, the score may be ranking rumours, mixed entities, or another region's company. Run vendor past performance retrieval first, then score. If retrieval returns nothing material, the honest output is a thin-history label, not a green rank.
A fourth slip sits next to these three: mixing this score with win probability scoring. Win odds are about whether you will land work. Delivery risk is about whether the bidder will finish the work you already intend to buy. Do not let a high win-probability story on a bid/no-bid memo substitute for first-party delivery cites at award.
What the panel owns after the note is written
The award panel still owns the decision. The model does not exclude, does not set a cut-off probability, and does not write the recommendation. Feed the ranked risk note, with cites, into the award recommendation report so price, programme, quality, and delivery history sit in one pack the minutes can record.
Record three things in the award minute: which bidders were scored from first-party history and which were labelled thin; which cited delays the panel treated as patterns versus one-offs; and that no bidder was removed solely because the score was missing or unflattering. If the panel awards against the rank, write why. That sentence is the quality control. Without it, the next package will treat the model as a gate.
Do not publish a delay probability for the winning bidder as if it were a measured programme result. The probability, if the model emits one, is an ordering device trained on past jobs. The live job will have its own variations, weather, and interfaces. The panel's job is to buy with eyes open, using the ranked note as evidence, not as a substitute for judgement.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first