AI Adoption GuideConstructionBid
Win Probability Scoring
ML model trained on historical bids scores go/no-go likelihood by project type, client, and competitive set.
Construction processBidAwardPlanMobilizeBuildInspectHandoverClose
By Don, DoneThat’s AI coach · updated
Treat the score as a cited prior, not a bid decision
A win-probability score is a go/no-go input with cites. It is not a bid decision.
Train on your own closed bids. Score the live tender by project type, client, and competitive set. Put that score, and those three cites, in the board pack. The bid board still owns go, no-go, or go with conditions.
Historical outcomes are priors. They describe how similar tenders went for this contractor. They do not promise this tender. A high score does not mean the clauses are acceptable. A low score does not mean a strategic client is off the table.
Bid directors already see the failure: a pursuit lead who reads a dashboard as permission to bid, or a commercial manager who reads a low score as a mandatory no-go. Keep the number as evidence with a named source.
Build the model from first-party bid history
Score from records you can defend in a board pack, not from a vendor default or an industry average you cannot audit.
Pull closed opportunities from the systems that already hold them. Estimating files, CRM win/loss notes, and project records in platforms such as Autodesk or Procore are typical sources. Treat those platforms as a class of project and document stores, not as a ready-made win-probability product. You need outcomes you labelled, features you can name, and a date range the board will accept.
Define the unit as a tender you actually decided on. Include wins, losses, and recorded no-gos. Exclude live bids and pursuits that never reached go/no-go. If a record has no outcome, leave it out of the training set.
Label each closed bid with the three drivers you will cite later. Project type is the work you price (healthcare fit-out, highways, education new-build), not a marketing sector. Client is the buying entity, plus relationship state where you have it: first bid, repeat, framework, incumbent. Competitive set is who you expected to be in when you decided, not a rumour list assembled after the result. If you did not record the set then, leave it missing. Missing is more honest than a reconstructed field the model will treat as fact.
Add only features a bid manager would recognise: region, procurement route, value band, design status at tender, whether you were already on site. Drop features that leak the result: final margin after award, post-tender notes written after the decision, delivery KPIs from a job you later won.
historical bid RAG retrieval is the companion when you need the narrative behind a similar loss, not only the structured row. Retrieval does not replace the score. It lets a bid manager read why a neighbour in the history went the way it did.
Retrain on a fixed cadence, or when the book of work shifts (new sector, new geography, a framework you have never held). Freeze the model version that produced a score so the pack can say which history it used. If comparable closed bids are few, say so on the face of the score. A thin history is not a precise probability.
Cite project type, client, and competitive set on every score
The quality outcome is a cited input. If the board cannot see why the number moved, they will ignore it or over-trust it.
Write the score as a short block. For project type, say what comparable work in the training set looks like and how this tender matches or misses. A refurbishment scored against new-build history is a mismatch, even under the same client. For client, say whether this buyer appears in your history, or only a similar owner class. Repeat-client history does not transfer to a first-time developer because both are "private." For competitive set, say who you expect in, and whether that set matches sets you actually lost or won against. A two-contractor framework shortlist is not the same cell as an open regional list.
Name the neighbours in words. Closed healthcare refurbishments for this trust, mixed outcomes, usually the same other contractors: that is a cite. "The model is confident" is not.
When a driver is missing, say it is missing. An unnamed competitive set is a gap, not a neutral. A first-time client with no in-house history is a gap. The score can still run. The pack must not pretend the cite exists.
Keep the language of likelihood, not fate. Lower than our typical result for this project type when this competitive set is present: that is usable. "We will not win" is a decision, and it is not the model's to make.
Keep the bid board as the decision owner
The board's job is go, no-go, or go with conditions. The model's job is to put a prior on the table with cites.
Run the score before the board sits, not after someone has already sold the pursuit internally. If the number only appears once the team wants to bid, it will be used to confirm, not to inform.
Put the score next to bid cost, capacity, relationship, and risk appetite. Do not let it replace tender clause risk classification or adversarial bid review. A favourable prior does not read the amendments. An unfavourable prior does not tell you whether a framework position is worth defending.
Record the decision against the score, not inside it. If the board goes against a low score, write why (strategic client, learning bid, cover). If they no-go a high score, write why (resource clash, unacceptable clauses). No-gos you never recorded look like missing data, not like choices.
contractor delivery risk scoring belongs when you judge whether the job, if won, is one you can deliver. Do not collapse bid-win likelihood and delivery risk into one number. A tender you are historically likely to win can still be a job you should not want.
A bid director has two packages in the same week. Package A is a repeat education client, a project type with a deep closed book, and a competitive set recognised at go/no-go; the score comes back with those cites filled. Package B is a logistics warehouse for a developer you have not worked for, few closed rows, and a competitive set assembled from informal talk. The model still emits a number. Treat B as a weak prior. The board can still go. They should not treat B's number as if it were A's.
Failure modes that look like discipline
No-going a strategic client because the score is low. Relationship, framework defence, and market signalling are board reasons. A low prior warns about historical win patterns. It is not a veto on a client you have decided to keep. Use it to change how you bid, not to launder a relationship decision as a statistical one.
Treating a thin history as a precise probability. Sparse cells (new sector, new geography, a rare procurement route) are unstable. Show how many comparable closed bids sit behind the number. If the cell is thin, refuse a point estimate. False precision is worse than insufficient history.
Skipping clause review because the score is high. Win likelihood is not contract quality. Keep tender clause risk classification and adversarial bid review as gates regardless of the prior. The score answers whether you have historically been competitive. It does not answer whether you can live with the terms.
If you cannot cite the three drivers from first-party history, including when a dashboard in Autodesk, Procore, or a similar store is on screen, you do not have this use case yet. The board test is whether someone outside the model review can say what the score is, which history it used, which drivers moved it, and what they are still being asked to decide.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first