AI Adoption GuideConstructionBid
Adversarial Bid Review
LLM stress-tests bid assumptions against scope and estimate, surfacing gaps and underpriced items before submission.
Construction processBidAwardPlanMobilizeBuildInspectHandoverClose
By Don, DoneThat’s AI coach · updated
What you get from a stress-test of the bid
An adversarial bid review does not produce a new price. It produces a gap list. Each row is an assumption in the bid, a cite in the scope, a cite in the estimate, and a status the estimator still owns.
The useful output is boring on purpose. For every flagged line you should open the extract and see the same clause, drawing note, or BOQ item the model used. If the extract does not contain that BOQ line, the cell stays empty. The model does not fill the hole with a guessed quantity, a provisional sum, or a typical rate from another job.
Treat the list as a check on completeness and underpricing, not a go/no-go. A long list can still be a bid you want. A short list can still hide commercial risk that belongs in clause review. The review does not change the submitted figure.
Pinning assumptions to scope, drawings, and the estimate
Start with three locked sources, not a chat about the project in general.
Scope is the tender pack you priced against: specifications, drawings, addenda, and the form of contract as extracted. Estimate is the current takeoff and pricing file: BOQ or CBS, rates, allowances, and notes. Assumptions are the statements the bid team already made in writing: inclusions, exclusions, qualifications, and allowed-for language in the cover letter or tender queries.
Run the comparison in that order. For each assumption, find the scope cite first. If the spec or drawing does not support the assumption, record that as a scope miss, not as a missing BOQ line. Then look at the estimate. If the BOQ already carries the work, the assumption is priced even if the wording is sloppy.
Flagging a priced item as missing is a common failure. It happens when the model matches on section titles and ignores a measured item two trades down, or when preliminaries absorb temporary works the architectural spec described under the envelope.
Keep platform context as context, not as a second source of truth. Many teams keep drawings, issues, and cost files in Autodesk or Procore products. Cite the extract you fed, not a live folder. If the extract is revision C and the live model is revision D, the list is only as current as revision C.
Feed takeoff that you already trust. If quantities are still being built from drawings, freeze automated BOQ takeoff from drawings before you ask the model to hunt for holes. An adversarial pass on a moving takeoff will re-flag the same measured work as not found every time the extract lags the spreadsheet.
Historical jobs are for the estimator. Historical bid RAG retrieval can remind you a similar envelope needed a mock-up. Put that on your checklist. Do not add a BOQ line unless the current extract contains the requirement.
Recording gaps with cites the estimator can check
Write each gap so a colleague who did not run the prompt can verify it in ten minutes.
Use a fixed row shape. Quote the bid team's assumption in their own words. Give a scope cite with document, clause or sheet, and a short excerpt. Give an estimate cite with BOQ or CBS code, or leave it empty. Name the gap type: unsupported assumption, unpriced scope, under-described rate, or conflict between two cites. Leave the estimator decision blank until a person fills accept, reject, or needs a query.
Empty estimate cites are allowed. They are required when the extract has no matching line. Do not substitute a sister code, a PC sum, or a round number to be confirmed. Inventing a provisional sum is the failure that looks most like helpfulness. It turns a completeness check into a pricing change the commercial lead never authorised.
Underpriced items are not missing items. If the BOQ has the curtain wall, but the rate note says supply only and the spec says supply, install, and seal, the gap is a rate-basis conflict with two cites. The review lists the conflict. It does not rewrite the rate or roll a new total.
Do not let the model merge gaps. Keep one assumption, one scope cite, and one estimate cite per row unless you explicitly attach multiple scope cites to the same estimate line. Collapsing three drawing notes into "prelims look light" is not a cite.
Example: a curtain-wall note the BOQ never carried
A bid lead is closing a mid-rise package. The architectural spec, in the glazing section, requires a full-size visual mock-up on site before production glass is released. The estimator's qualification list says facade as per drawings and specification. The BOQ extract has measured curtain-wall area, opening infills, and a preliminaries lump for site establishment. It has no mock-up line.
A useful gap row looks like this. Assumption: facade as per drawings and specification. Scope cite: glazing spec, mock-up clause, excerpt requiring a full-size visual mock-up on site before production glass. Estimate cite: empty. No BOQ line in the extract names a mock-up, sample bay, or visual prototype. Gap type: unpriced scope attached to a broad inclusion. Estimator decision: still blank.
The estimator now chooses: add a measured or lump item, qualify the mock-up out, or confirm it sits inside a prelim the extract did not label. The review does not pick among those three and does not insert a provisional mock-up so the sheet looks complete. If a prelim note already said includes specified mock-ups, reject the gap and fill the estimate cite. The first extract was incomplete; the bid was not automatically wrong.
If the same model also flags the measured curtain-wall area as missing because the BOQ title is aluminium framing rather than curtain wall, that second flag is noise. The area is priced. The estimator rejects it with the estimate cite. Priced work with odd names will keep appearing as holes unless you bind the review to codes, not to trade labels.
Calls the model must not make for you
Three failure modes show up on almost every first pass. Treat them as reject reasons, not as proof the review failed.
When the model flags a priced item as missing, bind the next run to BOQ codes and measured descriptions, not to heading text. Ask it to search the full extract before it writes not found. If it still flags a priced line, paste the cite and close the row. Do not add duplicate money.
When the extract has no line, the estimate cite stays empty. A PC, PS, or allow figure is a commercial instrument. Only the estimator creates one, and only after they decide the scope is real. The gap list can say the mock-up clause is unpriced in this extract. It cannot mint the sum.
Do not treat the review as a go/no-go. You can have zero gaps and still decline on programme, bonding, or liquidated damages. You can have a long list and still submit after qualifications. Win probability scoring sits downstream, and only after a person has accepted or rejected each row. Raw flags include priced items mislabelled as missing. They are not a bid/no-bid score.
The bid price does not move because the review ran. Totals change when the estimator accepts a gap and edits the estimate, or when they issue a qualification the commercial team signs. Until then, the submitted figure is the figure you already had.
What belongs next to this check, not inside it
Run clause tagging on the same extract so onerous risk is not mistaken for a missing BOQ item. A fitness-for-purpose clause is not an unpriced trade. Put that language through tender clause risk classification.
When the gap list is closed, every accepted row has an estimate cite or a written qualification. Every rejected row has an estimate cite that was always there. Empty still means empty. If the BOQ line was not in the extract, you either refresh the extract and re-run, or you price it yourself. The model does not finish the sheet for you.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first