Skip to main content
DoneThat

AI Adoption GuideConstructionInspect

Inspection Failure Prediction

ML predicts inspection failure probability for upcoming activities based on prior failure rates by crew, trade, and activity type.

Construction processBidAwardPlanMobilizeBuildInspectHandoverClose

By Don, DoneThat’s AI coach · updated

What the score is allowed to mean

A failure-probability score on an upcoming inspection is a risk flag with cites. It is not a failed inspection, not a punch item, and not authority to pull a crew off scheduled work. Quality still owns whether to add hold points, walk the work early, or leave the inspector's checklist as written.

The flag is useful only when you can open the prior failures it cites. Each cite should name the crew, the trade, or the activity type, and it should point at an inspection record from this project, or from another project only when that record is attached to the cite. A number without a cite is not a score.

Inspection logs and punch history already live in platforms such as Procore and Autodesk. The model reads that history. It does not replace the inspector who will sign the next report. If tomorrow's activity is flagged, decide what extra evidence quality wants before the inspector is called. Do not pre-mark the activity as failed.

How to score an upcoming inspection from cited fails

Start from the activity on the lookahead, not from a trade-wide memory. Name the inspection, crew, trade, and activity type the way they appear on the schedule and in the inspection template. Pull prior fails that match those keys.

For every cite, keep four things you can defend in a quality meeting: which inspection failed, which crew or trade or activity it belonged to, what the inspector wrote, and that you are allowed to use that project's record. If any of those four is missing, drop the cite. Do not fill the hole with a rate from a sister job, a national figure, or a story about how that crew "usually" works.

Score the upcoming inspection only from cites that survive that filter. A high score means the cited history is thick enough that quality should plan extra hold points. A low score means the cited history does not justify extra process. A blank score means there is no cited history for that crew, trade, or activity. Blank is a valid output. It is not a pass and it is not a fail.

Do not invent a failure rate. If Crew 4 has cited flashing fails on this job and nothing else that matches, the score cites those records. It does not become a percentage rounded up from another tower. If the only matching history is activity-type fails from a different waterproofing crew, say that. Do not relabel those fails as Crew 4's.

Pull tomorrow's inspections from the same place the superintendent already uses, match each one, and write the cites next to the activity before the morning huddle. If a lookahead constraint scan already lists the inspection as a constraint, the score is extra context for that constraint, not a second schedule.

Illustrative example, not a measured case: tomorrow's envelope inspection is window install on Grid D by Crew 4. The model returns a risk flag that cites prior fails on this project, sill-pan photos rejected on Level 3 for the same crew, and a flashing fail on the same activity type by the waterproofing trade two weeks earlier. It does not cite a rate. Quality opens both reports, checks that the inspector comments still apply to Grid D, and leaves the flag as a flag. Nobody writes "failed" on tomorrow's inspection. Nobody bars Crew 4 from the elevation.

Extra hold points quality still owns

A risk flag earns extra hold points only when quality decides the cites warrant them. The model does not add hold points. Quality writes them into the inspection plan the same way it would after reading punch history by hand.

Map hold points to the cited defects, not to a generic high-risk checklist. If the cites are sill-pan photos and flashing, the extra hold is those two details, with photos or a quality walk while the work is still open, before the inspector is called. If the cites are torque values on a different activity, do not invent a sill-pan hold because the scores looked alike.

Keep the scheduled inspection. Extra hold points sit in front of it. The inspector still owns pass or fail. The holds exist so quality can catch the cited failure mode while the work is accessible.

Tie each hold to the cite's grain. Crew-level cites put the extra hold on that crew's next occurrence of the activity. Trade-level cites apply to that trade's activity type even when the crew is new, and the note should say so. Activity-type cites must not be written as if a new crew personally failed the last inspection.

When the flag is low or blank, do not add hold points "to be safe." Extra process without a cite trains the field to ignore both the score and quality. Use the ordinary template and spend the extra walk on activities that have cites.

If the cited pattern is workmanship a skills review would also catch, send the superintendent to a deployed workforce skills gap analysis after the hold points are on the plan. The score does not become a training record.

When the record is empty

Empty stays empty. A crew with no inspection history on this activity does not get a borrowed rate. A new trade partner does not inherit another contractor's fail count. A first occurrence of an activity type on the job does not get a percentage copied from a different project unless that project's inspection records are attached as cites and quality accepts them.

Blocking a crew because they have no history looks like caution and acts like a blacklist. No history means you have nothing to cite. The inspection proceeds on the standard template. Quality can still walk the work. The superintendent can still pair the crew with a known foreman. None of that requires an invented score.

If someone wants to fill the blank from how this trade performed on the last job, require the cite: project, inspection ID, crew or trade or activity, inspector comment. Without that, the blank remains. If the cite exists and quality accepts the other project as in-scope, the activity is scored with foreign-project cites, labeled as such.

Do not treat empty as a clean bill of health. Absence of cited fails is not evidence of good work. Photo-based AI defect detection from site photos can still surface visible defects on the same elevation. Thin inspection history also does not prove the submittal is right; that belongs in a submittal spec compliance check, not in a made-up failure rate.

Write "no cited history" on the lookahead line so the blank is visible.

Misreads that turn a flag into a penalty

Treating the score as a fail puts predictions in the inspection log and sets crews arguing with a number instead of an inspector. Keep the flag in planning notes. Keep pass and fail on the inspector's report.

Blocking a crew with no history converts "we have not inspected them yet" into "they are high risk." That punishes the crews quality has not reached, which is the opposite of a cited-history method.

Inventing a rate from another project without a cite turns the flag into folklore. Sister-job stories and trade-wide memory are not cites. If Procore or Autodesk has the inspection record, open it and attach it. If it does not, leave the rate out.

Extra hold points are preparation. If hold-point photos are clean and the inspector still fails flashing, the fail is the inspector's. If the score was high and the inspector passes, the pass is the inspector's. Do not reopen a passed inspection because the model was pessimistic, and do not skip calling the inspector because the model was calm.

The note quality leaves next to the activity is the artifact: cited fails, extra hold points or none, and who will walk them. The score is only how those cites were found.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first