AI Adoption GuideConstructionClose
Lessons Learned Knowledge Base Extraction
LLM extracts lessons from RFI logs, NCRs, and delay events into a searchable RAG knowledge base for future projects.
Construction processBidAwardPlanMobilizeBuildInspectHandoverClose
By Don, DoneThat’s AI coach · updated
A cited lesson, or an empty field
The quality bar is a lesson that cites a specific RFI, NCR, or delay record, and that a reviewer can open that record and see the same cause, constraint, or decision. If the record does not contain a lesson, the field stays empty. The model does not invent a moral, does not close a lessons-learned workshop that never ran, and does not fill a blank with fluent advice.
Extraction is a filter. A paragraph that sounds useful but has no cite is not a lesson. A paragraph that restates the NCR narrative without stating what would change on a similar package is also not a lesson. Quality is not whether the model wrote something. It is whether the statement is supported by the cited record and would be safe to retrieve on the next job.
A project-controls or commercial lead should run three gates in order: extract with a cite, keep empty when unsupported, and confirm before the item is searchable. Skipping a gate is how a knowledge base fills with language that later retrieval will treat as fact.
What RFIs, NCRs, and delay files can support
Source files usually come from project systems in the Autodesk and Procore class, plus claim folders and structured exports. Read the fields and the narrative that belong to that record. Do not treat a neighboring email thread, a site diary, or a chat dump as part of the record unless it is attached as an official exhibit.
For each record, produce one extraction: source type, source ID, lesson text or an explicit empty, and a verbatim excerpt from the record that would justify the lesson. If the lesson is empty, give the reviewer a reason such as routine clarification, no process finding, or dates only. That reason stays in the review queue. It does not enter the searchable index. Do not merge an RFI, an NCR, and a delay into a single combined lesson.
From an RFI, a lesson is a coordination, specification, or design gap that forced the question, plus the answer that actually governed the work. Cite the RFI number. Do not promote a contractor preference the response never adopted. A color confirmation or a routine dimension check with no recurring issue yields an empty lesson.
From an NCR, a lesson is the recorded condition, the accepted disposition, and any process or design change the record itself names to prevent a repeat. Cite the NCR number. When the NCR already carries a cause code, keep extraction aligned with defect root cause classification. Do not invent a different cause so the lesson reads better. A one-off workmanship miss with no process finding stays empty.
From a delay event, a lesson is the event type, the notice or schedule evidence in that file, and a decision that changed sequence, only if the delay file states it. Cite the delay or event ID. Do not rewrite the event as a claim story. Chronology and entitlement belong with delay claim chronology assembly. If the log has dates and days but no cause analysis, do not write a cause.
The same empty rule applies across all three source types: no lesson in the record means no lesson in the knowledge base.
One NCR, one allowed extraction
A facade NCR records cracked sealant at a movement joint, photographs, a disposition to cut out and replace, and a note that the joint width on the approved detail was smaller than the sealant manufacturer's minimum. The allowed output cites that NCR and states that the approved joint geometry did not meet the product minimum, so a similar package should check joint width against the product data before shop drawings are released.
Not allowed: a lesson that quality culture was weak, that the subcontractor was unreliable, or that every sealant joint on the project should be redesigned. None of those claims is in the NCR. Writing them anyway is the first failure mode: a lesson the NCR does not support. Once indexed, those sentences retrieve against unrelated packages as if close-out had agreed them.
Apply the same test to an RFI that only confirms a finish and to a delay event that only records a late drawing with no agreed cause. Empty is the correct output. Do not complete the record so the dashboard looks populated.
Index for retrieval, not for a workshop pack
After a lesson has a cite, index it so a future estimator, planner, or package manager can retrieve it by trade, location type, cause, and source ID. The corpus is a RAG knowledge base: the lesson text, a short excerpt from the source, and citation metadata. The cite must survive retrieval. A hit without an RFI, NCR, or delay ID cannot be checked.
This is the same retrieval pattern as historical bid RAG retrieval: similar past language is a candidate, not an instruction. Confirmed lessons can later inform estimating benchmark update when they change a production rate, a repeatable risk, or a contingency assumption, and only after someone has accepted that the lesson is true of that record.
Do not index site-diary gossip. A diary line that the architect always delays glass is not a delay event. Indexing it teaches the next search that the architect is the cause even when the delay file never said so. That is the second failure mode: gossip in the corpus.
Keep each indexed chunk as one lesson, one cite, one short excerpt. Do not merge five RFIs into a theme before indexing. Theme-writing is workshop synthesis. This pipeline does not run a workshop. Splitting the lesson from the NCR number is how citations disappear in retrieval.
Confirm before anything is searchable
A person in project controls, commercial, or quality confirms the extraction before it enters search. The reviewer opens the cited RFI, NCR, or delay record, checks that the lesson is supported, and accepts, edits to match the record, or rejects to empty. The reviewer should reject to empty when the excerpt is not in the cited record, when the lesson names a person or company as the cause without a finding in that record, or when the text is a general slogan. Edits are limited to matching the record. The reviewer does not improve the lesson by adding knowledge from memory of the job.
Skipping confirmation is the third failure mode, and it does the most damage. An invented lesson looks fluent in a review queue and then appears in the next project's retrieval set as agreed close-out knowledge. A cited but unsupported lesson is the same error with an ID attached: the cite makes it look audited when it is not.
Keep status simple: extracted, confirmed, rejected-empty. Only confirmed records are searchable. Empty and rejected records stay out of the index. If the source system later revises the NCR disposition, re-run extraction and confirm again. Do not leave a stale lesson on a superseded record.
The close-out meeting can still happen. This pipeline does not replace it. It only turns records that already contain a lesson into cited, retrievable knowledge, and it leaves the rest blank so the next project does not inherit fiction.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first