AI Adoption GuideInsuranceQuote
Submission extraction and normalization
NLP parses ACORD forms, PDFs, broker emails, and attachments into structured rating fields for faster quoting.
Insurance processQuoteUnderwriteBindIssueBillServiceRenewClaim
By Don, DoneThat’s AI coach · updated
Quality means a cited rating field
A usable extract is not a filled rating worksheet. It is a set of rating fields where each populated value points to a source form field or an email passage, and every silent source stays blank until operations completes the file.
NLP can read ACORD applications, loss runs, statements of values, broker cover emails, and scanned attachments. Parsing is not the outcome. The outcome is a TIV, occupancy, construction, or limit an underwriter can open and trace. If the pack never states a building limit, the field stays empty. Nobody invents occupancy from a brochure.
Do not auto-rate from uncited fields. A model that guesses square footage to keep the policy admin system (PAS) happy has already failed, even if someone later corrects it. Completeness without a cite is not ready to rate.
This page is the quote-stage discipline for turning a messy inbound pack into structured rating inputs. Appetite screening still happens on appetite and complexity triage. Questions the pack never answered belong in conversational quote intake, not in guessed cells. Extraction here only claims what the submission itself said.
Load the full submission pack first
Load the entire submission before you run extraction. That means the ACORD or carrier application, every PDF in the email thread, the body of the broker email, and nested attachments. A pack split across a forwarding chain is still one submission.
Operations should treat the pack as a folder with a stable ID. The ACORD 125, 126, and 140 (or the carrier equivalent) are usually the spine. The email often carries exceptions, occupancy is warehouse, ignore the office code on the ACORD, or the requested limit is $5 million pending engineering. Those sentences are source text.
If you extract from the ACORD alone, you will miss the email override. If you extract from the email alone, you will miss form field labels that rating and audit expect. Load both. Keep original filenames and page numbers so a cite can say ACORD 140, Building 1, Occupancy, page 2, not "the PDF."
Do not start rating, triaging complexity, or pushing records into Guidewire, Duck Creek, Applied Epic, or Salesforce until the pack is assembled. Those platforms consume structured fields. They do not tell you which attachment was authoritative. Assembly is an operations job. Extraction runs on the assembled pack.
If the broker wrote that the SOV will follow, do not extract TIV from a prior-year spreadsheet in the thread. Park the file until the pack is whole.
Extract with a cite, or leave the cell empty
For each rating field the worksheet requires, the extractor either returns a value plus a cite, or it returns empty. Required fields that are empty block rating. They do not invite a default.
A cite names the artifact and the location: form name and field label, PDF page and region, or email timestamp plus the quoted sentence. Occupancy equals wholesale distributor, ACORD 140, Location 1, Occupancy. That is usable. Occupancy equals wholesale, with no pointer, is not. Treat an uncited populated field as a failure mode, the same as a wrong value. Route it to operations as incomplete, not as ready.
When the source is silent, leave the field empty. Empty is correct. Padding a blank construction class because most similar risks are masonry is inventing a rating input. Padding a limit because the SOV totals look close to a round number is inventing a limit. Both will price. Both will fail a file review.
Map extracted fields to the rating schema the PAS already uses. Do not create a parallel worksheet that underwriters retype. The destination can be Guidewire, Duck Creek, Applied Epic, Salesforce, or a rating workbook that later loads those systems. The quality rule does not change with the destination: cited or empty, then operations, then rate.
If OCR cannot attach a form field or readable passage to a number, drop the value and flag the page. A digit without a cite is how invented limits enter the file.
Occupancy on the ACORD disagrees with the broker email
A mid-market property submission arrives with an ACORD 140 that lists occupancy as office for a suburban building, a statement of values that describes warehouse and light assembly, and a broker email that says the tenant vacated the offices last quarter and the space is now used to store finished goods.
Extraction should not pick a winner. It should emit occupancy from the ACORD with a cite to that form field, occupancy language from the SOV with a page cite, and the broker's sentence with an email cite. The rating occupancy field the PAS will use stays empty, or sits in a conflict queue, until operations decides which source governs and writes one value with that decision recorded.
If the model writes warehouse because the email sounds newer, it has invented occupancy relative to the form of record. If it blends office warehouse so the field is not blank, it has created a value no source stated.
The same pattern applies to limits. An ACORD location limit of $2 million and a cover email asking to quote $5 million pending the engineering report are two cited facts. The rating limit is not $5 million until operations confirms the request and updates the field. Writing $5 million into the PAS because the email mentioned it, then sending that record to rating, is inventing a limit. That is the failure this page exists to stop.
This is an illustration of conflict handling, not a measured before-and-after. Volumes and forms will differ. Cite every claim, leave the PAS field empty while sources disagree, let operations complete.
Never treat the extract as already rated
A structured dump in the PAS looks like a quote file. It is not a quote until a human has cleared blanks and conflicts, and until uncited values are stripped or sourced.
Block auto-rate when any required field is empty or uncited. If year built has no cite, do not let a PAS default fill it. If occupancy is uncited, do not refill from last year's account. Last year is not this pack.
Do not let downstream enrichment overwrite a cited extract. Third-party building attributes can sit beside the submission on external data enrichment at quote. They do not replace a cited ACORD field unless operations chooses to. Mixing enrichment into the same cells as submission extraction erases the audit trail.
Vendor class does not change the control. Whether the rating engine lives in Guidewire, Duck Creek, Applied Epic, or Salesforce, the gate is the same: no rate from uncited fields, no invented limits, no invented occupancy. A workflow that "completes" the record so the API will accept it is the opposite of quality.
Train the queue so extract-complete is not the same status as file-complete. Extract-complete means every source document was read and every field is cited or empty. File-complete means operations has resolved blanks and conflicts. Only file-complete may enter rating.
Operations completes the file, then quoting proceeds
After extraction, operations works a queue of blanks, uncited values, and conflicts. Completing the file can be a broker reply, a clearer scan, a mapping from free-text occupancy to the carrier code list, or a documented underwriting decision in writing. Until that work is done, the submission is not in rating.
Work the queue by opening each cite and confirming the passage. When two cites disagree, record the governing source on the field. When the source is silent, ask for the specific blank, the occupancy field on ACORD 140 Location 1, not a general request for more information.
Once the rating record is cited and complete, appetite and complexity can use those fields without rereading the PDF stack. Unanswered questions still go through conversational intake instead of guessed cells. When the account later binds, the same citation habit applies to bind instruction extraction and system population.
The practitioner test is simple. Open a rated quote and pick a limit and an occupancy. If you cannot land on a form field or an email passage for each, the extract was treated as truth. If a field was empty in the pack and full in the PAS, someone invented it. Fix the gate, not the quote letter.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first