AI Adoption GuideHealthcareBill
Compliance audit risk scoring
LLM plus rules engine scores claims for documentation completeness and audit risk before submission, flagging gap items for the billing team.
Healthcare processAccessIntakeAssessDiagnoseTreatDischargeBillFollowup
By Don, DoneThat’s AI coach · updated
Flag a gap with a cite, or leave the list empty
A pre-submission compliance score is useful only when it points at a missing sentence in the signed note or at a written billing rule the claim fails. If the note already supports the codes and charges, the gap list stays empty. Billing still reviews. The number on the claim is not a hold, and it is not an audit rate.
Capture and coding decide what the claim might contain. This check asks whether the signed record actually backs that packet. Run it after automated charge capture from notes and diagnosis code suggestion have proposed lines, not instead of them.
Assemble the signed note, the claim, and the written rules
Load three inputs as one packet.
The signed clinical note for the encounter is the source of truth. Unsigned drafts, scribe scratch, and after-visit text the clinician did not sign do not count. If the EHR of record (Epic, Oracle Health, or another system) stores addenda for that visit, load those addenda with timestamps. If the note is not signed, stop. Route an unsigned record as a process exception. Do not score it as if it were the legal chart.
The claim is the set of lines that would submit: CPT and HCPCS, modifiers, ICD-10-CM, units, place of service, rendering and billing identifiers, and any condition codes already attached. If a 3M-class encoder, an in-house encoder, or an AKASA-class coding layer already suggested codes, treat those suggestions as candidate lines. They are not proof that the note supports the line.
The written rules are the ones billing compliance actually maintains: adopted LCD and NCD text, internal do-not-bill lists (procedure X requires finding Y in the note), query policy, and the organization's E/M or procedure documentation standard. An LLM without those rules will invent completeness. A rules engine without the note will flag codes it cannot read. Run both against the same packet.
Match identifiers before you score. The note encounter ID, date of service, rendering clinician, and place of service must be the same objects the claim will carry. A perfect cite from yesterday's admission note is still a wrong cite on today's clinic claim.
What a usable gap looks like on one claim
Every flag needs three parts: a cite from the signed note (or a clear statement that the required sentence is absent), the written rule that sentence fails, and an action a biller can take. Drop any flag that cannot fill those three.
Illustrative example, not a measured result. An outpatient follow-up is coded 99214 with E11.65 (type 2 diabetes mellitus with hyperglycemia). The claim also carries a hemoglobin A1c laboratory line. The signed note, in full, is: "DM follow-up. A1c pending. Continue metformin. RTC 3 months."
A usable medical decision-making gap would quote that note and say it does not document problems addressed, data reviewed, or risk in the terms your E/M standard requires. The written rule is that standard, not a guessed contractor selection rate. The action is to query the clinician for the missing elements, or to bill the level the note supports, then re-score.
A usable laboratory gap would quote "A1c pending" and apply the written rule for when a lab may ride this claim. If the rule requires an order and a result in this encounter's record, flag the line. If the lab belongs to a different encounter that already has an order, say that the problem is linking, not missing clinical text.
An unusable gap says "documentation incomplete" or "high audit risk" with no sentence and no rule. Billers cannot open a hole they cannot see. If the model cannot quote the note or name the rule, it outputs nothing for that issue.
Do not attach a made-up chance that a payer medical review or a Recovery Audit Contractor will select the claim. That figure is not in the note. Billing cannot act on it.
Complete claims stay silent
Empty output is the correct result when the signed note supports each code and the written rules are met. Do not force a finding so the score looks active.
Silence is not a waiver. Billing still reviews, especially for new providers, unusual code pairs, and claims near a policy change. Review then starts from a complete packet instead of a search for a missing paragraph.
Re-run the check when capture or coding changes a line. A complete result on yesterday's claim is stale once diagnosis code suggestion or charge capture adds a code the assessment never stated.
Watch for false completeness. The model may quote a sentence that is in the note but does not satisfy the rule, such as "discussed risks" when the policy requires a named risk and a decision. That is still a gap. Cite the sentence, cite the rule, and say why the sentence is not enough.
Keep the score advisory so billing can still release the claim
Do not wire the score to freeze the claim by default. A high score with a cite is a work item. A high score without a cite is a defect in the checker. An automatic hold turns noisy scores into a second, unofficial edit that nobody owns, and it delays clean claims that merely look risky.
Billing owns release. Compliance owns the rule set. The model proposes gaps. When a biller disagrees, they record why: the note was already sufficient, the rule was misapplied, or the wrong encounter was loaded. That disagreement is how the rule file improves.
After this check, a sibling tool may look at payer rejection patterns. That is denial prediction before submission, not a substitute for a documentation cite. If a denial arrives later, denial appeal letter generation should start from the same signed note and the same cited gap history, not from a recycled risk number.
Refuse three failure modes in design review: a gap with no sentence (kill it in QA; if a biller cannot move from the flag to the line in the note, the flag is noise), treating the score as a hold (keep it advisory unless a human-written rule already required that hold, such as a missing signature), and inventing an audit rate (no percents, contractor likelihood, or industry averages on this claim).
Score last, after charges and diagnoses exist
Run this after proposed charges and diagnoses exist. Charge capture from the note can create lines this check must then prove. Diagnosis suggestion can add an ICD-10-CM code the assessment never stated. Score last among those three so you are judging the packet that would actually drop to the payer.
Keep encoder output (3M-class or otherwise) and EHR workqueues (Epic, Oracle Health) as the systems that store notes and proposed codes. Keep autonomous coding vendors in the same class: they speed line creation. Completeness still means a cite in the signed note or a written rule, and empty when there is nothing to flag.
When you stand up the workflow, judge the checker on process quality, not on invented audit outcomes: whether flags include a verbatim cite, whether empty outputs match claims billing already judged complete, how long a biller takes to accept or reject a flag, and how often the wrong encounter was loaded. Those measures tell you if the tool is usable. They are not an audit rate.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first