Skip to main content
DoneThat

AI Adoption GuideHealthcareDiagnose

Diagnosis code suggestion

NLP reads the diagnostic impression from clinical notes and suggests ICD-11 codes for physician sign-off, flagging underdocumented specificity.

Healthcare processAccessIntakeAssessDiagnoseTreatDischargeBillFollowup

By Don, DoneThat’s AI coach · updated

A suggestion is unsigned until the physician signs

A diagnosis code suggestion is a candidate ICD-11 code plus the note sentence that is supposed to support it. It is not the encounter diagnosis. The physician still signs the code. Until that signature, CDI, coding, quality abstracting, and anyone estimating a working DRG should treat the field as a proposal.

A suggested code that cites the note sentence is useful. A blank that stays blank because the note is silent is also useful. A diagnosis the note does not state is not a suggestion. It is fabrication. Empty stays empty if the note is silent. Do not invent a diagnosis.

The failure that looks operationally convenient is treating the suggestion as already coded. A sidebar in the chart, an encoder pane, or a coding workbench can display a code as if it were done. Someone copies it into a query, a mortality review, or a working grouper. The attending later rejects it, or never sees it. You now have downstream artifacts that assert a condition the legal record does not contain. Reverse those artifacts. Do not confirm them against the model.

Epic, Oracle Health, 3M-class encoders, and AKASA-class coding automation all hold working lists of this kind. Group them as suggestion queues, not as the coded record. The vendor UI does not change the attestation rule. If your policy lets a coder accept on the physician's behalf, that is a separate control issue. This workflow assumes the physician is the signer.

Load the impression, then suggest only with a cite

Load the diagnostic impression first: the labeled impression, the assessment diagnostic sentences, or the closing diagnostic language your template actually uses. That span is what NLP may read for code candidates. HPI, ROS, labs, imaging, and the problem list are context for later flags. They are not diagnoses for this encounter unless the impression restates them as such.

If you ingest the whole note as if it were the impression, two leaks appear. Historical problem-list rows become this-visit codes. Listed possibilities from symptom-to-differential generation become confirmed diagnoses. Both invent a diagnosis the impression never made.

Run the how-to in this order:

  1. Isolate the impression span. If the template has no impression heading, take only the diagnostic sentences in the assessment, not the entire plan.
  2. Generate candidate ICD-11 codes from that span only.
  3. Attach a cite to each candidate: the verbatim sentence or contiguous clause that supports that code, shown next to the code.
  4. Drop any candidate that has no cite. A code with no sentence does not ship, not even as a low-confidence row.
  5. Hand the remaining list to the physician as unsigned suggestions.

A code with no sentence is the first failure mode to kill in review. It usually means the engine matched a string elsewhere in the chart, or generalized from similar notes, or read a problem list. There is nothing in the impression to show the physician. Delete it. If clinicians still believe the condition belongs on the encounter, query for language. Do not keep the orphan code as a hint.

Spoken documentation does not bypass this. If the impression was dictated through ambient clinical documentation and never written, you do not have a cite. Fix the note. Do not code from audio the chart does not contain.

Blanks and specificity flags are the correct incomplete outputs

When the impression does not mention a condition, leave that suggestion field blank. Do not backfill from an echo, a problem list, a prior discharge, or a habit of assigning a code to a typical presentation. A blank is a clinical documentation signal: query, or omit. It is not an invitation to the model to complete the thought.

Underdocumented specificity is the other incomplete output you should keep. The sentence may support a parent ICD-11 stem and not the child that would need acuity, etiology, laterality, episode type, or severity. Emit the parent if the cite carries it. Flag the missing axis in plain language. Do not pick the more specific child to look complete.

Inventing specificity is the next failure mode. The echo reports an ejection fraction. The impression says heart failure. Upgrading to a systolic or HFrEF code because the report exists is a documentation change the physician did not make. Either they add diagnostic language and you rerun suggestion, or the suggestion stays at the level the impression states. Reading the report as if it were the impression is how unspecified language becomes a precise code with no sentence that actually says it.

CDI queries belong on the missing language, not on the unsigned code. A query that the impression states acute heart failure without type or etiology, and asks for clarification if known, points at the note. A request to accept the suggested code points at the model. Only the first is a documentation query.

When the impression names the syndrome and nothing else

This is a constructed walk-through of the workflow, not a measured case.

Impression: Acute heart failure, volume overload. Continue IV diuresis. Echo pending.

Load that impression. Suggest only what those diagnostic words support, cited to the clause on acute heart failure and volume overload. Flag what the classification still needs and the sentence does not give: heart-failure type, chronicity (acute on chronic versus first presentation), and etiology. Leave those child codes blank. Do not emit HFrEF. Do not emit ischemic cardiomyopathy. Do not emit a post-echo etiology. The echo is pending. Filling those axes invents a diagnosis.

If a later addendum states that echo shows reduced ejection fraction, and that this is acute on chronic systolic heart failure, likely ischemic, reload the impression. Cite the new sentences. Only then offer the more specific ICD-11 children those sentences support. The first pass stays limited. The addendum is new diagnostic language, not a license to have been more specific the day before.

If instead the engine had emitted a precise systolic-ischemic code on day one, with no sentence, that row would be a code with no cite. Strike it even if the echo later agrees. Agreement after the fact does not create a cite that did not exist when the suggestion was made. Recode from the addendum.

After the signature, stop treating the queue as the record

Physician sign-off is the only transition from suggestion to diagnosis. Accept, edit, or reject each row with the cite visible. Accepted rows become the signed list. Rejected rows are gone. Edited rows are the physician's wording and code choice.

Until that action, do not feed suggestions into billing, quality abstracts, or CDI text that asserts the condition. After that action, stop maintaining a parallel suggested-codes column that can diverge from what was signed. The queue is a worklist. The signed list is the record.

Keep diagnosis suggestion out of charge acceptance. If notes are also mined for procedures or visit level, that is automated charge capture from notes. One combined accept click is how a rejected diagnosis still generates a charge.

Suggestions that leak into the chart without a cite, or at a specificity the impression never stated, are unsupported codes. Treat that leakage as you would any unsupported coded condition in compliance audit risk scoring. The scoring question is the same as the suggestion rule: is there a sentence, and did a physician sign.

Practical sequence, without a second system of record: load the impression, suggest only with cites, leave blanks when the note is silent, flag missing specificity without filling it, and treat physician signature as the coded diagnosis. Everything else is a draft.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first