Skip to main content
DoneThat

AI Adoption GuideHealthcareDiagnose

Lab result anomaly detection

ML flags abnormal patterns across lab panels in clinical context and surfaces actionable deviations to the ordering physician.

Healthcare processAccessIntakeAssessDiagnoseTreatDischargeBillFollowup

By Don, DoneThat’s AI coach · updated

A flag that cannot cite the result is not a quality signal

The quality bar for lab-result anomaly detection is a flag that names the result, the reference range or delta it is judged against, and the context field that made the pattern worth surfacing. If any of those three cites is missing, do not show a flag. An empty output on an ordinary panel is the correct result. A red tile with no result is not.

This workflow sits in diagnose, not in deterioration scoring and not in differential writing. Early warning score automation may use some of the same analytes. Symptom-to-differential generation may mention labs as supporting findings. Neither replaces a panel-level check that answers one question: is this set of results ordinary for this patient at this collection, or is there a deviation the ordering physician should see before they sign.

Do not invent a critical value to make the detector look decisive. Critical notification is a laboratory policy with call trees and read-back. The model does not own that list. The ordering physician still acts on whatever the detector is allowed to show.

Hospital stacks already hold the raw material. Epic and Oracle Health carry the order, the resulted values, and much of the encounter. Microsoft-class analytics layers and Roche-class middleware sit nearer the analyzer and the interface engine. Treat them as a class of sources and surfaces, not as a ranked shortlist. If the stack can color a badge and cannot quote the three cites, it is not ready for this use case.

Load the panel, the applied ranges, and the context field

Load every resulted analyte on the order: value, unit, collection time, result status (final, corrected, pending), and any instrument comment the lab already trusts (hemolysis, lipemia, icterus, dilution, clot). Then load the reference interval that applied when the result was released, not an interval the catalog updated later. If the lab uses delta checks, load the prior result the delta is computed from, with its timestamp, so the flag can cite the pair.

The context field is a short, named set of facts, not a chart dump. Age, sex, pregnancy status when known, care setting (ICU, clinic, dialysis unit), specimen type, and whether the order is a timed series are usually enough to decide whether an out-of-range number is interesting. A potassium of 5.6 mmol/L next to an adult interval is incomplete. The same number with hemolysis marked, or with a dialysis location, or with a prior 5.5 mmol/L from yesterday, is a different object.

Fail closed on a missing range. Do not substitute a textbook interval. Do not mint a critical. Hold the flag until the cite exists. Put that rule in the medical director's release criteria as one sentence: no cite, no flag.

Illustrative path, not a measured case. An outpatient basic metabolic panel returns creatinine 1.9 mg/dL. The adult interval printed on the result is 0.6-1.2 mg/dL. A prior creatinine three weeks earlier was 1.0 mg/dL. Context: 68-year-old, clinic (not dialysis), serum, ACE inhibitor on the active list, not a timed series, no hemolysis comment. The flag cites 1.9 mg/dL, the 0.6-1.2 interval, the delta from 1.0 mg/dL, and those context facts. It does not say acute kidney injury. It does not stage CKD. It does not page a critical the lab did not already define. The physician still decides: repeat, hold the drug, refer, or document why the value is expected.

If that same panel is inside interval, without a delta the lab would have held, and without an instrument comment, the detector writes nothing.

Flag with cites, or write nothing

Ordinary panels stay empty. Empty is not a miss. The miss is a worklist that lights up on every chronic, explained excursion until nobody reads the line.

Define ordinary before you tune a model. Ordinary means each resulted analyte is inside the applied interval, or the excursion is already covered by an in-policy exception the lab recognizes (documented chronic baseline, expected post-procedure shift, age-adjusted interval), and no existing delta-check rule would have held the result. When that is true, write nothing. Do not add a "no anomaly" badge unless the director asked for an audit trail. Physicians do not need a green tile for a normal CBC.

When you do flag, keep the payload to the three cites. Extra textbook prose about what hyponatremia "means" competes with the laboratory's own result comments. Keep the language at the level of the finding, the same way radiology AI triage and flagging should offer a worklist reason rather than a radiology report. This is also closer to adverse event signal detection than to an automated safety-event close: you are pointing at a pattern, not filing the event.

Route the flag to the least interruptive surface that still places the cites in front of the person who ordered the test. Epic, Oracle Health, Microsoft, and Roche-class middleware each expose different inboxes, result-review panes, and interruptive alerts. Interruptive alerts for numbers that are merely out of range are how the workflow dies.

Three failure modes that mimic quality

A red tile with no result. The interface shows "anomaly" or a severity color and does not quote the analyte, the number, the interval or delta, or the context. Physicians click through, find an ordinary panel, and stop trusting the row. Treat missing cites as a release blocker, not as a display bug. If the payload cannot be rendered, do not render the tile.

Treating the flag as a diagnosis. Someone copies the flag into the assessment, or a rule maps a creatinine delta to an AKI problem-list code. That turns a quality surface into a diagnostic act the detector was not released to perform. Keep flag text at the cite. If the service wants a differential, that work belongs with symptom-to-differential generation, a different owner and a different liability.

Inventing a critical. Uncertainty, a missing range, or an abnormal value that is not on the laboratory critical list must not produce a critical-value call. Criticals have notification, read-back, and documentation the lab already runs. An ML score is not that policy. If faster attention is required, hand the case to the existing critical or read-back path. Do not create a parallel "model-critical."

Keep this detector visually separate from early warning score automation. An EWS scores deterioration, often from vitals plus selected labs. Panel anomaly detection is a result-level quality check. If both fire on the same potassium, the physician should see two reasons, not one blended banner.

The ordering physician still acts

The laboratory medical director owns release: what is loaded, what counts as a cite, when empty is mandatory, and which failure modes block go-live. The ordering physician owns the response: repeat the test, change a medication, ignore with documentation, or escalate.

Do not close the loop inside the model. Do not auto-enter an order, auto-hold a drug, or auto-file a diagnosis from the flag. A suggested next step, if operations insist, sits behind a physician action and stays off the flag payload.

Review flagged and unflagged panels on a cadence the director can defend. Sample specifically for red tiles without cites, flags that read like diagnoses, and any invented critical language. Sample the empty outputs too. Those empties are the detector doing the job.

Open a flagged panel and read the flag aloud. If you can say the result, the range or delta, and the context, and you have not called a critical the lab did not already define, the quality outcome held. If you cannot, leave the tile off.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first