Skip to main content
DoneThat

AI Adoption GuideHealthcareFollowup

Care plan adherence monitoring

NLP plus classification tracks patient-reported adherence data from structured digital check-ins and flags deviations for nurse triage.

Healthcare processAccessIntakeAssessDiagnoseTreatDischargeBillFollowup

By Don, DoneThat’s AI coach · updated

Start with the locked plan, not the flag

Open the locked care plan before you look at a classifier output. The plan is the only legitimate comparison point. A check-in answer is a report, not a finding of non-adherence, until you have compared it with that plan and decided what the words mean for this patient.

Ambulatory follow-up lives in the days between visits. Structured digital check-ins fill that interval with scheduled questions: medication taken or not, diet and activity items, symptom scales, and a free-text box. NLP plus classification can map those fields onto the locked plan and put mismatches on a nurse list. The quality bar is narrow. A deviation names the mismatch and cites the check-in answer that produced it. If the submitted answers are consistent with the plan, the list stays empty. You do not invent a problem to fill the queue. You still triage every flag that appears.

Clinics already store the locked plan and the inbound check-in in products from vendors such as Epic, Oracle Health, Microsoft, and Biofourmis. Use those systems as the source for plan text and the payload the patient submitted. Do not treat a vendor dashboard as a substitute for a cited answer. The model does not rewrite the plan and does not add events the patient never reported.

Load the plan and the check-ins together

Work in one pass so you never score a check-in against memory.

Pull the locked plan as it stood when the check-in window opened. If the plan was amended after the patient submitted, do not score the old answers against the new instructions. Treat that as a version conflict in the chart, not as non-adherence.

Pull every structured check-in in that window: coded answers, numeric fields, and free text. Classification should see the same payload you can open. If a field is blank, it is blank. Do not impute a missed dose or diet breach from a blank unless the plan and questionnaire defined non-response that way.

Compare, do not narrate. For each plan element the questionnaire actually asks about, decide whether the submitted answer is consistent with that element. When it is not, write a deviation that includes a cite: the question label plus the patient's answer, or a short excerpt of free text. When it is consistent, write nothing for that element.

Leave the on-plan remainder empty. A clean check-in should produce an empty list for this workflow, even if you worry about the patient for other reasons. Other signals belong in their own queues, including remote patient monitoring with escalation, not in a fabricated adherence event.

Hand non-empty results to nurse triage. The classifier stops at cite and flag. It does not call the patient, change the plan, or document non-adherence in the legal record.

A deviation is a cited mismatch

If you cannot point to the check-in answer, you do not have a quality-grade flag.

NLP helps on the free-text box and on messy short answers such as "took it late", "out of town", or "ate out". Classification helps on coded items: yes or no, ordinal scales, and numeric thresholds written into the plan. Together they should name the plan element, why it looks off, and the exact answer. That cite is what makes the work auditable. Without it, a later reviewer cannot tell whether the model reacted to the patient or to its own prior.

Here is a single walk-through, not a measured case. A locked heart-failure plan in clinic follow-up includes daily weight, loop diuretic as prescribed, a two-gram sodium target, and instructions to report rapid swelling. The morning check-in shows a weight inside the agreed band and a yes on taking the diuretic. In free text the patient writes, "I ate soup from a can last night and my ankles look puffy." A valid flag names the diet and swelling items and cites that sentence. It does not add a missed water pill. The nurse reads the cite, calls, asks about the soup and the ankles, and documents what the patient confirms. If that same morning the check-in had been weight in band, diuretic taken, and "no swelling, regular meals," the list for this workflow stays empty.

When the patient actually reported a missed or delayed dose, outreach is a different job. Confirm first, then use medication adherence outreach agent. Do not auto-enroll someone from an uncited flag.

Empty is the correct on-plan result

This workflow fails if you hunt until something appears. The quality standard is not "find a problem." It is "do not invent non-adherence."

If coded answers match the plan and free text does not contradict them, there is no deviation to file. An empty list means the check-ins were on plan, not that the model failed to run. Lowering the threshold until the queue is never quiet trains everyone to ignore it.

Longitudinal risk is not this flag. Models that estimate worsening chronic disease belong in chronic disease trajectory modeling. A rising risk score without a cited check-in mismatch is not care-plan non-adherence. Keep the lists apart so a probabilistic trend cannot be charted as a skipped medication.

State a silent-field rule in one sentence you can defend. An honest rule: if the weight question is unanswered, do not create a diuretic deviation. A dishonest rule: if anything is blank, assume they are off plan. Blank is not a confession.

Do not ship flags that invent the story

Three failure modes show up in triage. Each one breaks the quality outcome.

The first is a flag with no check-in cite. The list says "possible non-adherence" or "diet risk" and points at the encounter, not at an answer. You cannot verify it. Reject it like a lab with no specimen ID. Do not call the patient to "confirm the AI." Call only after you can read the words or values you are confirming.

The second is treating the flag as non-adherence. Classification is a queue ticket. Non-adherence is a clinical, and often legal, characterization. If you copy the flag into the problem list or a quality registry without reviewing the cited answer with the patient, you have turned a hypothesis into a fact. The cite exists so you can ask a precise question: "You wrote that you ate canned soup and your ankles look puffy. What happened after that?"

The third is inventing a missed dose. The patient reported a diet slip, and the model adds "likely skipped furosemide" because swelling and missed diuretic often travel together in training text. Or the medication item is blank, and the model fills "no." Or yesterday's missed dose is replayed onto today's complete check-in. None of that is in the payload. If you need more history, look it up or ask. Do not chart a pill the patient never reported.

Plan drift is a cousin of invention. If education materials were rewritten for literacy, do not score old answers against new wording. That is an instruction change, not a behavior change. Keep wording work in patient instruction personalization instead of creating extra deviations.

Triage the cite, then decide

You remain the decision-maker. Read the locked plan element. Read the cited answer. Decide whether the mismatch is real, misunderstood, already resolved, or a questionnaire defect such as wrong language or someone else completing the form.

If the cite supports a real deviation, the next steps are ordinary nursing: clarify with the patient, assess risk, coach, involve pharmacy or the ordering clinician, and document what you confirmed.

If the cite does not support the flag, close it as false or uncited and say why in the feedback channel your program uses. If the queue was empty and the patient later reports skipped doses they never entered, that is under-reporting. Fix the questionnaire or the visit conversation rather than asking the model to invent doses.

The working habit is small. Load the locked plan and the check-ins. Flag only deviations you can cite. Leave on-plan check-ins empty. Triage every flag yourself.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first