Skip to main content
DoneThat

AI Adoption GuideHealthcareTreat

Clinical trial matching

RAG system matches patient profile against active trial eligibility criteria and surfaces qualifying trials to the care team at point of treatment, using tools like Tempus or TrialSpark.

Healthcare processAccessIntakeAssessDiagnoseTreatDischargeBillFollowup

By Don, DoneThat’s AI coach · updated

A match is a cited pair, not a ranked list

A qualifying trial match is a patient who meets a protocol inclusion row you can point to in the chart, with both cites visible to the care team at the point of treatment. A list of trial names without those two pointers is not a match.

Every criterion you mark as met names the chart field, note, lab, imaging report, or discrete result, and names the protocol inclusion or exclusion row it satisfies. If either cite is missing, that criterion is unmatched. If a required criterion is unmatched, that trial stays off the qualifying list. Empty stays empty. Do not invent eligibility.

Vendors in this class, including Tempus, TrialSpark, Epic, Microsoft, and similar trial-matching or EHR-adjacent tools, can retrieve active protocols and compare a patient profile against eligibility language. They do not enroll anyone. Treat the output as a cited worklist. A research coordinator still screens.

This is the same citation habit as evidence-based guideline retrieval: show the source row. Do not paraphrase it away.

Load the chart and the protocol before any comparison

Matching starts with two loads, then a comparison.

Load the chart as the record actually stands: diagnosis and staging, performance status, dated labs, prior lines of therapy, comorbidities, concomitant medications, and molecular or genomic results that are already resulted. Do not backfill a missing biomarker from a typical disease profile. If the marker is not in the chart, the marker is unknown.

Load the protocol as your site currently runs it: inclusion rows, exclusion rows, washout windows, required tests, cohort status, and the open amendment. A nationally listed study that is closed locally is not an active match for this patient today.

Only then retrieve. RAG here means the model searches eligibility text and chart passages and proposes alignments with pointers. It does not infer a missing lab because similar patients usually have one.

Pathology and genomic inputs should already exist as resulted documents, the same way pathology image analysis and pharmacogenomics-based drug selection leave a cited finding in the record before anyone acts. Trial matching is downstream of those results. It is not a substitute for ordering them.

Watch dates on both loads. A performance status from two years ago does not meet a protocol that requires ECOG at screening. A protocol PDF from a prior amendment does not govern the cohort that is open this week.

Dual-cite every criterion you claim is met

For each inclusion or exclusion the protocol treats as required, the output is a pair. The chart cite is the location of the value (field, document type, and date). The protocol cite is the inclusion or exclusion row, including the threshold or category the protocol uses. Both must be present before you mark the criterion met.

Take one coordinator pass, not a published study. A patient is in clinic with metastatic non-small cell lung cancer. The protocol requires ECOG 0-1, measurable disease, and a documented EGFR exon 19 deletion or L858R. The chart holds an oncology intake with ECOG 1 from last week, a recent scan that describes a target lesion, and a molecular pathology report with EGFR exon 19 deletion. The system may surface that trial with those three chart cites paired to the three protocol rows. If the molecular report is absent, EGFR stays unmatched and the trial does not appear as qualifying. The coordinator can order the test or chase an outside report. The model does not fill EGFR from histology.

Numeric rows need the same pairing. "ANC adequate" is not a cite. "ANC 1.8 from CBC on 12 Mar" against "ANC at least 1.5, inclusion 4.2" is a cite. If the CBC is older than the window the protocol names, the row is unmatched even when the number looks acceptable.

Exclusion rows use the same structure in the other direction. A listed CNS metastasis in a recent note is a chart cite against an exclusion for active CNS disease. Coordinators need the reason a trial dropped, not only the trials that survived.

Leave the field empty when a required row is missing

Unknown is not eligible. Unknown is not ineligible unless the protocol says that the absence of a result excludes the patient.

If a required inclusion has no chart field, leave that criterion empty. Do not write "likely met" or "assumed negative." Do not copy a value from a prior encounter the protocol would not accept. The qualifying list omits any trial that still has a required empty. Trials that failed on a cited exclusion can sit in a separate, labeled failed set so the reason stays inspectable.

The same empty-if-unknown rule shows up in smart referral matching: a suggested destination without the required sending data is not a completed match.

Washouts, line of therapy, and "no prior checkpoint inhibitor" are common empties. If the treatment history is incomplete, the prior-therapy row stays empty. Completing the med list is coordinator work. Generating a plausible history is not.

Local open status belongs here too. If the system cannot cite that the cohort is open at this site under the IRB-approved amendment, the trial is not a local qualifying match.

Coordinator screening is the enrollment gate

The list on the treatment screen is a work queue. It is not enrollment, not consent, and not a held slot.

The coordinator screens every surfaced trial against the live protocol, the source documents, and facts the model does not own: slot availability, competing studies, travel, consent language, and whether the treating clinician agrees the patient should be approached. Tools in this class can shorten the search. They cannot close the screen.

Hand the clinician a short packet: trial identifier, dual cites for required rows, empty rows that still need data, and the next action (order a lab, retrieve an outside genomic report, or do not approach). Keep failed trials visible with their exclusion cites so the team does not re-argue the same no.

Do not treat the vendor UI as enrollment status. Tempus, TrialSpark, Epic, Microsoft, and peer systems may show a match, a score, or a likely-eligible flag. Your source of truth remains the protocol row plus the chart field, confirmed by someone allowed to pre-screen. Discuss with the investigator, approach the patient, consent, and register only after that screen.

Three ways a "match" goes wrong

A match with no inclusion row. The system returns a trial because the disease name overlaps, or because an abstract mentions the tumor type, without pointing to a specific inclusion. Treat that as retrieval. Demand the row or remove the trial from the qualifying list. A confident title is not a criterion.

Treating the list as enrolled. A panel of studies on treatment day is easy to misread as "this patient is on a study" or "we already offered these." Until a coordinator screens, the investigator agrees, and the site registers the patient, the list is candidates only. Do not let orders, billing, or outside-hospital notes assume study status from a match widget.

Inventing a biomarker. The model infers HER2, MSI, a fusion, or a mutation because the cancer type often carries it, or because a note used a nearby term. That invented value can look like a dual cite when the protocol row is real and the chart cite is not. Require the chart cite to resolve to a resulted report. If the report is not there, the biomarker row stays empty.

Stale protocol PDFs, mixed units, and a never-drawn lab treated as negative fail the same way. Dual cites make those errors visible.

The usable output is narrow: qualifying trials with chart-and-protocol pairs, required gaps left empty, and a coordinator still between the list and the patient.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first