Skip to main content
DoneThat

AI Adoption GuideGovernmentFund

Cross-program fraud pattern detector

ML identifies fraud signatures appearing simultaneously across multiple programs and jurisdictions, surfacing coordinated schemes invisible to siloed review.

Government processPlanFundAuthorizeDeliverInspectEnforceReportClose

By Don, DoneThat’s AI coach · updated

What a cited overlap lead actually contains

A cross-program fraud pattern detector is useful only when its output is a lead an investigator can open. That lead names the programs and jurisdictions in the overlap, names the entity as it appears in those extracts, and names the matching field. If the overlap is not there, the page stays empty. The detector does not invent a fraud rate, and it does not substitute for the person who still has to open the case.

Siloed review is the reason this page exists. A housing rehabilitation file and a workforce stipend file can each look ordinary when a grants analyst reviews them in their own queue. Coordinated activity shows up when the same tax identifier, the same payment destination, or the same document fingerprint appears in more than one program at the same time, often across county or state lines that do not share a queue. The detector's job is to surface that coincidence with cites. It is not to declare that coincidence proved.

Treat a numeric score with no matching field as incomplete. A rank tells you the model was confident. It does not tell you which field to pull, which program extract to reopen, or which entity string to search. Investigators cannot work from confidence alone.

Load the program extracts you intend to compare

Load every extract you are willing to cite, and do not compare programs you did not load. The matching step can only speak about rows that are actually in the run. If a second jurisdiction's file never arrived, do not write that jurisdiction into the lead. Inventing overlap across programs you did not load is the fastest way to send investigators into a file that is not there.

Each extract should carry the identifiers the program actually uses: entity legal name and any DBA, tax identifier, award or application ID, award period, draw or payment dates, payee account or address fields your policy allows you to hold, and the document or invoice keys that sit on the payment packet. Align column names before you match. A vendor_id in one extract and a payee_tin in another are not the same field until you document that they are.

Agencies already keep these extracts in casework and analytics stacks from vendors such as Palantir, SAS, Microsoft, and Tyler. Treat those platforms as the systems of record and the places investigators will reopen the file. The detector is a comparison pass over extracts, not a replacement for those systems and not a claim about any one product's fraud module.

Scope the run in writing: which programs, which award years, which jurisdictions, and which entity types. A statewide housing program compared to a two-county workforce extract will only ever be able to cite those two. If you later add a disaster-recovery file, reload and rerun. Do not backfill the earlier lead with programs that were not in the extract list.

Export at the grain you will cite. If investigators open an award, extract award-level rows. If they open a draw, include draw IDs. Keep a run manifest that lists file name, program, jurisdiction, as-of date, row count, and the join keys you will allow. When a file fails validation, exclude it from the match set and record the exclusion. Do not silently match against a partial file and then speak as if the program was fully loaded.

Match on fields an investigator can reopen

Match with cites. For every hit, store the program names, the jurisdiction labels, the entity key as it appeared, and the field name plus the literal value that matched. Prefer fields an investigator can pull from the original packet: tax identifier, shared payee routing, repeated invoice or document number, identical supporting-attachment hash, or a contact identifier that your policy treats as in-scope. Fuzzy name matches without a second field are weak cites. Record them as weak, or drop them.

Pair this pass with neighboring controls instead of asking one model to do every integrity job. Document-level reuse belongs with the document forensics engine. Unusual draw timing inside a single award belongs with the drawdown anomaly monitor. Rule and eligibility collisions across programs are a different object from fraud signatures; keep those on the cross-program violation linker. Recipient history that is about concentration and repeat awards, not a simultaneous signature, belongs with the re-grant risk profiler.

Here is the one illustration. A county housing rehabilitation extract and a neighboring-county workforce stipend extract both contain the same employer identification number. The housing packet lists that number on the contractor payee line. The workforce packet lists it on the training-provider line. The matching field is the EIN. The lead cites Program A in County X, Program B in County Y, the EIN as it appears in both extracts, and the two award IDs. It does not add a third program from memory. It does not estimate how often this happens. An investigator opens both packets, checks whether the payee is in fact the same legal person, and decides whether the overlap is allowed stacking, a reporting error, or something that warrants a case.

If the model emits a score and the matching-field slot is blank, discard or quarantine that row. Do not pass it downstream as a lead. Simultaneous is part of the signature you are looking for. Record the award periods or draw windows that overlap. A shared identifier years apart may still be worth a review list, but it is not the same object as a signature appearing in more than one program in the same window.

Keep empty output empty

No overlap means no lead. Empty stays empty. Do not fill the blank with a similar name, a nearby ZIP code, or a prior year's file you did not load. A blank output is a successful run when the extracts do not share a citable field.

Write that rule into the queue so analysts do not complete empty rows. Downstream dashboards that require a count will pressure people to convert blanks into near-matches. Resist that. A near-match is a different product: a review list with an explicit weakness label, not a fraud pattern lead.

When extracts are incomplete, say so in the run log: missing jurisdiction, truncated award year, dropped document keys. Incomplete extracts shrink the set of claims you are allowed to make. They do not license invented overlap.

Treat the lead as intake, not a finding

The investigator opens the case. The lead is intake: programs, entity, matching field, and links back to the source rows. It is not a finding, not a questioned-cost amount, and not a referral letter. Treating the lead as a finding is a failure mode. It short-circuits corroboration, and it leaves you unable to explain why the overlap was allowed or why it was not.

Open from the cite. Pull the named programs, confirm the entity strings, and confirm the matching field in the original extracts before you request bank records, subpoenas, or a notice to the recipient. If the field does not reproduce, close the lead as not reproduced and leave the fraud question unanswered rather than implied.

Keep a short disposition set that does not pretend to be a fraud taxonomy: opened for investigation, overlap allowed under program rules, extract error, not reproduced, and no overlap. None of those dispositions is a fraud rate. Do not publish a rate from this queue. The detector measures whether a citable overlap appeared in the extracts you loaded. Investigators measure whether that overlap is a scheme.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first