Skip to main content
DoneThat

AI Adoption GuideGovernmentEnforce

Cross-program violation linker

Graph AI identifies entities with simultaneous active violations across multiple regulatory programs, enabling coordinated enforcement action.

Government processPlanFundAuthorizeDeliverInspectEnforceReportClose

By Don, DoneThat’s AI coach · updated

Load open cases and the identifiers programs already share

Start from active cases only. Closed, stayed, or proposed-but-not-docketed matters do not belong in the match set. Active means the program still treats the matter as open in its system of record: a pending penalty, an outstanding compliance order, an unresolved inspection package, or an open administrative case, using that program's own status codes.

Load, for each active case, the case ID the program uses in correspondence and the entity identifiers that program already stores. Typical keys are a facility ID, a legal-entity ID, a permittee ID, or a registered-operator number. Use the keys the source system already treats as canonical. Do not derive a new key in the linker.

Permitting, inspections, and enforcement records often sit in platforms from vendors such as Palantir, Microsoft, Accela, and Tyler, sometimes next to homegrown case databases. Treat those products as a class of systems of record. The linker reads identifiers they already hold. It does not replace their case files. If two programs use different identifier schemes, map only through an authoritative crosswalk your agency already maintains. If no crosswalk exists, those programs cannot share an entity key, and the linker must not invent one.

Refresh the extract, then freeze the batch you are about to match. A case that closed between extract and review should drop out of the active set before you treat a link as current.

Name collisions are not identifiers. Two "Main Street Plant" rows in different programs are not a shared entity. Two similar EINs with a transposed digit are not a shared entity. Fuzzy name match, address proximity, and officer overlap can inform a later human review. They are not inputs to the link field.

A cross-program fraud pattern detector looks for deceptive structures across programs. Simultaneous open violations are not the same question. Do not feed linker rows into a fraud workflow just because two cases share an entity key.

Emit dual cites or leave the row empty

Match only on the shared entity identifier. When two or more active cases from different programs carry the same canonical entity key, write a link that lists every participating case ID and that key. The minimum complete row is two case IDs plus the entity identifier. A row that lists one case ID and an entity, or two case IDs and no entity, is incomplete. Do not ship it as a link.

When an active case has no counterpart in another program, leave the link field empty. When two cases look related to a reviewer but do not share a key, leave the link field empty. Empty means no shared entity in this batch, not a blank waiting to be filled.

One working picture, not a measured result: an air program has open case AIR-1188 against facility key FAC-88021 for a missed stack test. A water program has open case WQ-441 against the same FAC-88021 for an overdue discharge-monitoring package. The linker writes AIR-1188, WQ-441, FAC-88021. That is the whole product. It does not say the violations share a cause, a statute, a deadline, or a recommended penalty. A third open case in a waste program against a different facility key stays unmatched, even if the plants sit on adjacent parcels and share a parent letterhead. The waste row stays empty.

If more than two programs hit the same key, extend the cite list. Do not collapse it into a cluster nickname. Coordinators need the case IDs they will look up, not a synthetic group label.

After emission, a person checks that each cited case is still active and that the entity key is the one each program actually stores. That check is clerical, not investigative. If a cited case has closed, drop it. If dropping it leaves a single remaining case, the link is gone. A one-sided remainder is not a link. Clear the field.

Three ways a row can look finished and still be wrong

A link with one case ID is the most common false complete. It usually appears when a job writes the seed case, fails to find a peer, and still stamps an entity key, or when a second case drops out after a status refresh and nobody clears the field. Treat a single-cite row as empty. Do not send it to a coordinator queue as a cross-program hit.

Treating the graph as a finding is the next failure. A dual-cite link says two open cases share an entity identifier. It does not establish knowledge, intent, a common course of conduct, or a violation of either program's rules. Investigators still prove their own elements. If a coordinator briefs leadership that the graph found coordinated noncompliance, they have overstated the output. The correct brief is that two programs have open cases on the same entity key and that the files have not been joined.

Inventing a related entity is the failure that is hardest to unwind. It shows up when a matcher cannot find a shared key and a reviewer adds an affiliate, a trade name, a site that used to be under the same permit, or a contractor who works at the facility. Those parties may matter in an investigation. They are not the shared identifier. Putting them in the link field creates a join you cannot defend. If you believe an affiliate should be in scope, open that question in the case file. Leave the linker empty until a real shared key exists.

Join is a coordinator decision, not a graph output

A complete dual-cite row is an invitation to talk, not an instruction to merge. Coordinators still decide whether to keep parallel tracks, share a contact calendar, run a joint inspection, or sequence notices. Statute, privilege, penalty policy, and staffing can all argue for staying separate even when the entity key matches.

Use the link to assemble the right people: the case owners, the program counsel who would have to agree on joint process, and whoever controls each docket. Walk the facts in each file on their own terms. If you later need a readiness check, run an evidence sufficiency assessor on each file separately. If you later need notice language, use an enforcement notice drafter after a person decides to notify. Do not skip from a graph row to a combined notice.

When you decide not to join, keep the dual-cite row. Non-join is not a data error. The link still tells the next reviewer that overlap exists and that someone already chose separate tracks. When you decide to join, record that decision in the case systems of record, not only in the linker. The linker does not file either case, and it should not be the only place a joint strategy lives.

Re-run the match when statuses change. A new active case on the same entity key extends the cite list. A closure shrinks it or clears it.

The quality test is narrow: every shipped link cites both (or all) case IDs and the shared entity identifier; every case without a shared entity stays empty; no invented overlap; no filing. If a row fails that test, it is not ready for a coordinator.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first