AI Adoption GuideGovernmentEnforce
Cross-program violation linker
Graph AI identifies entities with simultaneous active violations across multiple regulatory programs, enabling coordinated enforcement action.
Government processPlanFundAuthorizeDeliverInspectEnforceReportClose
By Don, DoneThat’s AI coach · updated
What a valid link contains
A valid link names both case IDs and the shared entity identifier. If you cannot cite all three, you do not have a link. The row is a coordination aid. It is not a violation finding, not proof of relatedness beyond the identifier, and not a filing in either program.
Empty is the correct output when two open cases do not share an entity key. Do not fill the gap with a parent company, a similar facility name, or a shared mailing address you noticed by eye. Those observations can go in a coordinator comment after human review. They do not belong in the link field.
The linker does not create cases, close cases, or attach charges. Programs keep their own files. Your job with this output is to see whether two live enforcement tracks already point at the same entity so you can later choose whether to coordinate.
A violation pattern classifier describes how a single case behaves over time. It does not establish a cross-program join.
Emit dual cites or leave the row empty
Match only on the shared entity identifier. When two or more active cases from different programs carry the same canonical entity key, write a link that lists every participating case ID and that key. The minimum complete row is two case IDs plus the entity identifier. A row that lists one case ID and an entity, or two case IDs and no entity, is incomplete. Do not ship it as a link.
When an active case has no counterpart in another program, leave the link field empty. When two cases look related to a reviewer but do not share a key, leave the link field empty. Empty means no shared entity in this batch, not a blank waiting to be filled.
One working picture, not a measured result: an air program has open case AIR-1188 against facility key FAC-88021 for a missed stack test. A water program has open case WQ-441 against the same FAC-88021 for an overdue discharge-monitoring package. The linker writes AIR-1188, WQ-441, FAC-88021. That is the whole product. It does not say the violations share a cause, a statute, a deadline, or a recommended penalty. A third open case in a waste program against a different facility key stays unmatched, even if the plants sit on adjacent parcels and share a parent letterhead. The waste row stays empty.
If more than two programs hit the same key, extend the cite list. Do not collapse it into a cluster nickname. Coordinators need the case IDs they will look up, not a synthetic group label.
After emission, a person checks that each cited case is still active and that the entity key is the one each program actually stores. That check is clerical, not investigative. If a cited case has closed, drop it. If dropping it leaves a single remaining case, the link is gone. A one-sided remainder is not a link. Clear the field.
Three ways a row can look finished and still be wrong
A link with one case ID is the most common false complete. It usually appears when a job writes the seed case, fails to find a peer, and still stamps an entity key, or when a second case drops out after a status refresh and nobody clears the field. Treat a single-cite row as empty. Do not send it to a coordinator queue as a cross-program hit.
Treating the graph as a finding is the next failure. A dual-cite link says two open cases share an entity identifier. It does not establish knowledge, intent, a common course of conduct, or a violation of either program's rules. Investigators still prove their own elements. If a coordinator briefs leadership that the graph found coordinated noncompliance, they have overstated the output. The correct brief is that two programs have open cases on the same entity key and that the files have not been joined.
Inventing a related entity is the failure that is hardest to unwind. It shows up when a matcher cannot find a shared key and a reviewer adds an affiliate, a trade name, a site that used to be under the same permit, or a contractor who works at the facility. Those parties may matter in an investigation. They are not the shared identifier. Putting them in the link field creates a join you cannot defend. If you believe an affiliate should be in scope, open that question in the case file. Leave the linker empty until a real shared key exists.
Join is a coordinator decision, not a graph output
A complete dual-cite row is an invitation to talk, not an instruction to merge. Coordinators still decide whether to keep parallel tracks, share a contact calendar, run a joint inspection, or sequence notices. Statute, privilege, penalty policy, and staffing can all argue for staying separate even when the entity key matches.
Use the link to assemble the right people: the case owners, the program counsel who would have to agree on joint process, and whoever controls each docket. Walk the facts in each file on their own terms. If you later need a readiness check, run an evidence sufficiency assessor on each file separately. If you later need notice language, use an enforcement notice drafter after a person decides to notify. Do not skip from a graph row to a combined notice.
When you decide not to join, keep the dual-cite row. Non-join is not a data error. The link still tells the next reviewer that overlap exists and that someone already chose separate tracks. When you decide to join, record that decision in the case systems of record, not only in the linker. The linker does not file either case, and it should not be the only place a joint strategy lives.
Re-run the match when statuses change. A new active case on the same entity key extends the cite list. A closure shrinks it or clears it.
The quality test is narrow: every shipped link cites both (or all) case IDs and the shared entity identifier; every case without a shared entity stays empty; no invented overlap; no filing. If a row fails that test, it is not ready for a coordinator.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first