Skip to main content
DoneThat

AI Adoption GuideGovernmentInspect

Violation pattern classifier

ML categorizes findings across inspections into violation typologies, surfacing systemic non-compliance across multiple sites or operators.

Government processPlanFundAuthorizeDeliverInspectEnforceReportClose

By Don, DoneThat’s AI coach · updated

A typology label that cites findings and the codebook

The classifier's quality output is a typology label that names a codebook entry and lists the finding IDs that support it. That is the whole product. It is not a compliance rate, not a ranked list of operators, and not an enforcement notice.

If the finding text does not match a code in the loaded codebook, the typology field stays empty. Empty is the correct quality result. A filled label that cannot point at a finding ID and a codebook row is not a classification; it is a theme someone typed.

An inspections analytics lead uses these cited labels to read across inspections: the same codebook entry appearing under more than one site or operator, each time backed by finding IDs. Pattern language only starts after those cites exist. Without them, "systemic" is a slide title, not an observation.

Inspection and case systems from Accela, Tyler, Palantir, or Microsoft already store findings, permits, and case files. Treat the classifier as a labeling pass over those finding records. Do not treat the vendor record as proof that a typology is true, and do not treat a typology as a substitute for the finding the vendor system already holds.

Load findings and the codebook, then label

Run the pass in a fixed order. Load the finding set first: each row needs a stable finding ID, the inspection it came from, the site or operator, and the inspector's text (or structured check result). Load the codebook second: each row needs a code, the statutory or rule citation the agency actually uses, and the matching language inspectors are trained to write. Do not label against a memory of last year's codes, a slide taxonomy, or a list a model invented.

Then label. For each finding, either attach one codebook entry and keep the finding ID on the label, or leave the typology empty. A typology that covers many inspections is a roll-up of those cited labels, not a new object. If you need a cluster across operators, the cluster still has to enumerate finding IDs and the same codebook entry. If you cannot name both, you do not have a pattern yet.

Findings can arrive from more than one capture path. A computer vision site auditor may emit visual findings that still need the same codebook test as a written checklist item. Vision output is another finding record. It is not a typology and it is not a notice.

Keep the codebook version on the batch. If the rulebook changed mid-quarter, labels from the old codebook and the new codebook are not the same field. Mixing them makes "repeat non-compliance" look like a pattern when it is a version collision.

Empty stays empty when no code matches

Unmatched findings stay unlabeled. Do not pick the nearest code, do not mint a house code, and do not write a free-text typology that sounds official. The statute named the codes in the codebook. If the finding does not match one of those entries, the classifier has nothing to cite.

Empty findings still belong in the extract. They tell you the inspector recorded something the codebook cannot absorb: an observation, a courtesy note, a condition outside the program, or language too vague to map. Dropping them hides coverage gaps. Forcing a code hides the same gaps behind a false typology.

When you brief leadership, report unlabeled rows as unlabeled. Do not fold them into an "other" typology. "Other" is an invented code.

Do not invent a match rate, a systemic rate, or a share of sites "in the typology." The quality outcome is the label (or the empty field), with cites. A percentage is a different product, and this classifier does not produce it.

Failure modes that poison the pattern view

Three failures show up as if they were analysis.

A typology with no finding ID. The label names a code, maybe even a codebook title, but nothing in the finding table can be opened. You cannot defend it in a file review, you cannot show an inspector what was mapped, and you cannot tell whether one bad sentence drove a "pattern" across operators. Discard the label or send it back to the pass. Do not chart it.

Treating the label as a citation. The codebook entry is the citation. The typology label is a pointer to that entry plus the finding IDs. If a later inspection report auto-drafter or a briefing quote uses the typology name as if it were the legal cite, the report is citing a classifier tag. Keep the statutory citation on the codebook row and keep the finding IDs next to the label. The tag is for grouping. It is not what you put in the record as the violation.

Inventing a code the statute never named. Models will offer tidy extra buckets: "general sanitation," "repeat culture of non-compliance," "multi-site risk." If those strings are not codebook entries, they are not labels. Adding them creates a shadow taxonomy that will not survive counsel, a hearing, or a public records request. If the program needs a new code, that is a codebook change with a version, not a one-off label in this batch.

A fourth, quieter failure: using the typology as if it were already linked across programs. Same-operator labels in one program are still one program. Cross-program identity is a separate step, handled by a cross-program violation linker that must cite its own records. Do not collapse two programs into one typology because the operator name looks the same.

A person still files the case

The classifier does not issue a notice, open a case, or close an inspection. After labels land, a person reads the cited findings, confirms the codebook entry, and files (or declines to file) in the case system. Pattern views can queue that review: the same codebook entry on two operators is a reason to open the finding IDs, not a reason to skip them.

When the file does need a notice, that draft is a different artifact. An enforcement notice drafter works from the finding record and the legal citation, not from a typology string. If the notice writer only has a label, stop and recover the finding IDs and the codebook row first.

Illustrative example, not a measured result. An analytics lead loads one week's food-facility findings for two operators. Finding F-4412 is "no hot water at the kitchen hand sink" and maps to codebook entry HW-01, hot water at handwashing, with F-4412 cited on the label. Finding F-4418 is "live cockroaches in dry storage" and maps to PEST-02, with F-4418 cited. Finding F-4420 is "dining-room floor tile worn." The loaded codebook has no worn-tile code for this program, so the typology on F-4420 stays empty. If HW-01 also appears on the second operator, the lead can say the codebook entry HW-01 is present on both operators and list the finding IDs. They cannot publish a sanitation rate, they cannot treat "HW-01" as the statutory text, they cannot add a homemade code for worn tile, and they cannot send a notice from the label. A reviewer still opens F-4412 and F-4418 and files whatever the program requires.

That is the quality bar: a typology that cites finding IDs and a codebook entry, empty when the finding does not match, and a human on the case.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first