Skip to main content
DoneThat

AI Adoption GuideGovernmentFund

Document forensics engine

Computer vision authenticates receipts, invoices, and certificates submitted as grant evidence, detecting fabrication and duplication.

Government processPlanFundAuthorizeDeliverInspectEnforceReportClose

By Don, DoneThat’s AI coach · updated

Cite the hash, the duplicate, or the artifact

A quality hold on grant evidence is useful only when it points at a file hash, a duplicate ID, or a named visual artifact. If the engine cannot cite one of those, the forensics field for that file stays empty. Empty is the right result for an ordinary receipt, invoice, or certificate.

The hold is a pause with a pointer. It is not a finding. You still decide whether the file is authentic, whether a duplicate is allowed, and whether an artifact is fabrication, a scanner quirk, or a phone photo of a crumpled original.

Hashing, duplicate indexes, and computer vision appear across government platforms as a class, including environments that already run Palantir, Microsoft, Adobe, and AWS. Treat that class as producers of cites. Do not treat a dashboard tile as an author of the award record. Do not invent a fabrication percent. A hold without a cite is noise, and a percent without a method is worse.

Hash the bytes, then compare the store

Hash the bytes that arrived. Do not hash a screenshot of the PDF viewer, a generated preview, or a print-to-PDF of a print-to-PDF. Those steps change the digest. You will miss a true duplicate or create a mismatch that is only an artifact of viewing.

Store the digest with the submission, the award, and the evidence type. Compare it to hashes already on this award. Then, where program rules allow, compare it to hashes from other awards. A collision produces a hold that cites the hash and the duplicate ID of the earlier file. That ID is the claim. Stop there.

A matching hash is not fraud. The same invoice PDF can be a legitimate shared cost, a double upload by a hurried applicant, or a second claim for one purchase. The engine cannot tell those apart. You can, once both files and the budget lines are in view.

When hashes differ, do not stop. Cropping, recompression, and a phone photo of paper will break a byte match even when the page is the same document. Structured fields such as invoice number, tax ID, or certificate ID, plus visual comparison, belong after the hash step. They still emit holds with cites, not scores.

File-level collision is next to, not the same as, a cross-program fraud pattern detector. That detector looks across awards for repeating entities and amounts. Forensics asks whether this file is the same bytes or the same page. Run the hash comparison before you treat a second invoice as new evidence. A drawdown anomaly monitor may later show two drawdowns against one purchase. The duplicate ID is the file-level reason those drawdowns belong in the same review.

Name the visual artifact or do not hold

Computer vision on receipts, invoices, and certificates should fire only when it can name what it saw: a cloned date field, repeated noise tiles, letterhead type that does not match the rest of the page, a security pattern that breaks under a copy-move, a QR or barcode that does not decode to the issuer printed on the certificate, line items that do not add to the rendered total.

The hold names the artifact and, when the interface allows, boxes the region. You look at that region. If the model cannot name the artifact, the output stays empty even if an overlay looks busy. A red tile with no artifact is not a cite.

Mess is not fabrication. Thermal paper fades. Staples punch holes. Phone photos skew and catch glare. Those files are ordinary. An artifact is a specific irregularity you could put in a letter to the applicant without waving at a heatmap.

Ordinary files stay empty

Most evidence will not collide and will not show a named defect. Leave those rows blank. Writing "looks authentic" or parking a confidence number on an ordinary file creates a record the next reviewer will over-read.

Forensics is not completeness. An evidence sufficiency assessor asks whether the packet has the required items. A complete packet can still contain a duplicate file. An incomplete packet can contain a perfectly ordinary scan. Keep the empty forensics field empty so sufficiency and authenticity do not collapse into one light.

A compliance pre-check generator can still fail a file that hashed cleanly if it is the wrong document type or missing an attestation. A clean digest does not mean the file satisfies the notice of funding opportunity.

Failure modes that waste the queue

Three mistakes show up when a forensics overlay is treated as the decision.

A red tile with no artifact. The overlay is alarming. There is no hash collision, no duplicate ID, and no sentence that names a defect. If you bounce the file on color, you reject legitimate messy scans and you have nothing the applicant can answer. Require a contestable sentence: which hash, which prior file ID, which region, and what was seen there. If you cannot write that sentence, clear the overlay and leave the field empty.

Treating the hold as a finding. Copying "fabricated" into the award record because the engine held the file skips the determination you are staffed to make. Close the hold with your language: authentic, duplicate allowed, duplicate not allowed, artifact explained by the original, or refer. The hold is a routing label.

Inventing that a receipt is fake because it looks messy. Mess is not a cite. If the vendor is readable, the date is readable, and the hash is unique, the file is ordinary unless you can point at a specific irregularity.

Do not let any of those three replace reading the narrative. Two identical PDFs can be fine. Two identical PDFs paired with two serial numbers in two narratives are not. The engine cites the duplicate. You read the story.

One collision, two paths for the reviewer

Illustrative path, not a measured case. Two awards in the same round each attach a supplier invoice PDF. The engine hashes both. The digests match. It emits a quality hold that cites the hash and the duplicate ID of the file already stored on the first award. It does not say the invoice is fake. It does not assign a fabrication rate.

You open both packets. On the first award the invoice supports a generator already in the budget. On the second, the same bytes are offered as proof of a second generator, and the narrative names a different serial. The hash tells you the files are the same. The narrative tells you they cannot both be true without explanation. You pause the second payment path, ask for an invoice that matches the second serial, and you leave fabrication language out of the file until you have a determination.

If instead the hashes match because one purchase is a documented shared cost across two line items, you clear the hold and note the share. The engine did the same work in both branches. You decided.

When a visual hold cites a cloned date on a training certificate, you still decide. Request issuer confirmation or a replacement scan. If the supposed clone is a JPEG block boundary on a zoomed photo and the date is consistent with the rest of the page, clear the hold. The file returns to ordinary. Empty stays empty.

Work the loop the same way every time: hash and compare, flag only with a cite, leave ordinary empty, reviewer decides.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first