Workpaper review LLM
LLM checks workpapers for completeness, sign-offs, tickmarks, and reference integrity.
Finance processPlanBudgetInvoiceCollectPayCloseReportAudit
By Don, DoneThat’s AI coach · updated
The hold cites a gap or it stays empty
A workpaper review LLM earns its place when the output is a quality hold you can clear, or silence. The hold names the missing sign-off, the tickmark that does not resolve, or the reference that does not land. If the paper is complete, the output stays empty. Empty is the correct result when there is nothing to hold.
Do not treat a generated comment as a finding. The model does not invent a tickmark gap because the file looks thin, because a similar binder last period had extra marks, or because a template checklist expects a mark this procedure never used. If the paper does not contain the defect, there is no hold.
The reviewer still signs. The model can queue a completeness pass. It cannot close review.
This sits next to other audit AI work that also has to stay conservative. A control-evidence retrieval agent can pull support. SOX control narrative drafting can draft language. Neither job, including this one, should fabricate a control or a gap that the file does not show.
Load the paper the reviewer would open
Start from the same packet a human reviewer would open, not a summary of it.
Load the workpaper itself: the lead sheet or procedure, the exhibits, the tickmark legend, the sign-off block, and the cross-references those items claim. If the binder lives in a workpaper or close system, snapshot the same package the reviewer would click through. Platforms in this class (AuditBoard, Workday, BlackLine, FloQast) hold workpapers, reconciliations, and close files in different shapes. If a tickmark points to an exhibit one click away, that exhibit is in scope. If a sign-off sits on a cover sheet the export dropped, the model will report a missing sign-off that is not missing.
Pin the review standard before you run anything. Name what complete means for this paper: which preparer and reviewer fields must be filled, which tickmarks must resolve to a source in the packet, and which references must resolve to a page, exhibit, or prior-period paper that is actually attached or linked. A cash confirmation paper and a journal-entry walkthrough do not share the same complete list. Do not run a generic workpaper checklist against every file.
Keep the model off the rest of the engagement unless you are checking a reference that points there. Pulling the whole binder in for context is how invented tickmark gaps appear: the model notices a mark used on another paper and reports it missing here.
Completeness checks that name the missing piece
Run a completeness pass, not a quality opinion.
Check sign-offs as fields and dates, not as a sense that someone reviewed this. If the preparer block is blank, the hold cites that block. If the reviewer block is blank on a paper you marked as ready for review, the hold cites that block. If both are filled, do not hold for weak review or a sign-off that looks rushed. That is not completeness.
Check tickmarks against the legend and the source they claim. A tickmark is complete when the mark appears, the legend defines it, and the cited source is in the packet. A hold for a tickmark names the mark, the cell or line it sits on, and the source that is missing or does not match. It does not say the sample is too small. It does not say the tickmark should have been a different letter.
Check reference integrity the same way. A note that says see WP 4.2 is complete only if WP 4.2 is in the packet or the link resolves. A footnote to an exhibit is complete only if that exhibit is attached. A hold for a broken reference cites the exact string and where it fails: dead link, missing exhibit, mismatched page, or a prior-year paper that was not carried forward.
Every hold should be traceable to a location a reviewer can open quickly: sheet, section, cell, exhibit name. If you cannot point at it, you do not have a hold.
One illustrative packet
A bank-confirmation workpaper lists three confirmations, a tickmark legend, and a sign-off block. The preparer has signed. Confirmation 2 is ticked to see Exhibit B. Exhibit A is attached. Exhibit B is not. The completeness pass should hold only that reference: Confirmation 2, tickmark to Exhibit B, exhibit not in the packet. It should not hold the preparer sign-off. It should not invent a missing reviewer sign-off if you loaded an in-process paper and did not ask for a final-file check. It should not hold Confirmations 1 and 3. Attach Exhibit B and rerun: that hold clears. If the paper was already complete, the rerun stays empty.
Clean papers get no hold
Leave complete papers empty.
A clean output is not a score, a looks good comment, or a generated review note that restates the procedure. Those leftovers become fake review evidence. Someone will paste them into the file. Someone will treat them as the review.
The failure mode that looks like diligence is holding a complete paper. The model wants to be useful, so it comments on formatting, suggests another tickmark, or flags that a rollforward could use more explanation. That is invention. Treat a complete paper as no row, no comment, no draft note. If the queue shows a hold, a reviewer should expect a cite. If the queue is blank, the paper goes to human review as complete on the mechanical checks, not as approved.
Inventing a tickmark gap is the same error wearing a completeness label. The legend has four marks. The paper used three because the fourth procedure did not apply. Holding the unused mark as missing creates a finding that wastes review time or, worse, gets cleared by adding a decorative tickmark that does not represent work.
Empty stays empty. Rerun after a real fix, not to squeeze a comment out of a clean file.
The reviewer still signs
The hold is a queue item. It is not the review.
Treating the hold as the review is the failure mode that lasts, because it feels efficient. The model listed three gaps. A senior initials the paper because the gaps were cleared. That skips what review is for: whether the procedure answers the assertion, whether the confirmation actually supports the number, whether the conclusion is the one the evidence earns. Completeness of sign-offs, tickmarks, and references is a gate. It is not professional skepticism.
The reviewer still opens the paper. They still decide if a cited hold is real (the export dropped Exhibit B) or noise (the exhibit sits in a linked reconciliation the snapshot missed). They still decide if a silent pass is enough to proceed to substance. They sign in the same place they always signed.
Use this pass late enough that the paper is meant to be complete, or label in-process papers so a blank reviewer block is not a hold. Use it before partner review so mechanical gaps do not consume partner time. Do not use it as a substitute for the first-level review of whether the work is right.
Related close and testing work has the same split. An account reconciliation agent can surface unexplained items. full-population transaction risk scoring can rank what to look at. A human still owns the conclusion. Same rule here: the LLM may cite a missing sign-off, a tickmark that does not resolve, or a broken reference. It does not sign, and it does not invent a finding so the file looks reviewed.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first