Skip to main content
DoneThat

AI Adoption GuideFinanceClose

Journal entry anomaly detection

ML flags journal entries by round numbers, late posting, unusual GL combinations, and user behavior.

Finance processPlanBudgetInvoiceCollectPayCloseReportAudit

By Don, DoneThat’s AI coach · updated

What the reviewer should receive

A useful flag names the journal entry and the pattern that triggered it. The cite has to be specific enough that a controller or internal-audit lead can open the document, see why it was raised, and decide whether to ask a question, request support, or do nothing.

Round number, late post, unusual GL combination, and user behavior are the four patterns in scope. If none of those patterns is present, the field stays empty. Empty is a completed result, not a miss.

The model does not post, reverse, or park the entry. Accounting still owns those actions. The model also does not invent a risk score or a percentage of anomalousness. A number that cannot be traced to a documented rule or threshold is not a quality output. Reviewers start treating it as authority, and the close then depends on a figure nobody can explain.

Signals that earn a cite

Round amounts are a frequent close signal, but they are not a fraud label. Payroll accruals, rent, and board-approved bonuses often land on round thousands because the contract is round. Flag the amount and the fact that it is round. Do not write "likely fraudulent" because the debit is round.

Late posting means the entry hit after the cutoff you defined for that entity and period. Cutoff is a policy choice. A model that uses a single evening hour for every book will flood shared-service centers that post in a later time zone. Cite the timestamp against the entity calendar, not against a generic clock.

Unusual GL combinations are pairs or sets of accounts that rarely appear together for that entity, book, or process. A revenue account clearing through a rarely used suspense GL is a combination worth citing. A standard AR, revenue, and tax posting is not unusual just because the dollar amount is large. The cite should name the accounts.

User behavior is about who posted relative to their normal role. A contractor who last posted two periods ago and now books a late true-up is a user pattern. A senior accountant who posts the same recurring close entries they always post is not. Cite the user and the departure: new poster, rare poster, or poster outside the process team. Do not cite "high-risk user" without the behavior.

These four signals stack. One signal on an otherwise ordinary entry can still be a valid flag if the cite is honest. Four signals on a documented recurring accrual can still resolve to leave the entry. The stack changes how fast you review, not whether accounting must reverse.

Coding quality upstream changes how noisy combination flags become. If posters guess at accounts, GL coding suggestion is the earlier control. Anomaly detection after the fact is a safety net, not a substitute for getting the code right on the first pass.

Running the pass on posted and draft entries

Load both posted journals and drafts that already have accounts, amounts, a poster, and a timestamp. Drafts without a GL pair or a user are not ready; incomplete rows produce invented combinations. Exclude auto-reversing close entries and system allocations only from an explicit list. Hiding them by default is how late manual true-ups get mixed into "the system always does this."

Run the pass twice: once when the draft book is mostly complete, so reviewers can still ask questions before lock, and once after the last manual post, so nothing late slips through uncited. Use the same rules both times. Changing thresholds between runs makes the second pass incomparable.

For each entry, emit either a cite or emptiness. A cite lists the document number or unique JE ID, the period and entity, and each matched pattern in plain language. Ordinary entries stay blank. Do not write "no issues found" on every row.

The reviewer acts. They open the cited entries, pull the support, and either accept the entry, request a correction, or ask accounting to reverse and rebook. Internal audit can sample the cited set and a small slice of blanks to test whether emptiness is being earned.

If the same unusual combination also breaks a reconciling item, send the reviewer to the recon pack rather than duplicating the investigation. The account reconciliation agent is the place that already holds the unexplained difference. The JE flag should point at the entry, not recreate the recon.

Illustrative path: Entity A's day-4 book includes JE 8841, a round hundred-thousand debit to a rarely used suspense account and a credit to product revenue, posted late in the evening by a contractor after the local cutoff, memo "true-up." The flag cites JE 8841 and four patterns: round amount, late post, unusual GL combination, user. The controller opens the entry, finds no calculation and no approval mail, and asks accounting to reverse. Accounting posts the reverse. The flag did not post it. If the contractor had attached the signed true-up schedule and the pair was the documented suspense path for that product line, the same four cites would still appear, the reviewer would note them, and the entry would stay. The difference is the support, not a score.

Failure modes that wreck the review

Flagging every round number as fraud is the fastest way to bury the pass. Recurring rent, rounded bonuses, and policy-driven accruals will dominate the queue. Reviewers then learn that round means noise. Keep the round-number cite, and look at process context. If you need less volume, restrict round-number flags to accounts or posters that are not already on a recurring template, rather than labeling the amount as fraudulent.

Treating the flag as a reverse is an operational failure. Close calendars assume humans post and reverse. If a reviewer accepts the model by reversing without reading support, you will reverse legitimate true-ups and skip the ones that needed a reverse but looked ordinary in the memo. The flag is a cite. The reverse is an accounting document with an author.

Inventing a risk percent is a documentation failure that becomes a control failure. A score nobody can recompute will not survive an auditor's how this was calculated question. If you cannot show the inputs, do not show the number. If you want ranking across a full population of transactions rather than a cite on a journal, that is a different job: full-population transaction risk scoring still needs a defined method. It is not a license to stamp a percent on a JE because the slide looks complete.

Other practical failure modes: using a single global cutoff for every entity; treating first-time posters in a new shared-service wave as inherently suspicious without a transition window; feeding unposted scratch journals that still have placeholder accounts; and storing flags in a side spreadsheet that never makes it into the workpaper. If the cite will be reviewed, it belongs in the pack the workpaper review LLM and the human reviewer will actually open.

What ERP and close platforms already cover

Finance teams already see pieces of this work in audit-analytics products and in ERP and close suites. MindBridge, Workday, SAP, and BlackLine sit in that class: some emphasize journal analytics, some the ERP record, some close tasks and reconciliations.

Whatever you use, the output still has to name the entry and the pattern, leave ordinary journals empty, and leave posting and reversing with accounting. If the platform only dumps a ranked list of journals with unexplained scores, you still owe the cite in the workpaper. If it already can tag late posts or unusual account pairs, use those tags as the cite rather than adding a second, undocumented score on top.

Do not buy a second tool because the first one does not detect fraud. Fraud determination is not the job. The job is a reviewable exception on a journal entry during close, with a human still deciding whether the entry stands.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first