Skip to main content
DoneThat

AI Adoption GuideFinanceAudit

Full-population transaction risk scoring

ML scores 100% of journal entries and transactions, replacing sampling-based testing, using tools like MindBridge or EY Helix.

Finance processPlanBudgetInvoiceCollectPayCloseReportAudit

By Don, DoneThat’s AI coach · updated

What belongs in the score

A complete score names the transaction and the risk factor. Round number, unusual account combination, late post, and posting user are acceptable cites. A rank with no factor is not a score. If the journal is ordinary, the field stays empty.

Empty is a valid result. Do not backfill it with a residual-risk percent, a "low" tile, or a coverage claim. Those figures are not what this procedure produces.

The scorer does not finish the test. Internal audit still selects the sample or the exception. The score is a pointer to a journal and a reason to look. It is not a finding, not a control, and not an opinion on residual risk.

Tools in this class (MindBridge, EY Helix, and scoring layers fed by Workday or SAP extracts) all attempt some form of population scoring. Treat them as a class. Differences in dashboards do not change the workpaper rule: cite the document and the factor, or write nothing.

Load the full population before you score

Sampling-based testing starts with a cut. This procedure starts with the extract.

Pull every in-scope journal entry and transaction for the period: posted and reversed, manual and automated, local and intercompany. If the file is already a sample, you are ranking a sample. You are not running full-population scoring.

Write the population definition before the model runs. List legal entities, ledgers, document types, and the date range. If you exclude recurring allocations, payroll cycles, or system-generated revaluations, list those document types as out of scope. Silent drops read as complete coverage later, which they are not.

ERP extracts are the usual source. Workday and SAP are typical. The scoring engine may sit in a specialist analytics platform. Keep the extract query, filters, and row count in the workpaper so a reviewer can rebuild the same population.

If a company code, feeder, or posting period is missing, stop. Scoring a partial file and labeling the run "full population" is how a coverage story gets invented. Completeness of the extract is a precondition, not a result of the model.

Match the extract to the trial balance or to the journal register before scoring. An account reconciliation agent can help confirm that subledger and GL movement tie out. That tie-out supports population completeness. It does not score transaction risk, and the two outputs must not be merged into one residual-risk number.

Cite the factor or leave the field empty

Run the scorer on the loaded population. For every line that trips a named factor, record:

  • journal or document ID
  • the factor (round number, unusual combo, late post, user)
  • enough to open the entry: amount, accounts, posting date, user ID

Ordinary lines stay empty. Empty means no named factor on this journal. It does not mean safe. It does not mean tested. It does not mean zero residual risk.

Scoring without a factor is a failure mode. A 0-100 index, a red-amber-green tile, or a sorted rank with no contributing reason cannot support an exception. If the tool only emits a number, do not paste that number into the workpaper as the rationale. Recover the factor from the model output, or treat the line as unscored and fall back to a documented sample.

Pattern detection that does not produce a cited risk factor belongs with journal entry anomaly detection. Use that when you need unexplained outliers. Use this procedure when every in-scope journal is scored and a non-empty result always names a factor.

Calibrate factor definitions in the program, not in the dashboard after you see the list. "Late" needs the close calendar. "Round number" needs the denomination and whether system allocations are in scope. "Unusual combo" needs the account-pair logic you will accept. "User" needs whether shared accounts and batch IDs count. Vague factors produce vague exceptions.

Month-end journals: round amounts, late posts, unusual pairs

One illustrative pass, not a measured engagement.

A single legal entity at period close. The extract is every posted journal in the month, including reversals. The scorer returns three kinds of non-empty rows: a posting to an unexpected P&L and balance-sheet pair (unusual combo), entries with posting dates after the close cutoff (late post), and even-thousand amounts entered by one user on a weekend (round number plus user). Each of those rows carries the document number and the factor. Routine depreciation, payroll, and allocation journals stay empty.

The audit lead does not treat the vendor's high-risk count as the conclusion. They take the unusual-combo item as an exception, take a judgmental sample of late posts, and schedule a walkthrough of the weekend round-amount entries with the controller. Empty journals enter the sample only if the lead adds them to look at ordinary processing.

No residual-risk percent is written. Nobody records that every transaction was tested. The population was scored. Testing is the investigation of the items the auditor selected.

Selection stays with the auditor

Decide the selection rules in the program, before the ranked list appears.

State which factors always become exceptions (unusual combo above an amount you set in the program). State which factors become a sample frame (late posts, round numbers). State whether empty journals get a small haphazard or random look, and why. If empty journals are not sampled, say the procedure is exception-only, not a substitute for testing ordinary processing.

After selection, investigate the way you would any exception. Inspect the journal. Pull the invoice, approval, or allocation support. Speak with the preparer when the cite is user or late post. A control-evidence retrieval agent can fetch attached support. Retrieval is evidence collection. It does not prove the score and it does not replace the factor cite on the scored line.

Treating the score as the finding is a failure mode. "High risk per model" is not a workpaper conclusion. The finding is what you observed: unauthorized poster, period-end dump into an unexpected pair, related-party combination, or a clean explanation that clears the item. Clearance belongs on the journal, with evidence, not as a lowered score.

When the file is drafted, a workpaper review LLM can check that every scored exception has a cite, a selection reason, and a disposition. Configure it to reject a file that reports a coverage percent the procedure never measured, or that records a risk score with no factor.

Do not convert a score into coverage

Full-population scoring is not a control. It is not continuous assurance by itself. It does not prove the remaining population is free of misstatement.

Inventing a coverage percent is a failure mode. Analytics platforms often display a figure for transactions analyzed. At best that describes extract completeness, and only if the extract was complete. It is not the share of risk addressed, not the share of the audit plan completed, and not residual risk. Do not copy it into the memo that sits next to the opinion.

Do not average scores. Do not roll flagged counts into a residual-risk percent. Do not write that sampling was replaced if you still selected a sample from the scored list. What you replaced is the first cut of the population. You still choose what to investigate.

MindBridge, EY Helix, and ERP-adjacent analytics on Workday or SAP data will differ in screens and in which factors they emphasize. None of that changes the file: name the transaction and the factor, leave ordinary empty, and let the auditor select the sample or exception.

If the tool cannot emit a factor, do not run this use case. Document a sample instead.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first