Skip to main content
DoneThat

AI Adoption GuideProcurementApprove

Approval anomaly flagging

ML detects unusual patterns in the approval queue, such as unusual amounts for a category or split-request attempts, and surfaces them for human review.

Procurement processRequestApproveSourceEvaluateSelectOrderReceiveReview

By Don, DoneThat’s AI coach · updated

A flag parks the row, not a finding

The job is to score each request sitting in the approval queue against your own history, hold the unusual rows, and put them in front of a controls reviewer. The model does not reject. It does not write an audit finding. A score is a reason to look, with the features that fired sitting on the row.

Most flags will be legitimate. A first buy for a newly onboarded plant, a planned overhaul, a calendar-driven renewal: all look unusual against a trailing median. If the first response is accusation, requesters and approvers will route around the queue, and you will lose the control.

The queue is the object. Split-PO detection watches purchase orders after they are cut. PO anomaly detection compares a PO to contract and history before it goes to the supplier. Those run later. If those are the only nets, the stamp on the requisition already happened.

Intake and source-to-pay suites such as Coupa, SAP Ariba, and Zip are where that queue usually lives. GRC workpaper tools such as AuditBoard are where a parked review can become an evidence file. Treat the model as a sidecar on the queue, or as a rules-plus-history layer in the same workflow.

Score four signals on the queue row

Score the open row, and keep the four features visible. A blended risk number that hides why it fired is how a reviewer clicks through.

Unusual amount for the category

Compare the request amount to the distribution for the same category node, company code, and site, over a trailing window long enough to include ordinary seasonality. Flag the high tail, and apply a dollar floor so you are not reviewing stationery that is $40 instead of $25.

Do not use a single global threshold. MRO at a mill and SaaS seats at HQ do not share a tail. If the category node is wrong, the comparison is wrong. That is a taxonomy problem, not an anomaly.

Unusual supplier for the category

Flag a supplier that rarely or never appears in that category at that entity, especially when the request is free-text rather than a catalog or contracted line. This is adjacent to off-contract spend prediction, which tries to catch the same pattern at intake. In the queue you are looking at a row that already routed.

Do not treat every first-time supplier as a hit. New legal entities, newly onboarded vendors with an open qualification file, and contracted panel members the site has not used yet will all look new. Down-weight or suppress those, or you will train reviewers to ignore the queue.

Split-across-days

Look across requests, not only inside one. Same requester or same cost center, same category, similar description or need-by, submitted inside a short window, each amount sitting under a DoA cut while the combined amount would have required a higher stamp.

That cluster is a queue hold. Do not auto-reject either row. Do not write it up as a split PO. The PO may not exist yet. If the cluster later becomes orders, split-PO detection is the order-stage counterpart. Here you stop the stamps until someone looks.

Self-approval patterns

Flag when the requester is the approver, when the named approver reports to the requester, or when a delegate mailbox is acting as both. DoA compliance enforcement answers a different question: is this person on the matrix for this amount. Self-approval can still be matrix-valid and still fail segregation of duties. Hold it.

A row that fires any of these must be ineligible for low-risk auto-approval, even if it sits under the usual dollar and catalog gates. Auto-approve is for boring rows. A flag means the row is no longer boring.

Worked example: two MRO requests under a $10,000 stamp

A split under a manager DoA looks like two ordinary approvals until you look across days. The walkthrough is illustrative, not a measured result.

A maintenance manager at a process plant needs bearings for a planned overhaul. Their manager-level DoA is $10,000. On Tuesday they submit $9,200 to Supplier A, a regular MRO vendor, mill cost center, need-by Friday. On Thursday they submit $8,800 to Supplier B, on the vendor master but unused at this site, same category, same cost center, same need-by. Each row routes to the manager stamp. Combined, the work would have gone to the plant controller. The Tuesday amount also sits in the high tail for mill MRO, which is usually hundreds to low thousands.

The model parks both rows for the procurement controls reviewer:

  • Unusual amount on Tuesday relative to mill MRO history, with the peer set of recent bearings requests.
  • Unusual supplier on Thursday: Supplier B never appeared in this category at this company code.
  • Split-across-days: same requester, same category, 48 hours, combined amount above the $10,000 DoA line that would have applied to one requisition.
  • Self-approval: the workflow's "department manager" is the requester, so both stamps would have been their own if nobody intervened.

The reviewer opens the work order. The overhaul is real, scheduled, and budgeted. The manager split the buy because two vendors quote different SKUs and they thought two requisitions were simpler, not because they were hiding the total. Supplier B is a contracted alternate the site had not used.

The outcome is not a fraud file. The reviewer combines the rows, routes the combined amount to the plant controller, and records that the split was process, not concealment. Treating the score as guilt would have been the failure.

First-time suppliers and month-end spikes will drown the reviewer

A calendar-naive amount model will park insurance renewals, quarter-end software true-ups, and year-end MRO stock-ups every time they hit. Those spikes sit against a trailing median, and they are how the close works. If the action on a flag is to block the row until a small controls team frees it, you will stall legitimate month-end spend and the business will demand the model be switched off.

Park with a service level, not a hard stop that only IT can lift. Allow a named controls owner to release with a reason. Teach the amount model the calendar: compare December to prior Decembers for that category, or exclude known renewal vendors from the tail, before you call the spike unusual.

First-time suppliers fail the same way. A new plant, an acquisition, a newly qualified converter, and a one-time specialist all look like an unusual supplier. Flag every one of them and the reviewer will stop reading. Require a second signal (free-text, off-panel, split, self-approval, or amount tail) before a first-time supplier alone creates a hold, unless the category is one you have already decided is sensitive.

Do not treat a flag as proven fraud. Internal audit cannot hang a finding on a percentile. The flag starts the workpaper. The documents, the cluster, and the interview close it, or they don't.

The reviewer writes the outcome, not the model

Give the controls reviewer the original request, the four feature values, the sibling rows in a split cluster, the DoA line that would have applied to a combined amount, and a short list of typical suppliers and amounts for that category at that site. They choose: release, combine and reroute, return to the requester, or open a workpaper.

Do not send the flag only to the original approver. If the pattern is self-approval or a split under their stamp, they are not a second pair of eyes.

Do not auto-reject. Rejection is a business decision with a named person and a reason the requester can contest. A model score cannot carry that.

If the reviewer opens a workpaper, the evidence is the documents and the cluster, not the percentile. Coupa, SAP Ariba, and Zip can hold the row in the queue. AuditBoard can hold the evidence file. Mixing "the model said fraud" into either place is how you get findings that will not survive a challenge, and a queue nobody believes.

Tune the volume to what the reviewing team can actually work in a week. Run silent against history before you hold live rows, and look at what would have fired at month-end and on first-time vendors. Switch off a signal that never produces a real reroute.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first