Skip to main content
DoneThat

AI Adoption GuideRetailBuy

Purchase Order Anomaly Detector

ML flags purchase orders with unusual quantities, costs, vendor terms, or delivery dates before they create excess inventory or margin leakage.

Retail processPlanBuyPriceStockSellFulfillReturnClear

By Don, DoneThat’s AI coach · updated

What this use case does

A purchase order anomaly detector reviews a draft or pending PO against recent buying history for the same SKU, vendor, category, and location before a buying controller releases it. The model compares quantity, unit cost, payment and freight terms, and requested delivery dates to patterns that usually hold for that item and supplier. When a field or combination of fields falls outside those patterns, the system returns a clear anomaly signal with the fields that drove it, so the controller can investigate before inventory or margin is locked in.

The detector does not rewrite the PO, choose an alternate vendor, or approve spend. It surfaces risk. The buying controller still decides whether to release, amend, or hold the order. That split keeps accountability with the people who own the open-to-buy and the vendor relationship, while the model does the repetitive cross-check that is easy to miss under volume pressure.

If the PO payload is incomplete, or if there is no usable history for the SKU–vendor–location combination, the detector returns empty output. It does not invent a baseline or guess from unrelated categories. Empty output means “not enough evidence to score,” not “no risk.” Controllers treat that as a cue to use the existing manual review path rather than as a clean bill of health.

When buying controllers use it

Controllers use this check at the last gate before release: after the PO is drafted from a buy plan, replenishment suggestion, or vendor quote, and before it posts to the ERP or EDI path. That timing matters because quantity and cost errors are cheap to fix as draft lines and expensive once receipt, accrual, and allocation begin.

Typical triggers include large seasonal buys, new or rarely used vendors, first-time SKUs on an existing vendor, promotions that change pack size or cost, and rush orders with compressed delivery windows. The same review also catches quieter failures: a unit-of-measure mismatch that multiplies quantity, a freight term flip that shifts landed cost, or a delivery date that clusters arrivals into a week the DC cannot absorb.

Related planning and sourcing steps sit upstream. Controllers often arrive here from an Automated Replenishment Buy Plan that proposed the buy, a Supplier Quote Comparison Agent that narrowed cost options, or a Private Label Product Brief Generator when private-label specs drive pack and cost expectations. The anomaly detector does not replace those tools. It checks that the PO that came out of them still looks coherent against live history.

Inputs the detector needs

The primary input is the pending PO: line SKUs, quantities, units of measure, unit costs or extended costs, vendor identifiers, payment and freight terms, requested and promised dates, and ship-to or DC identifiers. Supporting context includes recent closed and open POs for the same SKU and vendor, category-level cost bands where SKU history is thin, and known calendars for promotions or seasonal peaks when those calendars are available in the buying system.

History quality drives signal quality. Stable, high-volume SKUs with consistent pack and cost produce sharper thresholds. Slow movers, new items, and vendors with sparse receipt history produce weaker baselines. In those cases the detector either returns empty output or, when policy allows, returns only soft flags with explicit confidence notes so the controller knows the score rests on thin evidence.

The system should not silently fill missing fields. A PO without cost, without quantity, or without a resolvable vendor key is incomplete for this use case. Empty output is the correct response until buying completes the record.

How anomaly scoring works in practice

Scoring is comparative, not absolute. Quantity is judged against typical order size and recent velocity for that SKU and location, not against a universal “large order” rule. Unit cost is judged against recent invoices and PO costs for the same vendor–SKU pair, and against category peers only when same-SKU history is insufficient and policy permits peer fallback. Terms and dates are scored for sudden changes relative to the vendor’s recent pattern and for delivery clusters that conflict with known capacity or lead-time norms.

Multi-field combinations matter more than single outliers. A slightly high quantity alone may be seasonal. The same quantity with a cost increase, longer payment terms, and an earlier delivery date is a stronger case for hold-and-review. The detector should report the contributing fields in plain language so a controller can verify each one without opening a separate analytics workbook.

False positives are expected and acceptable if they stay reviewable. Controllers prefer a short list of explainable flags over a silent pass. False negatives are costlier: a released PO that later creates excess inventory or margin leakage. Empty output on missing history is preferable to a fabricated “normal” score that would hide those risks.

Human review and release

Human-in-the-loop is mandatory. The model flags anomalies; buying still releases the PO. Controllers use the flag list to confirm intent with the merchant or planner, correct unit-of-measure or cost typos, renegotiate terms, or split delivery dates before release. When the flag is intentional (for example a deliberate forward buy), the controller documents the reason and releases without treating the model as a veto.

Escalation paths stay inside existing buying governance. High-value or multi-DC anomalies go to senior buying or finance review. Vendor-term anomalies may loop in accounts payable or trade management before EDI goes out. The detector does not auto-reject or auto-amend. Its job ends when the anomaly packet is clear enough for a person to act.

After release, closed POs and receipts feed the history used on the next review. That feedback loop is how thresholds stay current as assortments, pack structures, and vendor pricing change. Controllers should expect empty or soft results again when assortment resets wipe comparable history, until enough new POs close to rebuild a baseline.

Outcomes buying teams should expect

When this review is in place, unusual quantities, costs, vendor terms, and delivery dates are caught before they become excess inventory or margin leakage. Controllers spend less time scanning every line the same way and more time on the small set of POs that actually deviate from pattern. Merchants and planners get earlier questions while the buy is still malleable.

The operational success criteria are practical: fewer post-release cost corrections, fewer surprise overstocks tied to oversized POs, clearer audit trails for intentional exceptions, and consistent empty-output behavior when data is missing instead of spurious green lights. The detector does not guarantee perfect buys. It reduces the chance that an anomalous PO ships without a human second look.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first