AI Adoption GuideProcurementReview
Maverick spend detection
ML detects spend occurring outside contracted suppliers or approved channels and quantifies the financial impact for CPO reporting.
Procurement processRequestApproveSourceEvaluateSelectOrderReceiveReview
By Don, DoneThat’s AI coach · updated
What maverick spend detection does
Maverick spend is purchasing that bypasses contracted suppliers, preferred channels, or approved catalogs. It shows up as invoices, P-card charges, marketplace orders, or free-text PO lines that do not map cleanly to an active agreement. For a CPO analyst, the job is not to chase every one-off buy. It is to find patterns that leak negotiated value, weaken compliance, and distort category plans.
An ML-assisted review workflow scores transactions and suppliers for off-contract likelihood, then estimates the financial gap versus contracted rates or channel terms where comparable pricing exists. The model proposes candidates and impact ranges. Category managers and finance decide what is truly maverick, what is an allowed exception, and what to remediate.
When spend lines, supplier master data, or contract coverage are missing for a period or category, the system returns empty output for that scope rather than inventing matches or leakage figures.
Signals that distinguish off-contract activity
Reliable detection depends on joining several operational sources: accounts payable invoices, purchase orders, P-card and T&E feeds, supplier master records, and the contract repository (including effective dates, SKUs or service scopes, price lists, and preferred-channel rules). Matching quality matters more than model complexity. Fuzzy supplier names, incomplete item descriptions, and stale contract end dates create false positives if left unchecked.
Useful signals include:
- Supplier not on the contracted or preferred list for the category and business unit
- Item or service description that maps to a contracted catalog SKU or rate card but was bought elsewhere
- Channel markers such as marketplace, spot buy, or non-approved punch-out when a catalog path exists
- Price variance versus contracted unit price for comparable units of measure
- Repeated small buys that avoid PO thresholds yet aggregate above a material threshold
- Buyer or cost-center patterns that consistently route around a named preferred supplier
The model should treat “no active contract in this category” differently from “contract exists and this buy ignored it.” The first case is a coverage gap for sourcing. The second is maverick behavior for compliance and savings recovery.
How the model scores and quantifies impact
Scoring typically combines a classification layer (likely on-contract, likely off-contract, or inconclusive) with an impact estimate. Classification can use supervised labels from prior audit findings, rules as hard constraints (for example, mandatory suppliers for regulated categories), and similarity features between invoice text and catalog or contract language.
Impact quantification should stay conservative. Where a contracted price and comparable unit of measure exist, estimate leakage as quantity times (actual unit price minus contracted unit price), clipped so negative “savings” do not mask true maverick volume. Where only a preferred supplier exists without a clean price benchmark, report volume and count of off-contract lines without fabricating a dollar savings claim. Where match confidence is low, mark impact as unestimated and push the record to human review.
Outputs that help CPO reporting usually include:
- Ranked list of suspect suppliers and categories by estimated leakage and transaction count
- Trend of maverick share of addressable spend over the review period
- Concentration by buyer, cost center, and business unit
- Confidence and data-quality flags per line or supplier cluster
Empty or incomplete inputs (missing AP extract, absent contract effective dates, or unmapped supplier IDs) should yield no ranked list and no portfolio-level leakage total for the affected slice. Partial periods should be labeled as incomplete rather than annualized without disclosure.
Human review and ownership
The model flags. Category and finance act. That split keeps auditability and avoids automated “savings” claims that cannot survive a challenge.
A practical review loop looks like this:
- Analyst opens the ranked queue filtered by category and confidence band.
- Category manager confirms whether the supplier or channel was authorized (exception, dual source, regional gap, or true maverick).
- Finance validates price benchmarks and whether the spend is addressable for reporting.
- Confirmed maverick cases feed remediation: buyer coaching, catalog enablement, contract amendment, or policy escalation.
- Label outcomes return to the training or rules set so the next cycle reduces noise.
Human-in-the-loop is mandatory for policy exceptions, related-party suppliers, emergency buys, and categories with incomplete catalogs. The system should never auto-write off a supplier relationship or auto-post a savings journal entry from a detection score alone.
Reporting for the CPO without overstating savings
CPO packs need three views: where maverick spend sits, how large the addressable gap is under stated assumptions, and what actions closed prior findings. Separate “detected off-contract volume” from “validated leakage versus contract.” Mixing those numbers is how reports lose credibility with finance.
Good reporting hygiene:
- State the spend base (AP only, AP plus P-card, exclude capex, and so on)
- Disclose match confidence and the share of spend left unscored
- Show validated versus estimated impact in different columns
- Track remediation status (open, disputed, closed with catalog fix, closed with policy exception)
- Avoid presenting model output as booked savings until finance agrees the baseline and period
When data for a category or legal entity is missing, omit that slice from totals and state the exclusion. An empty section for “Category X: insufficient contract or spend coverage” is preferable to a zero that looks like perfect compliance.
Failure modes and operating constraints
Common failure modes include treating all non-catalog spend as maverick when the catalog is thin, double-counting the same invoice across PO and non-PO feeds, and scoring expired contracts as active. Another frequent issue is supplier hierarchy: a subsidiary invoice may look off-contract when the parent holds the agreement.
Operate with clear constraints:
- Require minimum fields (supplier identifier or resolvable name, amount, date, and category or GL mapping) before scoring a line
- Gate portfolio KPIs on a documented coverage threshold for contracts and spend feeds
- Recalibrate after major ERP, catalog, or supplier-master changes
- Keep an appeals path so buyers can document legitimate exceptions without fighting a black box
Maverick spend detection earns trust when it is narrow about what it claims, loud about data gaps, and deferential to category and finance judgment on confirmation and remediation.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first