AI Adoption GuideProcurementRequest
Spend category auto-classification
ML classifies each request into a procurement taxonomy, such as UNSPSC, using request text and historical data.
Procurement processRequestApproveSourceEvaluateSelectOrderReceiveReview
By Don, DoneThat’s AI coach · updated
A suggestion is not a recode
Auto-classification should return three things on every request: a suggested category, a confidence, and a reason a buyer can reject. It should not silently overwrite the code that will ride into approval, the purchase order, and the spend cube.
A wrong node is not a cosmetic label. It chooses which buyer sees the request, which contract the policy pre-check at intake thinks applies, and which roll-up spend analytics and savings tracking will treat as last year's baseline. IT hardware filed as office supplies makes office supplies look large, IT hardware look small, and every later savings number in those families untrustworthy.
Speed, here, means the requester and the category owner do not hunt the tree for a leaf they cannot name. The model proposes UNSPSC or your internal equivalent. A named human still owns the taxonomy: which nodes exist, how deep they go, and which families a model is not allowed to settle.
This job usually sits next to intake and spend cubes in suites such as Coupa, SAP Ariba, and Jaggaer, and in spend-analytics platforms such as Sievo. None of them remove the need for a suggestion, a confidence, and a person who can say no.
Suggest from the request text and from history
Classify at intake from the text the requester wrote and from labels a buyer already accepted.
Start with the description, item lines, and any supplier named on the request. If intake is a sentence rather than a catalog pick, run free-text request parsing first so you classify nouns and quantities, not a paragraph. Requester org, cost center, and amount are supporting signals. They are not the category. A laptop can hit an IT expense account and still be coded as office supplies in the procurement tree.
Historical requisitions and POs are the training labels. If those labels mix IT hardware and office supplies on the same laptop wording, the model will learn the mix. Clean a sample of high-volume families before you train. Do not publish a clean cube by letting the model guess the rest.
Predict at the depth your reporting uses. If category managers report at UNSPSC family, do not force a commodity code so the field looks complete. A family with a reason is usable. A guessed eight-digit leaf is fake precision.
Set a confidence threshold the category owner chose. Above it, show the suggested node and the reason (matched wording, similar historical buys, or the supplier's usual family) and let the buyer confirm in one action. Below it, do not write a leaf. Route the request to the buyer who owns that part of the tree, with the model's top suggestions still visible as suggestions.
Persist two values: the stored category the process will use, and the model suggestion. If they differ, that disagreement is a work item.
One laptop filed as office supplies
A facilities coordinator types "laptop and docking station for a new hire, same as last time" and picks office supplies because that is the punchout they always open. The following is illustrative, not a measured result.
The model should read those nouns, look at prior buys with that wording, and suggest information technology (UNSPSC segment 43), not office equipment and supplies (segment 44). The reason should name the nouns and the historical IT hardware POs, not "similar requests." Confidence should drop if this coordinator's only labeled history is office supplies, or if the description is "computer supplies" with no model name.
If the form keeps the requester's office-supplies pick because the field was already filled, the suggestion is decoration. The laptop spends as stationery. A demand aggregation signal keyed off the stored code will not cluster it with other IT hardware. The IT category manager will not see a consolidation conversation.
If you silently recode it to IT hardware and then tell finance that IT hardware spend appeared this quarter, you invented a category shift. The coordinator did not change behavior. You changed the label.
The requester's pick still wins if the form keeps it
Most intake forms already have a category before the model runs. Punchout defaults, a familiar family, or a cost-center GL map fill it. That stored pick is what approvals, routing, and the cube will use unless you change the write path.
If the model suggests IT hardware and the stored field remains office supplies, you annotated the request. You did not classify it. Downstream jobs that key off the stored code, including policy pre-check at intake, will still treat a laptop as stationery.
Decide who may change the stored category: the requester with the suggestion in front of them, the buyer on the low-confidence queue, or a taxonomy owner on a recode queue. Keep the requester's original pick; you will need it when someone asks why the cube moved.
GL mapping is not a taxonomy. Evaluate the model against the procurement tree you report, not against the account string.
A guessed code is not a savings baseline
Savings reports assume last year's category is true. If you backfill a model's guesses into history, this year's same-family price or volume is compared to a fiction.
That failure shows up in two places. Spend analytics and savings tracking will treat a recode from office supplies to IT hardware as a mix shift or a sourcing win. It is neither. Market price benchmarking will look up the wrong peer set: a laptop priced as a stapler, or a stapler priced as a laptop, depending which way the guess went.
Keep a recode ledger separate from the savings numerator. When a buyer or taxonomy owner changes a node, record the old code, the new code, who changed it, and whether the change is a label correction or a different thing bought. Only human-accepted codes, including historical rows you are willing to defend, belong in the baseline you publish.
Do not let a low-confidence guess become the "should have been" category in a business case. If the model is unsure, the baseline is unknown. Say that.
If you classify to UNSPSC and then map into an internal L1-L3 tree, the map is another place a guess becomes a published number. Own the map the same way you own the nodes.
Low confidence goes to a buyer; rejects go in a log
Low confidence is a queue, not a quieter suggestion. The owner of that queue is the buyer or category manager for the families in play, not a generic intake clerk clearing a list.
Show them the request text, the top suggestions with reasons, and the requester's original pick. They accept the suggestion, choose a different node, or send the request back for a usable description. Capital versus expense, dual-use items, and anything with tax or regulatory meaning stay on that queue even when confidence looks high.
Every reject belongs in a log: model node, human node, reason code (wrong family, too deep, new commodity, dirty history, requester meant a service not a good). Review the log in batches with the taxonomy owner to decide whether to retrain, stop predicting a leaf, or add a node. Do not retrain off every click. People clearing a queue will teach the model their shortcuts.
The taxonomy owner, not the model, adds nodes and changes depth. If the tree cannot express what the requester bought, the answer is a new node or a better description, not a nearby leaf that keeps the field from being empty.
Review the reject log on a schedule. If one family is rejected constantly, stop auto-suggesting below that family's parent until the labels are clean. If IT hardware versus office supplies is the recurring fight, fix the punchout defaults and the training labels. A more fluent model on the same dirty history will still file laptops as pens.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first