AI Adoption GuideProcurementSource
Market price benchmarking
AI pulls real-time market and historical pricing data to build a should-cost model before RFQ issuance.
Procurement processRequestApproveSourceEvaluateSelectOrderReceiveReview
By Don, DoneThat’s AI coach · updated
A should-cost is a named range, not a guess
Before the RFQ goes out, you need a should-cost range you can defend: low, mid, high, and the drivers at each end. The output is not a single heroic number. It is not a chart labeled "market" with no series name.
A model that invents a commodity index is not a benchmark. If you cannot open the source (Fastmarkets kraftliner 42 lb delivered US East, LME aluminum cash, BLS PPI for that commodity, your last PO file, a contracted lane rate), you do not have a should-cost. You have a vibe.
Keelvar, SAP Ariba, Coupa, and Jaggaer are the class of sourcing suites where events and bid files already live. PO history usually sits in the same suite, or in the ERP behind it. Assemble the range before RFP/RFQ auto-drafting publishes the pack. Use that same range as a reserve, then as an outlier challenge when bids return. Those platforms are not ranked here. None of them substitutes for naming the index.
The cost outcome is the range in the event file before anyone is invited. Do not book savings against it. A gap to last year's award is a question, not a result.
Pull indices, last POs, and freight into one stack
Build three layers. Keep them separate so a movement in one does not silently rewrite the others.
Named indices. Commodity series that actually price the material in this spec. Metals: LME cash or the contract you buy against. US containerboard: Fastmarkets kraftliner, unbleached 42 lb, open market, delivered US East, and semi-chemical medium, 26 lb, same geography. Recycled content: a Fastmarkets OCC (old corrugated containers) assessment for the region you buy in. BLS PPI is a secondary public check, not a stand-in for the grade. Write the series name, unit, geography, Incoterm of the print, and the date you used. If the category has no named series, say so. Do not invent a packaging index from unrelated PPI lines.
Last PO prices. Same spec, same Incoterm, same or similar volume band, from ERP or P2P history. Strip expedites, samples, credits, and one-off display packs. Keep plant and currency. Convert with the FX source finance already uses, dated. Last PO is not the market. It is what you paid, which may have been high, low, or a favor. It is still the strongest internal signal, and the one teams skip.
Freight. Lane cost to the ship-to. Fuel where the contract uses a published diesel series (US EIA weekly retail on-highway diesel is a common one). Accessorials that actually apply. A mill-gate index plus a DAP RFQ with no freight line is not a delivered should-cost. Contracted 3PL rates are the primary. A DAT-style spot print is a sanity band only.
Then add conversion you can name: last PO converting for this RSC, mill processing for this alloy. Do not invent a conversion index. Show the stack as a range: index band over the lookback you chose, last comparable awards, freight high and low for the lanes you will use. Put the drivers next to the dollars: linerboard up, diesel down, Southeast lane tighter than Midwest.
Pick a lookback that matches how this market moves. Some resins need weeks. Stable converting can use many months. Write the dates. A large call-off is not comparable to an emergency buy.
If spend category auto-classification has the row in the wrong node, the last-PO cohort is garbage. Fix the category before you trust the average. Total cost of ownership modeling is a later stack (yield, damage, inventory, warranty). Should-cost for the RFQ is delivered unit price for the spec on the event. Do not mix them into one number.
Illustrative walkthrough: corrugated RSC before the RFQ
The walkthrough is illustrative, not a measured result.
A food manufacturer is re-sourcing RSC cartons, 32 ECT, C-flute, two-color print, DAP to two plants: Midwest and Southeast. The RFQ goes out next week. Three converters are on the panel. This week's job is a should-cost range in the event file, not a target to email the incumbents.
The stack: Fastmarkets kraftliner 42 lb and semi-chemical medium 26 lb, delivered US East, with print dates in the file; twelve months of DAP unit prices for this RSC spec, split by plant, with expedites and a display shipper removed; contracted truckload to each plant, fuel tied to EIA diesel. Midwest POs sat in a tight band. Southeast POs ran higher in two months after a converter dropped a lane.
The file shows two delivered ranges and the drivers. In this illustration, Midwest is $860-$980 per thousand cartons DAP. Southeast is $940-$1,120. The gap is the lane and those two high PO months, not a slogan that "the South is expensive."
What they almost did: they published one heroic should-cost for "corrugated, US" on the RFQ cover. Southeast freight disappeared. The Midwest converter looked expensive. The point became a political number. They also treated a scraped Amazon listing for similar-looking boxes as the market. Retail pack quantity, channel, and print are not an industrial DAP carton. That number sat below every serious bid.
The useful file had two ranges, named series, PO cohort rules, and freight by plant. That file is what the reserve is set from.
A laptop priced as a stapler is not a benchmark
If the classifier files a notebook under office supplies, last-PO should-cost for that node looks like pens and staplers. The RFQ reserve will be nonsense, and every IT bid will look like a gouge. That is a taxonomy error, not a clever blend.
Do not run should-cost on a node until a person has confirmed the category, especially at taxonomy boundaries: IT hardware versus office supplies, corrugated converting versus folding carton versus packaging merchants, MRO bearings versus a generic maintenance code. A reject log on the classifier is cheaper than an event priced off the wrong cohort.
If you cannot name a series for a bespoke assembly, do not fake one. Use last comparable POs, a teardown you actually did, and freight, and label the holes. Negotiation position generation can still use a wide band. It cannot use a fake index.
Set the reserve, then challenge bid outliers
Set the RFQ reserve from the high end of the delivered range for that ship-to, after you freeze the spec the RFQ will publish. The reserve is an internal ceiling, not a number on the supplier pack. A bid above it is a flag, not an auto-reject. If every bid lands above it, the stack is wrong or the spec on the event is not the spec you costed.
After bids land, use the same drivers to challenge outliers. A bid well below the low end is not a win until you check spec, Incoterm, and whether freight is missing. A bid well above the high end gets a written challenge: index print, last PO band, freight lane. That write-up feeds negotiation positions, and the decision whether to run a reverse auction at all. An auction start price pulled from a scraped listing, or from last year's PO with no index movement, is a bad reserve with a countdown clock.
Do not email the should-cost to the bid list. Do not put the mid-point on the RFQ as a "target price." You will buy at that number or worse.
Record the range and the sources in the event file before invitations go out. After award, compare outcome to the band. Persistent bias (always over, always under) is a model problem or a spec problem. It is not a savings percentage. Do not invent one.
The job is finished when a category lead can open the event and see a delivered range per ship-to, each driver named, the reserve set from that range, and a rule for what happens to a bid that sits outside it.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first