AI Adoption GuideProcurementReceive
Receipt quality scoring
ML aggregates delivery accuracy, damage rates, and timing compliance per supplier into a running quality KPI.
Procurement processRequestApproveSourceEvaluateSelectOrderReceiveReview
By Don, DoneThat’s AI coach · updated
Score your goods receipts, by legal entity
Receipt quality scoring is a running KPI built from goods receipts your warehouses posted, keyed to the legal entity that shipped and the ship-tos you actually use.
The vendor-master name on the PO is not the unit of analysis. Sister plants, billing companies, and trading names often share a group brand. If you score the brand, you mix docks, mix company codes, and then cannot tell a plant manager which receipts are theirs.
A score you cannot cite back to GR numbers is not a quality KPI. It is a tile. Suites such as SAP Ariba, Coupa, and Jaggaer, and the ERP that posted the receipt (often SAP), will show a supplier performance view if you put numbers there. They do not decide what counted. You do: which vendor number, which window, which events, and when the file is too thin to score.
Do not substitute a third-party rating, a supplier's own delivery claim, or last year's bid score. Those files answer other questions. This one answers what landed on your dock.
Roll short, wrong, damage, and late into a cited series
Each posted GR is one dated observation. Label it before you average anything.
Keep four event types separate until a category rule combines them:
- Short: received quantity below the PO or delivery note after open partials are excluded.
- Wrong: material, grade, or pack that does not match the PO line.
- Damage: crush, leak, contamination, or similar, with the photo or note already attached by delivery discrepancy classification.
- Late: the date your inbound process actually records, after the confirmed date the buyer held. Do not invent a second clock.
One GR can carry more than one event. A late load can also be short. Count both. Each event stores the GR number and, when it exists, the discrepancy record. That citation is what you take to a supplier meeting. Without it, the meeting becomes whose extract is right.
Combine events only with a written category rule. Timing may weigh more for ingredients that stop a line. Damage may weigh more for glass or fresh goods. Paperwork errors belong in their own bucket so they are not treated as product defects. Use a dated window, not a lifetime blend that never forgets a bad week from years ago.
Work the file in this order:
- Pull posted GRs for the review window, keyed by vendor legal entity, ship-to, and category.
- Label each GR as clean, or as short, wrong, damage, late (more than one allowed).
- Keep the GR number on every event, plus the discrepancy record when there is one.
- Apply the written category weights and emit a dated KPI point. Do not fold the whole history into one lifetime average.
- Hold the score back below the minimum GR count. New suppliers stay in that unscored set.
Quantity shorts that already failed three-way match automation should appear here as shorts, not as a second unexplained match-fail reason. If AP codes them only as invoice exceptions, the quality KPI will look cleaner than the dock.
Receiving errors will hit the score. Miscounts, GRs posted to the wrong PO, and damage photographed late all mark the supplier. Sample a slice of defect events against the physical receipt before you publish internally, and again before suppliers see the number. Strip or flag confirmed dock errors.
A model can classify the event and refresh the series as receipts post. It cannot move the KPI without a document pointer. If it cannot cite a GR, it does not count.
Thin files stay unscored, including new suppliers
A thin file is insufficient evidence, not a high score.
Set a minimum GR count in the window, for that legal entity and the ship-tos in scope. Below that count the supplier is unscored. Show that state. Do not let a blank cell fill with a default.
The failure mode is treating no receipts as a perfect score. It shows up when a vendor record exists, invoices have been paid, and receiving posted nothing, or posted against another number. An empty file is a missing control. It is not 100% quality.
New suppliers stay unscored through onboarding and first live receipts. Early loads are when you should watch the events, and when a percentage will swing on a single pallet. Watch the events. Do not publish the percentage.
If one supplier ships several company codes, keep a series per legal entity. A roll-up is allowed only when the entity list stays visible. A clean plant must not hide a dirty one.
Sister plants under one parent get two KPIs
Picture a packaging category lead at a food manufacturer. Cartons come from one supplier group. Two plants ship.
Plant North is the entity on most POs for Site A. Receipts are complete. Crushed corners show up on some loads; the dock photographs them. Plant South is a sister legal entity shipping to Site B. Someone later merged South's vendor number into the parent during a master-data cleanup. South's inbound deliveries miss the confirmed date more often than North's. North's do not.
Score the parent and both problems appear at once. North inherits South's late receipts. South's records, if any still post under the old number, drop out of the entity you pay. The quarterly pack then reports one group number. Neither plant can act, because the cited GRs mix two companies and two docks.
Keep the KPI on the vendor number that matches the ship-from legal entity on the GR. If master data merged them, split the score back to the receipts, not to the surviving name. The parent can sit on a group watchlist. It does not receive a blended quality grade that no plant manager owns.
That is also why this KPI does not belong inside proposal scoring against RFP criteria as a bid score. An RFP scores claims in a document. Receipt quality scores what landed. A bidder with no GRs in your ERP stays unscored here, even when the proposal reads well.
Awards need the event file, not the tile
A running quality KPI supports a quarterly business review and a sourcing file. It is not, by itself, an award.
The failure mode is using the number to cut or grow share, or to grant a preferred badge, while the GR list is missing, not reviewable by the supplier, or not scoped to the entity that would ship the new volume. The supplier cannot audit the score. You cannot show which shorts, wrongs, damages, and lates moved it. The award then rests on a tile.
Before the score is used in an award, a quarterly review, or a volume argument, require:
- The dated series for that legal entity and ship-to set.
- The event file: GR numbers, event type, and the discrepancy record or photo where it exists.
- Sample size on the face of the score. Unscored stays unscored. A barely qualified file is labeled as such.
- Dock errors found in sample review stripped or marked, not left in as supplier defects.
Use historical performance retrieval to pull this same file at the next sourcing event, not a screenshot of last quarter's rank. Keep supplier risk continuous monitoring on news, filings, and ESG events. That signal is different. A distressed vendor can still post clean receipts. A stable vendor can still crush cartons. Do not blend risk and receipt quality into one health number.
SAP Ariba, Coupa, Jaggaer, and SAP will display whatever you load. Configure the entity key, the citations, the window, and the unscored state. The default supplier performance view is not automatically that rule.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first