AI Adoption GuideProcurementEvaluate
Historical performance retrieval
RAG-based retrieval surfaces past delivery, quality, and dispute records for shortlisted suppliers from prior contracts.
Procurement processRequestApproveSourceEvaluateSelectOrderReceiveReview
By Don, DoneThat’s AI coach · updated
Your own receipts are the only performance file that counts
For a shortlisted supplier, historical performance retrieval means assembling this legal entity's file from your contracts: goods receipts, quality notes, claims, and scorecards you already ran. The output is a dated pack the evaluation panel can open. It is not a web profile of the brand, and it is not the case study in the bid.
A retrieval-augmented search can look across the systems where those records already live, then return citations. It cannot know what arrived at another customer's dock. If you have never received from this entity, the honest pack is empty. Write no history. Do not fill the gap with a parent-company award, a brochure, or a synthesized score.
Keep this pack off the bid score. Proposal scoring against RFP criteria still scores what they wrote. This page is about whether your operations already know them.
Search the legal entity you pay, plus affiliates you have received from
Brand-name search is the fastest way to retrieve the wrong company. The name on the bid may be a converting plant, a trading affiliate, a European parent, or a distributor that shares a logo. Receipts posted against one billing company tell you nothing about a sister plant you have never received from.
Before you run anything, pin identity from supplier master:
- The vendor number you pay, plus the tax or VAT identifiers you already store.
- Known affiliates you have actually transacted with, listed so the search can include a sister plant that ships to you under a different billing company.
- Sister companies you have not received from, listed as exclusions. Same brand is not shared history.
If the index returns a quality escape, a late ASN, or a claim, read the vendor number on the source document before it enters the pack. Retrieving a different affiliate with the same brand is the failure mode panels remember. Once they see a Midwest rejection attributed to a plant that never shipped to you, they will not trust the next pack.
When ownership changed, keep the old entity's records and print the dates. New management does not erase the goods receipt. It also does not inherit a sister company's file automatically.
Assemble receipts, late deliveries, claims, and dated scorecards
Ask for four document classes, not a single performance score.
Goods-receipt quality. Accepted versus rejected quantity, inspection or nonconformance notes, and the PO and material they attach to. Receipt quality scoring is how those events become a running KPI later. For evaluation you want the events, with dates, not a lifetime average with no citations.
Late deliveries. Promised date or ASN versus goods-receipt date, by PO line. Delivery discrepancy classification types those events (short, late, damaged, wrong item) so a quantity mismatch in AP is not treated as the same problem as a dock delay. Do not ask the model to invent an on-time in-full (OTIF) percentage. If your ERP already computes one, and you can open the report, the period, and the vendor number, cite that report. If you cannot, show the dated receipts.
Claims. Quantity short, quality, freight, and commercial disputes on this vendor number. Include open and closed. A closed claim is still history.
Prior scorecards. Cards from supplier performance cycles you actually ran, with the period and the KPI definitions used that cycle. An older plant-survey card on "responsiveness" is not a substitute for current receipts.
Source-to-pay suites such as SAP Ariba, Coupa, and Jaggaer are where supplier records, receipts, and (where you run supplier performance) periodic scorecards typically sit. Enterprise search such as Glean can retrieve across those stores and adjacent contract or ticket systems when the user's permissions allow. None of them creates history that was never posted. If goods receipt was never captured, retrieval has nothing to find.
Build the pack so a category lead can audit it in one sitting:
- Identity header: legal name, vendor number, affiliates included, affiliates excluded.
- An event list: GR, NCR, claim, or scorecard ID, date, and a document number that opens the source.
- A hard line that says no records found when a class is empty. Silence is not a pass.
Spot-check before the pack reaches the panel. Open several source documents. If one is the wrong affiliate, a bid attachment, or a claim with no GR behind it, stop and fix the identity map. Panels stop using retrieval after one wrong citation.
Worked example: incumbent converter versus a plant with no file
The walkthrough is illustrative, not a measured result.
A food manufacturer is re-bidding corrugated for two plants. The panel has shortlisted two converters. Converter A is the Midwest incumbent. Converter B is a Southeast plant the category has never used. Converter B's bid includes a delivery case study from a sister company in the same group, written for a different customer.
Retrieval is scoped to each legal entity on the bid, plus affiliates this company has actually received from.
For Converter A the pack returns dated goods receipts: accepted loads, rejected lots with inspection notes on crush and print, receipts posted after the promised date, and a closed quality claim from the prior year. A quarterly scorecard from the last supplier-performance cycle is attached, period printed on the card. The category manager can read the dates. Nothing in the narrative is turned into an OTIF.
For Converter B the search against that vendor number and known affiliates returns no goods receipts, no claims, no scorecards. The sister company's case study sits in the proposal library. It is kept out of the performance pack because those receipts are not yours. The pack states no history with this legal entity. The panel can still score the bid, run ESG supplier screening on the certificates they attached, and treat operational performance as unknown.
If the search had been run on the brand string alone, it would have pulled Converter A's rejections onto Converter B's row. That pack would have been worse than no retrieval.
Missing history is unknown; a thin file is not a high score
A new supplier has no history. Write that. Do not score them as neutral, and do not score them as excellent because nothing bad appeared.
A thin file is the same problem in a smaller envelope. A couple of receipts from years ago, or a single claim with no GR behind it, is not a clean record. It is insufficient evidence. Putting a high score on a thin file is how a panel confuses "we did not look" with "they never failed."
Do not let a sales case study stand in for your receipts. A customer logo in the proposal is not a goods receipt at your dock. If the only "performance" the index can find lives in the bid, that is proposal content. Score it with the RFP. Leave the history line blank.
Do not let the model emit an OTIF, PPM, or quality percentage unless you can open the report, the period, and the vendor number that produced it. Retrieval's job is the file. Any score you publish is a separate, dated calculation you already own.
Dispute and legal records often should not be visible to everyone on the panel. Keep retrieval permission-aware: the category lead and legal may see claim narratives; the wider panel gets a dated count and a pointer, or a redacted summary. Supplier risk continuous monitoring watches the same supplier after award. This pack is the pre-award slice of that file, frozen for the event.
If the first packs you show contain a wrong affiliate, an invented percentage, or a brochure quoted as a receipt, stop. Fix identity and citations before you run the rest of the shortlist.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first