Skip to main content
DoneThat

AI Adoption GuideRetailBuy

Supplier Quote Comparison Agent

LLM extracts supplier quotes, normalizes commercial terms, flags deviations, and produces a buyer-ready comparison table for faster sourcing decisions.

Retail processPlanBuyPriceStockSellFulfillReturnClear

By Don, DoneThat’s AI coach · updated

What this agent does

Retail buyers often receive supplier quotes in mismatched formats: PDFs, spreadsheets, email attachments, and portal exports. Comparing landed cost, MOQ, lead time, payment terms, and Incoterms by hand is slow and easy to misread. A supplier quote comparison agent uses an LLM to extract line items and commercial terms from each quote file, normalize them onto a shared schema, flag deviations from the RFQ or category baselines, and produce a buyer-ready comparison table.

The agent accelerates the analysis step. It does not award business. The buyer still reviews the table, checks flagged items, and decides which supplier wins.

When a quote file cannot be read (corrupt PDF, scanned image without usable text, password-locked workbook, or unsupported encoding), the agent returns empty output for that file and records a clear failure reason. It does not invent line items or fill gaps with guesses.

Inputs the agent needs

The agent works from the RFQ package plus each supplier response:

  • RFQ baseline: requested SKUs or specs, target volumes, delivery windows, Incoterms preference, payment terms, packaging, and any mandatory clauses
  • Supplier quote files: PDFs, XLSX/CSV, or structured exports attached to the bid
  • Supplier identity: name, vendor ID if known, currency, and quote validity date
  • Category baselines (optional): historical unit costs, typical MOQs, and lead-time norms for the same or similar items
  • Scoring weights (optional): how the buyer prioritizes unit price vs. lead time vs. service level for this event

Without a readable RFQ baseline, deviation flags are weaker: the model can still align fields across quotes, but it cannot reliably mark “off RFQ” terms. Prefer structured RFQ fields when available; free-text RFQs still work if the agent can parse the required commercial dimensions.

How extraction and normalization work

For each readable quote, the agent extracts candidate fields: item identifiers or descriptions, quantities, unit prices, currency, MOQ, pack size, lead time, ship-from location, Incoterms, payment terms, validity window, and surcharge or freight notes. It maps supplier-specific labels (“net 60,” “FOB Shanghai,” “carton of 24”) onto a shared vocabulary so rows can sit side by side.

Normalization includes:

  1. Currency alignment: convert quoted amounts to a buyer-selected comparison currency when rates and as-of dates are provided; otherwise leave native currency and mark conversion as pending
  2. Unit of measure: express price per consistent UOM (each, case, kg) using pack-size clues in the quote
  3. Lead time: map calendar days or “ships in X weeks” into a comparable day count
  4. Term codes: map payment and Incoterms variants to standard codes with the original phrase retained as evidence

Ambiguous fields stay marked as ambiguous rather than forced into a false precision. The comparison table shows source snippets or page references so the buyer can verify extraction before relying on a cell.

What the comparison table surfaces

The output is a matrix keyed by RFQ line (or normalized item) with one column per supplier. Typical columns include unit cost, extended cost at the RFQ quantity, MOQ, lead time, Incoterms, payment terms, quote validity, and a short notes field for freight, tooling, or exclusivity language.

Alongside the matrix, the agent emits:

  • Deviation flags: terms that differ from the RFQ (wrong Incoterms, longer lead time than requested, MOQ above target, missing mandatory SKUs)
  • Cross-supplier outliers: prices or lead times that sit far from the peer set for the same line, labeled as outliers for review, not as automatic rejects
  • Coverage gaps: RFQ lines with no quote, or quotes that only partially cover the assortment
  • Read failures: files that produced empty output, with the reason (unreadable, encrypted, unsupported format)

Flags are advisory. A higher price with better payment terms or a nearer DC may still be the right award. The table is there to make those tradeoffs visible in minutes instead of hours of spreadsheet stitching.

Buyer review and award control

Human-in-the-loop is mandatory. The model builds the comparison; the buyer still awards.

Recommended review loop:

  1. Confirm every quote file either contributed rows or appears in the read-failure list
  2. Spot-check high-impact cells (unit price, MOQ, lead time) against the source PDF or sheet
  3. Resolve ambiguous extractions before ranking
  4. Apply commercial judgment: risk, quality history, capacity, and relationship factors the agent does not own
  5. Record the award decision and rationale in the sourcing system of record

Do not auto-submit POs or auto-notify losers from model output alone. Downstream systems may ingest the approved table, but award authority stays with the buyer or category manager.

Operational guardrails

Treat quote documents as commercially sensitive. Restrict who can run the agent, retain access logs, and avoid sending full quote packs to tools outside your approved boundary. Prefer redacting unrelated attachments before ingestion.

Keep prompts and post-processing deterministic enough that a re-run on the same files yields the same table structure. Version the RFQ baseline ID and quote file hashes on each run so auditors can see which inputs produced which comparison.

When quality is low (many empty outputs, many ambiguous cells), stop and fix inputs: re-export the workbook, request a text PDF, or ask the supplier for a structured quote template. Empty output on unreadable files is correct behavior; fabricating a comparison from partial OCR is not.

Used this way, the agent shortens quote turnaround in retail buy workflows without substituting for sourcing judgment. Buyers spend time on award decisions, not on aligning columns across five incompatible quote formats.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first