Skip to main content
DoneThat

AI Adoption GuideProcurementEvaluate

Proposal scoring against RFP criteria

LLM extracts vendor claims from submitted PDFs and scores each against weighted evaluation criteria, producing a structured comparison matrix, using tools like Inventive AI or Zycus AutoScore.

Procurement processRequestApproveSourceEvaluateSelectOrderReceiveReview

By Don, DoneThat’s AI coach · updated

Suggested scores are a starting pack, not an award

The job is a suggested score against this event's published criteria, with the passage and page the vendor actually wrote. The evaluation panel still owns every number. The model does not award.

Each cell should carry the claim in the bidder's words, a page or exhibit cite, and a suggested score only where that claim maps to the published score-band descriptor. Where the bid is silent, evasive, or only markets itself, the cell stays blank. Do not invent a mid-band score because the executive summary sounded confident, and do not highlight a winner on the extract.

Sourcing and evaluation suites in the Inventive AI, Zycus, Jaggaer, and Coupa class can pull claims from submitted PDFs and lay them into a comparison matrix. Treat that class as extraction and arithmetic on the issued weights. The panel still scores. In public and regulated procurement, an unexplainable machine score is a challenge waiting for a file.

Run internal contradiction detection before anyone scores. A bid that promises around-the-clock coverage in the summary and business-hours coverage in the SLA schedule should be clarified, not averaged into a neat 4.

Load the scorecard that went to bidders

Load the evaluation workbook issued with this RFP: criterion names, weights, pass/fail gates, and the score-band descriptors bidders saw. Do not load last year's leftover template.

Leftover criteria fail quietly. A facilities event last year scored on-site print and mail. This year's pack dropped that line. If the model still has the old column, every vendor is silent on print, every cell gets a guessed mid-score, and the weighted total punishes people for not answering a question you did not ask.

If a criterion is not in the issued pack, it is not scored. If the panel later wants a new question, that is a clarification to all bidders, not a private column.

Compliance requirement extraction is the sibling job on the shalls: certifications, statutes, named policies. Scoring the quality of a technical answer is not checking whether a mandatory certificate is present. Keep those lists apart so a missing certificate is a gate, not a 2 on quality systems.

Cite the page or leave the cell blank

For each vendor PDF, and for the pricing workbook only if a commercial criterion sits on the issued scorecard:

  1. Map each published criterion to the sections that could answer it.
  2. Extract the vendor's claim in their words, with page, exhibit, or clause.
  3. Match the claim to the published score-band descriptor, for example named roles, ramp weeks, and a cutover plan versus a proven-methodology slogan.
  4. Propose a score only when the bid actually answered. If it did not, leave the cell blank and tag it unanswered.
  5. Do not backfill from marketing language, from another vendor's answer, or from what you remember they delivered last time.

Historical performance retrieval is a separate pack for the panel: SLA misses, claims, transition pain on the incumbent. It is not a license to fill a silent cell. Last year's good mobilization does not answer this year's unanswered mobilization question.

Silence is silence. The panel can issue a clarification, score a zero under the published rule for non-response, or hold the cell. The model does not get to pick which. Price belongs where the issued scorecard put it. If commercial is a separate envelope, do not let a technical extract bleed unit rates into a quality score.

Three bids against one issued pack

This is a worked example with made-up names, not a case study.

A city lets a three-year facilities management RFP. The issued scorecard is:

  • Mobilization and cutover (25 percent), scored 0 to 5 against named roles, a week-by-week ramp, and building-by-building cutover.
  • Staffing model (20 percent), scored 0 to 5 against named FTEs by site and skill.
  • Reactive SLA (20 percent), scored 0 to 5 against stated response and restore times by priority.
  • Waste and recycling plan (15 percent), scored 0 to 5 against a site-level diversion method.
  • Mandatory living-wage commitment (pass/fail). Fail means ineligible, not a low weight.
  • Price is a separate envelope, opened after technical.

Northline writes a mobilization chapter (pages 18 to 24) with a named transition lead, a six-week ramp, and a building schedule. Staffing is a table of FTEs by site. The SLA schedule states P1 restore in four hours. Waste is one sentence in the executive summary: we take sustainability seriously. Living-wage: a signed schedule matching the city's floor.

Harbour FM has a glossy summary claiming around-the-clock coverage and zero disruption. The technical appendix (page 41) states helpdesk hours 07:00 to 19:00 on weekdays. Mobilization is a stock methodology with no named roles. Staffing is right-sized teams. Waste is a corporate brochure. Living-wage: signed.

Pike Services answers mobilization and staffing with cites. SLA times sit only in the pricing workbook, not the technical pack. Waste has a site-level plan. Living-wage: the bid is silent; the boilerplate says we comply with all applicable laws.

What the extract should do:

  • Northline: suggested scores on mobilization, staffing, and SLA, each tied to pages. Waste cell blank (unanswered), not a 3 because the summary used the word sustainability. Pass on living-wage with the signed schedule cited.
  • Harbour: the coverage contradiction already flagged before scoring. Mobilization and staffing stay low or blank against the descriptors. Do not propose a strong SLA score from the summary while page 41 says weekday hours. Pass on living-wage.
  • Pike: mobilization and staffing suggested with cites. SLA: if the issued rule says technical envelope only, the cell is blank until a clarification, not filled from the price file. Living-wage: fail, not a 2 that can be averaged off.

What they almost did: scored last year's leftover print-room column (all blank, all mid-scored), gave Harbour a 3 on waste from the brochure, and published a ranked matrix with Harbour in second because the weighted average hid the SLA contradiction.

A mandatory fail is not a low weight

Pass/fail gates stay binary. If living-wage, security clearance, insurance, or a named certification is mandatory in the issued pack, a silent or non-conforming bid is ineligible. Do not convert the fail to a 1 out of 5 and let staffing and price pull the weighted total over the line.

That averaging failure is how a panel discovers a winner who never passed the gate. The matrix will look complete. The award will not survive a challenge that asks where the commitment is.

ESG supplier screening is the same kind of gate when policy says so: a screening outcome, not a quality score invented from a brochure. Do not fold screening into the weighted technical average unless the issued scorecard published it that way.

After the panel owns criterion scores, combining them is a different job. Multi-criteria decision matrix uses the panel's numbers, not the model's first pass. Do not treat the extract's weighted total as the award.

The matrix is an exhibit, not the award

Inventive AI, Zycus, Jaggaer, and Coupa are a class of places the PDFs, scorecard, and matrix already live. Use the suite as the file: submissions, criterion columns, cites, suggested scores, blank cells, and the panel's overwritten numbers. Do not stand up a second scoring spreadsheet that diverges the day after moderation.

The panel still scores independently or in a moderated session, using the extract as the pack, not as the vote. Two evaluators at 2 and 5 on the same cited passage is a discussion, not a number to average in silence. The award recommendation in the evaluation report is theirs. The matrix is an exhibit.

Do not publish a ranked AI score to stakeholders. Do not let a project sponsor treat the highlighted row as the decision. In a challenge you will need the issued criteria, the cited passages, the panel's scores, and the reason a blank cell became a zero or a clarification. The tool filled it is not a reason.

Trial on one completed event first: load the issued scorecard, extract with cites, compare suggested cells to what the panel actually scored, and look hard at blanks, leftover columns, and any mandatory that received a numeric score. Then use it live as a starting pack.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first