AI Adoption GuideITSelect
Vendor shortlist scoring engine
LLM scores vendors against decomposed weighted criteria, including security, TCO, integration fit, and support, from RFP responses.
IT processPlanSelectDeployProvisionSupportUpgradeReplaceRetire
By Don, DoneThat’s AI coach · updated
What you get from a scoring engine
An LLM scoring engine turns decomposed, weighted evaluation criteria into a reviewable grid: one row per vendor, one column per criterion, each cell tied to a specific RFP response row and criterion ID. The engine reads vendor submissions (structured tables, narrative answers, attachment excerpts you have already normalized) and assigns a score or leaves the cell empty when the response is silent. Speed comes from not re-reading forty-page PDFs for every security, TCO, integration, and support question. Judgment stays with the committee. The sheet supports shortlisting; it does not pick a winner or compute a rolled-up weighted total unless your team explicitly builds that step elsewhere.
Use this when you have an RFP in flight, a fixed criteria matrix from legal and architecture, and a calendar that does not allow three rounds of manual cross-checking before the evaluation workshop.
Load RFP responses before you score
Start with a single response index per vendor. Each indexed row needs a stable ID, the source document, page or section reference, and the raw text the model will cite. If you ran an upstream pass with the RFP response summarizer, import those rows rather than re-parsing originals. Split wide answers: a vendor may answer "integration" in three places (API appendix, implementation plan, reference architecture). Index each fragment separately so citations stay precise.
Map every scoring column to one or more criterion IDs from your matrix. Weights live in the matrix metadata; the scoring engine reads them for context but does not emit a composite rank. Typical IT sourcing columns include security controls and certifications, three-year TCO assumptions (often cross-checked later in the total cost of ownership calculator), integration fit against your integration catalog, and support model (SLA tiers, escalation, named TAM). Vendors in this class (ServiceNow, SAP Ariba, Coupa, and analyst frameworks such as Gartner's market guides) arrive with different response shapes: some lead with product modules, others with SI partnerships. Your index normalizes that variance so the scorer compares answer text to criterion text, not brochure structure.
Lock the criteria version for the scoring run. If architecture adds a net-new row mid-evaluation, re-run only affected vendors and columns; do not silently overwrite cells from an earlier criteria set without marking the run ID.
Score cells with mandatory row cites
For each vendor and criterion, the model returns: score (or blank), short rationale, criterion_id, and response_row_id. A valid cell always includes the row cite. If the model produces a score without a row ID, treat it as a failed cell and re-prompt or hold for human review. That failure mode is common when a vendor implies an answer across sections; the fix is tighter indexing or a narrower criterion prompt, not accepting an orphan score.
Scoring prompts should mirror committee language. Example criterion: "Describe SSO support for SAML 2.0 and SCIM provisioning; state whether included in base license." The engine matches indexed text, extracts whether SAML and SCIM are both addressed, and scores against your rubric (e.g., fully documented, partial, not addressed). Illustrative run: four vendors, twelve criteria, forty-eight cells. Vendor B's cell for criterion SEC-04 shows "partial" with cite VENDOR-B / row 47 / Security appendix p.12 because SCIM is mentioned only in a footnote. Vendor C's cell for the same criterion is empty with rationale "no indexed row mentions SCIM or SAML provisioning" because their response discusses LDAP only. Empty is correct; do not backfill from marketing sites or analyst summaries.
Keep rationales one to three sentences. They are briefing notes for workshop prep, not audit prose. Link the run to your selection decision audit trail so each cell's cite survives challenge weeks later.
Leave blanks when the response is silent
A blank cell means the indexed RFP material does not substantiate a score for that criterion. Silence is not a zero unless your rubric explicitly maps "no answer" to the lowest band and the committee agrees that rule applies before scoring. Many teams use blank vs. "not addressed in RFP" as a flag for clarification questions rather than an automatic penalty.
Do not let the model infer from vendor reputation, prior contracts, or category knowledge. ServiceNow may be strong on workflow; Coupa on spend; SAP Ariba on supplier networks; Gartner may describe market positioning. None of that belongs in a cell unless the vendor stated it in this RFP response row. If integration fit requires a named connector and the response omits it, blank stays blank until the vendor amends the submission or the committee accepts external evidence through a separate, logged process.
Blanks also surface criteria that vendors interpreted differently. If half the grid is empty for TCO-02 because vendors bundled pricing narrative elsewhere, that is a indexing or RFP template problem, not a reason to invent scores.
Shortlist in committee; do not auto-award
The scoring engine output is input to a shortlist conversation, not a contract trigger. Export a workshop view: vendors as rows, criteria grouped by theme, blanks highlighted, cites expandable. Facilitators ask "do we accept this partial on SEC-04 for Vendor B?" and record decisions in the audit trail. Shortlist rules stay human: top N by domain coverage, mandatory pass on must-have criteria, or tiered proceed vs. clarify vs. drop.
Explicitly forbid two anti-patterns in your runbook. First, treating the highest numeric sheet as an award: without a committee-approved weighting rollup, column scores are comparable only within a criterion, not as a single leaderboard. Second, inventing a weighted total inside the tool: weights inform prioritization of review time, not an automatic rank. If the business wants a weighted model, build it in a controlled spreadsheet or governance-approved calculator with signed weights; do not let the LLM silently multiply scores.
After shortlist, downstream work (negotiation, contract risk extraction, final BAFO) references the same criterion IDs so nothing scored in workshop contradicts what legal reviews in the contract.
Failure modes to block before sign-off
Review the grid for these before you call the evaluation complete:
Scores without row cites. Any populated cell missing response_row_id is invalid. Re-run or clear the cell.
Sheet treated as award. If stakeholders quote "Vendor A won the scoring engine," reset expectations: the engine accelerates reading, not decision rights.
Invented weighted totals. Reject any export that shows a composite rank unless it came from a separate, versioned weighting artifact the committee approved.
Hallucinated vendor capabilities. Cross-spot check high-impact criteria (security, data residency, exit terms) against the cited row text, not the rationale alone.
Criteria drift. Mixed criterion versions in one grid invalidate comparison; partition runs by criteria_version.
Silent blanks filled by assumption. Empty cells should trigger clarify-or-exclude workflow, not quiet downgrade.
When these checks pass, you have a defensible shortlist package: fast first pass, traceable cites, honest gaps, and a committee still accountable for who proceeds to negotiation.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first