Skip to main content
DoneThat

AI Adoption GuideGovernmentFund

Proposal scoring assistant

LLM scores and ranks funding proposals against decomposed weighted criteria, producing ranked summaries for review panels.

Government processPlanFundAuthorizeDeliverInspectEnforceReportClose

By Don, DoneThat’s AI coach · updated

Dual-cite scores the panel can audit

A usable criterion score names two things: the published criterion from the Notice of Funding Opportunity (NOFO), and the proposal paragraph that is supposed to meet it. If either cite is missing, the cell is not a score. It is a note that ranking has already started to drift.

The quality bar is dual citation, not a number that looks official. The assistant may propose a criterion-level score only when it can point at both sources. The review panel still records the official score. Nothing in the ranked summary replaces the panel's sheet, and nothing in the ranked summary should be copied into the official score column.

This page is for the secretary who builds the scoring packet: criteria extracted from the NOFO, one proposal at a time, ranked summaries that a panel can walk through without guessing where the language came from. Completeness of the application package is a different job. Run that check with the application completeness checker before you ask anyone to score.

Load the NOFO rubric before the proposal

Start with the published criteria, not the narrative. Pull the evaluation section from the NOFO, keep the weights as published, and decompose compound criteria into the units the panel will actually score. A criterion that mixes capacity and past performance needs two rows if the NOFO scores them separately. A criterion that already lists numbered subfactors should become numbered rows, not one blended comment.

Load those rows into the scoring sheet first. Then load the proposal. Order matters because the model will otherwise invent a rubric from the applicant's headings. Applicant section titles are not the NOFO. If the proposal uses "Organizational capacity" and the NOFO scores "Staffing plan" and "Fiscal controls" as separate items, you score the NOFO rows. You do not merge them because the PDF had one heading.

Where the application arrived through a Grants.gov-class portal, or sits in Salesforce Government Cloud, Microsoft, or Accela, treat those systems as the filing cabinet. They hold the PDF, the attachments, and the official NOFO identifier. They do not hold a panel score until a reviewer enters one. Do not ask the assistant to read the system of record for a score that does not exist yet.

If the NOFO is still in draft, stop. Scoring against an unpublished rubric produces ranks the panel cannot defend in a debrief. Statutory fit, eligibility, and required forms belong in a compliance pre-check generator pass, not in the merit rows. Mixing those checks into merit scoring is how an ineligible proposal gets a high rank and a silent merit criterion gets a guessed number.

Score each row with two citations

For each criterion row, the assistant should return four fields: the criterion identifier and quoted text from the NOFO, the weight as published, a proposed score only if evidence exists, and the proposal cite (section heading plus paragraph, page, or attachment filename). Dual cite means both the criterion quote and the proposal quote are present and specific enough that a panel member can find them without searching the whole file.

Walk one criterion at a time. Do not ask for a total first. Totals hide empty rows. After a row is filled, check the two quotes before you move on. If the criterion quote is a paraphrase instead of the published sentence, replace it. If the proposal quote is a restatement of the criterion with no applicant language, treat the row as silent.

Illustrative example: a workforce training NOFO weights employer partnerships at 20 points and credential attainment plan at 15. The proposal's partnerships section names three employers and attaches letters. The assistant can propose a criterion score for partnerships because it can quote the 20-point criterion and the paragraph that lists the letters. The credential section describes classroom hours and does not mention a credential, an accreditor, or an attainment target. That row stays empty. The ranked summary can still list this proposal among others on the partnership criterion. It cannot invent a mid-range credential score so the spreadsheet looks complete.

After every criterion has been attempted, produce the ranked summary the panel asked for: proposals ordered by the assistant's filled rows only, with weights applied only to scored cells, and a visible count of blank criteria per proposal. A rank that omits the blank count is not ready for the room. A rank with no cite on any winning row is not a rank. Send it back and re-run the empty-cite rows before anyone sits down.

Silent criteria stay blank

Empty stays empty if the proposal is silent. Do not fill a silent criterion with a middle score, a partial-credit guess, or language recycled from a different section. Silence is evidence. A three or a five in that cell is not.

The failure mode is cosmetic completeness. A secretary under time pressure sees a blank, worries the panel will think the packet is unfinished, and accepts a hedged paragraph that implies the missing element. That hedge is how a silent criterion becomes a middle score, and how a rank with no real cite gets into the packet. If the model cannot quote a proposal paragraph that addresses the criterion as published, the score field stays blank and the note field says silent, with the criterion quote still attached so the panel sees what was looked for.

Do not let the model infer partnerships from a job-placement sentence, or infer a credential plan from a list of courses. Adjacent topics are not the published criterion. The panel can decide that the adjacent language is enough. The assistant cannot make that decision by filling the cell.

Blank rows are the panel's job, not the model's. Reviewers decide whether silence is disqualifying, a clarification item, or a scored zero. The assistant does not choose among those, and it does not pre-fill a zero unless the NOFO itself says that absence scores zero and the packet already records that rule on the row.

The panel still scores

Treat the assistant output as a reading aid: ranked summaries, dual cites, blank flags. Copy those into the packet the panel will use. Do not paste assistant numbers into the official score column in the grants system of record.

The second failure mode is treating the assistant score as the panel score. Once those numbers sit in the same column, later exports, debriefs, and closeout files will treat them as the panel's judgment. Keep two columns or two sheets: assistant (advisory, cited) and panel (official, signed). If your Grants.gov-class portal, Salesforce Government Cloud, Microsoft, or Accela record only has one score field, the panel field is the only one that belongs there. Print the assistant ranks on paper or in a sidecar file the chair can discard after deliberation.

A rank with no cite is the error that looks most finished. A sorted list of applicants with scores and no criterion quotes, no paragraph quotes, and no blank flags is a ranking you cannot walk a chair through. It also cannot survive a debrief question about why proposal B sat above proposal C. Re-run criterion by criterion until every non-blank score has both cites.

When the panel has scored, the secretary's job shifts from scoring assistance to recording the decision. Do not go back and reconcile panel numbers to the assistant ranks. Divergence is expected. Allocation of remaining funds across awards is a later step. The resource allocation optimizer and scenario comparison matrix are for that stage, after official scores exist.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first