Skip to main content
DoneThat

AI Adoption GuideConsultingSell

Proposal Red-Team Reviewer

LLM stress-tests a draft proposal against likely evaluator objections and scoring criteria, surfacing gaps before submission.

Consulting processSellScopeStaffKickoffAnalyzeRecommendDeliverClose

By Don, DoneThat’s AI coach · updated

Score the draft against the sheet, do not rewrite it

A red-team pass exists to name which scored lines you would lose on, not to produce a second draft. Evaluators mark against a published matrix. The moment your prompt says "improve this proposal" or "make it more compelling," you have left that job and hired a ghostwriter.

Proposal platforms such as AutogenAI and Responsive sit in this class. A general LLM on the RFP scoring matrix (Copilot in Word, a custom GPT with the rubric pasted in) can run the same attack. Ask for objections mapped to criterion IDs, with the evidence used and the evidence missing. Do not ask the tool to rewrite the proposal.

The output is a list of score risks. A partner still decides which to accept, which to answer, and which mean you should not have bid. If you have not decided to bid, stop. Run a bid/no-bid opportunity classifier first. Stress-testing a no-bid is a weekend you do not get back.

First cuts from a RAG proposal draft generator often sound finished. Red-team them as if they were already submitted.

Paste the scoring criteria or the model invents the evaluator

Without the published scoring sheet in the prompt, the model fabricates objections that sound like an evaluator and are not. It will complain about tone, "lack of vision," or missing "innovation" because those words appear in generic bid-writing advice. Evaluators do not mark those lines unless they are on the sheet.

Feed the model, in this order:

  1. Evaluation criteria, weights, mandatory pass/fail gates, and published scoring descriptors (what a 3 looks like versus a 5).
  2. Questions or requirements mapped to those criteria, with IDs if the RFP numbered them.
  3. The draft, sectioned the way it will be submitted.
  4. Constraints that still affect the score: page limits, font, mandatory forms, annex rules. A model that does not know the page cap will demand more proof. A model that does not know a form exists will not notice you skipped it.

If the buyer only published high-level criteria, say that in the prompt. Ask the model to mark uncertainty rather than invent a detailed rubric. Recurring loss themes from win/loss pattern synthesis belong in a second pass ("we lose when we cannot name two in-sector references"), not as a substitute for this buyer's sheet.

A finding that cannot point to a criterion ID, a weight, or a mandatory gate is a nitpick. Discard it, or park it for color.

Walk a first-cut ops bid through the published weights

A scored pass on a library first cut shows gate failures a style review never finds. Take a mid-size operations practice answering a manufacturer RFP: 16-week operating-model redesign, four workstreams (decision rights, management forums, span of control, shared-services split). The buyer published this sheet:

  • Methodology and approach, 30
  • Relevant experience, 25
  • Team and availability, 20
  • Commercial and price realism, 15
  • Social value, 10
  • Mandatory: two comparable manufacturing references in the last 36 months, named, with permission to contact

The prompt: for each criterion, quote the descriptor, quote the draft, mark pass, score-risk, or fail, and name missing evidence. No rewrites.

Findings worth a war-room slot

Relevant experience (25) plus the mandatory gate. The draft cited a 2023 food-manufacturer operating-model job (in-sector, inside 36 months) and a "confidential financial-services transformation." The second reference is not manufacturing, is not named, and cannot be contacted. That is fail-the-gate risk, not a credentials polish item.

Methodology (30). The RFP asked for a week-by-week plan against four named workstreams. The draft offered a proven-playbook narrative and a 12-week Gantt from an older job. It never mapped weeks to the shared-services split. An evaluator scoring "extent to which the approach addresses the stated workstreams" has nothing to mark. Remap the plan, or accept a low score on the heaviest line.

Commercial (15). The fee table answered price. The high-score descriptor asked for price realism: assumptions, what happens if the split takes longer, and which costs sit with the client. The draft was silent. That is unanswered criteria, not a pricing-strategy debate.

Lines to delete from the red-team output

The same pass also said the executive summary was "too passive," that the vision "lacked ambition," and that adding a digital twin would "differentiate." None of that is on the sheet. Digital twin is invented scope. Passive voice is color. Kill those lines so the partner does not spend a review cycle on them.

If the model starts producing replacement paragraphs, stop it. A reviewer that rewrites becomes a co-author, and color review turns into an argument with the tool instead of with the buyer.

Keep findings that name the criterion, the gap, and the fix

A finding earns a place on the punch list only if a partner can act on it before submission. Useful output has a fixed shape:

  • Criterion ID and weight. "C2, 25 points, plus mandatory reference gate."
  • Evaluator instruction. Quote the descriptor. Do not let the model paraphrase it into something easier to attack.
  • What the draft says. Quote or cite the section. If the evidence is absent, say absent.
  • Score implication. Fail gate, likely low band, or unanswered.
  • Smallest fix. Swap the ineligible proof point. Add the assumption table. Map weeks to workstream 4. Not "strengthen the narrative."

Nitpicks look like this: more energy, more innovation, fewer adverbs, "consider adding a case study" with no eligibility test, grammar on a section that already fails a gate. Grammar can wait. A failed mandatory cannot.

Cap the list. Ask for the objections most likely to cost points, ordered by weight times gap severity. A long dump is how a proposal team learns to ignore the tool. Keep the risks that would actually move a war room, plus any mandatory fails.

Unstated delivery assumptions that would blow up after award belong in an assumption gap detector on the SOW, not in a bid red team. If they appear in the approach, flag them as credibility or score risk, then stop.

For recommendations rather than bids, use the recommendation adversarial stress-tester. Do not mix those prompts.

Color, relationships, and the room still belong to the partner

The model cannot see the evaluation panel, the incumbent, or the last time your firm annoyed this buyer. Color review is the human pass: voice, what you are willing to claim, and which score risks you will carry because the oral or the relationship will carry them.

Keep the sequence honest:

  1. Proposal manager: coverage and gates. Every scored requirement has an answer. Mandatories are met or escalated. Page limits and forms are intact. Triage the red-team list: keep, fix, or discard as invented.
  2. Practice lead: proof points and method. Ineligible experience dies here. So does a workplan that does not match how you deliver.
  3. Commercial partner: price, assumptions, and which risks to accept. The model does not get a vote on margin.
  4. Color reviewer, usually the bid partner. Voice, emphasis, and what you will say in the room. They read the remaining score risks. They do not owe the model a rewrite.

Design for the failure modes you will actually hit:

  • No rubric in the prompt. Invented evaluator. If someone ran Copilot on the Word file with "review this proposal," throw the comments out and rerun with the sheet pasted in.
  • Strategy not locked. You get a long list and no direction. Lock win themes, team, and commercial shape before you attack the prose.
  • Night-before run. You get anxiety. You cannot swap a reference or remap a workplan overnight. Run while those can still move.
  • Score-chasing scope. The model adds workstreams, analytics, or social-value padding you cannot deliver. That is how you win a bid you will regret.
  • Politics unlabeled. The model will not know the incumbent wrote the specification, or that one evaluator hates consultant jargon. The partner adds those objections after the scored list, labeled as not on the sheet.

The red team is doing its job when the war room argues about unanswered criteria and ineligible proof, and the color pass is short. If the war room is arguing about adjectives, you fed the model the draft and forgot the sheet.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first