AI Adoption GuideNonprofitReport
Report Accuracy Adversarial Review
LLM cross-checks claims in a draft report against source data, flagging overstated results or unsupported assertions.
Nonprofit processPlanFundOutreachDeliverMeasureReportStewardRenew
By Don, DoneThat’s AI coach · updated
Overview
A grants or evaluation lead reviewing a draft report needs a second pass that is hostile to unsupported wording. An adversarial accuracy review treats every result, attribution, and completeness claim as something that must be traced to source data, then lists mismatches for staff to fix before anyone submits.
This is not a rewrite of the report and not a replacement for program judgment. The model flags overstated results and assertions that the attached sources do not carry. Staff still decide what to change, what to qualify, and what to send.
What an adversarial accuracy review does
Funder, board, and evaluation reports mix tables, narrative, and inherited phrasing. Under deadline pressure, a sentence that felt true in a staff meeting can outrun the spreadsheet: a pilot becomes “the community,” a one-year change becomes a trend, or a correlational finding becomes “the program caused.” Those slips are easy to miss when the same team wrote both the analysis and the prose.
The review walks the draft as a skeptical reader with the source pack open. It extracts discrete claims (numeric results, comparisons to target or baseline, statements about who was reached, causal or contribution language, completeness of the evidence) and asks whether the cited or attached data actually supports each one at the strength the sentence implies.
Related work sits nearby. A Performance vs. Target Variance Report is often one of the sources this check should read. A Quant-Qual Impact Narrative is a common place overclaiming appears, because qualitative color is easy to write as if it were population-level proof. Funder Report Section Drafting may have produced the draft under review. The accuracy pass comes after drafting, not instead of it.
Output is a structured flag list, not a polished second draft. Each flag should name the claim, point to the passage, state what the sources show (or fail to show), and classify the issue so a reviewer can triage: overstated magnitude, unsupported attribution, missing denominator or period, cherry-picked slice, or assertion with no source at all.
Inputs the check needs
The review is only as good as the pair of artifacts it is given: the draft report and the source data that is supposed to underwrite it. Sources typically include monitoring extracts, survey or administrative tables, finance or output trackers, evaluation annexes, and prior-period figures if the draft claims change over time.
Staff should attach the version that actually informed the draft, not a later dashboard refresh, unless the report is being updated to that refresh on purpose. If the draft cites a table, that table (or an extract with the same filters, period, and population) needs to be in the pack. If the draft cites a qualitative finding, the coded notes or evaluation excerpt that support it need to be in the pack. Vague “see MEL folder” pointers are not a source.
The model should not invent a substitute evidence base. When the draft report is missing, or when source data is missing, the review returns empty output. A half-pack is not an invitation to “do the best you can.” Empty output is the correct failure mode: it tells the reviewer that the check did not run, so they must not treat silence as a clean bill of health.
If sources exist but are incomplete relative to the draft’s claims, the run can still proceed, provided both a draft and a source pack are present. Incomplete coverage then becomes flags (claim has no matching source), not a reason to fabricate numbers or to assume the missing file would have confirmed the sentence.
How claims are tested against sources
The useful unit is a claim, not a paragraph. A paragraph can contain a supported output count and an unsupported outcome story. The review should split those.
Numeric claims are checked for value, unit, period, geography, and population. “Eighty-seven percent of participants improved” is a different claim from “eighty-seven percent of survey respondents who completed both waves improved on this scale.” If the source table uses respondents, completers, or a subgroup, the prose must not widen the group. Rounding that changes the story (for example, treating a small-n percentage as a round, confident result) should be flagged even when the underlying count is real.
Comparison claims need the same comparator the source uses. Beating last year, beating target, and beating a control or comparison group are different statements. A variance table that shows shortfall against target does not support a sentence that only reports year-on-year improvement as if the grant were on track.
Attribution and contribution language is checked more strictly than description. Sources that show association, staff observation, or participant testimony do not automatically support “the grant caused,” “we achieved,” or “the model works.” The review should flag causal or credit-taking verbs when the pack only supports “participants reported,” “outputs were delivered,” or “outcomes moved during the period.”
Completeness claims (“all sites,” “the cohort,” “no serious incidents”) are checked against coverage in the data: missing sites, unreturned surveys, truncated periods, and suppressed small cells. If the source cannot speak to a site or a month, the draft should not.
The model is not the records system of record. It does not correct the spreadsheet. It maps language to evidence and lists gaps.
Flags staff should expect
Overstated results are the most common class: the number is nearby in the pack, but the sentence is stronger than the row. Examples include stretching a pilot site to the whole portfolio, treating a short window as a year, or reporting a midpoint as if it were the final figure.
Unsupported assertions are claims with no matching source, or with a source that addresses a different question. A success story about one household does not support a county-level outcome sentence. A process milestone (curriculum delivered, stipends paid) does not support an impact sentence.
Watch for inherited boilerplate from previous reports. Language that was accurate last cycle can become false when targets, samples, or geographies changed. The review should treat last year’s PDF as a draft input only if it is explicitly in the source pack and the claim is about that prior period.
Also watch for silent selection: highlighting the one indicator that moved while the variance source shows several that did not, or quoting the favorable qualitative theme without the dissenting codes in the same evaluation excerpt. Adversarial review is the right place to surface that imbalance. Staff may still choose to lead with a positive result; they should not be surprised that the pack contains the rest of the picture.
Each flag should be specific enough to edit from. “Tone down the outcomes section” is not actionable. “Paragraph 4 states a 22-point gain for all enrolled youth; Table 2 shows that gain for the 61 youth with paired assessments” is.
What remains a human decision
The model flags. It does not submit. A grants or evaluation lead still owns correction, internal sign-off, and the file that goes to the funder or board.
Some flags are mechanical and should be fixed: wrong period, wrong denominator, a percentage that does not match the table. Others are judgment. How strongly to claim contribution, whether to keep a story as illustration rather than proof, and how to describe mixed results are program and relationship decisions. The review should not auto-rewrite those sentences into a house style. Staff edit, or they document why a flagged sentence stays.
False positives will happen. A source may support a claim through a defined indicator that the model did not align, or through a footnote the pack included as an image. The human reviewer dismisses those flags with a note, rather than treating the list as a score.
The same person who drafted the section should not be the only person who clears flags if the report is high-stakes. A second staff reader, using the flag list as an agenda, is the intended loop: machine extraction of mismatches, human decision on wording, human submission.
Do not run this review as a substitute for evaluation design. If the sources cannot support an impact claim, the fix is to qualify the report, not to ask the model to sound more confident.
When the review returns nothing
Empty output is required when the draft report is missing or the source data is missing. In that state there is nothing legitimate to cross-check, and a generated “no issues found” comment would be misleading.
Empty output is also the right response if the run cannot access both artifacts after a failed upload or a blank extract. Staff should treat empty as “did not review,” attach the draft and the source pack, and run again.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first