AI Adoption GuideNonprofitMeasure
Measurement Plan Validity Review
LLM adversarial review identifies confounders, missing indicators, and validity threats in a draft evaluation framework.
Nonprofit processPlanFundOutreachDeliverMeasureReportStewardRenew
By Don, DoneThat’s AI coach · updated
Why review a draft measurement plan before fieldwork
A complete-looking measurement plan can still fail to support the claims a board, funder, or program team will later want to make. An adversarial review treats the draft evaluation framework as an argument. It asks whether the indicators, comparison logic, and data sources can speak to the outcomes named in the theory of change.
Nonprofit evaluation leads use this review when a program team or consultant has produced a draft framework and the organization needs an independent stress test before locking indicators, sampling, and reporting claims. The evaluation lead owns design. The model is a critic, not the methodologist of record. It flags threats, missing constructs, and plausible alternative explanations so a human can decide which risks to accept, which to mitigate, and which to rewrite.
Time this review after a theory of change exists and before instruments go into the field. It is not a substitute for watching data once collection starts; that work lives with Dataset Quality Monitoring. It is also not a forecast of results. Early signals from incomplete outcome data are a separate task: Mid-Program Outcome Prediction.
Required inputs and empty output
Run the review only when three artifacts are present. You need a draft evaluation framework that states questions, claims, and intended uses. You need a named set of indicators with data sources and timing. You need a program description that states who is served, what is delivered, over what period, and in what operating context.
If any of those is missing, the output must be empty. A critique invented from a slogan, a logic-model heading, or a funder template without indicators will invent threats that do not apply and miss the ones that do. Outcomes listed without measures, or a survey draft without a comparison strategy, are not enough. Those packets invite the model to complete the design, which it must not do.
Optional context helps only after the three artifacts exist. Useful additions include the audience for findings, known constraints such as sample size, access, ethics, and staff time, prior evaluations of the same program, and any comparison or contribution story the team already plans to tell. Do not treat optional context as a license to fill gaps in the required packet.
A useful critique is specific. Each finding should cite the claim or indicator it attacks, name the validity threat in evaluator language, explain the mechanism, and suggest a design or claim change without rewriting the framework. Group findings by whether they break a primary claim, weaken a secondary question, or are residual limitations. Vague notes such as "consider selection bias" with no path through this program's intake are not review output.
Confounders and comparison logic
Ask first whether observed change could reasonably be credited to the program. History (a policy shift, seasonal demand, a peer program launching nearby), selection (who enrolls versus who is eligible), maturation, attrition, and instrumentation changes are ordinary threats in nonprofit work. The model should name each threat in plain language, tie it to a specific claim or indicator, and say what would have to be true for the threat to matter.
Check whether a comparison is specified at all. Before-after designs without a credible counterfactual, matched comparisons that ignore unobserved selection, and participant versus non-participant stories that collapse eligibility, motivation, and dosage into one contrast are common. The critique should not prescribe a gold-standard design the organization cannot field. It should state what the current design can and cannot support.
Treat dosage and fidelity as validity issues, not only as implementation notes. If the framework treats "enrolled" as "received the intervention," outcome movement can be credited to a service that many participants barely received. Missing process or implementation indicators leave outcome claims uninterpretable. The model should ask whether the plan can distinguish non-receipt, partial receipt, and full receipt, and whether those groups are mixed in the outcome tables the team intends to publish.
Missing indicators and construct coverage
A plan can carry many metrics and still miss the construct. If the named outcome is stable housing and every indicator is a point-in-time shelter exit, the framework under-represents duration, returns, and housing quality. If the named outcome is youth wellbeing and the only measures are attendance and grade promotion, program process is standing in for the construct.
Map each evaluation question to indicators and mark the gaps. Typical holes include no baseline, no harm or unintended-effect indicators, no equity disaggregation where the theory of change implies differential effects, and no measure of the mechanism the program claims to use. Flag mono-method bias when a construct rests on a single survey item or a single administrative field.
Catch indicator-claim mismatch. Output counts presented as outcome evidence, satisfaction items used as effectiveness, and short-horizon measures used to stand in for long-term change weaken the argument even when the numbers are clean. Pulling quotes or results from reports does not repair a weak plan. Use Outcome Evidence Extraction only after the constructs themselves are worth extracting.
Claim language, ethics, and proportionate design
Statistical conclusion threats matter when the plan already announces analyses the data cannot carry: underpowered subgroup claims, a long list of unadjusted comparisons, or noisy administrative fields treated as precise. Construct validity threats matter when the operational definition does not match the language used with funders. External validity threats matter when the sample is a convenience cohort and the intended use is a statement about the population.
Separate problems the team can fix in design from problems that require more modest claims. Adding a comparison group, a follow-up wave, or dropping an overclaimed outcome are design moves. Switching from attribution to contribution, or from proof to learning, are claim moves. The model should not invent precision the data cannot support, and it should not demand experimental language for a developmental evaluation.
Ethics and burden sit inside validity, not beside it. Instruments participants cannot complete, indicators that require identifiable sensitive data the organization cannot protect, and measurement that changes the service itself are design defects. Name them as such so the evaluator can drop, redesign, or gate those measures before anyone is asked to provide data.
Using the critique without handing over the design
Read the output as a structured challenge list, not as a rewrite. For each flagged threat, decide to mitigate in design, monitor as a limitation, or change the claim. Keep a short decision log so later reporting can say which threats were accepted.
Do not paste model language into the evaluation plan. Translate accepted points into indicator definitions, sampling notes, analysis caveats, and use statements. If the model proposes a new indicator, the evaluator specifies the operational definition, data owner, frequency, and quality rule, or rejects it. The model does not add indicators to the live framework, assign data owners, or approve instruments.
Re-run the review after a material revision, not after every wording tweak. Stop when remaining threats are known, documented, and proportionate to the intended use of findings. The standard is a plan that can honestly support the claims you will make, not a plan with no remaining risk.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first