AI Adoption GuideNonprofitPlan
Logic Model Assumption Stress-Test
LLM adversarial review challenges causal assumptions in a draft logic model against the published evidence base.
Nonprofit processPlanFundOutreachDeliverMeasureReportStewardRenew
By Don, DoneThat’s AI coach · updated
What this use case does
Logic Model Assumption Stress-Test takes a draft logic model (or an equivalent causal chain from activities through outputs to outcomes and impact) and runs an adversarial pass against a defined evidence base. The model’s job is not to rewrite the program design. It is to surface which if–then links look weakly supported, over-specified, context-blind, or internally inconsistent, and to show why.
A strategy or evaluation lead typically uses this after a first draft exists and before the model is locked for a proposal, board packet, funder report, or evaluation plan. The stress-test is most useful when the team has already named intended outcomes and needs a disciplined check on whether the causal story holds under scrutiny.
Related planning work often sits nearby: Theory of Change Generation helps produce the narrative structure; this page focuses on pressure-testing the assumptions that narrative depends on. Stakeholder input and demand signals can inform later revisions via Stakeholder Feedback Clustering and Service Demand Forecasting, but they do not replace an evidence-based assumption review.
When to run it (and when not to)
Run the stress-test when you have both of the following:
- A draft logic model with identifiable links (for example, activity → output → short-term outcome → intermediate outcome → long-term impact), including any explicit assumptions the team already wrote down.
- An evidence corpus you are willing to treat as the review standard: peer-reviewed studies, systematic reviews, evaluation reports, government or foundation evidence summaries, or your own prior evaluations with clear methods notes.
Do not run it as a substitute for program design workshops, community consultation, or evaluator judgment. An adversarial model can challenge wording and cite tensions in published literature. It cannot decide which trade-offs your organization accepts, which outcomes matter most to the people you serve, or how to weigh conflicting evidence in your specific context.
Skip or defer the pass when the draft is still a vague aspiration list with no causal structure, when evidence sources have not been selected, or when the only “evidence” available is marketing material or unsourced claims. In those cases the right next step is to complete the draft and assemble sources, not to generate a false sense of rigor.
Inputs, workflow, and outputs
Inputs (required). The draft logic model in a structured form (table, linked statements, or annotated theory of change) plus the evidence set (documents, bibliographic records, or excerpts the organization has approved for use). Optional but useful: population and setting notes, dosage or intensity assumptions, implementation constraints, and known risks already flagged by the team.
Workflow. Parse the model into discrete causal claims. For each claim, attempt to locate supporting, qualifying, or contradicting material in the evidence set. Flag assumptions that are implicit (unstated but required for the chain to work), contested in the literature, dependent on conditions the draft does not mention, or mismatched to the stated population or delivery model. Produce a challenge memo organized by link, not a polished rewrite of the full model.
Outputs. Expect structured findings such as: challenged link; assumption stated in plain language; evidence status (supported, mixed, unsupported in the provided corpus, or not addressed); severity (blocks the chain vs. weakens confidence); and suggested clarifying questions for the program lead. The system should not silently invent new outcomes or “fix” the model. Where the corpus is silent, say so.
Empty output rule. If the draft model is missing, unreadable, or lacks identifiable causal links, return no stress-test findings. If the evidence sources are missing or empty, return no stress-test findings. Partial runs that speculate without a corpus create false confidence; empty output is the correct failure mode.
What gets challenged in practice
Adversarial review focuses on the load-bearing joints of the logic model, not on wordsmithing activity lists.
Causal leaps. Does “workshops delivered” reliably imply “behavior change,” or does the draft skip mediators such as practice opportunity, supervision, incentives, or environmental barriers? The challenge should name the missing step and ask whether evidence in the corpus supports that leap for a comparable population and dosage.
Context transfer. Evidence from one setting (for example, urban clinic-based delivery) may not transfer to another (rural outreach, school-based, or digital-only). The review should call out transfer risk when the draft cites or assumes results from mismatched contexts without adaptation assumptions.
Dosage and fidelity. Many outcome claims depend on intensity, duration, staff skill, and fidelity. If the draft promises outcomes associated with intensive models while specifying light-touch activities, that contradiction should be flagged.
Attribution and contribution. Logic models sometimes imply that the program alone produces community-level impact. The stress-test should push for contribution language where the evidence base shows multi-actor pathways, secular trends, or selection effects.
Negative and null findings. A useful adversarial pass surfaces studies or evaluations that found null or adverse effects under conditions similar to the draft. Suppressing those references defeats the purpose of the review.
Measurement readiness. If an outcome is central to the chain but the draft has no feasible indicator, data source, or timing for observation, the challenge should note that the assumption is currently untestable, not that it is therefore false.
Ownership, human review, and quality bar
Program and evaluation leads retain ownership of the logic model. The model challenges; people decide. Treat every finding as a prompt for revision, escalation, or documented acceptance of residual risk, not as an automatic edit.
A practical review loop looks like this: run the stress-test; triage high-severity challenges with the design team; revise links or make assumptions explicit; re-run only on changed sections if needed; record decisions (revised, accepted with caveat, deferred pending evidence). Funders and boards should see the owned model, not a raw adversarial transcript, unless transparency of the challenge process is itself a deliverable.
Quality depends on corpus quality. A narrow or outdated evidence set produces narrow challenges. A corpus that excludes null results will under-challenge optimism. Teams should document what was in scope and what was excluded so later readers can interpret the stress-test’s limits.
This use case supports the quality outcome in the plan stage: fewer unsupported causal claims survive into proposals and evaluation frameworks, and residual uncertainty is visible rather than buried in diagram aesthetics. It does not certify that a program “works.” It reduces the chance that a polished logic model conceals a weak middle.
When the draft and evidence are ready, the stress-test earns its place between first draft and lock-in. When either input is absent, produce nothing and wait for the team to complete the prerequisites.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first