Skip to main content
DoneThat

AI Adoption GuideNonprofitPlan

Theory of Change Generation

LLM drafts a theory of change narrative from program parameters and target population descriptors.

Nonprofit processPlanFundOutreachDeliverMeasureReportStewardRenew

By Don, DoneThat’s AI coach · updated

What This Use Case Produces

Theory of change generation turns structured program inputs into a readable causal narrative. The model proposes how activities are meant to produce interim results, how those results connect to longer-term outcomes, and which assumptions sit under each link. The output is a draft for critique, not a finished strategy.

Program designers use this when they already know the program shape (who is served, what is delivered, over what timeframe) but still need a coherent written account of how change is expected to happen. Funders, boards, and evaluation partners often ask for that account early. Drafting it from scratch is slow when the designer is juggling service design, partnership maps, and budget constraints at the same time.

The model can assemble a first-pass narrative that mirrors the language of the inputs you provide. It can surface missing links (an activity with no stated pathway to an outcome, or an outcome with no supporting activity). It cannot invent local context you did not supply, and it cannot decide which trade-offs your organization will accept.

Ownership stays with the program lead. The draft is a starting text for review with field staff, community partners, and evaluation colleagues. Until that review happens, treat the narrative as provisional.

Inputs You Must Provide

Empty output is the correct response when either program parameters or target population descriptors are missing. Do not ask the model to invent a theory from a program name alone.

Program parameters should include, at minimum:

  • Goal or problem statement in plain language
  • Core activities or service components
  • Delivery setting and dosage (intensity, duration, frequency) when known
  • Time horizon for expected change
  • Constraints that shape design (eligibility rules, staffing model, partnership roles, regulatory limits)

Target population descriptors should include:

  • Who is eligible or prioritized, and who is explicitly out of scope
  • Relevant demographics or situational characteristics that affect how the program is supposed to work
  • Barriers the program claims to address
  • Assets or existing supports the pathway assumes (family networks, referral partners, prior services)

Optional but useful inputs include existing logic models, prior evaluation findings, and known risks (attrition, referral leakage, seasonal demand). When those are present, the draft can cite them as assumptions rather than as proven facts. When they are absent, the narrative should mark causal links as hypothesized.

If either the parameter set or the population description is incomplete, return no narrative. A partial theory that quietly fills gaps with generic nonprofit language is worse than a blank page, because it hides what still needs design work.

How Drafting Should Work

Run generation only after the required inputs are present and internally consistent enough to describe one program, not a portfolio of unrelated services.

  1. Normalize the inputs. Restate activities, populations, and intended outcomes in a stable vocabulary so the narrative does not drift between synonyms mid-draft.
  2. Propose a causal spine. Order short-term, intermediate, and longer-term results so each step is a consequence of the previous one, not a restatement of the activity list.
  3. Attach assumptions to links. For each major link, name what must be true for the link to hold (participation, dosage completion, partner follow-through, external conditions).
  4. Flag weak or missing links. Call out activities without pathways, outcomes without mechanisms, and population claims that do not match eligibility rules.
  5. Produce a readable narrative. Prefer short paragraphs over dense matrices. Keep jargon minimal unless the inputs already use a funder’s required terms.
  6. Stop at the draft. Do not auto-approve, publish, or treat the text as the organization’s official theory.

The model should stay faithful to the supplied parameters. If the inputs describe a six-week workshop series for caregivers of school-age children, the draft should not expand into a multi-year systems-change agenda. Scope creep in the narrative creates false confidence and misaligns later measurement plans.

When inputs conflict (for example, an outcome that requires sustained case management while parameters describe a one-touch referral), the draft should surface the conflict instead of smoothing it over.

Human Review and Ownership

Program leads remain accountable for the theory of change. Review is not cosmetic editing. It is a design decision about which causal story the organization is willing to defend.

A practical review checklist:

  • Does every major activity appear in the pathway, and does every major outcome have a mechanism?
  • Do population descriptors match who will actually enroll under current eligibility and outreach practice?
  • Are assumptions stated as assumptions, not as evidence?
  • Would frontline staff recognize this pathway as how the work actually runs week to week?
  • Are risks and failure points named early enough that monitoring can watch them?

Involve people who see different parts of the program. Field staff catch dosage and engagement realities. Community partners catch referral and trust assumptions. Evaluation colleagues catch whether interim results are observable. Leadership catches whether the narrative matches strategy and resource limits.

Revision cycles should change the theory when practice or evidence disagrees with the draft. Do not preserve elegant language that no longer matches operations. After material program changes (new eligibility rules, cut dosage, added partners), regenerate or rewrite rather than patching outdated paragraphs.

How This Fits Adjacent Planning Work

Theory of change generation sits upstream of stress-testing and demand planning. A clear narrative makes assumption testing sharper, because reviewers know which links to probe. Clustering stakeholder feedback can then be mapped onto specific pathway steps instead of sitting in an undifferentiated comment pile.

Demand forecasting is related but distinct. Forecasting estimates how many people may need or seek a service. Theory of change explains how serving them is supposed to produce change. Mixing the two creates muddled documents: volume projections dressed up as causal claims, or causal stories that ignore capacity. Keep population descriptors aligned across both workstreams so the theory and the demand model describe the same people.

Use related pages when you move from drafting to critique:

Limits, Misuse, and Safe Defaults

This use case fails when treated as automatic strategy. Common failure modes:

  • Missing inputs still get a fluent draft. Prevent this with a hard empty-output rule when parameters or population descriptors are absent.
  • Generic sector language replaces local design. Reject drafts that could apply to any program in the domain without referencing your supplied activities and population.
  • Assumptions are written as facts. Keep hedging language where evidence is thin, and require leads to approve any stronger claim.
  • The narrative outruns operations. If staffing, partnerships, or dosage cannot support a link, revise the theory or the program, not the wording alone.
  • Version drift. Store the inputs that produced each draft so later readers know what the model saw.

Safe default behavior: draft only from complete inputs; label the output as a draft owned by the program lead; require human approval before the narrative is shared with funders or published externally; regenerate after material input changes rather than silently editing prior prose.

Used this way, theory of change generation shortens the blank-page phase and makes causal claims inspectable earlier. It does not replace judgment about what your program can honestly claim.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first