AI Adoption GuideConsultingScope
Scope Risk Classifier
LLM classifies each scope element by delivery risk, including ambiguity, dependencies, and client readiness, against past project patterns.
Consulting processSellScopeStaffKickoffAnalyzeRecommendDeliverClose
By Don, DoneThat’s AI coach · updated
Score the written lines, not the missing ones
Classify each already-listed scope element for how likely it is to blow up delivery: ambiguous wording, waiting on the client, data that may not exist, an environment that may not be ready, and political conditions the text pretends are settled. Output is a reason attached to a line, not a new clause and not an hour count.
This is not an assumption gap detector. That pass hunts for what the SOW never said. This pass scores what it already said. Mixing them produces a wall of "risks" that are really missing sentences. Write those first, then classify what remains.
This is also not effort estimation from historical actuals. Hours are a different question. A high-risk line may deserve a wider range later. It does not become a padded midpoint here. Using risk labels to inflate every estimate is a commercial move dressed as delivery hygiene.
Run it in a general LLM on the SOW text: Copilot in Word, Claude, or ChatGPT, with your firm's risk types in the prompt and, if you have them, short notes from comparable overruns. Professional services automation in Kantata or Planview is a reasonable place to store accepted scores in a risk log. Those products do not ship this classifier. Do not treat a PSA risk field as model output.
Pull similar closed work with a comparable past scope retriever and use overrun reasons as calibration. Superficial industry match is not a pattern.
Ask five questions of every listed element
Work line by line through in-scope items, client responsibilities, and named dependencies. Skip background and boilerplate. Require a label and a one-sentence reason on each axis. "High" without a reason is not a classification.
- Ambiguity. Can a delivery lead staff this without a follow-up call? "Support the replatform," "optimize the operating model," and "enable stakeholders" are high. Mitigation is rewrite, not a bigger contingency.
- Client dependency. Does the line wait on a named client action with a date, or on "timely access" and "reasonable availability"? Unnamed owners and unbounded review cycles are high. Mitigation is a named person, a response time, or stop-work if the input does not arrive.
- Data. Does the work treat extracts, quality, or historical depth as settled because the SOW wrote them as fact? "Five years of clean orders" as a given is high if nobody has seen a sample. Mitigation is a sample before kickoff, a quality clause, or a phase gate.
- Environment. Does delivery need a sandbox, VPN, non-prod that resembles prod, or access through a third-party host? "Work in client staging" with no owner for refresh and access is high. Mitigation is an environment checklist as a client obligation, or work in your shop until staging exists.
- Political readiness. Does the line depend on alignment the document asserts ("steering committee monthly," "single accountable sponsor") without evidence that those people can actually decide? This axis is the weakest in the model. Score it unverified unless you have a human note. Do not let fluent governance language become a green score.
Cap what you act on. The engagement manager picks a small set of lines to reword, rephase, convert to a client obligation, or take to the partner as a walk-away. If the prompt says "find all risks," you will get a red SOW.
Illustrative SOW: a twelve-week commerce replatform
This is an illustrative scenario, not a case study.
A mid-market retailer has a draft for twelve weeks of advisory around a commerce replatform: current-state of merchandising workflows, a target architecture on paper, integration design for the existing PIM, a cutover checklist, and "change enablement for store and digital teams." Client responsibilities already in the draft: weekly product extracts, staging that mirrors production, and a monthly executive steering committee.
- PIM integration design is high on ambiguity. "Existing PIM" does not name the product, catalog count, who owns attribute mapping, or whether the design includes a proof in staging. Rewrite toward a named PIM, a bounded catalog count, and a design pack with a review date.
- Weekly product extracts is high on client dependency and data, even though the line is already written. The owner is a function, not a person. Cadence is stated; quality and schema are not. Keep the line. Add a named data owner and a first-extract sample before week two, or a pause if the sample fails.
- Staging that mirrors production is high on environment. Mirroring is a claim. Who provisions it, who refreshes data, and whether payment and identity are stubbed is unsaid on this line. Classify the written claim. Do not turn it into a new assumption list; that is the other pass. Mitigation: "advisory design only until staging is demonstrated," or access and refresh as a dated client obligation.
- Change enablement for store and digital teams is high on ambiguity and unverified on political readiness. The phrase is a slogan: not town halls, a training pack, or two workshops. The model often red-flags politics here because two teams are named. That is a guess. The partner still has to say whether merchandising and e-commerce share a cutover date.
- Monthly executive steering committee looks like governance and often scores low on a generic prompt. Treat political readiness as unverified. A committee on the page is not a decision right. If the engagement manager knows the CIO is leaving, that note belongs in the human overlay. The model will not see it.
A sloppy run paints every line red, invents a "data quality workstream" that was never in scope, and quotes hours. None of that is this job.
A full-red sheet is not a ranking
A classifier that marks most of the SOW as high has not ranked anything. Delivery leads stop reading. Partners start using "high risk" as permission to pad.
Force discrimination:
- Require at least one axis that is not high on every line, or require the model to rank lines against each other inside this SOW.
- Act on a handful of lines, not the whole list. Reword, rephase, or add a named client action. Leave ordinary consulting uncertainty alone.
- Separate client-wait risk from wording risk. A crystal-clear client obligation can still be high dependency. The fix is commercial (stop-work, date, named owner), not more adjectives.
- Recheck at close: which flagged lines actually hurt, and which red ink was noise. Feed that into the next prompt as examples. Without that loop you are running generic caution.
If you cannot get a short, ranked list a partner will argue with, stop. You do not yet have a classifier. You have a highlighter.
Put accepted scores where people will see them
Once a partner accepts a score, it can live next to the work: on the SOW comment, in the engagement workbook, or later in a Kantata or Planview risk log. The PSA is a filing cabinet. It does not classify. Do not copy model output into the log without an accept/reject step.
Widen hour ranges for the few accepted high lines when you estimate. Do not raise the midpoint because a paragraph sounded scary.
After signature, incoming asks are a scope creep detector problem. During delivery, velocity and issue trend are a workstream delivery risk predictor problem. Do not keep re-scoring the original SOW as if that were live delivery risk.
Political readiness stays with the people who know the account. The model reads "sponsor," "steering," and "aligned leadership" and will either treat theater as readiness or red-flag every client. Both failures are normal. Write "unverified, needs partner note" when the text is only governance language. The partner adds what the PDF cannot contain: a sponsor who cannot overrule IT, two functions that do not share a date, a procurement freeze that is not in the SOW.
The system is working when a few written lines change (clearer artifact, named client action, a phase gate) and the rest of the SOW stays ordinary. It is failing when every element is red, hours move because of adjectives, or a steering-committee sentence is treated as proof that the client is ready.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first