AI Adoption GuideConsultingDeliver
Scope Creep Detector
LLM compares incoming client requests and meeting notes against the original SOW and flags out-of-scope items in real time.
Consulting processSellScopeStaffKickoffAnalyzeRecommendDeliverClose
By Don, DoneThat’s AI coach · updated
A cited flag is the product; a commercial email is not
Compare each incoming client request and meeting note to the signed statement of work. If the ask sits outside in-scope or inside an exclusion, flag it with the clause quoted. The engagement manager then absorbs it, prices a change order, or refuses. That is the whole job.
The model does not write to the client. It does not open a change request in the PSA. It does not tell the team to stop. A flag that mails the sponsor a commercial paragraph is a relationship incident dressed up as governance. Keep the queue inside the firm.
Scope creep is rarely one dramatic request. It is a string of small ones that never got written down: an extra site, a live report, a third review round, "while you're in there." If nobody maps those asks to the signed text while they are still fresh, week ten becomes "we thought that was included."
This is a quality check on what you are about to do. A flag that cannot quote a clause is a hunch. Treat hunches as comments, not as findings.
Run the comparison on the signed text and the incoming ask
A model that only sees the latest Slack thread will invent a scope. A model that only sees the SOW will miss what the client just asked for. Attach both: the signed in-scope and exclusions (the PDF or Word that legal closed, not the proposal deck), and the incoming artifact.
Incoming artifacts are ordinary: an email from the sponsor, a ticket in Jira, Asana, or monday.com, a recap from a meeting-to-action-item agent, a transcript from Gong. Microsoft Copilot can sit in Word or Outlook and run the same comparison if you paste the signed clauses and the ask. None of those tools is the system of record for scope. The signed SOW is.
Require a fixed output shape or the flags will be essays:
- One-sentence restatement of the ask
- Verdict: in-scope, out-of-scope against a named in-scope clause, or inside a named exclusion
- Quoted language from the signed document, with section or exhibit
- Why the ask does not fit that language
- Suggested options for the engagement manager only: absorb, change-order, refuse. Do not pick one.
Cap the first pass. A weekly digest of a handful of flags will get read. A firehose of "possible misalignment" will not. Keep flags that would add a deliverable, a site, a system, an extra review cycle, or ongoing operational work. Drop tone, politeness, and "they asked a clarifying question."
Do not compare the ask to what this firm usually does, or to last year's similar project. "We usually include a dashboard" is not a clause. If the signed text is silent, that is a gap you should have caught in scoping, not a license to treat custom as included.
Route every flag to the engagement manager, or to the partner they have deputized. Never to the client, never to the full delivery Slack channel as a gotcha.
Illustrative week: a Tableau dashboard in the steering notes
Illustrative only. No measured outcome.
A fourteen-week operating-model engagement is in week six. The signed SOW in-scope is current-state workshops, process maps, a future-state design, and an implementation roadmap. Exclusions name "build or operate reporting," "system configuration," and "ongoing PMO services." Hypercare and a live data product are not on the page.
In steering, the sponsor says, "While you're in the data, can you also stand up a weekly ops dashboard in Tableau so leadership can see the baseline?" The analyst hears a helpful add-on and opens a ticket in Asana: "Stand up weekly Tableau dashboard." Gong has the utterance. Nobody re-reads Appendix B.
Run the model on the signed SOW plus the ticket and the transcript excerpt. A useful flag looks like this:
- Ask: weekly Tableau ops dashboard for leadership, sourced from current-state data.
- Verdict: inside exclusion "build or operate reporting"; not covered by in-scope process maps.
- Quote: Exhibit B, Exclusions: "Firm will not build or operate reporting or dashboards."
- Why: a live weekly Tableau workbook is operated reporting, not a process map appendix.
- Options for the EM: absorb a one-page static snapshot in the next readout (goodwill, not a live dashboard); raise a change order for a reporting workstream; refuse and point at Exhibit B.
The engagement manager absorbs the static snapshot, writes "live Tableau out of scope" in the working file, and tells the analyst to close the Asana ticket. They do not email the sponsor a change-order template that same afternoon. The sponsor asked in a steering meeting. An auto-generated commercial note would read as a snub.
The model will not know that the sponsor is using "dashboard" to mean "put the process map numbers on one slide." That is ordinary consulting judgment. If the flag cannot tell a slide from a product, the EM has to.
Ordinary judgment is not creep; a verbal "we will also" still is
If you flag every extra hour of thinking, engagement managers mute the channel. Choosing which six people to interview when the SOW says "stakeholder interviews," rewriting a slide for a different audience, adding a breakout inside a paid workshop: that is delivery judgment. It is not a new deliverable. Prompt the model to ignore method choices inside an already-named artifact.
The opposite miss is worse. A verbal "we will also" in Gong is easy to skip because it is informal, late, and unrepeated. "Can you also sit in on the Tuesday ops huddle for a couple of months" is ongoing work. A meeting-to-action-item agent may capture it as a task with an owner and never check the SOW. The creep detector has to see that task, or the huddle becomes unpaid operating cadence.
Vague signed language produces junk. "Support the transformation" and "as required" will flag everything and nothing. Tighten the SOW before you automate the check. If the in-scope list cannot support a clause citation, stop. You are asking the model to litigate a document that was never specific enough to litigate.
Do not auto-raise a change order, and do not skip the assumptions pass
A flag is not permission to create commercial paper. Auto-raising a change order in Jira, Asana, monday.com, or the PSA, or auto-sending a rate-card email, trains the client to stop telling you what they need. Batch flags. Look at the cumulative picture, not each ask in isolation. Raise the pattern at the next milestone conversation, unless a single ask would consume the week.
Until a change is priced, signed, and recorded, the roster does not move. Use agentic re-staffing on scope change only after an approved change order exists, not after a detector flag. Accumulated unabsorbed flags belong in workstream delivery risk predictor as a delivery-risk input: work that is happening without a clause behind it.
This detector cannot recover an assumption you declined to write down. If data-extract fitness, named client FTE, or an iteration cap never made it into the signed SOW, every related fight will look like creep or like you under-scoped. Run assumption gap detector before signature. If the draft started from a SOW drafter from discovery transcript, the verbal "we will also" should already be an open item, not a week-six surprise.
If partners will not have the milestone conversation, do not spend cycles refining prompts. The flags will pile up. The team will still do the dashboard.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first