AI Adoption GuideConsultingDeliver
Client Satisfaction Sentiment Monitor
LLM continuously analyzes client emails and meeting transcripts for satisfaction signals and escalation risk patterns.
Consulting processSellScopeStaffKickoffAnalyzeRecommendDeliverClose
By Don, DoneThat’s AI coach · updated
Route every alert to the engagement manager, never to the client
The monitor exists to give the engagement manager a few hours of warning, not a scorecard the client, the steering pack, or the account dashboard can see. If a packet would embarrass the firm if someone forwarded it, it does not belong in a shared channel.
A useful output is a private note: who said what, on which retained thread or call, which pattern it matches, and what the engagement manager might do this week. A useless output is a client-facing "relationship health" tile, a weekly happiness number on the status pack, or an automated mail that tells the sponsor you noticed they "seem frustrated." Do not auto-email the client. A human still decides whether to call, absorb extra work, raise a change order, or wait.
Escalation on a live job often shows up before dates slip. Pair this inbox with the workstream delivery risk predictor rather than treating tone as a substitute for missed milestones. After the last invoice, stop scraping the live mailbox and hand the relationship to the post-delivery relationship health monitor.
Score the pattern, not the tone
Sentiment scores are weak on consulting mail. Polite clients write politely while they escalate. Direct cultures sound negative to a model trained on US corporate English. Sarcasm, dry humor, and a non-native accent in a transcript are ordinary noise. Treat a single "negative" sentence as unusable.
Look for sequences a human would already worry about:
- The same ask appears twice or more with no firm answer from the team: status, access, a named deliverable.
- The economic sponsor goes quiet: delayed replies, skipped standups, a deputy sent in their place, calendar holds cancelled without a new date.
- A new legal, procurement, or audit address appears on threads that used to be working-level.
- Language shifts from "can we" to "please confirm in writing," "per the SOW," or "our counsel."
- Attendance on the steering call thins while the working sessions stay full.
Those patterns are behavioral. They survive sarcasm. They also overlap real delivery fights. An unanswered ask that is actually a new request belongs with the scope creep detector, not with a mood label. An unanswered ask that was already owned in last week's notes belongs with the meeting-to-action-item agent: the relationship risk is often that the team dropped a commitment, not that the client "turned negative."
A weekly "happiness score" for the account is a failure mode. Teams start writing warmer emails to the model. Partners argue about the number. Nobody calls the sponsor. Drop the score. Keep a short list of open patterns, with evidence and an owner on the consulting side.
Only read mail and recordings you already have permission to keep
Run the model only on mail and recordings the firm already retains under the engagement's mailbox and recording rules. Do not expand the harvest to make the monitor more complete.
Mailbox rules, in practice:
- Shared engagement mailboxes and the named project aliases the SOW already contemplates.
- Client emails the team is allowed to file in the matter folder.
- Not personal inboxes, not partner-to-sponsor side threads marked private, not internal venting on a team channel, not HR or ethics mail that happens to mention the account.
Recording rules, in practice:
- Transcripts from calls where recording was disclosed and the client did not refuse.
- Not a secretly captured hallway recap. Not a transcript of a meeting the client asked not to record. Not a vendor bot that joined a call the client never approved.
If the relationship would not survive the client finding out you scored their mail, do not build the pipeline. That question is the gate. A quietly unhappy client still writes politely. Reading private threads will not fix that, and it will create a worse problem if counsel ever asks how the file was assembled.
Readiness of the client organization is a separate check. A sponsor who was never staffed to govern the work will look quiet from day one. Use the client organizational readiness classifier at kickoff so you do not treat an empty calendar as a sudden sentiment drop.
Illustrative example: happiness stayed green after the sponsor left
The walkthrough below is illustrative: realistic firm details, no measured results.
Engagement manager at a 90-person advisory firm. Sixteen-week ERP process-redesign workstream inside a larger systems program. Weekly steering with the client's VP of Operations, plus two working sessions. Volume on the order of a few dozen client emails a week and two recorded calls.
What they tried
A weekly pulse: the model scored every client email and transcript from negative to positive and posted a "happiness" tile in the internal status pack. Conversation tools already produced transcripts. The tile consumed those files plus the shared mailbox.
What broke
Week nine, the VP stopped joining steering and sent a director. Replies got shorter. The happiness score stayed mildly positive because the director wrote "thanks" and "looks good." Week eleven, legal appeared on a thread about data extracts. The model tagged one sarcastic "great, another workshop" as strongly negative, and the engagement manager spent a day on a non-event. Week twelve the VP asked, in writing, for a pause. Nobody had called them. The team had been managing a number.
What they changed
They killed the score. The monitor now emits a packet only when a pattern fires: unanswered ask repeated, sponsor absent from two consecutive steers, new legal or procurement cc, or "in writing / per SOW" language on a working thread. Packets go to the engagement manager's private channel, with quotes and links to the retained file, never to the client, never onto the steering pack. The engagement manager still decides whether to call. They did not auto-mail the VP. They stopped reading a partner's private side thread with the CIO.
Keep conversation tools as sources, not as the monitor
Gong, Otter.ai, Fireflies, and Microsoft Copilot sit in the conversation-intelligence and meeting-notes class. Use them, where the client has agreed to recording, as the place transcripts already live. Do not treat their meeting notes or call summaries as a consulting satisfaction metric. Do not build a firm-wide ranking of vendors. The monitor is a thin layer: allowed mail plus allowed transcripts, pattern rules, a human inbox.
A packet worth sending contains:
- The pattern name (repeated unanswered ask, sponsor quiet, new legal cc, written-confirmation language).
- Two or three quotes from retained files, with date and thread or meeting title.
- Whether the same item is already on the action log or looks like new scope.
- A suggested next step for the engagement manager, worded as a question, not as a mail to the client.
Trial it on one live job, in shadow, for a few weeks:
- Confirm mailbox and recording consent on that job before any model runs.
- Show packets only to the engagement manager until false positives are boring.
- Drop patterns that fire on sarcasm, accent, or a single blunt sentence.
- Only then let a partner see a digest. Still never the client.
Judge success by whether a human called earlier than they would have, with a specific thread in hand. Judge failure by a weekly score, a packet that quotes a private thread, or a client who learned they were being scored.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first