Skip to main content
DoneThat

AI Adoption GuideConsultingAnalyze

Multi-Stakeholder Interview Transcript Analyzer

LLM extracts themes, contradictions, and sentiment signals across multiple stakeholder interview transcripts simultaneously.

Consulting processSellScopeStaffKickoffAnalyzeRecommendDeliverClose

By Don, DoneThat’s AI coach · updated

One-by-one summaries bury the contradiction

The useful output is a contradiction map across roles, with the quotes that make each disagreement real. A folder of per-interview recaps is the wrong artifact. Each person can sound consistent while operations, finance, and the plant describe incompatible worlds.

That false consensus is the quality failure. Speed of synthesis is a side effect. A fast "the organization believes X" that erases a live split is worse than a slower read that names who disagrees.

Transcripts often arrive from tools in the Otter.ai, Fireflies, Gong, and Microsoft Copilot class, each with a per-meeting recap. Use the transcript as the source file. Do not treat the recap as the analysis. A recap is one conversation. One interview is notes. Neither can show a split across roles.

Load the full set, then ask for disagreements first

Do not loop "summarize this interview" across the folder. Load every diagnostic transcript in one pass, with speaker, role, site or function, and date on each file. If a file has no role, or still says Speaker 2, fix the labels first. A nameless speaker cannot appear on a contradiction-by-role map.

Keep the team's huddles and any sales-cycle recordings out of the set unless you are studying the sale. Mixing those files with diagnostic interviews invents a consensus that existed only in the pitch or in the project room. Use the same discussion guide so silence is meaningful. If only the plant was asked about night counts, "HQ never mentioned night counts" is not a finding.

Ask for three layers, in this order:

  1. Shared claims: a statement in two or more interviews, each with a supporting quote and a role label.
  2. Contradictions: two or more roles asserting incompatible facts or causes about the same object, each with a quote. This is the headline.
  3. Role-only claims: a statement in only one role. Keep these. Do not promote them to "the organization."

Themes come after the contradiction map. Ask for themes first and the model will cluster "inventory" or "culture" and hide that the quotes under the cluster disagree.

Report every disagreement with the quotes, not a paraphrase. "Operations and finance disagree on inventory accuracy" cannot survive a partner review. You need the plant manager on skipped night counts and the controller on write-offs sitting inside policy.

Configure attribution to match what interviewees were told. If you promised nothing attributable to them, the output is role and site at most, and you still strip quotes only one person could have said. If you promised named, on-the-record comments, names can stay. Write the attribution rule into the prompt the same way you wrote it into the interview opener.

Sentiment scores are optional and weak. Treat a confident score on a bad transcript, an accent, or another language as a flag to re-listen, not a finding.

Before anything reaches the partner, the lead interviewer (or the engagement manager who sat in) marks each top theme and each contradiction: heard it, did not hear it, or stretch. The model does not override the person who was in the room.

If these interviews are meant to feed a later SOW, keep the set out of the SOW drafter from discovery transcript until the contradiction map has been reviewed. A draft that papers over a live split is an assumption gap you would be signing.

Illustrative example: Meridian Foods, week three

This is a worked example with made-up names, written to show the cuts, not a case study with results.

Sam is the engagement manager on an operations diagnostic at Meridian Foods. Week three produced eight interviews across HQ operations, the Westbrook plant, finance, IT, warehouse nights, and a union steward. Transcripts arrived from Otter.ai and Microsoft Copilot. Two Fireflies huddle files and Gong recordings from the sales cycle stay out of this set.

Sam's analyst ran each file through a recap prompt. Every recap mentioned inventory, service levels, and alignment. The partner asked for the story. It looked like inventory accuracy was the issue, and the client already agreed.

Then Sam loaded the eight diagnostic transcripts as one set and asked for contradictions first.

Priya, VP Operations: the ERP cannot support cycle counts at the SKU grain the business needs. Tomas, plant manager: the system is fine; night shift does not complete the counts already on the calendar. Mei, controller: write-offs sit inside the policy band, so inventory is not a finance problem. Anders, night warehouse supervisor: they skip counts when outbound is late, and they keep a paper workaround HQ has never seen. IT: cycle-count transactions post; they have no ticket that looks like an ERP defect. The union steward talks about overtime on outbound, not about counts.

That is not one theme called inventory. That is five incompatible explanations of the same object. The per-interview dump hid it because each person was consistent with themselves.

Sam's attribution rule from the interview opener was role and site, not names. Anders's paper-workaround quote is still identifying: there is one night supervisor at Westbrook. It has to be rewritten as "a front-line warehouse role at Westbrook described an undocumented workaround" or dropped from anything that leaves the team. Priya's ERP claim can stay as HQ operations; several people at HQ could have said it.

The interviewer who sat with Anders marks heard it. The interviewer who sat with Mei marks heard it, and she was calm, not frustrated, which kills a sentiment label the model had printed on her transcript. English is Mei's third language. The transcript is choppy. A negative score on that file is noise.

The sample is HQ plus one plant, days plus one night supervisor. No other site, no customers, no 3PL. Sam can say Westbrook and HQ disagree on cause. Sam cannot say Meridian believes X.

Identifiable quotes, confident sentiment, and a sample that is not the firm

Identifiable quotes that break the promise. Unique stories, numbers, complaints, and quotes that only one seat could have said will identify the speaker in a role-only report. Read the output as a hostile reader who knows the org chart. If a quote could only have come from Anders, it is named, whatever the label says. Honor the opener you read in the room, not the model's citation style.

Confident sentiment across accents and languages. Transcript quality drops with overlapping speech, accents, and languages the dictation model was not built for. Do not put "finance sentiment: negative" on a slide. If something sounded off, re-listen. If you cannot re-listen, drop the score.

A skewed sample treated as the organization. Eight interviews at HQ and one plant are eight interviews at HQ and one plant. The model will still write "stakeholders agree" because agreement is a cheap sentence. Force the output to name who was in the sample and who was not.

Do not treat this as a meeting-action extractor. Decisions and owners from a working session belong in a meeting-to-action-item agent. That is a different set and a different promise.

Ship the contradiction map, then hang themes on the tree

The artifact the partner should see is a short table of disagreements by role, each with quotes that survive the attribution rule, plus a one-line sample note. Themes are the second page. Each theme maps to a branch on the MECE hypothesis tree generator. If a theme has no branch, it is a new hypothesis or it is noise. If a branch has no quotes, it is still a hypothesis, not a finding.

Contradictions that are really "we will not do this" or "we do not have the people" are readiness, not operational fact. Pass those to the client organizational readiness classifier as political feasibility and capacity, separate from the fact pattern.

Do not auto-resolve the disagreement. The engagement manager decides what to test next: a floor walk at night, a write-off extract, a system demonstration. The model can suggest tests. It does not pick a single root cause so the deck can look decided.

If the output is twelve summaries with a list of themes, you ran the wrong job. If you later write a SOW from this discovery, run the assumption gap detector against the draft. A SOW that assumes the ERP is the constraint, when Westbrook says the constraint is night-shift practice, is a dispute you scheduled yourselves.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first