AI Adoption GuideConsultingAnalyze
Agentic Research and Synthesis Agent
Multi-step agent runs web and internal document research and produces structured findings briefs for consultant review, using tools like McKinsey Lilli.
Consulting processSellScopeStaffKickoffAnalyzeRecommendDeliverClose
By Don, DoneThat’s AI coach · updated
Lock the question before the agent plans any search
An agent will research whatever you type. Fast synthesis of the wrong question still burns week one. Lock the question in a sentence the engagement manager will defend. Pull the live branches from a MECE hypothesis tree generator. Each branch becomes a search step with an expected source class: a filing, a trade article, a prior-firm memo, a client price list.
Do not start from "what is happening in this market." That prompt produces a readable brief that does not answer the engagement. The agent will plan its own steps. Those steps are useful only if they stay inside a question the team has already agreed to.
For each branch, write three things before any tool runs:
- The claim you are trying to support or kill.
- Where a true answer would live (10-K MD&A, licensed trade data, a prior diligence memo, a dealer survey in the data room).
- What counts as a miss: no source, a source that does not contain the claim, or a source nobody on the team can open this week.
Then let the agent plan search steps inside those constraints. Multi-step research is a sequence: find the competitor's aftermarket language, then the SKU families, then whether the client's channel saw the same shift. Multi-step without a locked question is a longer wrong brief.
If two branches would be confirmed by the same paragraph, they are not exclusive. Fix the tree before you spend the corpus. Ask the agent for a search plan, not an answer. Review the plan against the branches. Kill steps that wander into adjacent industry news.
What a week-one brief looks like on a distributor margin problem
This walkthrough is illustrative. A PE-backed specialty chemicals distributor. Gross margin is under pressure. Leadership wants to know whether the squeeze is national-account price giveaways, mix toward lower-margin OEM blends, or a new importer in two SKU families. You are not writing the recommendation this week. You are building the fact base for those three branches.
Locked question: which of those three mechanisms is actually moving gross margin?
Search plan:
- Public: competitor 10-Ks and earnings transcripts for private-label, import, and aftermarket language; trade press on new entrants; any customs or market data the firm already licenses.
- Firm: prior chemicals diligence memos and aftermarket studies, found through a comparable past scope retriever and loaded for this engagement with a project knowledge base ingestion agent. Use only documents this team is allowed to see.
- Client: board packs and price files. If they arrive as PDFs, run them through an unstructured document extraction pipeline into this engagement's corpus, not into a firm-wide index.
Ask the agent to return a table, not a narrative: branch, claim, quote, locator, corpus label (PUBLIC / FIRM / CLIENT), status. Prose paragraphs bury citations and make spot-checks slower.
A usable finding names the branch, states a claim, includes a quote that appears in the source, and gives a locator a human can open: a URL, a data-room path, or a firm-knowledge ID. Status stays unverified until an analyst opens it.
A bad finding looks like this: "The industrial cleaners import market grew 14% last year (Statista, 2025)." The link opens a real Statista page about a different chemical category, a login wall, or a page that never states 14%. That is a confident fabricated statistic wearing a real-looking citation. It survives a skim. It does not survive a partner who clicks.
If a search step returns nothing, record a miss. Do not fill the cell with a nearby market story. Empty is information. Adjacent is contamination.
Do not fold interview color into the desk-research brief. Stakeholder stories belong in a multi-stakeholder interview transcript analyzer. Mixing them hides whether a claim came from a filing or from one plant manager.
Run public, firm, and client corpora as separate searches
The tools that do this work sit in one class: McKinsey Lilli, Perplexity Enterprise, AlphaSense, Glean, and Hebbia. They are not interchangeable. Some are built around the open web and published filings. Some are built around the firm's own documents and workspace. Some sit on a client data room. Treat them as research agents that can plan steps and return cited synthesis. Do not rank them. Do not assume any of them will refuse to mix corpora unless you configured that boundary.
Run three labeled searches:
- PUBLIC: web, filings, licensed research.
- FIRM: prior engagements, practice notes, expert decks, scoped to what this team may see.
- CLIENT: this engagement's data room only.
The leak failure mode is simple. An agent with a global firm index will cite Client B's pricing study while writing Client A's week-one brief, because both are chemicals distribution. That is a confidentiality break, not a helpful analog. Engagement-level access is a precondition, not a later cleanup.
A prior-firm memo can still be the wrong analog even when access is legal. A 2019 study of commodity solvents does not answer a 2026 question about two specialty SKU families. Label the analog as analog. Do not let the brief present it as this client's fact.
If you cannot technically separate the corpora, do not run internal search until information security says the boundaries hold. Public-only research is slower to make firm-specific. It will not put another client's numbers in this deck.
A finding without an openable source and a quote is not a finding
Drop any row that cannot produce all three:
- A locator a human on the team can open this week.
- A short quote or table excerpt that actually appears at that locator.
- The branch it belongs to.
A citation that does not contain the claim is a fabrication with a costume. Agents do this when the question is quantitative and a nearby source is almost right. Nearby is not good enough for a fact pack. The year, geography, and product definition have to match the branch, not merely the industry noun.
Spot-check is not a sample of titles. Open the source. Find the quote. Confirm the number, if any number is claimed. If the source is paywalled and nobody on the team has access, the finding is blocked. Leave it out or mark it blocked. Do not paraphrase around the wall.
Do not let the brief average across sources into a new statistic the sources never stated. Synthesis can group claims. It cannot invent a market size by blending three incompatible figures.
When the source is a table extracted from a PDF, keep the page pointer. Extraction error on a margin rate is a credibility event once it reaches a client slide. Verify the figure against the page, not against the model's confidence.
Spot-check the brief before it leaves the team
The output is a structured findings brief for consultant review. It is not a client deliverable. Nothing from it reaches a working-session slide until a person has opened the sources behind the claims that will be seen.
Review in this order:
- The engagement manager reads the locked question against the brief. If the brief answered "market trends" instead of the three mechanisms, you researched the wrong question quickly. Throw the brief out. Do not decorate it.
- An analyst opens every finding that will appear in a fact pack or a slide, not a sample of the ones that look important.
- Someone who did not write the prompt checks the CLIENT and FIRM labels. Anything that smells like another engagement gets escalated, not edited into softer language.
What you should have at the end of week one is a shorter fact base on the live branches, with sources the team can defend. A branch may already be dead. The tree may need a redraw. That is the reason to compress desk research: you reach the kill tests sooner. You do not skip reading.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first