AI Adoption GuideConsultingClose
Project Knowledge Base Ingestion Agent
Agent classifies, tags, and indexes project IP into the firm's knowledge graph for retrieval by future engagements.
Consulting processSellScopeStaffKickoffAnalyzeRecommendDeliverClose
By Don, DoneThat’s AI coach · updated
Connecting the project site is not ingestion
Do not point enterprise search at the project SharePoint, Teams site, or client-year folder and call that a knowledge graph. Ingest stripped method only. Connecting the whole project SharePoint is how confidential exhibits become search answers.
The speed is real: the next engagement should retrieve a demand-review cadence or a cutover war-room kit without pinging the last partner. That only works if the index holds method. The connector will also index the plant forecast workbook, the named org chart, the exhibit the client marked confidential, and the deck that still has their logo. Retrieval will treat those files as answers.
Products in this class (Glean, Hebbia, SharePoint, Notion, McKinsey Lilli) retrieve against whatever you connect. They will not refuse a confidential exhibit, invent a problem-type taxonomy, or copy an engagement wall into the index unless you set them that way. None of these is a ranking. Pick the system your knowledge managers will keep clean. Then do not feed it the site.
Ingestion is a controlled write after someone has decided what is allowed to be found. It is not "turn on the connector and close the job."
Extraction and dual review are the gate
Do not write to the graph until a reusable IP extraction agent has stripped method from client exhibits, and until two humans have signed that pack.
Extraction decides what is firm IP. Ingestion decides how that IP is found. If you ingest first, you are indexing the exhibit under a method-shaped filename, and the dual-review signatures are theater.
The same two reviewers who cleared extraction must clear the ingest write:
- The engagement manager confirms the item is the firm's to reuse (IP clause, client limits, conflict walls) and that the pack still carries no identifying structure.
- The knowledge manager confirms classification, problem-type tags, and that the write will inherit the engagement's access rules rather than defaulting to everyone in the practice.
If either says no, the item stays in the project archive. Do not "index it privately and fix tags later." Later does not happen, and the embeddings are already live.
A lessons learned synthesizer draft is input to extraction, not a document to ingest. Postmortems name people, quote the client, and describe failures in language that identifies the account. Pull a reusable operating cadence if one is there. Leave the named executive and the blame in the project file.
Classify the object, then tag the problem, never the client
Classification says what the object is. Tagging says when a future team should find it. Do both on the stripped pack, not on the raw site.
Use a small classification set the practice will actually maintain:
- Method. Workshop kits, blank templates, facilitation guides, model structures with dummy inputs. This is the only class that is a candidate for the reusable graph.
- Exhibit. Client data, org charts, process maps, screenshots, anything that would identify the account if it left the team. Never chunked.
- Working paper. Drafts, redlines, comment threads, scratch analysis. Stay in the project store.
- Record. Signed SOW spine, delivered-hours join, close note. These may feed a comparable past scope retriever. They are not open method.
Extraction already forced method versus confidential. Ingestion still classifies because a cleared method pack can contain more than one object (a workshop kit and a model template), and because records and working papers keep showing up in the close folder. Filename is a weak signal. If the classifier is empty, you are tagging files, not objects.
Tag method by problem type the practice names: "S&OP exception-ladder diagnostic," "ERP recovery after a failed cutover," "shared-services operating-model redesign." Industry and date are filters. Artifact kind is already in the classification.
Do not tag by client name. A client-name tag teaches people to search "Cedarline" instead of "S&OP exception ladder." It also turns retrieval into an account dump. Staff who should never open that folder will query the name and get hits. Problem type is the find key. Client identity, if it must exist, lives behind the engagement ACL as metadata the cleared team can see, not as a public facet.
Index stripped method only, and copy the engagement wall
Chunk only the text and structure the reviewers cleared. Do not embed the original PowerPoint, the workbook with the client's sheets still inside, or a "clean" PDF that still has their footer. If the method needs a figure, redraw it with generic labels.
Copy the engagement's access rules into the graph. If a person could not open the file in the source system, they cannot retrieve a chunk, a snippet, or a "similar work" card that still contains it. Everyone in the practice is the wrong default. Conflict walls and client limits on reuse apply to embeddings the same way they apply to a shared drive.
Default the retrievable payload to:
- Problem type, industry, date, and a short method summary.
- The stripped artifact, for people the ACL allows.
- A pointer back to the project store for anyone who already had access and needs the full file. A pointer is not a back door. The source system still enforces the wall.
Named-entity stripping is not enough. An identifiable forecast grain, a plant layout, or a sentence that could only describe one manufacturer is still an exhibit. If you cannot describe the method without the client's facts, it is not ready to index.
Later jobs will trust this corpus. A RAG proposal draft generator will paste whatever it retrieves. A stakeholder onboarding brief synthesizer will drop another client's name into a joiner pack if that name sits in the index. Hygiene at ingest is cheaper than a leak in a draft.
When retrieval returns another client's exhibit
This is a worked example with made-up names, written to show the cuts, not a case study with results.
Rafi is the knowledge manager. The firm has just closed an S&OP diagnostic at Cedarline Manufacturing: demand-review cadence, an exception ladder the team adapted, and a plant-level forecast workbook the client marked confidential. The partner's instruction is simple: connect Glean to the project SharePoint so next year's operations work can find this.
The site holds the reusable exception-ladder kit, Cedarline's SKU-by-plant forecast file, a named planner org chart, the final deck with Cedarline's capacity calendar, and a lessons-learned note that names the operations VP who blocked a freeze. Connecting the site indexes all of it.
The whole site became the index
Months later a partner scoping a discrete manufacturer's demand-planning diagnostic searches "S&OP" and "exception ladder." The nearest file is Cedarline's forecast workbook. Embeddings like the vocabulary. That is indexing the whole project site, then calling the result knowledge.
The tags were the client
The same library was tagged "Cedarline." Analysts search the manufacturer they remember, not the problem. The facet works as an account dump. People who should never see Cedarline's folder now have a path that looks like search instead of like opening a restricted drive.
Retrieval returned another client's exhibit
The partner running the new diagnostic is not on the Cedarline account. The hit still includes the SKU-by-plant file, or a "clean" exception-ladder PDF whose example row still has Cedarline product codes. That is retrieval returning another client's exhibit. Entity stripping missed the codes because they look like SKUs, not names.
What Rafi should have done: wait for extraction to pull the exception-ladder steps and a generic example row. Dual-review with the engagement manager against the IP clause and Cedarline's reuse limits. Classify the workbook and org chart as exhibits, the deck as a mix of working paper and exhibit, the lessons-learned note as a record that is not ingestible as written. Tag the cleared method "S&OP exception-ladder diagnostic," with industry and date as filters. Index that pack only, with Cedarline's engagement ACL copied so a partner outside the wall gets a problem-type hit they are allowed to see, or nothing, never the forecast file.
The system is doing its job when the next demand-planning diagnostic retrieves a generic exception ladder without anyone opening Cedarline's site, and when a partner who could not open that site cannot retrieve it as a snippet either.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first