Skip to main content
DoneThat

AI Adoption GuideSalesProspect

ICP lookalike account discovery

Embedding search over firmographic data finds accounts matching the closed-won profile, using tools like Clay or 6sense.

Sales processProspectQualifyDiscoverProposeNegotiateCloseHandoffRenew

By Don, DoneThat’s AI coach · updated

Train on closed-won, not the whole CRM

Lookalike account discovery raises list quality only when the seed is accounts you closed. Seed it with the whole CRM, every opportunity, or every SQL, and you clone targeting mistakes at scale.

The mechanism is embedding search: represent each company as a bundle of firmographic and stack attributes, then retrieve other companies that sit near your closed-won set in that space. Clay, 6sense, and ZoomInfo sit in the account-discovery class that searches a commercial universe for companies near a seed. Salesforce is where the seed, suppression, and territory live. Mixing those jobs is how lookalikes leak into the wrong patch.

Closed-won is the quality bar because those accounts already passed budget, champion, legal, and implementation. Training on every opportunity encodes who you chase, including deals you lost and deals that stall. The model then ranks more of the same, and sales ops reads that as coverage.

Even closed-won needs a filter. Drop product lines you no longer sell, founder-relationship deals you cannot repeat, and a whale that would dominate the embedding. If you have too few wins to describe a profile in writing, write the ICP as a hypothesis instead of fitting a model to a handful of logos.

A similarity rank is a retrieval order: closer to the seed than the next row. It is not a win probability. If a tool prints a match figure, use it to sort. Do not forecast pipeline from it.

Embed firmographics and stack, not famous logos

Encode the attributes that made the win possible, not the brand on the closed-won record.

Firmographics that usually carry signal: employee band, industry or sub-vertical, HQ country, and, when it actually showed up in wins, public versus private. Stack signals that usually carry signal: CRM, ERP, cloud, identity, or an integration category you recorded on closed-won, not a category you wish they used.

Logo cloning is how teams invent an ICP they cannot sell. Three enterprise logos in a mid-market book will pull more companies like that brand. Those wins often came from a board intro, a tiny division, or a services bench you no longer have. The list then fills with accounts that fail your security process, sit below your ACV floor, live in a geo you do not cover, or run a stack you do not support.

Write the seed as a profile a rep would recognize. Then walk the top of the list with a few AEs before anyone sequences it. If they say they would never win that account, the embedding matched the logo, not the deal.

Exclude lost deals, pipeline, and out-of-territory accounts

Hygiene on the seed and hygiene on the output are different jobs. Both happen before BDRs see a name.

Keep closed-lost out of training. Those rows resemble deals you failed. Including them pulls more of the same objections, budget misses, and wrong titles.

Keep current pipeline out of training and out of the output. Open opportunities already have an owner. Rediscovering them as lookalikes creates double-touch and a fake sense that the TAM grew.

Keep customers, recent closed-lost, named accounts, and partners off the output even when they match. Salesforce, or whatever you treat as system of record, is the suppression source. A quarterly CSV is not.

Humans still own territories. The model proposes candidates. Sales ops maps them onto patches, named-account lists, and do-not-prospect rules. An account that matches closed-won but sits on someone else's named list, in a partner book, or in a geo you do not cover is not a BDR lead. Auto-assigning lookalikes across patches starts intra-team fights and burns accounts that were being worked quietly.

If you cannot state the territory rule in one sentence, do not push the list into a sequence or into an autonomous BDR agent.

Illustrative example: Apexline rebuilds the seed

The following is a made-up but realistic rebuild, not a case study and not reported results.

Apexline sells IT service management to mid-market operations teams: four BDRs, one RevOps owner, Salesforce as CRM. They turned on lookalike search in the same class as Clay, 6sense, or ZoomInfo, and seeded it with every opportunity from two years because more rows felt safer.

The first list failed in three ways. Top rows included logos they had already lost (too large; legal never cleared). A block of names were open opportunities AEs already owned. The rest was a long dump of software companies in a broad employee band, because that was the shape of the opportunity object, not the shape of closed-won. BDRs received hundreds of accounts overnight, could not research them, and either sprayed a generic sequence or ignored the list.

They rebuilt:

  1. Closed-won only, last 24 months, one product line. They dropped the single enterprise win that came through a board member.
  2. Attributes: 150-800 employees, US or UK HQ, operations or IT buyer, plus a ticketing or ITSM category they had actually captured on wins.
  3. Suppression against Salesforce: customers, open opps, closed-lost in 18 months, named accounts, partners.
  4. Sales ops, not the model, mapped remaining candidates onto existing territories. Anything outside a patch stayed off the BDR queue.
  5. Each BDR got a weekly slice small enough to check the site, the stack claim, and whether a real trigger existed.

Fit-only names did not go straight to outreach. They waited for a timing signal from buying-signal trigger monitoring or a pricing-page visit from website visitor de-anonymization.

Hand BDRs a researchable slice, not a TAM dump

A lookalike list is a quality filter, not a volume machine. Flooding BDRs with accounts they cannot research recreates the problem you were trying to fix: activity on names nobody understands.

Cap what leaves sales ops. Annotate each row with the matching attributes (size band, industry, stack), Salesforce status, and the territory owner. If a BDR cannot say in one minute why this account resembles a closed-won, the list is not ready.

Fit is not intent. A close lookalike with no hiring, no funding, no stack change, and no site visit is still a cold account. Use lookalikes to decide who is allowed on the list. Use triggers and inbound identity to decide who is worked this week.

Before a name becomes a meeting, run the same disqualification rules you already use. A disqualification recommender is the right next gate if BDRs are booking lookalikes that fail ICP on title, geo, or product fit. Once an account is in the funnel, predictive SQL conversion score is a separate model. Do not reuse the lookalike rank as that score.

Review weekly with a few AEs: which new lookalikes they would actually take, which cloned a logo they cannot sell, and whether volume is exceeding research time. If the list only grows, you trained on the CRM again.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first