AI Adoption GuideHRSource
Multi-source semantic candidate search
AI search across open-web profiles ranks candidates by skills, trajectory, and fit beyond keyword matching.
HR processPlanSourceSelectHireOnboardDevelopRewardExit
By Don, DoneThat’s AI coach · updated
Rank by trajectory and skill evidence, not keyword overlap
Semantic candidate search ranks people by how their skills, career trajectory, and role fit line up with a requisition, not by whether a boolean string hit a job title. Keyword search still belongs on exact licenses, certifications, and location filters. It fails when the same skill is named several ways, when a title drifted, or when the strongest person is adjacent rather than identical.
The output you want is a shortlist. Each person is tied to a licensed profile record and to at least one skill named on the req. If the system cannot name that source and that skill, the row stays empty. Rank is a reading order for a sourcer. It is not permission to contact.
Tools in this class include LinkedIn Recruiter, Eightfold, Phenom, and SeekOut. Treat them as licensed talent databases and recruiter search products, not as cover for pulling profiles from the open web. Contract, seat, and data-use terms decide what you may query. If a source is not licensed for this search, do not run it. Do not stitch a profile from public pages. Do not fill the gap with a guessed match.
If the req itself is biased or vague, semantic rank will amplify that. Clean the posting first with an inclusive job description optimizer so the skills you search are the skills you actually need.
Load the requisition before any search runs
Start with the req, not with a talent pool. Pull must-have skills, nice-to-haves, seniority, location or remote rule, and hard constraints such as clearance, license, language, or work authorization. Store those as structured fields. Pasting the whole posting into a free-text box produces fluent ranking and weak evidence.
Map each must-have to the vocabulary your licensed sources actually store. A req that says "distributed systems" should also list the concrete signals those databases index: language, framework, domain, years in role. You are not inventing skills. You are translating the hiring manager's language into fields the licensed index already contains.
Set the empty-result rule before you search. No licensed profile identifier, no row. No req skill cited, no row. Do not let the model substitute a similar skill the manager never asked for.
People already in your ATS are a separate licensed pool. Run ats candidate rediscovery against those records with the same req fields. Do not merge ATS rows with external rows unless both sides carry a source cite the sourcer can open.
Search licensed talent databases only
Query only sources your company has licensed and approved for recruiting: recruiter seats, talent CRM indexes, and vendor talent graphs under contract. LinkedIn Recruiter, Eightfold, Phenom, and SeekOut sit in that class when your organization has paid access and a permitted use. An approved internal talent database counts. A public profile page found in a browser does not.
For each query, log which licensed source ran, which req skills were used, and which filters applied (location, seniority, current-company exclusions). That log explains the shortlist and proves you did not search an unlicensed source.
If a vendor index returns a person but the payload has no profile URL or record ID your seat can open, drop the person. A name and a title without a licensed cite is not a candidate. Do not look them up on the open web to complete the row.
Semantic similarity can surface people whose titles do not match the req title. That is the point. Keep the evidence bar the same: the licensed profile must still show the req skill, in their own words or in a structured skill field the source exposes. Trajectory means the sequence of roles on the profile, not a story the model writes. If the profile lists adjacent work, say that. Do not upgrade it to a strong fit.
If coverage is thin, widen the licensed query with adjacent titles or related skills the manager approved, or stop. Thin coverage is a sourcing problem for the sourcer, not a prompt to invent people.
Cite the source and the skill, or leave the result empty
Every kept match has two citations, visible to the sourcer without opening a page they cannot access.
The first is the licensed profile source: product name, record or profile identifier, and a link or in-product handle that opens under your existing seat. The second is the req skill: the skill string from the requisition, plus the snippet or structured field on that profile that supports it.
If either citation is missing, the result is empty. Empty is correct. Do not backfill from memory, from another unlicensed site, or from a résumé the candidate never submitted to you.
Do not attach a fit percent. Rank order can be useful as a look-at-these-first sort when the sourcer knows it is a sort, not a score. A generated fit percentage is not evidence. It cannot be audited against the profile, and it trains people to skip the cite.
After the sourcer accepts a person into review, structured resume scoring can score materials you are allowed to hold: an application, an inbound résumé, or a profile export the vendor permits. Scoring is a later step on licensed documents. It is not a substitute for the source-and-skill cite on the search hit.
Illustrative path: a sourcer loads a staff backend req that lists Kubernetes, Go, and on-call ownership as must-haves. Licensed search returns three people. Person A has a licensed profile listing Kubernetes and Go, plus a recent role that mentions on-call rotation. The row cites the vendor record and those three req skills, with snippets. Person B ranks higher on similar titles, but the licensed payload has no snippet for Kubernetes, Go, or on-call. Discard the row. Person C is a fluent narrative from a public page the company does not license, with no vendor record ID. The row stays empty. The sourcer opens Person A in the licensed tool and decides whether to contact.
Keep outreach in the sourcer's hands
Search does not send. No auto-email, no bulk first-touch from the ranking job. A person appearing on a ranked list is not consent and is not a decision.
The sourcer reviews the cite, opens the licensed profile, checks do-not-contact flags and existing pipeline status, then chooses to reach out, pass, or save. If they reach out, they own the message. A personalized outreach sequence generator can draft from the same req skills and the cited snippets. It still waits for a human send.
Treating rank as a send burns a market. A high semantic rank can mean overlapping vocabulary, not that the person is open, eligible, or worth a sourcer's credibility. Volume outreach from an unreviewed list also collides with people already in process.
Failure modes that look like progress
A match with no licensed-source cite. The UI shows a name, a plausible title, and fluent English. There is no record ID or seat-openable profile, or the claimed source is a public page outside contract. Delete the row. Do not hunt the person down to attach a cite after the fact.
Treating rank as a send. The list is sorted, so the top names go into a sequence. Rank is a reading order. Contact is a separate act with a named sourcer, a licensed profile in view, and a check against existing candidates.
Inventing a fit percent. A number appears because the prompt asked for confidence. No one can point to the profile field that produced it. Remove the number. Keep the skill snippet and the source cite. If you need a later score, run structured scoring on documents you are allowed to store, and label it as a document score, not a person-fit score from search.
Other quiet failures: substituting a related skill the req did not list; merging two people with similar names; using trajectory language the profile does not support; searching a source after the license lapsed because the integration was still connected. Each of these should fail closed: empty result, logged source, no outreach.
When the pipeline is working, a sourcer sits down with a req, runs licensed search, reads a short list of cited matches, and contacts only the people they chose. Quality is the cite plus the skill, not the length of the list.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first