Skip to main content
DoneThat

AI Adoption GuideEducationRecruit

Look-alike Prospect Discovery

Embedding similarity identifies external prospects whose profiles resemble high-value enrolled student cohorts.

Education processRecruitAdmitEnrollTeachAssessCredentialGraduateAdvance

By Don, DoneThat’s AI coach · updated

A named look-alike cites the enrolled features that produced it

A look-alike prospect is a named person already in your inquiry or search file whose recorded profile resembles a named enrolled cohort on features you can list. The usable output is the person, the cohort label, and the feature cites. It is not a yield number, not an admit recommendation, and not permission to contact someone the file does not contain.

Enrollment offices already store inquiries, applicants, and enrolled students in CRM and SIS tools such as Slate, Technolutions, Salesforce, EAB, and Ellucian. Look-alike discovery reads those records and ranks people who resemble a cohort you defined. A recruiter still decides outreach priority. If the prospect record is thin, leave the match empty. Filling the cell so the list looks complete is a quality failure.

Define the enrolled cohort before you match anyone

Name the enrolled group first. "High-value" only means something once you say which enrolled students count: program, entry type, term, campus, and the recruiting outcome you care about, such as enrolled students who persisted through first census. Do not point the model at every enrolled student and hope geography does the rest.

Write the cohort as a filter you could reproduce next term. Include academic interest or major family, entry path (first-year, transfer, adult), start term, and academic markers that exist on both the enrolled side and the prospect side. Exclude features that only appear after admit, such as deposit date or housing assignment, unless you are matching a later-stage list. Those fields describe students you already won, not people still in the search file.

Keep yield out of the cohort definition. Similarity to last year's enrollees is not the same task as predictive yield scoring. Yield models estimate who among known inquirers or admits is likely to enroll. Look-alike discovery only says this person's recorded profile is close to a cohort you named. Mixing the two produces a fake certainty: a high similarity score treated as if it were a high chance of enrolling.

Illustrative example, not a measured case. A transfer recruiter for a pre-licensure nursing pathway defines the cohort as students who enrolled as transfers into the BSN for a recent fall, with nursing as the intended major, a prior-credit band the SIS actually stores, a campus of interest, and an inquiry source the CRM recorded. The feature cites on any look-alike must be those fields, not "lives nearby" alone. A zip-only match would also surface a neighbor who inquired about business and never mentioned nursing. That person is not a look-alike of the BSN transfer cohort, even if the drive time is identical.

Match only features the prospect file actually holds

Embedding or nearest-neighbor matching is only as honest as the overlapping fields. For each candidate, require a named feature set on both the enrolled cohort and the prospect row: intended major or academic interest, entry type, intended term, campus or modality, inquiry source, and any academic band you already store (prior credits, test-optional flag, high school or college type). Cite those features on the output row. If a field is blank on the prospect, do not impute it from the cohort average.

Do not match on zip, census tract, or high school CEEB as the sole signal. Geography is a weak proxy for academic intent and recreates last year's travel pattern instead of last year's enrolled profile. Use location as one feature among several, and only when the prospect file actually has it.

Do not source a person the inquiry file does not contain. Purchased names and outside directories are a different project, with different consent and suppression rules. This process returns a named look-alike inside the file you already hold. If the nearest embedding points at someone who is not in that file, drop them. Do not create an inquiry so the match can exist.

Vendor platforms differ in how they store interests, sources, and custom fields. Treat Slate, Technolutions, Salesforce, EAB, and Ellucian as systems of record you will query, not as interchangeable scoring engines. Map cohort filters to the fields those systems actually populate. A similarity job that reads empty custom fields will match on whatever is left, often zip and name, or return noise.

After a recruiter accepts a look-alike for contact, personalized outreach generation should use the major, term, and source you already cited, not a generic "students like you" line that hides the match logic.

Recruiter review owns outreach; similarity does not

Every look-alike row needs a human pass before it hits a call list or a journey. The reviewer checks three things: the person is in the inquiry or search file, the cited features are true on that record, and the cohort still matches the campaign you are running. Wrong major, stale intended term, or a suppressed record is a reject, regardless of embedding distance.

Similarity is not an admit decision. Closeness to enrolled students does not mean the person will meet academic requirements, submit a complete file, or pass holistic review. Those questions belong with holistic application scoring once an application exists. Using look-alike rank as a stealth admit score also trains staff to skip people who look unlike last year's class.

Similarity is not yield. Do not print a conversion percentage next to a look-alike. You do not have an outcome yet. If leadership wants a yield estimate, run that model on the right population and keep the two scores in separate columns.

Admissions still owns outreach priority. The model can sort. It cannot assign territories, counselor caseloads, or travel. A recruiter may demote a strong look-alike who already has a counselor of record, and may promote a weaker match who requested a visit this week. That override is the process working, not a failure of the embedding. Do not drop look-alikes straight into an automated journey. Similarity is a review queue, not a send trigger.

Do not fold this list into melt work. Enrollment melt prediction applies to students who already committed. Mixing pre-inquiry look-alikes with melt scores confuses stage and invents urgency.

Empty is the correct result when the record is thin

A thin prospect file is a row missing the features the cohort was built on: no academic interest, no entry type, no intended term, and maybe only a name, a zip, and a source of "search." Matching that row will over-weight geography or invent a profile. Neither is a quality outcome. Return empty, and send the record to data hygiene or a short required-field capture before it re-enters the job.

Empty also applies when the enrolled cohort is too mixed to cite. If you cannot name the features that define "high-value" without pointing at the entire enrolled population, stop. Split the cohort in the SIS extract, then rematch. A campus-wide embedding with no program slice will recreate zip-only behavior even if the method looks sophisticated.

Document the cites on every non-empty result: cohort name, feature list, and the prospect values for those features. That record is what a recruiter audits. Without it, similarity is an unexplained rank that staff will treat as yield or admit advice.

Run the job on a schedule that matches inquiry-file changes, and re-check suppressions, counselor assignments, and intended term each time. A look-alike whose term flipped or whose record was merged is not still a look-alike until you re-cite the features.

The quality bar is simple. You can name the person, name the enrolled cohort, and point to the overlapping features that justified the match. If you cannot, the cell stays empty, and outreach priority stays with admissions.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first