Skip to main content
DoneThat

AI Adoption GuideEducationGraduate

Employment Outcome Tracking

NLP classifies alumni employment status by field and level from public and third-party sources to reduce manual survey effort.

Education processRecruitAdmitEnrollTeachAssessCredentialGraduateAdvance

By Don, DoneThat’s AI coach · updated

What the classified status is allowed to contain

A usable employment-outcome record is a classified status, a field and a level when the evidence supports them, and citations to the source records that produced that classification. Status means employed, not employed, continuing education, or not enough evidence. Field and level use the codebook institutional research already publishes for first-destination or alumni follow-up, not the wording on a public profile.

This record is not a polished job title. It is not a NACE first-destination row until IR codes it as one. It is not permission to replace a survey the graduate already submitted. IR still owns the official outcome. Model output is a draft with receipts.

Empty is a valid state. If sources conflict, or if they are too thin to support field or level, those classified fields stay empty. Filling the gap with a guessed title is how an office publishes employment it cannot defend in an audit or a trustee packet.

Career offices often want a sentence they can print: software engineer at a named firm. That sentence is the failure mode. Titles on public profiles are marketing copy. They do not map cleanly to occupational taxonomies, and they change faster than a reporting cycle. Store the raw title on the citation when a source contains one. Do not promote it to the outcome.

When IR has confirmed a draft, the status can inform later counseling work such as career pathway matching. Until confirmation, it should not appear in published outcomes, program pages, or any extract that looks official.

Ingest only the sources IR has already permitted

Begin with a written allow-list and a capture protocol. Typical allowed inputs are the institution's first-destination or alumni survey, registrar identity keys, and third-party or public records IR has contracted or approved. Those may include state unemployment-insurance wage files where lawful, licensed alumni directories, employer pages the institution is allowed to collect, and professional profiles policy treats as public.

Not every page a crawler can reach is an allowed source. FERPA, vendor contracts, and research-ethics rules decide what may sit next to a student identifier. Ellucian, Salesforce, Anthology, and Workday commonly hold the student, CRM, advancement, and HR-adjacent person records that supply those identifiers. Treat that class of campus systems as identity and prior-response sources. They do not turn a headline into a NACE category. Pull keys and existing survey values from them. Do not write a model class back into the official outcome field until IR confirms.

Keep source classes in separate ingest buckets. Survey responses carry collection date, instrument, and a respondent flag. Public and third-party observations carry a record id, capture date, and a short excerpt. Wage-file or licensed-directory hits, if you have them, stay in their own bucket with the license terms attached. Mixing those streams before classification is how a scraped headline overwrites a graduate's own answer.

Identity resolution belongs in ingest, not in a later cleanup. Match on the keys IR already trusts. Ambiguous matches (common names, missing graduation year, employer-only hits) do not enter the classified set. They remain in a review queue or they stay out. An unmatched scrape is not evidence.

Schedule ingest on IR's calendar, not on whenever a vendor refresh lands. A new scrape does not by itself create a new official outcome.

Classify field and level, then cite every record you used

Run classification only on allowed, identity-matched observations. The model proposes a status, and when the evidence is consistent, a field and a level from your codebook. Every proposed value carries cites: source record ids, excerpts, and dates. A class without a cite is not a class. It is a guess.

Do not ask the model for a job title. If a source contains a title string, keep that string on the citation as raw evidence. The classified output still uses field and level. Publishing the raw title as the outcome is how teams invent employment they cannot stand behind.

LinkedIn-class profiles are noisy. Headlines concatenate aspirations, contract gigs, and employer brands. They are not NACE first-destination categories. A profile that says the person is seeking roles in data science is not employed in data science. A profile that lists three employers without dates is not a timeline. If counselors need labor-market context, keep that in job market signal monitoring as a separate feed. Do not fold market signals into the graduate's official status.

Illustrative example: a biology bachelor's graduate has three allowed rows. In June the first-destination survey marked employed, healthcare, non-research. In August a professional profile lists Research Associate at a named clinic. The same week's clinic staff page lists the same name under patient services with no research title. Classification may cite the survey for status and field. It must not emit Research Associate as the outcome. It should not upgrade the person into a research occupation on the strength of the profile headline. If IR policy says public sources may only fill survey non-response, the public hits are unused because the survey already occupies the row. If policy allows corroboration, the cites still point at the survey and the staff page, and any specialty the profile alone asserts stays empty.

Leave disputed and thin files empty

Conflict means two allowed, identity-matched sources that disagree on status, field, or level after your coding rules. Thin means a single weak observation: an undated profile, a name match without employer confirmation, a directory listing with no role, a wage hit without occupation. In both cases the disputed classified fields stay empty.

Empty is not a prompt to impute employed. Empty means the office will not represent a fact it cannot cite. Downstream reports should count empties as missing, the same way they count survey non-response, unless IR already has a published imputation method. This workflow does not invent one.

Do not average the sources. Do not pick the more flattering title. Do not prefer the most recent scrape by default. Recency without a collection protocol is whichever crawler ran last. Write the conflict into the review packet: source A excerpt and date, source B excerpt and date, the fields left blank, and a reason code (conflict or thin).

A related failure mode is treating a confident-looking profile occupation as if it were a NACE or CIP-SOC mapping. Crosswalks are maintained. Profile text is advertising. If you cannot map the evidence onto the codebook without guessing, leave field and level empty even when employed versus not employed is clear.

IR confirms before a draft becomes the official outcome

Classification produces a review queue, not a file IR can publish. IR inspects the proposed status, the cites, and the empty reason codes. Confirmation is a recorded action: who approved, on which date, against which codebook version. Only then may the official outcome field update, and only through the coding process IR already uses.

Never overwrite a survey response with a model class. If the graduate answered, that answer remains unless the documented IR protocol allows a later correction, such as a respondent amendment. Public and third-party sources fill non-response and, where policy allows, corroborate. They do not silently replace the person.

After confirmation, career services can use the status in placement counselor co-pilot work. Advancement may later consume a stable employment flag in something like an alumni giving propensity model. A draft class is not a giving feature and not a counselor fact. Pushing unconfirmed titles into those systems creates a second outcomes database IR cannot correct.

Operate as a cycle. Ingest on the schedule IR sets. Classify with cites. Hold conflicts and thin files at empty. Confirm in IR. Export only confirmed rows. If a later source would change a confirmed row, it opens a new review. It does not auto-write. That is how you cut manual survey chase without replacing the survey with a guess.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first