AI Adoption GuideHRSelect
Hiring decision bias audit
Adversarial model reviews offer and reject patterns by demographic group, surfacing drift before regulatory or reputational exposure.
HR processPlanSourceSelectHireOnboardDevelopRewardExit
By Don, DoneThat’s AI coach · updated
What the audit is allowed to return
A hiring decision bias audit returns a drift flag, or it returns nothing. The flag names the decision cohort (job family, level band, location grouping) and the vintage of the offer and reject file it read. It does not name a new hire, flip a reject to an offer, or print a demographic rate.
If the export is thin, the result stays empty. Sparse files are a data-quality problem for talent ops, not a reason to impute a protected-class rate. An empty output is correct when the model cannot cite who was decided, when, and from which file.
The quality bar is citation, not intensity of language. A flag that cannot point at the cohort and vintage is not an audit result. Treat it as a failed run. Rerun only after the file can support a cite.
Protected-class analysis stays human-owned. The model can surface that offer and reject patterns in a cited cohort look different from the prior vintage of the same cohort. People decide whether that difference is noise, a pipeline mix shift, a process change, or something that needs counsel.
The flag must repeat, in plain text a reviewer can paste into a ticket: the cohort identifier you supplied, the decision window of this file, whether the pattern that moved was offers, rejects, or both, and a pointer to the slice used (requisition ids, date bounds, and row counts as counts of rows). Row counts are inventory, not a disparity statistic. If any cite field is missing, the output is invalid even if the prose sounds confident.
Files you load, and what stays blank
Load the offer file and the reject file for one decision vintage, not a rolling all-time dump. Applicant tracking systems such as Greenhouse, Lever, Ashby, and Workday all store those outcomes. The audit does not depend on which of those products produced the export. It depends on each row being a decision with a stable candidate id, requisition, stage, outcome (offer or reject), decision timestamp, and only the demographic attributes your policy allows you to hold for this review.
Do not invent columns. If the export has no demographic field, do not infer group from names, photos, schools, or interview notes. Leave the attribute blank and stop. If a row has an offer date but no requisition id, keep it off the model input or park it on a human exception list. The model does not guess the job.
Pair the files. An offer-only extract cannot show reject patterns. A reject-only extract cannot show who received offers. If one side is missing, the result stays empty.
Define the cohort before you run. Same job family, same level band, same location grouping, same decision vintage. Mixing intern rejects from last spring with staff offers from this week produces a flag no one can interpret. Write the cohort name on the ticket before the CSV hits the model, so the flag can only echo what you already chose.
Who appears in these files is already shaped upstream. Resume ranking and structured screens concentrate the slate long before an offer or reject is written. Keep this review after structured resume scoring and ai screening interviews. Do not use it as a replacement for those controls.
Freeze the export. Live boards rewrite stage history as recruiters recode candidates. A flag may only cite a frozen file. A reject reopened after the pull belongs in the next vintage.
Flagging offer and reject patterns with a cite
Run the adversarial pass only after the cohort name and vintage are written down. The model compares offer and reject outcomes across demographic groups inside that cohort, against the prior vintage of the same cohort when you have one. The only allowed machine output is a drift flag that repeats the cite. A score, a traffic light, or a suggested hire is out of scope.
If the prior vintage is missing, the flag must say the baseline is absent. Do not fabricate one. The first vintage of a new location, a new level, or a post-cutover org (for example after moving among Greenhouse, Lever, Ashby, or Workday) is thin on the baseline side. Empty is allowed.
Row counts stay counts. A valid flag names date bounds and requisition ids and can report how many offer rows and reject rows sat in the cited slice. Filling a blank with an adverse-impact percentage, a four-fifths shorthand, or a rounded gap is inventing a statistic the file did not contain. Do not enrich a flag with an industry benchmark. Benchmarks are someone else's file.
Illustrative example, not a measured result: A talent-ops lead exports one month of staff-engineer offers and rejects from the company's ATS. They label the cohort "Staff SWE, US remote, January vintage" and load both files. The model returns a drift flag that cites that cohort and vintage and points at the January-dated staff-SWE offer and reject slices it used. It does not print a group rate. The lead opens the cited slice, confirms every row belongs, then reads process notes: panel composition, a missing required interviewer, a mid-month hiring freeze. If the January reject file had been too small to support a cite, the flag would have stayed blank and the lead would have logged insufficient data instead of asking for a percentage.
The talent-ops review, not a reverse decision
Talent ops owns the review. A DE&I lead may sit with them. Recruiting ops may pull ATS records. Counsel may be looped when the question is legal exposure. None of those steps is a model action, and none of them is an automatic change to a living offer.
Confirm the cite first. Right cohort, right vintage, rows that actually match. If the pointer is wrong, the flag is void. Only then look at process: Did sourcing change the slate? Did interviewers change? Did copy from an inclusive job description optimizer ship in the same window? Did comp recommendation per offer guidance change who signed? Offers and accepts are different events. Mixing accepts into the offer file can look like offer-pattern drift when it is an economics effect. Keep them on separate vintages unless you are explicitly auditing accepts.
Do not auto-change a hire decision. An outstanding offer stays outstanding until a human withdraws it under written policy. A reject stays a reject until a human reopens the requisition. Using the flag as a reverse decision (advance everyone in one group, hold everyone in another) is a second uncited process, not a correction. The audit does not have the facts to re-score a person. Do not wire it to the ATS as a write-back.
When the file is thin, write "insufficient data" in the review log and close the ticket. Do not resubmit the same CSV with a prompt that asks for a number. Empty stays empty. Re-run on the next frozen vintage.
Failure modes that look like rigor
A flag with no cohort cite. If the output says there may be bias and does not name the requisition family, level, location grouping, and vintage, discard it. You cannot reproduce it, brief counsel on it, or compare it to the next vintage.
Treating the flag as a reverse decision. Flipping outcomes to balance a chart creates a new biased process and still leaves the original file unexplained. Make the cited pattern visible, then change sourcing, panels, rubrics, or timing, then rerun the next frozen vintage.
Inventing a disparity percent. If the model, a spreadsheet, or a vendor tile fills a blank with a gap figure, throw the number out. Your export did not contain that statistic. Publishing it is its own regulatory and reputational problem, because you cannot show how it was computed from the cited rows.
Other quiet failures: joining demographic data from a system that is not authorized for this review; comparing offers from one ATS to rejects from another across a cutover as if they were one vintage; leaving candidate names in the prompt; letting a diversity-insight dashboard replace the cited flag. Greenhouse, Lever, Ashby, and Workday are sources of decision files, not the owner of the protected-class review.
When a run is empty, the next action is a better export or a tighter cohort, not a more aggressive prompt.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first