AI Adoption GuideHRDevelop
Calibration bias detector
Compares ratings across managers and demographic groups, flagging drift and inconsistency before calibration meetings.
HR processPlanSourceSelectHireOnboardDevelopRewardExit
By Don, DoneThat’s AI coach · updated
What the flag is allowed to say
A calibration bias detector is a quality check on the rating file you will take into the meeting. It compares ratings across managers and across demographic groups your counsel already allows you to use. It does not emit a new rating, a recommended score, or a disparity percent.
The only successful outputs are a flag and a blank. A flag cites three things: the rating file you loaded, the grouping rule you applied, and the vintage of that file. If a cell is too thin to interpret under that rule, the cell stays empty. Empty is not a finding of fairness. Empty means the detector did not have enough people to speak.
People analytics still writes the brief. Calibration chairs still decide in the room. If the detector has already changed a rating in the system of record, it has left its job.
Load the rating file and the grouping rule
Start with exports you can name out loud. Pull the current-cycle ratings from the system that holds them. Performance suites such as Workday, Lattice, and Culture Amp, and people-analytics extracts such as Visier, all produce files a detector can read. Treat those vendors as a class of sources, not as a ranking and not as a substitute method. Do not assume any of them already applied your grouping rule, already stored the vintage you need, or already ran this check.
Confirm the rating file before you group it. Check that the cycle label, the employee IDs, and the rating scale match the meeting you are about to run. If two exports disagree, stop and pick one vintage. Do not blend files.
Pair the rating file with a grouping file from HRIS. Use only consented demographic fields counsel has already approved for this purpose. Those fields live on the employee record. Do not infer protected class from names, photos, email prefixes, Slack display names, or free-text comments. If the field is not in HRIS and not consented, it is not an input.
Write the grouping rule so a chair can audit it: which fields, which comparison set (for example each manager's distribution against same-level peers in the same function), and which minimum cell size your analytics team already uses. The detector does not invent a cell-size policy. It applies the policy you loaded.
Here is the one pass that matters. An engineering talent lead preparing a mid-year calibration loads the current rating export and the HRIS grouping extract counsel already signed off, gender and job level only, nothing inferred from names. They apply the team's existing rule: compare each manager's ratings to same-level peers in the function, and leave any cell below the agreed headcount blank. That is the setup. There is no model that discovers a protected class, and no need to type a disparity percent into the slide.
Upstream narrative still comes from humans and from the performance review draft generator. The detector reads ratings, not prose. Do not feed draft language into the grouping file.
Cite the vintage on every flag
Run the comparison only after both files are loaded and named. For each manager cut and each allowed demographic cut, apply the grouping rule. Where the rule marks inconsistency against the comparison set, write a flag. Where the cell is below the loaded minimum, write nothing.
Every flag must be reconstructable next week. Name the rating file, the grouping rule, and the vintage, meaning the as-of date or cycle label on that file. A chair who cannot tell which snapshot produced the flag cannot use it.
A flag with no vintage is the first failure mode. "Manager B looks harsh" on a slide, with no file name, no cycle label, and no grouping rule, is a rumor. Do not present it. Reload the named file or drop the row.
When a cell is too thin, leave it empty. Do not fill it with a rounded percent, a borrowed rate from a larger org cut, or a qualitative "likely skewed." Do not invent a disparity percent for a thick cell either. The flag is that the distributions diverged from the comparison set in a way your loaded rule marks as inconsistent. The brief can say direction in words (this manager's ratings sit higher or lower than the peer set; this group's ratings sit apart from the rest of the level) without minting a statistic the file does not support.
If you want drift across cycles, load two named vintages. Comparing this cycle to "how we usually rate" without a prior file is not a vintage-backed flag.
Brief the room. Do not write ratings back.
People analytics owns the pre-read. The brief lists flagged cells, the three cites, and the blanks. It does not rewrite ratings in Workday, Lattice, Culture Amp, or Visier. It does not pre-fill the calibration grid with "corrected" scores.
Treating the flag as a rating change is the second failure mode. An analyst writes an adjusted rating back before the meeting so the grid looks clean. The chair then rubber-stamps a decision the detector never had authority to make. Reverse that write. Restore the original ratings. Bring the flag as a flag.
Use the flag as a reason to open evidence already in the packet: goals, work samples, and the review draft. A manager-consistency flag is a prompt to inspect, not a prompt to overwrite.
Keep money decisions on their own rails. A rating flag is not a pay-equity finding and not a merit recommendation. When those questions come up, point chairs to continuous pay-equity monitoring and the merit cycle agent rather than stretching this detector to answer them.
Hiring-time bias work is a different file and a different meeting. The same discipline (cite the source, do not infer class from names, do not auto-change an outcome) is how a hiring decision bias audit should run. Do not mix hiring scores into a performance calibration flag.
What you refuse to compute
Inventing a disparity percent is the third failure mode, and it travels. A thin cell, or a grouping rule that never asked for a rate, still shows up on the slide as a precise gap. That number gets quoted in the room and in follow-up email. If the detector did not compute a counsel-approved rate from a thick enough cell, do not type one. Leave the cell blank or describe the inconsistency without a percent.
Inferring class from names or photos is a related refusal even when no percent appears. A grouping built from first names, profile pictures, or display names is not an HRIS-held, consented field. Throw that grouping out. Reload from HRIS, or run ungrouped manager-consistency flags only.
After flags are cited and thin cells are blank, the remaining work is the work you already do. Read the packets. Hear the managers. Let the chair record the rating the room agrees. Keep the cites in the calibration notes so a later cycle can compare vintages on purpose. Keep blanks blank in the notes too, so nobody fills them from memory a week later.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first