Skip to main content
DoneThat

AI Adoption GuideHRSelect

AI screening interviews

Voice or chat agent conducts adaptive first-round interviews and scores responses against the role rubric.

HR processPlanSourceSelectHireOnboardDevelopRewardExit

By Don, DoneThat’s AI coach · updated

What a usable score packet contains

A first-round AI screen is a quality artifact only when each score names a rubric item and quotes the candidate answer span it rests on. If the candidate never addressed that item, the cell stays empty. The recruiter still advances or rejects. The agent does not auto-reject, and it does not invent a fit percent.

The packet is what you review, not a colored badge on a profile. A second person should open the same row, read the cited words, and be able to agree or disagree with the score. If they cannot find the span in the transcript, the score is not evidence. It is a claim.

Voice agents and chat agents both fit this shape. Adaptive follow-ups are fine when they still target a loaded rubric item. A fluent tangent is not a new competency, and it should not mint a new row.

Keep this packet separate from structured resume scoring. Resume scoring decides who is invited into the first round. The interview packet records what you actually heard after they joined. Folding both into one rank hides which evidence is missing.

Load the rubric before the first prompt

Configure the role rubric before the greeting. Map every planned probe to a named item. A dimension with no probe will not produce a citeable answer, so it should not produce a score.

Write items in the hiring manager's language. "Walks an angry customer down before offering a workaround" can be probed and cited. "Great communicator" cannot. Broad labels invite the model to score tone and file it under a competency nobody defined.

The same rubric should travel into later rounds. If a live interview copilot scores a different trait list, the debrief becomes an argument about which document is real instead of a decision about the candidate.

Interview and applicant-tracking products (HireVue, Greenhouse, Lever, Ashby, and similar systems) will store the recording, the score, or the application. Treat them as a class of stores. Do not assume the export already pairs a rubric row with an answer span. If you only receive a single number per candidate, insert a packet review before anyone advances or rejects.

Lock the scale in the configuration. If the team scores 1 to 4 with behavioral anchors, the agent scores 1 to 4. A 0 to 100 "match" is a different instrument. It is also the usual path by which a fit percent appears without a weighting rule.

Name who may edit the rubric on a live req. If sourcers change probes mid-week, Tuesday's packets are not comparable to Thursday's. Freeze the rubric for the req, or version it and store the version identifier on the packet.

Interview with cites on every score

The agent runs an adaptive first round in voice or chat. It may ask a follow-up when an answer is thin. After the session, score each rubric item only from the transcript or from a timestamped clip you can quote.

Each scored row needs three fields: the rubric item, the score on the agreed scale, and the answer span. The span is the candidate's words, not a paraphrase the model prefers. A fluent rewrite with no quote is how you end up with a score that has no answer cite.

If the product can highlight audio, store the clip bounds next to the text. A later reviewer should land on the same seconds and hear the same claim.

Do not score from the agent's recap. Recaps drop hedges, reorder events, and invent structure. Open the transcript, find the span, then assign the score.

Async video interview analysis is a sibling, not a replacement. The prompt is fixed and the candidate records once. An adaptive agent can ask again. Both still require a rubric cite and an answer span. Neither should auto-reject from a model score.

Support screen, one item left blank

A recruiting lead is filling a support seat. The rubric includes policy accuracy and de-escalation. The agent asks for a time the candidate handled an angry caller. The candidate explains the refund form, the approval limit, and which field to complete. They never describe what they said to the person first, or how they lowered the temperature.

The packet cites the refund-form span under policy accuracy and scores that item. De-escalation stays empty. The agent does not treat a calm tone as de-escalation, does not write "likely a 3," and does not fill a fit percent so the profile looks finished.

The blank is the useful signal. The recruiter can take de-escalation to a live round, or choose not to advance because a required item produced no evidence. Either path is a recruiter decision.

Silent answers stay empty

Empty stays empty. Dead air, a skipped question, a one-word "yeah," or an answer that belongs on another row does not authorize a score.

The row to throw out is a score with no answer cite. That is not a polite incomplete. It is an unauditable claim. Invalidate the row. Do not reject the candidate because the row is broken. Rescore from the transcript, or mark the dimension as not observed.

Do not fill cells to satisfy a completeness meter. Three cited scores and two blanks are more honest than five scores with two of them unsourced.

If audio drops, the chat disconnects, or the candidate asks to pause, mark the session interrupted. Missing minutes are not low performance.

Language access follows the same rule. A candidate working through an interpreter, or one who asked for extra time, can still produce citeable spans. Score those spans. Do not treat slower delivery as a missing competency.

The recruiter still advances or rejects

The packet informs a human. Auto-reject on a screening score turns a missing cite, a slow start, or a bad probe into a closed door. The agent does not own the outcome.

Do not invent a fit percent. A percent smuggles in weights, a missing-data rule, and a threshold nobody wrote down. If leadership wants a composite, put the weights on the rubric document and compute only from cited rows. Blanks stay blank. They become zeros only when the hiring manager documented that rule before the first screen.

When the recruiter advances or rejects, the note should point at the packet: which cited scores carried the decision, which blanks were acceptable for this role, and what the next interviewer must still test. That file is what a hiring decision bias audit can actually inspect. A stack of unexplained "no" labels is not.

Hold the cited-evidence rule across sources. Referrals and inbound applicants get the same packet standard. If one path skips the agent, write that in the file. Do not treat a missing packet as a pass.

Watch the three failure modes as a set. A score with no answer cite is invalid evidence. Treating the score as a reject hands the outcome to the model. Inventing a fit percent hides missing rows behind a number that looks decisive. If you cannot export rubric item, answer span, and blank-for-silent, you may still use the agent as a structured conversation. You do not yet have a quality screen you can defend in a debrief.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first