AI Adoption GuideHRSelect
Async video interview analysis
Transcribes one-way video answers and scores responses against a competency rubric for high-volume roles.
HR processPlanSourceSelectHireOnboardDevelopRewardExit
By Don, DoneThat’s AI coach · updated
Usable output is a cited rubric score, not a percent
A one-way video score is usable only when it names a competency from the rubric and points at the transcript span that supports it. If the model cannot show that span, the field stays empty. A recruiter still moves the candidate.
High-volume select work fails when the pipeline treats a number as a decision. A competency percent with no quote is an invention. Auto-reject on that number is a process error. The clip is evidence. The rubric is the standard. The recruiter is the decision.
This page is for the select stage after someone already passed a screen. It is not structured resume scoring, and it is not a live conversation. It is the recorded answer to a fixed prompt, scored against the same competencies you would use in a structured interview.
Load the clip and transcribe before anyone scores
Start with the file, not the model. Pull the one-way recording from whatever recorded it. HireVue often holds the async session. Greenhouse, Lever, and Ashby typically hold the requisition, the stage, and the packet the recruiter works from. Treat those products as a class of interview and ATS surfaces. Do not assume a special analysis feature exists in any of them.
Confirm the clip is the right candidate, the right prompt, and a complete take. If the candidate restarted, score the take you told them would count. If two prompts live in one file, split them so each answer maps to one rubric row.
Transcribe next. The transcript is the scoring surface. Muted audio, a language the job was not set to, a cut-off file, or a candidate holding a script the speech model cannot see are empty-score cases, not low scores.
Do not score from a summary of the clip. A summary already discarded the spans you need to cite. Keep timestamps when you have them. Keep the original file.
A recruiter who opens the packet should see the prompt text next to the transcript. Without the prompt, a fluent answer to the wrong question looks like a strong competency hit.
Score the answer against the rubric with spans
Lock the rubric before the first clip in the req. Each row is a competency the hiring manager already signed: de-escalation, ownership, safety awareness, schedule reliability, or the equivalent for this role. Each row has observable anchors, not adjectives. "Gives a next step the customer can hear" is an anchor. "Sounds confident" is not.
For each row, the scorer either writes the rating your scale already uses (does not meet, meets, exceeds), plus the span, plus the rubric item name, or leaves the row empty because no span supports a rating.
Do not invent a competency percent. There is no honest way to turn one 90-second answer into "communication 74%." Percents imply a sample and a calibration you do not have. They also invite ranking people against each other on noise.
One illustrative pass, not a case study. Role: high-volume contact-center associate. Prompt: "A customer says their order is late and they want a refund. Walk us through what you would say and do." Rubric rows: de-escalation, ownership, policy awareness.
The candidate says: "I'd pull up the tracking first and tell them exactly what I see, then apologize for the wait. If the policy allows a replacement, I'd offer that before I talk about a refund, and I'd give them a time they can expect an update."
A usable ownership score cites that span and the ownership row, and rates it against the anchor: names a concrete next step the customer can hear. A usable policy-awareness score cites "if the policy allows" and notes the candidate deferred to policy instead of inventing a refund. If the transcript never shows de-escalation language, that row stays empty. Empty is not a zero. Empty means the clip did not produce evidence for that item.
Kill this in review: a score with no transcript cite. "Ownership: strong" with no quote is not reviewable. A second recruiter cannot disagree with a ghost. If the model cannot attach a span, it must not emit a rating.
Empty output when the clip cannot be scored
Unusable clips are common in high-volume async: car noise, audio that dies after twenty seconds, the wrong language setting, a freeze after the prompt, a file that is a wall.
Leave the score empty. Do not down-rank "poor communication." You did not observe communication. You observed a file problem, a setup problem, or unclear instructions.
Empty still advances to a recruiter. The recruiter can request a retake, move the person to a live interview, or close the loop with a human screen. The model does not choose among those.
Write the empty reason in the packet: "audio unusable after 0:18," "wrong prompt," "no citeable span." That note is for the recruiter, not a score. If only one row has a citeable span, ship that row. Do not average empty with scored. Averaging manufactures a competency percent by another name.
The recruiter advances; nothing auto-rejects
The packet shows cited ratings, empty rows, and the transcript. The recruiter decides whether the candidate moves, retakes, or waits for a live interview. No ATS threshold should auto-reject on an async video score.
Treating the score as a reject wrecks the workflow. A "below 3, disposition" rule turns a noisy single-prompt sample into a hard out. It also hides empty-clip cases inside a fail bucket, so you never fix instructions or audio.
Do not invent a competency percent and sort the req by it. Two candidates who both "meet" ownership on different spans are not 61 versus 67. Sort by stage, recruiter, or time in stage, not by a number the rubric never defined.
When you later audit who advanced, you need the cite and the human decision. That is the evidence trail a hiring decision bias audit depends on: what was scored, what was empty, who moved the candidate, and whether empty clips were treated as fails.
Keep this distinct from ai screening interviews. Screening decides whether someone enters the process. Select-stage async scoring decides whether a recorded answer produced rubric evidence. Mixing those jobs collapses a chat screen and a video score into one unexplained rank.
Where this sits next to screening and live interviews
Use async video analysis when the prompt is fixed, the rubric is shared, and volume makes live first-round interviews the bottleneck. Warehouse, retail, and contact-center roles are the usual fit because every candidate gets the same scenario.
Do not use it as a personality read, an accent filter, or a stand-in for a work sample you could assign in text. Talking about a packing SOP is weaker evidence than a short task. Handling a heated customer in speech is closer to the job, if you score content against anchors and not warmth of tone as a proxy.
Live interviews still matter. A live interview copilot is notes and probes in a two-way conversation. Do not paste an async score into the live room as a hiring recommendation. Hand the interviewer the cited spans and the empty rows so they know what was never observed.
Resume scores stay upstream. Structured resume scoring can decide who gets the async invite. Do not blend it with the video rating into one "fit" number.
The recruiting lead owns the bar: rubric locked to prompts before the req opens; clip loaded and transcribed; each row cited or empty, never a percent; unusable media empty with a reason; recruiter action required to advance or decline; ATS dispositions are human, not model labels.
If a vendor dashboard only emits a score and a color, do not store that color as the decision. Paste the span, the rubric item, and the recruiter's call. If you cannot get a span out of the tool, score manually from the transcript.
Audit a sample of packets: every non-empty rating has a span and a rubric name; every empty rating has a reason; no candidate was auto-rejected from this score; no packet contains a competency percent the rubric did not define.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first