Skip to main content
DoneThat

AI Adoption GuideHRSelect

Structured resume scoring

LLM rates each decomposed job requirement separately and produces an explainable rubric score per candidate.

HR processPlanSourceSelectHireOnboardDevelopRewardExit

By Don, DoneThat’s AI coach · updated

Score each requirement against a cited resume span

A structured resume score is a row per job requirement, not a verdict on the person. For every decomposed requirement, the model either quotes a resume span and scores that span, or it leaves the cell empty. A recruiter still screens. Nothing in this step auto-rejects a candidate, and nothing in this step invents a fit percent.

The quality bar is the cite. If a reviewer cannot find the quoted text in the resume, the score is invalid. If the resume never speaks to the requirement, emptiness is the correct output, not a guessed zero and not a padded "partial."

This sits after you have a frozen, atomic requirement list and before a human decides who gets a conversation. Applicant tracking systems such as Greenhouse, Lever, Ashby, and Eightfold already store the job and the resume. Use them as the system of record for those two artifacts. The scoring pass is a rubric layer on top, not a replacement for the screen.

Load the req and the resume, nothing else

Load two inputs: the decomposed requirements for this req, and the resume text you will score.

Split stacked language into atomic rows before the model runs. "Python in production, Kafka, and on-call ownership" is three requirements. If the posting is still a paragraph of nice-to-haves, clean the source job first. inclusive job description optimizer is the right place to tighten language so you are not scoring against vague or exclusionary bars.

Use the resume as submitted. Do not concatenate a LinkedIn scrape, a sourcer note, or a comment from a prior req. A cite must be something the candidate put on the page. Recruiter-authored text dressed up as resume evidence is how you get a confident score that the candidate never wrote.

Pass the requirement text, any must-versus-nice flag you already maintain, and the resume. Do not pass a target score, a "this is a knockout" instruction, a diversity objective, or a request for overall fit. Those side instructions are how blank cells become invented competence.

Looking up people who already applied to a related role is a different job. Keep it off this call and on ats candidate rediscovery.

Run one score per requirement, with a cite

Score requirements independently. Do not ask the model to "read the resume and rate the candidate."

For each requirement, require four fields: requirement id, a score on a scale you froze for this req (for example 0-3, or meets / partial / strong), the exact resume span, and a one-line rationale that only restates what that span shows. If the model cannot produce a span a recruiter can find with search, the score is blank and the rationale states that the resume is silent.

Do not average the rows. Do not add an "overall" row. The recruiter reads the grid.

Keep the scale definition in the prompt identical for every candidate on the same req. If "3" means demonstrated production ownership for person A, it means the same bar for person B. Changing the scale mid-slate makes the grid incomparable.

After the grid is filled, a human opens the resume and the scores together. The model's work ended at cite-or-blank.

One pattern, not a measured case. Requirement: "Shipped a production payments API." Resume span: "Led checkout API at Acme; PCI scope; 2022-2024." The model scores that row, quotes that line, and stops. Next requirement: "On-call owner for a service with a defined SLO." The resume never mentions on-call, SLOs, or pager. That cell stays empty. The recruiter still screens. They can ask about on-call in the first conversation. They do not treat the blank as a fail, and they do not fill a 2 because "payments people usually carry a pager."

Empty stays empty when the resume is silent

Silence is not a zero. A zero claims the resume addressed the requirement and the evidence was weak. A blank claims there was nothing to score.

Collapsing blanks into zeros is how a rubric becomes a reject list without a human. Candidates omit on-call, clearance, or version pins for many reasons: length, what a prior employer allowed in writing, or a resume that leads with outcomes instead of tools. The job of this step is to mark missing evidence. Inferring the missing skill is a different act, and it is not scoring.

If a score arrives with no resume cite, drop the score. A number plus a paraphrase ("appears strong in distributed systems") is not structured scoring. It is a vibe with a digit. Validate before the grid hits a recruiter: no span, no score.

The same rule blocks invented fit percents. A percent is an aggregate. Aggregation hides which rows were cited and which were empty. Once a percent match sits on a profile, people will sort and cut on it. The outcome you want is a requirement score that cites a resume span. A percent is a different object. Do not emit it.

Recruiter screening owns the next action

The recruiter screens with the grid open, not after a numeric cutoff.

Cited rows tell them what is already on the page. Blank rows become questions, not disqualifiers. They can advance someone with blanks if the cited rows are the ones that matter for this hire, or they can hold for a clarifying screen. That call is the work. The model does not make it.

Do not map a low score or a blank to auto-reject, auto-archive, or an automatic stage skip. Greenhouse, Lever, Ashby, and Eightfold can hold custom fields and scorecards. Store the rubric there as display and audit trail. Stage changes stay on a human action.

If a blank should become an interview prompt, hand it to the screen as a question, not as a no. ai screening interviews is the next step for turning silence into a structured ask.

When you later inspect a loop for systematic tilt, a cited grid is usable because a reviewer can see which bars were scored from text and which were empty. A composite is not. Pair the grid with hiring decision bias audit instead of asking the same model call for a fairness number.

Failure modes that undo the rubric

A score with no resume cite. If you do not throw it out, recruiters will trust the digit. You cannot show a hiring manager what was scored. Invalid row.

Treating the score as a reject. The moment a 0 or a blank archives the candidate, you automated a screen you said a recruiter still owns. Blanks hit career-changers, short resumes, and anyone whose last title does not echo your requirement wording.

Inventing a fit percent. Summing, averaging, or "overall fit" puts back the single rank structured scoring was meant to remove. If a stakeholder wants a rollup, show count of cited rows and count of blanks. That is coverage of the resume, not a ranking of the person.

Prompt bleed. Asking the same call to rank "culture add" or flag "overqualified" contaminates the cites. One job per call: this requirement against this resume.

Unstable requirement text. If the req changes after half the slate is scored, cites still point at the old bar. Freeze the list for the batch, or rescore everyone when it changes.

Operating loop: load the decomposed req and the resume, score each requirement with a cite, leave silence blank, recruiter screens. Empty stays empty. No auto-reject. No invented fit percent.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first