Offer-acceptance probability model
Predicts close rate per offer from candidate behavior, compensation delta, and pipeline history to prioritize recruiter effort.
HR processPlanSourceSelectHireOnboardDevelopRewardExit
By Don, DoneThat’s AI coach · updated
What this score is allowed to claim
The model produces a quality signal on a live offer: a scored brief that cites candidate behavior on this requisition, the compensation delta versus the packet you are sending, and the vintage of the comparable pipeline history. That is the entire claim. It is not a close-rate percent, not a staffing capacity figure, and not a decision.
If those three cites cannot be named, do not emit a number that looks like a prediction. A recruiter still owns the close. Empty stays empty when history is thin. Nothing in the ATS should auto-withdraw from this field.
Applicant tracking systems such as Greenhouse, Lever, Ashby, and Workday already store the offer object, the activity timeline, and prior requisitions. Treat them as one class of system of record. The model reads those records. It does not invent outcomes the records do not support, and it does not rank those vendors.
Load the packet, the activity, and dated comparables
Work from the offer you would actually send, in the order a closer would audit it.
Pull the current packet: role, level, location or remote policy, compensation mix, start date, and constraints the candidate already put in writing. Pair the money with the comp recommendation per offer so the compensation delta is measured against a documented recommendation, not against a remembered band or a spreadsheet tab someone renamed last quarter.
Pull behavior that already exists as ATS events for this candidate on this requisition: response latency after each stage, questions about leveling or equity, reschedules and no-shows, take-home completion, requests for a written offer before a verbal. Behavior means dated activity the recruiter can open. It does not mean a personality label, a "hungry" comment, or a vibe from a debrief.
Pull comparable offer history: same role family, similar level, same location policy, with offer-out and resolution dates you can stamp. History without a date is not usable history. Write the vintage of the cohort next to the records: when those offers went out, when they resolved, and which leveling framework was in force.
Keep localized offer letter generation on a separate track. Letter localization is how the packet reaches the candidate in the right language and employing entity. Probability scoring is whether the close is even in reach. Combining them yields a polished letter on an offer nobody has reasoned about.
Do not score from a dashboard tile that has already collapsed those inputs. If the recruiter cannot open the packet, the activity rows, and the dated comparable set, the model is not ready. Assign who loads what: compensation partner owns the recommendation used for delta, recruiting coordinator owns activity hygiene in the ATS, recruiting lead owns the comparable-set definition and the vintage stamp.
Emit cites, delta, and vintage, or emit nothing
Score only after the three inputs are attached to the offer record.
The output should be a short brief a closer can argue with:
- Behavior cites: which events, on which dates, from which ATS activity rows.
- Compensation delta: recommended package versus this offer, and which components moved (base, equity, sign-on, location differential).
- History vintage: which prior offers counted as comparable, and the date range of that set.
A score with no vintage is a failure mode, even when the number looks confident. Old competing-offer patterns, old time-to-accept norms, and old leveling maps do not travel. If the comparable set is from a different hiring climate or a superseded framework, the cite must say so, or the field stays blank. Shipping a confident score on undated comps is worse than leaving the field empty, because it looks like diligence.
Do not invent a close-rate percent. Relative language tied to this cohort and these cites is honest. A percentage chance of acceptance is made-up precision. Those percentages get pasted into weekly staffing reviews and then treated as headcount. They are not headcount. They are also not a reason to tell a hiring manager the seat is "probably filled."
Walk through one offer as a method, not as a measured result. A staff engineer packet is out. The candidate has replied to scheduling the same day, asked two written questions about refresh grants, and has not named another employer's package. The compensation delta versus the documented recommendation is a small base lift and no extra equity. Comparable staff offers in this office last resolved under a previous remote policy, and the set is small enough that a recruiting lead would not want to generalize from it. The model should list those comparables, stamp how old that set is, and mark confidence as limited by age and mix. It should not print a close-rate percent, tell the recruiter to stop working the candidate, or move the ATS stage.
Blanks, missingness, and adjacent-role leakage
Thin history is normal for new levels, new locations, first hires on a team, and roles you have not offered in a long time. Treat thin history as an expected state, not as a data-quality incident to paper over before Friday's review.
When the comparable set is too small, too mixed, or too old to cite, leave the probability field empty. Empty is a correct quality outcome. Filling it with a score that has no vintage, or with a borrowed rate from another function, trains recruiters to ignore the field.
Do not backfill from adjacent roles without saying you did. A product manager close pattern is not a data-science close pattern. A backfill in one city is not a greenfield hire in another. If you must show something, show the missingness: no comparable offers in the dated window. Then stop. Do not average in a company-wide acceptance figure. This model is not allowed to mint that figure.
Candidate rediscovery is a different workflow. If this offer dies, ats candidate rediscovery is how you find them later. It is not a reason to invent a probability so the requisition looks healthier this week.
Scoring can also encode who receives "likely to accept" language. Apply the same discipline you would on a hiring decision bias audit. If the model is mostly picking up response speed, interview polish, or how hard someone negotiates, it may be scoring presentation and comfort with process, not acceptance risk. Put the limitation in the audit trail. Do not hide it inside a number.
Operational checks before you trust a filled field:
- Every non-empty score has a vintage stamp and a comparable set the recruiter can open.
- Compensation delta points at the current recommendation, not at an expired band document.
- Behavior cites are events on this candidate, not comments from a hallway debrief.
Closers decide; nothing auto-withdraws
The recruiter owns the close. The model does not auto-withdraw, auto-expire, or auto-reroute the requisition because a score looked weak.
Treating the score as a withdraw is the failure mode that shows up when someone wants the field to do something in the ATS. A low score or a blank means put a closer on the next conversation and check whether the comp recommendation per offer still matches what went out. It does not mean the ATS should move the candidate to declined.
No workflow, stage change, or expiration job may fire from the probability field. No SLA may require a numeric score before an offer can go out; thin history must still be allowed to offer. Recruiter notes remain the source of truth for stated intent, and staffing plans do not consume a close-rate percent this model did not produce.
Use a scored, cited, vintage-stamped brief to prioritize closer time where the conversation is live but fragile, and where the packet is expensive and the field is blank. Do not starve an offer because the model stayed empty.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first