AI Adoption GuideHROnboard
AI roleplay skill coach
Simulated customer or colleague scenarios let new hires practice difficult conversations with structured feedback.
HR processPlanSourceSelectHireOnboardDevelopRewardExit
By Don, DoneThat’s AI coach · updated
Cite the turn and the rubric item, or leave the cell empty
A scored roleplay is a packet a manager can open after the session. For every item you mark, name the scenario turn the evidence came from and the rubric row it was judged against. If the new hire never produced a turn that could be scored on that row, the cell stays empty. Empty means there was nothing to judge. It is not a zero, not a pass, and not a proficiency percent. The same cite-or-blank rule applies to a colleague scene as to a customer scene.
Performance and onboarding systems in the same class as Lattice, Workday, and 15Five can record that practice was assigned or completed. Use them as the place the manager already looks for whether conversation practice ran. Do not treat them as the scoring engine.
Three failure modes appear as soon as teams skip that bar. Feedback with no turn cite and no rubric item is a comment, not a score. Treating a finished session as a pass is a completion flag. Printing a proficiency percent the rubric never defined is a fabricated summary.
Load the scenario pack and the scoring sheet before anyone speaks
Load two artifacts together: the scenario (who the counterpart is, what they want, what constraints apply, how the scene opens and closes) and the rubric (one observable behavior per row, plus what evidence looks like). The new hire's brief, the model that plays the counterpart, and the scorer must see the same pair. If those three see different sheets, the cites will not line up.
Keep the scenario small enough that every turn can be pointed at. A billing dispute with a named product, a stated amount, and a short turn budget is enough. A vague difficult-customer prompt with no facts forces invented plot that cannot be cited.
Map each rubric row to something a listener could timestamp. "Acknowledges the customer's stated problem before offering a fix" can be cited to a turn. "Sounds senior" cannot. If a row cannot survive a cite, cut it or rewrite it until it can.
Who gets which pack is a routing problem, not a scoring problem. Tie the scenario to the role using the same map as role-based learning auto-assignment. Put the session on a calendar week from the 30-60-90 plan generator so the manager is not inventing a date in Slack.
This is not a new-hire conversational assistant. That assistant answers where the laptop stipend form lives. Roleplay is the hire saying the hard line out loud, with a counterpart who will not let them skip the uncomfortable turn. It is also not an ai personal tutor on product facts. The tutor can check whether they know the refund policy. The roleplay checks whether they can say the policy to someone who is already angry.
Score a billing-dispute practice as it happens
A new support hire is in week two. The counterpart is a customer charged twice for an add-on after a failed downgrade. Refunds over a threshold need a supervisor code the hire does not have. The scene opens on the customer's first line and ends when the hire has a next step the customer accepted, or the customer has asked for a transfer.
The rubric has four rows L\u0026D and the manager agreed on before the session:
- Names the double charge and the add-on in the customer's words before proposing a fix.
- States what they can do in this call and what needs a supervisor, without promising a refund they cannot issue.
- Offers a concrete next step with a time expectation.
- Confirms the customer understands the next step, or hears the request to escalate, before closing.
Turn 1 is the customer's opening. The hire has not spoken. Rows 2, 3, and 4 stay empty.
Turn 2, the hire says they are sorry the customer is upset and asks how they can help. That is not row 1. Row 1 needs the double charge and the add-on in the customer's words.
Turn 4, the hire restates two charges for the add-on after the downgrade failed, and repeats the ticket ID. Cite row 1 to turn 4. Do not also mark rows 2 through 4 from the same sentence. They were not evidenced.
Turn 6, the hire promises a full refund on the spot. Cite turn 6 against row 2 as not met, with the quote. Do not convert that miss into a percent. Do not pass the scenario because they were warm on turn 2.
Turn 8, they walk back the promise and name the supervisor path and a same-day callback window. Cite turn 8 for row 3 if the next step and the time expectation are both present. If the customer talks over them and never confirms, row 4 stays empty. Do not infer confirmation from a polite okay that was about something else.
When the scene ends, the manager opens a table of rubric row, status (met, not met, or empty), turn number, and a quote. If a row is empty, the manager decides whether to re-run or coach live. The system does not declare the hire proficient.
Mark only evidenced turns; keep silent cells blank
For each rubric row, search the transcript for a turn that could satisfy it. If you find one, write the turn index and the quote. If you find a turn that contradicts it, write the turn index and the quote and mark not met. If you find neither, leave the cell empty.
Do not fill empty cells with implied, probably would have, or a blend from other rows. New hires stall, over-apologize, or wait for the counterpart to change the subject. Filling those stalls with a guessed percent teaches the dashboard, not the person.
Ban feedback that cannot be pointed at a turn. Work on your tone, with no timestamp, is not coachable in a short review. Turn 6 promised a refund you cannot issue (rubric row 2) is coachable. If the model cannot produce that pair, do not show a score. Show the transcript and ask the manager to mark it. A missing cite is a pipeline failure, not a soft-skill result.
Never auto-pass on session complete, time spent, or a claim that the customer seemed calmer. Those are not rubric items unless you wrote them as observable rows in advance, and seemed calmer cannot be cited cleanly. Completion belongs in the HR system of record. Skill evidence belongs in the cited table.
Do not invent a proficiency percent. Four rows with two met, one not met, and one empty is not half, two-thirds, or three-quarters. The empty row was not tested. Averaging tested and untested items ships a number that cannot survive a manager question. Do not treat any rollup as a pass.
Give the packet to the manager; do not file a pass
The manager still coaches. The model's job ended when the table was filled with cites and blanks. The manager sits with the hire, replays the cited turns, and chooses the next practice: the same scenario with a tighter open, a harder counterpart, or a live shadow on a real call.
Use empty cells as the agenda, not as a failing grade. An empty row 4 means they never got to confirmation. That might be because the counterpart steamrolled, because the hire closed early, or because the scenario was too short. The manager can tell which.
Do not file a pass in Lattice, Workday, 15Five, or a peer system unless a human marked a decision you defined, such as ready for supervised live calls. Storing that a roleplay ran is fine. Storing a certification because a session object exists is auto-pass in another database.
If you need a second run, keep the same rubric. Change the scenario facts if you want variety. Keep the behaviors stable until the manager is done coaching that skill. Schedule the review close to the session. Keep the packet to a table, quotes, and blanks. Essays from the model hide missing cites.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first