Skip to main content
DoneThat

AI Adoption GuideHRDevelop

Performance review draft generator

LLM synthesizes peer comments, goals, and self-assessment into a manager review draft.

HR processPlanSourceSelectHireOnboardDevelopRewardExit

By Don, DoneThat’s AI coach · updated

A sourced sentence is the quality bar

A performance review draft generator is useful only when every claim in the draft is traceable to a peer comment, a goal record, or the employee's self-assessment. If those sources do not support a sentence, that sentence stays blank. The manager still writes the review and still submits it. The model does not file anything, and it does not invent a performance claim to make the packet look complete.

People-ops and HRBP leads should treat this as a writing aid with a hard quality rule, not as an author of record. Review platforms such as Lattice, Culture Amp, 15Five, and Workday already hold the comments, goals, and self-assessment you need. The generator reads those packets and proposes language. It does not replace the manager's judgment, the employee's voice, or your calibration process.

Watch for a fluent paragraph that sounds like a review and cites nothing. That paragraph is not a draft. It is an unattributed performance claim.

Load peer comments, goals, and the self-assessment first

Do not prompt a model with a name, a rating, and "write the review." Load the packet first.

Pull the peer comments that belong to this cycle, including skip-level and cross-functional notes if you collected them. Keep attribution with each comment so a later sentence can point at a specific source, not at "peers generally."

Load the goals that were in force for the period, with status notes the manager or employee already recorded. Do not treat a goal title as evidence of delivery. Treat the recorded progress, blockers, and outcomes as the source.

Load the self-assessment the employee wrote, including examples they chose and areas they flagged. The draft may quote or paraphrase that material. It may not contradict it with a claim the packet does not contain.

If a field is empty in the system of record, it is empty for the model. Missing peer comments are not a reason to infer skill from tickets, pull requests, or calendar density. That is a different workflow, skills inference from work artifacts, and mixing it in here is how invented claims sneak into a review.

Confirm the packet is for the right person, cycle, and manager before generation. Wrong-cycle comments are a common source of false narrative, and they look authoritative once they sit in prose.

Draft with cites, and leave blanks when sources are silent

Generation should produce candidate sentences, not a finished letter.

Each sentence that asserts something about performance must carry a cite: which peer comment, which goal record, or which self-assessment passage it rests on. A cite is a pointer a manager can open, not a decorative footnote. If the model cannot attach a pointer, it must leave that sentence blank rather than smoothing over the gap.

Use blanks on purpose. A blank next to "collaboration" when no peer mentioned collaboration is the correct output. Filling it with "works well with others" is inventing a performance claim. Filling it from the manager's private opinion is also inventing a claim relative to this packet, even if the manager later chooses to write that opinion in their own voice.

Keep the manager's observations in a separate layer, after they have seen the sourced draft. That split shows what the packet supports and what the manager is adding.

Do not auto-submit. The draft stays in a working state until the manager edits, accepts, or discards language and then submits through the same path they use today. Treating a generated draft as submitted is a failure mode. It skips the human who is accountable for the record.

A sentence with no source cite is the other failure mode to scan for on every sample. If a sentence cannot show its comment, goal, or self-assessment, delete it or convert it back to a blank.

Cited draft versus invented draft

Alex is an IC. Peers describe Alex as the person who unblocked a delayed vendor integration and documented the workaround. One goal is marked complete with a note that the integration shipped in the cycle. The second goal has no progress note. The self-assessment talks about the vendor work and asks for clearer priority-setting from the manager. No peer comment mentions mentoring. No source mentions missed deadlines.

A sourced draft might read:

  • Peers described Alex as the person who unblocked the delayed vendor integration and documented the workaround so the next team did not repeat the outage. [peer comments: Jordan, Sam]
  • The vendor integration goal is recorded as complete for this cycle. [goal: vendor integration]
  • Alex's self-assessment focuses on the vendor work and asks for clearer priority-setting. [self-assessment]
  • Mentoring: (blank)
  • Delivery against the second goal: (blank)
  • Missed deadlines: (blank)

An invented draft might add that Alex is a strong mentor, consistently hits every deadline, and exceeded both goals. None of those sentences have a cite in this packet. They are performance claims the model made up to look complete. That output fails the quality bar even if the prose is polished.

The manager can still write about mentoring if they observed it. They write it as their observation, not as something the packet proved. They can also use the blanks as a prompt to collect a missing comment or to leave the topic out of this cycle's review.

After the draft exists, you still need human review of ratings and narrative against the rest of the org. A fluent, cited draft can still sit at the wrong level relative to peers. Pair this workflow with a calibration bias detector before ratings lock. Do not treat a clean draft as a calibrated rating.

The manager writes and submits

Hand the sourced draft to the manager as editable text with cites visible. Ask them to keep, rewrite, or drop each sentence. Ask them to fill blanks only with language they can stand behind, or to leave topics out.

The manager submits in the review tool. The generator does not click submit, does not advance the workflow, and does not notify the employee on the manager's behalf. If your stack can technically post the document, keep that action off. Auto-submit is how a draft becomes an official record before anyone accountable has read it.

Keep compensation language out of this draft. Merit, bonus, and equity decisions are a separate packet. Point managers who skip from prose to money at the merit cycle agent rather than stuffing pay recommendations into a review paragraph the packet never supported.

When the review is filed, development follow-through is also separate. If the sourced draft and the manager's additions point to a ramp or reset, a 30-60-90 plan generator can take the agreed themes after the conversation, not before the review exists.

People-ops should sample drafts during the cycle, not after close. Look for sentences without cites, blanks that were filled with generic praise, and any draft that moved to submitted without manager edits. Those three are how quality dies in a busy review week.

Guardrails you can run this cycle

Publish the quality rule next to your review timeline: a draft sentence cites a peer comment, a goal, or a self-assessment; empty stays empty; the manager writes and submits; nothing auto-files; nothing invents a performance claim.

In the tool path, require the packet to load before generation is enabled. If peers, goals, or self-assessment are missing, show that they are missing. Do not hide absence behind a full page of prose.

In the output path, require a cite on every generated sentence. Reject or highlight any sentence the model emits without one. Blank fields are valid. Uncited sentences are not.

In the submit path, keep the human click. Log that a draft was generated, but do not treat generation as completion. Completion is the manager's submit.

Remind managers that citing a self-assessment is not the same as agreeing with it. They can quote the employee and then write a different conclusion in their own section, as long as that conclusion is clearly theirs.

If you do only one thing this cycle, enforce empty-stays-empty. Completeness pressure is what produces invented claims. A shorter, cited draft plus a manager's own paragraphs is a review. A long, unsourced letter is a liability, even when it feels helpful.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first