Skip to main content
DoneThat

AI Adoption GuideHRHire

Contract clause review

LLM compares vendor or amended employment contracts against approved policy and flags deviations.

HR processPlanSourceSelectHireOnboardDevelopRewardExit

By Don, DoneThat’s AI coach · updated

Dual-cite or do not flag

A complete flag names the contract clause and the approved policy clause. Anything less is not a flag. It is an ungrounded comment, and comments do not go into the counsel packet.

The quality bar is strict. The model compares vendor paper or an amended employment contract to the policy counsel already approved. When a topic diverges, the row cites both sides. When the texts match on that topic, the row stays empty. Empty is a valid result. The system does not invent a deviation to fill space.

Counsel still approves every packet. The model does not sign, does not send for signature, and does not convert a flag list into clearance. A dual cite can still misread materiality, miss how two clauses interact, or cite the wrong policy version. Those are counsel judgments.

This check belongs after pay and letter language are largely locked. Numbers and bands are handled in comp recommendation per offer. Jurisdiction-specific letter text is handled in localized offer letter generation. Clause review is the later pass that the legal text still matches policy once those pieces sit in the same packet.

Assemble the signing copy and the signed playbook

Load two sources, then run. The first is the contract as it will be signed: vendor paper, a redline of your template, or an employment agreement marked up by the other side. The second is the approved policy as a versioned document, not a paraphrase from chat or a slide. If the playbook is stored in a CLM such as Ironclad, export the signed template and the exception list counsel last approved. If hiring ops stores offer attachments in Workday or Greenhouse, pull the files that will actually attach, not a leftover PDF from a prior req.

Record identifiers on the review. Filename plus date, envelope ID, or document ID for the contract. Policy ID plus effective date for the playbook. When a flag looks wrong, the first question is which two texts were compared.

Do not add a third "context" file that is not the approved policy. Prior deals, sales decks, and side emails invite the model to treat custom history as the standard. The standard is the policy you loaded.

Keep screening facts out of this comparison. If the contract mentions background checks, you are still comparing contractual language to policy. The contents of a screening report are a different extract, covered in background-check risk extraction.

Stop the run if either source is missing, unlabeled, or known stale. A comparison against last year's playbook will produce dual cites that are internally consistent and operationally wrong. Counsel should see the policy version on every flag so they can discard a bad packet.

Example: 90-day termination against a 30-day rule

Suppose a preferred vendor returns an amended master services agreement. Your playbook requires 30-day termination for convenience, mutual ownership of jointly created work product, and a liability cap of twelve months of fees. The vendor draft keeps the cap and the IP split. It changes termination to 90 days and adds a one-way indemnity for any claim arising from use of the services.

A complete termination flag cites both sides: Contract section 12.2 (90-day termination for convenience) against Policy Term-04 (30-day termination for convenience). A complete indemnity flag cites the new indemnity paragraph against the policy clause that limits indemnities to mutual IP infringement and bodily injury. The IP and liability-cap rows stay blank. Those clauses match.

An incomplete flag says termination is "longer than we usually accept" and names no policy section. That line cannot be audited. Send it back. A flag that cites only the contract is a highlighter, not a policy comparison. Counsel should not be asked to guess which playbook rule the model had in mind.

The other failure is invention. The model claims the liability cap is too low even though it matches Term-09. That invented deviation creates false work and trains reviewers to distrust the queue. HR ops must not "fix" blank rows by prompting the model to find more issues. If both texts match, the row stays empty.

This is an illustration of the output shape, not a measured outcome. Do not treat the section numbers as a template you copy into production without mapping them to your own playbook.

Empty rows are matches

Reviewers often read a blank as a failed run. It is not. The instruction is: flag only when both cites exist; otherwise leave the field empty. A packet with two dual-cited flags and a page of blanks is a successful comparison if those two topics are the only mismatches.

Route the packet so blanks do not block counsel, and so counsel still reads the contract. Blanks mean no dual-cited mismatch on the checklist topics. They do not mean the agreement is approved. Unlisted topics, interactions, and local law still sit with the human.

Do not treat the flag list as a sign-off. If operations sends a DocuSign envelope because the model found only two items and legal is busy, the control has failed. Flags are a worklist. Signature remains a human act, including when the worklist is empty.

Keep the checklist aligned with the policy source. Adding data-processing, insurance, or audit-right topics after the fact, without loading matching policy clauses, is another path to invented deviations. Expand the playbook first, then the checklist.

Hiring-process fairness is a separate trail. Whether decision-makers applied criteria consistently belongs in hiring decision bias audit, not in clause comparison.

Human approval stays outside the model

Employment counsel, or the designated reviewer, decides materiality, exception, and whether to send a redline. The model does not rank vendors and does not choose a CLM. Workday, Ironclad, Greenhouse, and DocuSign are where files, offers, and signatures already live. Park the comparison output next to those systems. None of them replaces dual-cite review, and this page does not assign them features they may or may not have.

When counsel agrees with a flag, they grant an exception or they counter. When they disagree, they record why: policy misread, clause interaction, or a commercial override. That record improves the next prompt and the playbook. It is not a reason to auto-accept the model's view next time.

Do not auto-sign. Do not let a workflow move Ironclad to approved, confirm an offer in Workday, or send a DocuSign envelope from a blank comparison or from an unreviewed flag list. The human approval step is the control. The model is a citing assistant.

Employment agreements follow the same rule as vendor paper. An amended offer contract that changes garden leave, IP assignment, or post-termination restrictions still needs a dual cite against the employment policy. "What we did on the last exec hire" is not the policy.

The operational test is whether a reviewer can open a flag and jump to both clauses without asking the model to explain itself. If yes, keep the workflow. If no, the output is not ready for counsel.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first