Skip to main content
DoneThat

AI Adoption GuideLegalAssess

Playbook deviation report

Scores how far each clause deviates from internal playbook thresholds, producing a ranked issue list.

Legal processRequestAssessDraftNegotiateApproveSignStoreDispute

By Don, DoneThat’s AI coach · updated

Overview

A playbook deviation report compares incoming contract language against your organization's approved playbook thresholds and returns a ranked list of clauses that fall outside acceptable bounds. The output is a quality assessment artifact: it tells counsel which provisions diverge from internal standards, how far they diverge, and which divergences matter most. It does not replace judgment about what to accept, reject, or negotiate.

Playbooks encode what your legal team has already decided is acceptable, preferred, or prohibited. Vendor paper rarely matches those defaults on first pass. Manual comparison clause by clause is slow and inconsistent across reviewers. A deviation report automates the comparison step so counsel starts from a structured issue list instead of a blank redline.

What the report produces

The primary output is a ranked issue list. Each entry represents one clause (or defined span within a clause) that fails one or more playbook rules. Ranking reflects deviation severity and, where configured, business impact weighting. High-severity items surface first so limited review time goes to the provisions most likely to block signature or create downstream exposure.

The report is scoped to assessment, not execution. It scores and orders; it does not auto-accept language, send counter-proposals, or update clause libraries. That separation keeps the workflow auditable and leaves negotiation strategy with counsel.

When no playbook is configured for the document type or jurisdiction in scope, the report returns empty. An empty result is intentional: without baseline thresholds, deviation scoring has nothing to measure against. Teams should treat an empty report as a configuration signal, not a clean bill of health. Pair an empty playbook state with missing clause detection or third-party paper summarization if the goal is coverage review rather than threshold comparison.

How deviation scoring works

Scoring begins after the contract is parsed into clause-level units. Each unit is evaluated against applicable playbook rules: numeric caps (liability limits, notice periods), enumerated allowed values (governing law, dispute forum), boolean requirements (mutual indemnity, audit rights), and textual patterns (prohibited carve-outs, non-standard definitions).

A deviation score quantifies distance from the playbook's acceptable range. Exact mechanics vary by rule type. A liability cap set at two times annual fees when the playbook ceiling is one times annual fees produces a measurable gap. A governing-law clause naming a non-approved jurisdiction may score as a categorical miss rather than a gradient. Rules can carry weights so that, for example, data-protection gaps outrank minor payment-term drift.

Multiple rules may fire against the same clause span. The report typically presents the highest-severity finding per span for readability, with secondary rule hits available in detailed views where vendors support them. Conflicting rules (one rule requires language another forbids) should surface as configuration errors during playbook maintenance, not as silent failures at review time.

Ranking combines severity tier with optional business context. A medium-severity limitation-of-liability deviation on a high-value strategic deal may rank above a high-severity force-majeure gap on a low-spend order form if the playbook assigns deal-tier multipliers. Document the weighting model so reviewers trust the sort order.

Fields in each deviation record

Every listed deviation should be traceable back to source text and playbook authority. Minimum viable records include:

Clause span. The exact location in the document: section number, heading, or character offset range depending on parser fidelity. Spans must be stable enough for counsel to jump to the provision in the source file or CLM viewer. Vague references ("indemnity section") erode trust.

Playbook rule ID. The internal identifier for the rule that triggered the finding. Rule IDs tie the report to your playbook version, change logs, and approval history. When playbooks update, historical reports remain interpretable if rule IDs persist or map cleanly across versions.

Severity tier. A discrete label (commonly critical, high, medium, low, or a vendor-specific scale aligned to your taxonomy). Tiers reflect pre-negotiated risk appetite, not model confidence alone. A clause can match playbook language poorly yet sit in a low tier if the business has explicitly accepted that posture for the contract type.

Optional enrichments worth configuring when available: suggested fallback language from the playbook, comparison to your standard paper, links to related findings (e.g., indemnity and limitation-of-liability pairs), and cross-references to clause risk classification or jurisdiction risk flags when those assessments ran on the same document.

Counsel workflow and negotiation boundaries

The deviation report is input to decision-making, not the decision. Counsel uses the ranked list to prioritize redlines, delegate low-risk accepts, and prepare negotiation briefs for business owners. Items marked critical or high typically warrant pushback or escalation; medium and low items may be candidates for accept-with-note or batch counter-language from playbook fallbacks.

Business stakeholders often ask for a single "pass or fail" score. Resist collapsing the report into one number without context. A contract can satisfy most playbook rules yet contain one critical deviation (unlimited liability, missing data-processing terms) that blocks execution. The ranked list preserves that nuance.

Negotiation choices stay with humans: trade one deviation for another, accept risk with documented approval, or walk away. The report should never imply that playbook compliance equals legal sufficiency or that deviation equals unacceptability. External law, counterparty leverage, and deal timing all sit outside playbook math.

For inbound third-party paper, run deviation scoring after third-party paper summarization when reviewers need orientation first, or in parallel when the team already knows the document shape. Use missing clause detection alongside deviation reporting: absence of required clauses is a different failure mode than presence of sub-threshold language.

Vendor capabilities

Four vendors commonly support playbook-style deviation analysis in enterprise legal workflows. Capabilities differ in playbook authoring, rule expressiveness, and CLM integration depth.

Ironclad embeds playbook logic in its contract lifecycle platform. Playbooks drive workflow routing and approval as well as deviation surfacing during review. Strength: tight coupling between playbook maintenance and live negotiations. Consider when playbook and CLM are already centralized on Ironclad.

LegalOn focuses on AI-assisted review against customer playbooks and market benchmarks. Emphasis on reviewer UX and explainable hits. Strength: fast time-to-value for teams that want playbook deviation lists without full CLM migration.

Kira (Litera) offers structured extraction plus comparison workflows historically strong in due diligence and large corpus review. Playbook-style checks often layer on extracted fields and clause taxonomy. Strength: high-volume document sets and repeatable extraction schemas.

Litera (broader platform, including Kira) pushes playbook and precedent content through review and drafting products. Deviation reporting may appear in Compare, Contract Companion, or Kira projects depending on SKU. Strength: organizations already standardized on Litera for drafting and compare.

Evaluate vendors on rule ID stability in exports, severity tier configurability, empty-state behavior when playbooks are missing, and whether deviation records include precise clause spans usable outside the vendor UI.

Operational prerequisites

Reliable deviation reporting depends on playbook hygiene more than model choice. Maintain versioned playbooks per contract type and jurisdiction. Align rule IDs with jurisdiction risk flag triggers so geographic exceptions do not fight generic rules. Retire or remap rules when legal positions change; stale rules produce false deviations and train reviewers to ignore the report.

Define severity tiers with worked examples. "High" for uncapped liability should mean the same thing across reviewers and automation. Calibrate tiers against a sample set of historical deals counsel already classified.

Log playbook version on every report generation. When counsel disputes a finding six months later, you need to know which threshold set applied.

Treat an empty report after playbook configuration as a parsing or scope failure worth investigation: wrong document type tag, unsupported file format, or rules gated on metadata that was not supplied.


Related assessments: Clause risk classifier · Missing clause detection · Third-party paper summarization · Jurisdiction risk flag

[REDACTED]

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first