Skip to main content
DoneThat

AI Adoption GuideLegalAssess

Clause risk classifier

Classifies each clause as red, amber, or green against a risk rubric and outputs rationale per flag, using tools like Kira or LegalOn.

Legal processRequestAssessDraftNegotiateApproveSignStoreDispute

By Don, DoneThat’s AI coach · updated

Overview

A clause risk classifier reads each provision in incoming contract text, scores it against a defined risk rubric, and returns a red, amber, or green flag with a short rationale. The point is not to replace legal judgment. It is to give counsel a structured first pass so review time goes to the clauses that matter most.

The classifier works best when your organization already has a playbook or risk matrix: acceptable language, fallback positions, and hard stops. Without that rubric, the tool has nothing authoritative to score against and should return empty output rather than invent thresholds. When the rubric exists, each flag must cite the exact clause span and the rubric rule that triggered it. That traceability is what turns a color code into something a reviewer can trust or override in seconds.

What the classifier produces

For every clause the system identifies, output includes three elements: a risk color, the matched text span, and a rationale tied to a specific rubric rule.

Red means the clause hits a hard stop or falls outside approved bounds. Examples include uncapped liability, one-sided indemnities, or assignment rights that violate firm policy. Amber means the language is negotiable but off-playbook: it may be common in market paper but conflicts with your preferred position or needs a defined fallback. Green means the clause aligns with the rubric as written, including approved carve-outs and defined thresholds.

The rationale field is not a summary of the clause. It states which rubric rule fired and why the span matched. A strong output looks like: "Red: Limitation of liability excludes indirect damages without cap; rubric rule L-04 requires mutual cap at fees paid in prior 12 months." A weak output paraphrases the clause without naming the rule. Quality review should reject outputs that lack both span citation and rule reference.

If no rubric is loaded for the contract type, jurisdiction, or counterparty tier, the classifier returns nothing. That is correct behavior. Scoring against a generic "best practices" model without your firm's positions creates false confidence and wastes counsel time unwinding bad flags.

Where it fits in contract intake

Clause risk classification sits early in assess-stage review, usually after the document is parsed and before deep redline work begins. It pairs naturally with third-party paper summarization when counsel needs orientation on unfamiliar templates, and with missing clause detection when the question is not "how risky is this clause?" but "is a required clause absent entirely?"

It complements, rather than duplicates, playbook deviation reporting. Deviation reports typically compare draft language to approved fallback tiers across the whole agreement. A clause risk classifier applies the rubric at provision level and assigns severity colors for triage. Some teams run deviation analysis first for negotiation strategy, then use per-clause flags to prioritize which deviations to fight. Others run classification on first receipt of counterparty paper to decide whether the deal warrants senior review at all.

Jurisdiction risk flagging handles a different axis: governing law, forum, regulatory exposure, and cross-border enforceability. A clause can be green on substance but amber on jurisdiction. Workflows that combine both checks reduce the chance that counsel approves acceptable commercial terms in an unacceptable legal frame.

How classification runs in practice

Most implementations follow the same pipeline. The document is ingested (PDF, Word, or CLM export), segmented into clauses or defined sections, and each segment is evaluated against the rubric for that agreement type.

Rubric design determines output quality more than model choice. Effective rubrics define rule IDs, trigger patterns (keywords, structural tests, or exemplar clauses), severity mapping, and approved alternatives. Rules should be testable: two reviewers reading the same span should agree on whether L-04 applies.

Segmentation affects precision. Tools differ in how they split merged sections, schedules, and defined terms. Mis-segmentation produces flags on partial sentences or merged topics. Counsel should spot-check segmentation on the first few runs for each template family.

Scoring applies rubric rules to each span. Hybrid systems combine pattern matching for clear playbook hits with language models for paraphrased or non-standard wording. The rationale must still anchor to a rule ID, not model intuition alone.

Human review closes the loop. Reviewers confirm, downgrade, or upgrade flags and optionally feed corrections back into rubric tuning. The classifier is a triage layer; sign-off remains with licensed counsel.

Vendor capabilities

Four platforms commonly support rubric-driven clause classification or close equivalents. Capabilities overlap, but emphasis differs.

Kira is widely used for due diligence and contract analysis at scale. Its strength is structured extraction plus user-defined smart fields and checklists that map cleanly to red/amber/green rubrics. Large matter teams often build clause libraries and scoring projects that export span-level hits with rule metadata for downstream review.

LegalOn centers on playbook alignment for in-house teams. Its review flows compare incoming clauses against curated playbooks and surface deviations with suggested fallback language. The product experience is oriented toward per-clause guidance rather than bulk diligence grids, which suits assess-stage triage on vendor paper.

Ironclad embeds AI-assisted review inside CLM workflows. Classification and playbook checks typically run at intake or during workflow stages, with flags attached to the record counsel already uses for approval. The value is less standalone analysis than keeping risk signals inside the system of record.

Litera (including Compare and related contract intelligence offerings) leans toward document-centric review: comparison, precedent alignment, and emerging AI review features across Word and matter tools. Teams already on Litera for redlining often extend into rubric-style checks without changing primary authoring surfaces.

When evaluating vendors, ask whether flags include verbatim spans and rule IDs, whether rubrics are versioned per contract type, and what happens when no rubric matches. Also confirm export format for audit: counsel needs to show why a clause was red on a given date, not just that a model said so.

Quality bar and governance

Treat these as minimum acceptance criteria for production use.

Traceability. Every non-green flag names the rubric rule and quotes or offsets the clause span. Screenshots or PDF highlights are not enough for audit if the underlying rule reference is missing.

Empty when unconfigured. No rubric means no score. Do not fall back to vendor default "risk models" without explicit governance approval.

Counsel decides. The classifier informs priority; it does not approve contracts. Workflow design should make override easy and should log overrides for rubric improvement.

Calibration. Run parallel review on a sample of agreements quarterly. Measure false positives (amber/red on clauses counsel would accept) and false negatives (green on clauses counsel changes). Tune rules before expanding to new contract types.

Consistency with related checks. Align severity definitions with your deviation report tiers and jurisdiction flags so reviewers see one coherent risk language across tools.

Limitations worth planning for

Classification quality drops on heavily amended PDFs, scanned images, and non-standard headings. OCR noise and table layouts can fragment spans. Non-English or mixed-language agreements need rubrics and models scoped to those locales.

Market paper often uses different labels for the same risk (e.g., "consequential damages" vs. "indirect losses"). Rubrics should include synonym groups or embedding-based matchers, with human spot checks on edge cases.

Highly negotiated one-off deals may outpace static rubrics. For those matters, treat classifier output as advisory only or maintain a separate "strategic deal" rubric with fewer hard reds.

The classifier accelerates assess-stage review when rubrics are clear, outputs are auditable, and counsel retains final say. Invest in rubric maintenance first; tool selection second. Without that foundation, red, amber, and green are just colors on a screen.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first