Skip to main content
DoneThat

AI Adoption GuideFinanceInvoice

Dispute-likelihood scoring

ML predicts which invoices will be challenged based on customer and line-item history.

Finance processPlanBudgetInvoiceCollectPayCloseReportAudit

By Don, DoneThat’s AI coach · updated

What a dispute-likelihood score is allowed to do

A useful score tells billing which invoices look like they will be challenged, and why, before the bill leaves the queue. The why is the cites: the customer history and the line-item pattern the model actually used. Billing still sends. Collections still works the dispute if one arrives.

The score is a quality signal, not a collections decision and not a cash hold. A high score means a person should look at the draft, the backup, and the wording. It does not mean the invoice sits until the score cools down. Treating the score as a hold that never sends is a failure mode.

Empty is a valid result. If the customer is new, the account has too few billed periods, or the line mix has no prior analogue, the score stays empty. Do not invent a historical dispute percent to fill the gap. A made-up rate is worse than a blank, because reviewers will treat it as measured history.

Systems that already hold billed records (Salesforce, Workday, HighRadius, Tabs) can supply invoices, activity, and sometimes contract context. They do not replace a score that cites its evidence, and they should not be asked to mint a dispute rate when the record is thin.

Load billed history the customer actually received

Start from invoices that were billed, not from quotes, not from unsent drafts, and not from a reconstructed typical dispute rate for the segment. Pull customer identity, invoice dates, line descriptions, amounts, tax and freight treatment, credit memos, and the outcomes of prior challenges when those outcomes exist in the record.

Match lines the way the customer would recognize them: SKU or charge code, description family, recurring versus one-off, and whether the line sat on a usage true-up or a fixed fee. A customer who routinely challenges usage true-ups is not the same as a customer who challenges tax on professional services. Collapsing both into one account-level risk number hides the pattern you need to cite.

Do not backfill missing periods. If three of the last twelve months never billed, those months are not zero-dispute months. They are missing. Scoring as if the gap were clean history is the same class of error as scoring a first-time customer as if they had history.

Run this load next to, not instead of, an invoice pre-flight check. Pre-flight catches structural defects (wrong entity, missing PO, tax on a tax-exempt account). Dispute-likelihood scoring asks a different question: even if the invoice is structurally complete, does this customer's history say they will still challenge these lines?

If the invoice was generated from a contract, keep the generation trail attached. A line that came out of contract-to-invoice generation still needs the billed-history check. Contract fidelity and dispute likelihood are not the same signal.

Score with cites, then suppress thin customers

The model output billing can use has three parts: a score, a reason-code family, and the cites. Cites name the prior invoices or line-item patterns that drove the result. A usable cite looks like this: this customer challenged usage overage lines on the last two invoices that carried overage, and this draft includes an overage line in the same charge family. That is enough for a reviewer to act. A silent score with no cites is not.

One walk-through: a reviewer opens a draft for a long-running subscription account. The score is elevated. The cites point to two prior invoices where the customer disputed a platform overage line that used a different unit description than the order form, and to a credit memo that reversed that line after a week of email. The current draft repeats the same unit description. The reviewer aligns the description to the order-form unit, attaches the usage backup, and sends. No dispute rate was invented. The score did not block send. Collections was not asked to pre-work a dispute that had not happened.

Suppress thin customers before the score is shown as a number. Keep the rule qualitative and strict: a first invoice, a first invoice in a new charge family, or a customer whose billed history does not include the line types on this draft. In those cases the field is empty and the screen should say why (insufficient billed history, no matching line-item pattern). Do not substitute a segment average.

Scoring a first-time customer as if they had history is the failure mode that looks sophisticated and is wrong. There is no customer pattern to cite. There may be a product pattern across other customers. That is a different model, and it must not be labeled as this customer's dispute likelihood.

Do not emit a historical dispute percent unless that percent is computed from this customer's own closed challenges over a defined billed window, and that window is shown. If the window is too short or the challenges are not coded, omit the percent. Inventing one to make the score look complete is a quality defect.

Billing reviews the draft, then sends

Put the score on the invoice that is about to go out, in the same place billing already reviews amounts and backup. A queue of high likelihood, empty because thin, and the rest is enough. Reviewers need the cites on the same screen as the lines.

Review is allowed to change line descriptions, quantity presentation, attached evidence, tax or freight treatment when the history says those are the disputed fields, and a short note in the invoice message that points to the backup. Review is not allowed to hold the invoice indefinitely because the score is high, zero out a legitimate charge to dodge a predicted challenge, or overwrite empty with a guessed rate so the dashboard looks populated.

After send, the score does not become a collections strategy. If the customer disputes, collections works the dispute with the same evidence the reviewer attached. If the customer pays, the outcome feeds the next history load. If the customer promises a date, that belongs on a promise-to-pay tracker, not back into the pre-send score as if a promise were a challenge.

Do not route high-score invoices into a quieter dunning path just in case. Adaptive dunning sequences start after the bill is outstanding. Mixing pre-send dispute likelihood with post-due dunning trains the system to nag a customer for a bill you already suspected they would reject, without first fixing the bill.

Keep the score out of jobs it cannot do

The score does not decide credit, does not replace a collections agent, and does not prove the invoice is wrong. It only ranks which drafts deserve a human pass because this customer's billed pattern looks like a challenge.

Keep three failure modes on the runbook. First, scoring a first-time customer as if they had history: empty stays empty. If you want a new-logo checklist, write a checklist. Do not dress it as this customer's likelihood. Second, treating the score as a hold that never sends: set a review SLA. If the reviewer misses it, the invoice still sends with whatever backup is on the draft. Unreviewed is not unbilled forever. Third, inventing a historical dispute percent: if you cannot show the numerator (closed challenges), the denominator (billed invoices in the window), and the customer identity, you do not have a rate.

Re-score when lines change in review. A description fix can retire the cite that drove the elevation. Do not leave the old score on a rewritten invoice.

Feed outcomes back on a delay. A dispute opened this week is not yet a pattern. Wait until the challenge is coded (invalid bill, pricing, quantity, tax, scope) before it becomes history the next score is allowed to cite.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first