Skip to main content
DoneThat

AI Adoption GuideHospitalityReview

Complaint root-cause classifier

Classifies negative reviews by operational category, such as housekeeping, F&B, or noise, to feed department-level KPIs.

Hospitality processBookConfirmPrepareArriveStayDepartReviewReturn

By Don, DoneThat’s AI coach · updated

What you are allowed to write on the record

The quality outcome is a category label that cites the review sentence supporting it. For a rooms-division manager that usually means buckets already used in the weekly ops huddle: housekeeping, food and beverage, noise, maintenance, front desk, or another department you already hold accountable. The record stores the label and the span. It does not store a generated review score, a severity number, or an automatic debit against a department.

If the guest never states a cause, the category and the cite both stay empty. Empty is a finished result. Filling a bucket so a chart has fewer holes is how housekeeping, F&B, or noise absorb reviews they did not earn.

This step is narrower than theme extraction. A review sentiment and topic extractor can tell you the comment is negative and which topics appear. Root-cause classification answers one operational question: which department-shaped cause did the guest actually write down, and which sentence proves it.

Reputation inboxes and property systems such as Revinate, TrustYou, Oracle Hospitality, and Mews are a class of places the text already lives. Load the comment from the system you already use. Do not treat the classifier as a replacement for those products, and do not assume any of them already returns a cited operational category.

Load the review before you pick a department

Start by loading the full review body, not a truncated dashboard line and not last week's pre-tagged department. You need the sentences in order so a cite is a contiguous span a night manager can re-read without opening a second tool.

Stay facts you already have (room type, arrival date, outlet names) are background for a human check. They do not authorize a label when the text is silent. A guest who stayed in a connecting room is not automatically a noise complaint.

If the same comment also tripped a review alert monitor, treat the alert as a timing signal only. An alert says look now. The classifier says this sentence names this operational cause, or it says nothing.

Do not attach a model-invented 1-to-5 or 0-to-100 complaint score while you load. If the guest entered a rating in the source system, keep that rating as the guest's rating on the review record. Inventing a second number and calling it quality is a failure mode, not extra insight.

Assign a category only with a cited sentence

Assign a category only when you can quote the sentence, or a short contiguous clause, that names the cause. The cite is the quality gate. A category with no sentence cite is the first failure mode: housekeeping tagged because the stay "felt messy" to the reader, while the guest wrote only about a late checkout argument. That label will flow into a housekeeping KPI the text does not support.

Work the steps in this order:

  1. Read the loaded review once for any operational cause the guest actually stated.
  2. Pick the primary category that matches that statement, using your property's existing department list.
  3. Store the category together with the exact span you relied on.
  4. If you cannot highlight a span, leave category and cite blank.
  5. Hand the record to operations. They decide whether it counts toward a department KPI, a coaching note, or a discard.

One illustrative pass, not a property result. Guest text: "Breakfast was slow both mornings and the coffee was cold. The bed was comfortable and the view was fine." The classifier may return food and beverage, citing "Breakfast was slow both mornings and the coffee was cold." It must not stamp housekeeping, invent a review score, or deduct points from the restaurant automatically. Comfortable bed is not a second root cause in this workflow. You are classifying the complaint, not writing a balanced report card.

If the guest names two distinct operational failures in two sentences, your local rule should already say whether you store one primary category or two labeled rows. Either rule is an ops choice. Every stored category still needs its own cited span. A second label with no cite is still the uncited-category failure.

Empty stays empty when the guest never names a cause

Negative language without a cause is common: "Worst stay in years. Never again." That is sentiment. It is not a root-cause category. The correct output is empty category and empty cite.

Pressure to put it somewhere is how departments absorb reviews they did not earn. Do not map vague disappointment onto the team that is currently over target, onto the last complaint you remember, or onto a default rooms bin.

A mid-stay complaint predictor may have flagged the reservation while the guest was in-house. That prediction is a different artifact. Do not copy its suggested department into the post-stay classifier when the published review never states a cause. Prediction and classification fail in different ways. Merging them hides both.

Blanks also apply when the only complaint is something the property cannot operationalize as a department cause, such as a rating with no text. Do not manufacture a maintenance or noise label to keep the row complete.

Department KPIs still belong to operations

The second failure mode is treating the category as if it were already the KPI. A count of housekeeping labels is a feed. Housekeeping's actual KPI remains whatever the rooms-division manager already owns: inspection results, reclean rates, verified linen comments, guest-recovery actions, or the metric your brand requires. The classifier does not auto-assign a department penalty, a bonus deduction, or a ranking.

Use the label this way:

  • Roll cited categories into a weekly department review as a candidate list.
  • A supervisor opens the cite, confirms the sentence, then decides whether the item counts.
  • Uncited or blank rows stay out of the numerator.
  • Disputed cites go back to a human, not into an automated debit.

The third failure mode is inventing a review score and ranking departments by that score. Root-cause classification does not produce a 2.4 for F&B or a 67 for noise. Guest-entered ratings stay guest-entered. Do not generate a second number and call it quality.

When a cited category is confirmed, a personalized response drafter can use the same span so the public reply names the issue the guest wrote. Drafting a reply is downstream. It must not change the category or fill a blank.

Keep the rooms-division loop in charge of counting

Who signs off stays with operations. The night manager, rooms-division manager, or designated reputation owner accepts or rejects each cited label before it hits a department dashboard. Engineering and F&B managers see items that survived that check, with the sentence attached, so they are not arguing with a naked tag.

Do not wire the classifier to payroll, contractor scorecards, or automatic work-order penalties. A noise label citing an ice machine is a prompt to inspect the corridor, not a system-generated write-up. A housekeeping label citing stained carpet is a prompt for the HK manager to verify, not an automatic reclean fee on an invoice.

If comments arrive from more than one place, a Revinate or TrustYou inbox versus notes sitting next to a reservation in Oracle Hospitality or Mews, normalize to one review record before you classify: same text, same stay identifier, one field for category and cite. Duplicate ingest is how the same sentence becomes two KPI hits.

Re-run classification only when the source text changes, such as an edited public comment or a concatenated follow-up. Re-running to fill blanks so a chart looks complete is how silent reviews get fake causes.

The durable rule fits on a huddle board: category plus cited sentence, or empty. Operations still owns whether it counts. No invented score. No automatic department penalty.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first