Skip to main content
DoneThat

AI Adoption GuideHospitalityReview

Review sentiment and topic extractor

NLP extracts themes, staff mentions, and sentiment scores across all platforms into a unified reputation dashboard, using tools like Revinate or TrustYou.

Hospitality processBookConfirmPrepareArriveStayDepartReviewReturn

By Don, DoneThat’s AI coach · updated

Quality is a cited theme, not a reputation score

The extractor is doing its job when each theme points at the exact review span it came from. That cited theme is the quality outcome. Staff mentions and sentiment scores sit on those spans as supporting fields. They do not replace the cite, and they do not become a property KPI.

Reputation suites such as Revinate, TrustYou, Oracle Hospitality, and Mews already ingest guest comments from OTAs, post-stay surveys, and direct channels. Treat them as a class of licensed sources. The extractor reads text those systems are allowed to store and writes rows a reputation analyst can open next to the original review. It does not replace the vendor dashboard, and it does not invent a house score those products already own.

If the model cannot quote the words that justify a label, the label does not ship. If the guest never mentioned a department, that department stays empty. Operations still decides which columns the unified reputation dashboard shows, and which vendor-native scores remain on screen.

A theme is an annotation of a quote. Tallying how often "slow breakfast" appears this week can help you decide what to read first. Promoting that tally to a KPI, then ranking outlets against it, turns a text label into a fake operating metric. Leave scoring to the licensed platforms.

Load licensed reviews and nothing scraped

Start from the property's licensed export or API. Required fields on every row: property code, review id, platform, language, full text, stay date when the source provides it, and any guest-facing rating the source already stored. Keep that source rating as a passthrough. Do not recompute it from adjectives.

Unify platforms at ingest so one grid can filter OTA, survey, and direct comments. Unification means shared columns and a shared property key. It does not mean averaging scores across vendors, and it does not mean pasting unlicensed social screenshots into the same table.

Drop records with empty bodies. A star rating with no text is not a theme source. Same-day volume spikes belong on the review alert monitor. The extractor's job is the cited theme row, not the pager.

When a review will also become a complaint file, keep this pass narrow. Theme labels can later feed a complaint root-cause classifier. Do not rewrite themes to match a root-cause taxonomy while you are still extracting.

If the license lapses, or the only copy you have is a screenshot, stop. Uncitable text is not an input.

Extract themes and staff mentions from quoted spans

For each licensed review, propose zero or more themes, zero or more staff mentions, and a sentiment polarity per theme span. Do not attach one global sentiment to the whole stay when the paragraph mixes praise and complaint. Every theme must include the exact sentence or character span it cites. Every staff mention must include the name or role as written, plus the clause around it.

Reject a theme that has no span cite. "Housekeeping friction" with a null quote is a failed row, even if the rest of the review sounds unhappy. A rooms director cannot act on a label they cannot click. The same rule applies to people: do not infer "the night manager" when the guest wrote "front desk."

Sentiment lives on the span. Store mixed polarity as separate theme rows, each with its own cite.

Here is one illustrative pass, not a case study. A guest writes: "Check-in took forever and Marco at night audit kept apologizing, but the room was quiet and the pillows were great. Never saw the pool."

A valid extract records three things and leaves the rest blank. Theme: check-in wait, cite "Check-in took forever," sentiment negative. Staff mention: Marco, night audit as written, cite "Marco at night audit kept apologizing." Theme: room quiet and bedding, cite "the room was quiet and the pillows were great," sentiment positive. Pool stays empty. "Never saw the pool" is not a pool-quality theme. Do not invent "pool unavailable," and do not output a pool score.

An invalid extract invents an overall reputation score, labels "F&B underperformance" with no F&B sentence, or treats check-in wait volume versus last week as the outcome of this job.

After extraction, reputation or the duty manager covering it reviews rows before they publish. The model does not write the dashboard.

Leave blanks when the review is silent

Silence is a valid result. Most reviews name one or two moments and skip the rest of the hotel. Do not backfill housekeeping, F&B, spa, or location because those columns exist.

Empty is not the same as a failed parse. If the model could not read the language, mark the run skipped and keep the raw text for a later pass. If the model read the text and found no span for a topic, leave that topic blank. Do not collapse those two states into one null.

Do not translate a missing theme into a neutral sentiment of zero. Neutral sentiment is a claim about words that exist. A blank is a claim that the words were not there. Inventing a review score to fill the hole, whether by averaging span polarities or by blending OTA and survey stars, is out of scope.

When ops asks for department coverage, the honest line is which departments were cited this period and which were not mentioned. That is a quality statement about the extract. It is not a league table.

Operations owns the dashboard the model only annotates

Sample new rows before they become tiles: every high-severity staff mention, a slice of mixed-polarity reviews, and any theme tagged with low cite confidence. Confirm the quote is in the review, the label is not stretched, and no score was generated.

Approved rows land in the unified dashboard operations already uses. That dashboard may still show scores native to Revinate, TrustYou, Oracle Hospitality, or Mews. Those scores stay vendor-owned. The extractor only adds cited themes, staff mentions, and span-level sentiment beside the original text.

Do not treat approved themes as KPIs once they are on the board. They are filters for which quotes to open. If "reduce check-in wait mentions" becomes a target, teams start arguing with the label instead of the stay.

Response writing is downstream. After a cited theme is approved, a personalized response drafter can use the quote. It must not thank departments the guest never named.

Cross-property views come last. A cross-property reputation benchmarker can show which cited themes appear at which hotels. It still cannot turn theme counts into a KPI or invent a house reputation index. If a sister property's extract is empty for spa, that spa column stays empty.

Stop the pipeline when output rows cannot be clicked back to source text. Quality is the cited theme. Everything else is annotation around it.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first