Skip to main content
DoneThat

AI Adoption GuideHospitalityReturn

Return propensity scorer

ML model scores lapsed guests by return likelihood using recency, frequency, spend, and trigger events.

Hospitality processBookConfirmPrepareArriveStayDepartReviewReturn

By Don, DoneThat’s AI coach · updated

What the score is allowed to claim

A return propensity score ranks lapsed guests by how likely they are to book again, using the stay history you actually hold. It is not a conversion forecast, not an approval to send, and not a substitute for a campaign brief.

The operational output is three fields: a score, the factors that moved it, and the vintage of the history those factors came from. Recency, frequency, spend, and trigger events (a cancelled stay, a membership expiry, a rate inquiry after checkout) can all contribute. None of them, on their own, is a reason to fire an email.

If the model cannot name why a guest scored high, or cannot say when the last stay record was observed, marketing cannot tell whether the score is still true. A number without those two attachments is a ranking for a slide, not a signal for a CRM queue.

Property CRM, PMS, and revenue stacks such as Revinate, Oracle Hospitality, Mews, and Duetto already store stay and booking fields this kind of model needs. Treat them as a class of history sources. The work is the same: export dated stay history, score what is complete, leave the rest blank.

Load stay history before you score

Pull a guest-level stay file first. Minimum useful fields: last stay date, stay count inside a defined lookback, spend in that same window, booking channel, rate or package type, cancellation and no-show flags, and trigger events the property already records (loyalty status change, membership lapse, abandoned hold, post-stay complaint).

Define lapse on a calendar rule, not on instinct. A guest whose last checkout sits outside the property's return window is lapsed. A guest at 60 days may still be in a normal booking cycle. Write the rule so this score and the win-back segment prioritizer share the same clock. If those two disagree on who is lapsed, you will rank people the campaign is not allowed to touch.

Deduplicate on a stable guest key. Folio merges, email changes, and loyalty-number collisions split one person into a thin record and a fat one. Score the merged history. Otherwise the model will reward the incomplete copy or punish the complete one at random.

Do not mix live booking abandonment into this file. Guests who never finished a checkout belong with booking abandonment recovery. Training on unfinished purchases and then scoring finished-but-lapsed stays contaminates both recency and trigger events.

One illustrative pass, not a measured result: a city hotel exports two years of stays for guests whose last checkout is older than the property lapse window. The CRM lead joins last stay date, stay count, room-night spend, and two trigger flags (membership expired, post-stay complaint). Guests with a complete join get a score. Guests who only have an email and a single night with no dated stay do not. Nobody is added to a send queue at this step.

Cite factors and stamp the vintage

Write three fields back to the CRM for every scored guest: the score, the cited factors, and the history vintage.

Cited factors are the inputs that moved this guest, not a global feature dump. A usable cite looks like: last stay 11 months ago, three stays in the lookback, spend concentrated in shoulder dates, membership expired last quarter. "High propensity" is not a cite.

Vintage is the date of the newest stay or trigger record the model saw. If the last checkout in the extract is March and you score in September, vintage is March. Marketing needs that gap on the record. A high score on stale history is a hypothesis about an old guest, not evidence about a current one.

Store the lookback next to the vintage (stays used: the window ending on the vintage date). Recency, frequency, and spend only mean something inside a window. Without it, "frequent" might mean frequent years ago and silent since.

Do not emit a conversion percent. Propensity is an ordering among lapsed guests you can describe. A percent would mean you already measured campaign response for this score band, on this property, in a named period. If that test has not run, you do not have that number. Inventing one turns a ranking into a fake forecast and it will be quoted in budget meetings as if it were observed.

Leave thin history blank

Empty stays empty. If stay count is missing, last stay date is missing, or spend cannot be tied to a stay, do not impute a midpoint and do not borrow a segment average.

Thin history is ordinary hotel data, not an edge case: one complimentary night, a group booking under a company folio, a cash desk stay with no profile match, a loyalty number that never attached to the PMS. A score on those rows is a guess with a decimal.

Blank is an operational state. Show "unscored, history insufficient" and the reason: no last-stay date, spend not joinable, vintage older than the allowed freshness window. That queue is for data repair. It is not a cheaper win-back list.

If vintage is missing, invalidate the score even when the number looks precise. A score with no vintage cannot be compared to last week's run and cannot be trusted after a PMS cutover or a CRM re-extract.

Known loyalty events should not be laundered into a stay-pattern score. Points expiry and tier-year resets belong with loyalty milestone trigger agent when the event date is known. Mixing them in, then citing only recency, frequency, and spend, hides the real trigger.

Marketing still chooses the campaign

The score orders who is more likely to return among guests you can describe. It does not choose the offer, the channel, the holdout, or the send time.

Marketing or CRM still picks the campaign: which score band enters a test, which creative, whether to suppress complaint flags, whether to wait because vintage is stale. The hyper-personalized re-engagement generator can draft copy from cited factors. It must not send because a threshold was crossed.

Treat "score as send" as a process failure. Auto-enrolment from a cutoff will include guests whose top cite is a complaint, guests whose vintage predates a renovation, and guests who already booked on another channel that has not landed in the CRM.

Use the score to order work. Fresh vintage, clean cites, and a complete stay join go to the first test cell. Mid scores wait. Blanks stay out. After a campaign, record response against the score version and vintage you used. Judge the next ranking on bookings you observed in scored cells versus a holdout, not on a conversion percent nobody measured.

Failure modes that look like a working model

A score with no vintage. The dashboard looks complete. Rankings jump because the extract date changed, not because guests changed. Stop any send that depends on that run. Backfill vintage from the extract timestamp or re-score from a dated file.

Treating the score as a send. Operations wants a list by Friday. The top band dumps into the ESP. Complainers, already-booked guests, and imputed thin records go out on the same template. Keep the score in the CRM. Build the audience in the campaign tool with human suppressions.

Inventing a conversion percent. Someone asks what lift to expect. A slide gets a number because empty feels worse than precise. Refuse it. Until a holdout test exists, the honest artifact is an ordered list with cites and vintage.

Source-system drift is quieter. If Revinate, Oracle Hospitality, Mews, or Duetto (or the warehouse that lands their exports) changes a stay-status code, frequency and recency shift without a retraining ticket. Re-check field definitions when a score distribution jumps between runs.

The desk test is simple. Open one high-score guest. You should see why they scored, when the history was current, and that a person still has to choose whether they get a campaign. If any of those three is missing, the score is not ready to use.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first