Skip to main content
DoneThat

AI Adoption GuideRetailReturn

Return Reason Classification

LLM classifies free-text return reasons into structured defect, fit, description, delivery, and expectation categories for merchandising and supplier action.

Retail processPlanBuyPriceStockSellFulfillReturnClear

By Don, DoneThat’s AI coach · updated

What return reason classification does

Return reason classification turns free-text customer explanations into a structured label set that quality, merchandising, and supplier teams can aggregate. Shoppers often write short, messy notes such as “too small,” “color looked different,” or “box crushed.” Those notes rarely match the coarse dropdown options on a return form, so analysts either ignore them or spend hours reading samples by hand.

An LLM reads each free-text reason (and any linked order or product context you choose to pass in) and assigns one or more categories from a controlled taxonomy: defect, fit, description mismatch, delivery damage or delay, and expectation gap. The model does not change refunds, dispositions, or catalog content. It produces structured classifications that humans use to decide what to investigate next.

Empty or near-empty reason fields should yield no classification. If the text is too thin to support a confident label, the system returns no category rather than guessing. That keeps dashboards from treating silence as a fake “other” spike.

Why free-text reasons matter for quality work

Dropdown-only reasons compress nuance. “Wrong size” and “doesn’t fit” can mean different things: sizing chart error, stretch after wash, or style that runs small across an entire style family. Description complaints often point at photography, material callouts, or incomplete specs. Delivery issues belong with logistics partners, not with the product team that owns fit.

Quality analysts need counts by SKU, style, vendor, and time window that map to those distinctions. Without structured categories, merchandising cannot tell whether a spike is a supplier defect wave, a size-curve problem, or a PDP that overpromises. Classification is the bridge between narrative returns and operational queues.

Human judgment still owns the response. The model labels; buyers, quality leads, and supplier managers decide whether to open a vendor claim, adjust size charts, revise imagery, or leave the assortment alone.

Taxonomy and labeling rules

A practical taxonomy stays small and action-oriented:

  • Defect: manufacturing fault, broken component, stained or worn-out-on-arrival item, missing part.
  • Fit: size, cut, proportion, stretch, or wearability relative to expected size.
  • Description: color, material, dimensions, features, or imagery that do not match what arrived.
  • Delivery: damage in transit, late arrival that drove the return, wrong item shipped, packaging failure.
  • Expectation: subjective disappointment that is not clearly a defect, fit miss, or false description (for example, “not as nice as I hoped”).

Multi-label output helps when a customer writes “small and the seam came apart.” Prefer the most actionable primary label when the text supports only one strong reading. When evidence is weak, leave the field empty instead of forcing “expectation.”

Prompt and schema design should forbid inventing product facts that are not in the reason text or approved context fields. Classification confidence can be returned as a score for sampling review, but low-confidence or empty-text cases should not auto-populate analytics as hard facts.

Inputs, outputs, and empty-output policy

Typical inputs:

  • Free-text return reason (required).
  • Optional structured reason code from the return portal.
  • Optional product identifiers (SKU, style, size, color).
  • Optional order and fulfillment metadata (carrier, ship method, delivery status).

Typical outputs:

  • Primary category and optional secondary categories from the taxonomy.
  • Short rationale grounded in the customer text (for auditor sampling).
  • Empty classification when the reason is blank, emoji-only, or too vague (“didn’t like it” with no further detail may still map to expectation if your policy allows; “n/a” or “.” should not).

The empty-output rule is a quality control, not a bug. Analysts should track the rate of unclassifiable reasons separately. A rising empty rate often means the return UI is discouraging detail, not that classification is broken.

How quality and merchandising use the labels

Once reasons are structured, analysts can:

  • Rank SKUs and vendors by defect-class volume and rate, not by raw return count alone.
  • Separate fit spikes from description spikes so size-curve work does not get mixed with PDP fixes.
  • Route delivery-tagged returns to logistics or packaging reviews instead of supplier quality meetings.
  • Sample high-volume expectation returns to see whether marketing copy is creating avoidable regret.

Workflows stay human-led. A weekly quality review might pull the top defect-classified styles, open a small inspection sample, and only then raise a supplier corrective action. Merchandising might freeze a size run after confirming fit classifications against size-chart feedback, not because a model score crossed a threshold.

Pair classification with related controls carefully. Fraud risk scoring and policy compliance checkers answer different questions (abuse and refund eligibility). Reason classification answers “what went wrong with the product experience?” Do not collapse those signals into one label.

Limits and operating guardrails

Models misread sarcasm, bilingual notes, and telegraphic slang. They may over-weight the first clause in a long complaint. They cannot see the physical item; a customer saying “defective” may be describing preference, not a true fault. Treat classifications as triage labels for sampling and dashboards, not as proof for chargebacks or vendor penalties.

Operational guardrails:

  • Require a minimum text length or substance check before classifying.
  • Keep a human review sample (for example, stratified by category and SKU volume) to measure agreement and drift.
  • Version the taxonomy; changing labels without a migration plan breaks trend charts.
  • Never auto-trigger supplier deductions, assortment kills, or PDP rewrites from a classification alone.
  • Log empty outputs and low-confidence cases for UI and prompt improvement.

When those guardrails hold, return reason classification gives quality analysts a repeatable way to turn noisy return notes into categories merchandising and suppliers can act on, while leaving final decisions with the people accountable for product and partner outcomes.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first