Return rate by category: what's healthy, what's a product problem, what's fraud
"Our return rate is too high" is one of the most common complaints in e-commerce and one of the least useful statements. The number only means something in context. Context that most merchants never compile, because they look at a single store-wide average and compare it to a benchmark from a trade publication that lumps their category in with fourteen others.
A 20% return rate on dress shirts is fine. A 20% return rate on ceiling lights is a crisis. A 20% return rate on a specific SKU that just spiked from 6% is neither. It's a signal you need to read correctly before you react.
Below is the framework for reading it. Benchmarks, decision tree, what each pattern usually means, what to do about it.
The honest benchmarks
Published benchmarks vary wildly depending on who's collecting the data and how they define a return. The numbers below are calibrated against a mix of public industry reports, Shopify merchant aggregates from various platforms, and pattern data we see across our own customer base. Reasonable estimates, not gospel.
| Category | Healthy return rate | Notes |
|---|---|---|
| Apparel (general) | 20% to 30% | The outlier category. High rates are baseline |
| Apparel (basics: tees, socks, underwear) | 5% to 12% | Simpler sizing, lower expectations |
| Apparel (dresses, suits, occasion wear) | 30% to 45% | Fit-hard categories, plus wardrobing exposure |
| Footwear | 20% to 35% | Sizing variance between brands is worse than apparel |
| Accessories (bags, wallets) | 8% to 15% | Color/style decisions, not fit |
| Jewelry (fashion) | 10% to 20% | Photo-vs-reality mismatch is real |
| Jewelry (fine) | 4% to 10% | Higher consideration purchase, lower return |
| Consumer electronics (general) | 8% to 15% | DOA, specs-don't-match, compatibility |
| Small appliances | 10% to 18% | Installation frustration is a major driver |
| Home goods (furniture) | 6% to 12% | Freight cost deters casual returns |
| Home goods (décor, lighting) | 4% to 10% | Color-match is the big driver |
| Beauty (unopened) | 4% to 8% | Most can't be returned opened |
| Beauty (trial-size, sample programs) | 15% to 25% | Different economics |
| Food / grocery / consumables | 1% to 4% | Returns are rare. Leakage is in substitutions |
| Pet supplies | 3% to 8% | Goods-fit-pet is a narrow failure mode |
| Sporting goods | 10% to 20% | Size/fit on apparel-adjacent, low on gear |
| Books, media | 2% to 6% | Almost nobody returns books |
Overall rates. Your category-specific rate should land within one of these bands. When it doesn't, the diagnostic question is always the same: why this category specifically, and why now?
The three patterns to separate
Once you know whether your category-level rate is elevated, the next question is what kind of elevation. Three patterns, and they look different in the data.
Pattern 1: product problem
A single SKU or closely related group of SKUs has a return rate significantly above its category baseline. The spike is recent. Started in the last 30 to 90 days. Return reasons cluster around a specific complaint. "Runs small," "didn't fit as pictured," "stopped working after X uses," "color different from photos."
A product problem. A fabric changed, a supplier substituted a component, the size chart didn't update when the fit did, photography lighting shifted, the product description created an expectation the item doesn't meet.
What the data looks like:
- 1 to 3 SKUs account for most of the spike
- Return reasons concentrated in one or two categories
- The spike coincides with a batch, a supplier change, or a photo refresh
- Customer complaint sentiment is consistent. They all felt the same thing was wrong
What to do: fix the product or the listing. Not a fraud problem. Not a policy problem. The return rate is correctly signaling that something about the item isn't matching expectations. Until you address that mismatch, every return policy change you make will cost you legitimate customer trust.
Pattern 2: fit problem
A category has an elevated rate (say apparel running at 35% against a baseline of 22%) but the elevation is distributed across many SKUs, not concentrated. Return reasons are dominated by sizing. "Too small," "too large," "sizing inconsistent."
A fit problem. Either your size chart is wrong, your model photos are misleading, your fit varies between batches, or your customers have learned to bracket because they don't trust the size chart.
What the data looks like:
- Elevation distributed across many SKUs, not one or two
- Sizing-related return reasons dominate
- Return rate doesn't decline for repeat customers (they're bracketing, not learning)
- Multiple sizes of the same SKU in a single order is common
What to do: improve fit information. Model measurements in reviews, fit indicators (runs small / large / true), better size charts, a fit-finder tool. This is where to spend engineering and content effort, because every percentage point you take off the fit-driven return rate is a meaningful margin recovery.
Pattern 3: fraud / abuse pattern
A category has an elevated return rate that's concentrated not by SKU but by customer. A minority of customers produce a majority of returns. Return reasons are inconsistent or skew toward benign categories. "Changed my mind," "didn't like," "other." Returned items show signs of use, or arrive in conditions that don't match the claimed reason.
Abuse. Wardrobing, bracketing at its abusive end, coordinated return fraud.
What the data looks like:
- 2% to 8% of customers produce 30% to 50% of returns in the affected category
- Return reasons are benign and varied rather than specific
- Repeat-customer return rate is high and doesn't decline over time
- Items return with subtle wear, missing tags, non-matching serial numbers
- Sometimes address clustering or shared payment methods across accounts
What to do: the hardest category to address, because broad policy changes hurt your legitimate customers more than they constrain the abusers. The effective response is customer-level pattern detection. Identify the specific customers whose pattern crosses your threshold and apply friction selectively. Blocklist after escalation, stricter documentation requirements, limits on bracket sizes, refusal to accept returns without clear tag presence. All applied to the 5% with the bad pattern, not the 95% who have nothing to do with it.
The decision tree
When your category return rate spikes, work through it in order.
- Is the spike recent and concentrated in specific SKUs? Product problem. Audit the SKUs.
- Is the spike distributed across the category but concentrated in fit-related return reasons? Fit problem. Audit your size chart and fit information.
- Is the elevation concentrated in a minority of customers with generic return reasons and high repeat-return rates? Abuse pattern. Customer-level detection.
- None of the above, a general elevation across SKUs, customers, and reasons? Usually a seasonal shift, a promotion-driven acquisition of different-than-normal customers, or a shipping carrier issue producing damaged arrivals. Check external factors before changing internal policy.
The most common mistake is treating pattern-3 problems with pattern-2 responses (tightening policies across all customers when only some are abusing them), or treating pattern-1 problems with pattern-3 responses (blocklisting customers when the actual problem is that your product has a defect).
The category-specific traps
A few category-level gotchas worth knowing.
Apparel bracketing inflates category benchmarks. The 20% to 30% baseline for apparel includes a meaningful amount of legitimate bracketing. A store trying to match the lower bound of that range might just be a store where bracketing has been suppressed, often by a return policy that's costing more customers than it's saving returns.
Electronics returns often cluster on opening-experience failures. A 12% return rate on a gadget might mean 12% of customers couldn't figure out the setup, not that the product is broken. Investment in onboarding docs and unboxing experience reduces this rate more than any policy change.
Furniture and freight items are under-reported in return rate. Customers often keep freight items they don't love because returning them is expensive and time-consuming. A 6% return rate on furniture isn't 6% satisfaction. It's 6% of the actually-dissatisfied customers who were motivated enough to return despite the friction. If you want a real dissatisfaction rate, look at review sentiment, not return rate.
Beauty returns are suppressed by policy, not happiness. Most beauty stores don't accept opened returns. The return rate is near zero, but the dissatisfaction rate is higher than that implies. If your beauty business has a high repeat-customer rate, great. If it doesn't, the returns-disallowed number is hiding a problem.
The takeaway
Return rate by category is a diagnostic, not a verdict. The same number means wildly different things in different categories, and within a category, the same number means different things depending on which customers and which SKUs are producing it.
Before reacting to an elevated rate, separate the three patterns. A product problem, a fit problem, or an abuse pattern. They need different responses, and applying the wrong response usually makes things worse. A product problem fixed with a policy change drives legitimate customers away while the underlying product stays broken. An abuse pattern fixed with a product change throws money at a defect that doesn't exist.
The stores that handle this well do one thing consistently. They look at return-rate data at the SKU-and-customer intersection, not the store-wide roll-up. The roll-up is the least actionable number in your dashboard, and yet it's usually the only one anybody looks at.