Blog
September 10, 2026

How to Measure the Conversion Cost of Sizing Uncertainty

Size guide opens, fit questions to support, and size returns show what sizing doubt costs. How to join them by product type and test fit information honestly.

Aaron
Aaron
10 mins read

Sizing uncertainty is costing you conversions when fit behavior piles up on the products that underperform: shoppers open the size guide, switch sizes, ask support whether an item runs small, and later return it as too small or too large. Measure it by joining size guide events, support tickets, and return reasons by product type, then test clearer fit information and track both view-to-cart and size-related returns.

Returns desk with stacks of printed support emails, a highlighter, and folded jeans under a yellow tape measure

Every “does this run small?” email is a data point, if someone counts it. Editorial image in disposable-camera style.

Most fashion teams know sizing hurts conversion. Few can say by how much, or on which products, which makes the fix hard to justify: new measurements, reshoots, and rewritten fit notes across a catalog all cost time.

The useful news is that sizing doubt leaves a trail. Shoppers open the size guide, flip between sizes, email support, and send items back with a reason attached. Join those trails and you can put a rough number on it, one product type at a time.

This guide is written for whoever owns returns or customer experience and wants evidence before asking for budget.

Analyze size guide events

Shopify’s standard pixel events cover product views and cart additions, but not what happens between them. Nothing in the standard set records a size guide opening or a size change, so you publish those yourself.

Shopify’s Web Pixels API lets theme code publish custom events with Shopify.analytics.publish, and any custom or app pixel can subscribe to them. Whatever you pass as event data arrives in the subscriber’s customData field. Shopify asks you to prefix custom event names so they cannot collide with standard ones, and merchants cannot publish standard events.

// In a script tag in theme.liquid, when the size guide opens
Shopify.analytics.publish('my_store:size_guide_opened', {
productId: sizeGuideButton.dataset.productId,
productType: sizeGuideButton.dataset.productType,
selectedSize: sizeGuideButton.dataset.selectedSize,
});
// In a custom pixel
analytics.subscribe('my_store:size_guide_opened', (event) => {
const { productId, productType, selectedSize } = event.customData;
// Forward to your analytics tool here.
});

Publish a second event when the shopper changes the selected size, with the old and new size, and you have the three numbers that matter per product:

MetricHow to calculate itWhat it tells you
Size guide open rateProduct view sessions that opened the guide, divided by product view sessionsHow often the page left size unanswered
Size switching rateProduct view sessions with two or more sizes selected, divided by product view sessionsHow often shoppers felt between sizes
View-to-cart rateProduct view sessions with a cart addition, divided by product view sessionsThe outcome you are trying to move

One trap is easy to fall into. Shoppers who open a size guide are more interested than shoppers who do not, and also more unsure, so comparing their conversion rate with everyone else’s measures who they are, not what the guide did. Compare products with each other instead. A product with a high guide open rate and a low view-to-cart rate, next to products of the same type and price with lower fit signals, is a suspect worth a closer look.

Support inboxes hold the clearest evidence of sizing doubt, because shoppers write it in their own words. It is easy to answer these questions one by one and never count them.

Tag every pre-purchase question about size or fit with the product and product type. Six tags cover most of it:

  1. Does it run small or large?
  2. I am between sizes, which should I pick?
  3. Can you send the measurements?
  4. How long is it on someone my height?
  5. How much does it stretch?
  6. What size is the model wearing?

Then normalize. Divide each product’s size questions by its product views and multiply by 1,000, so a bestseller does not look worse just because more people see it. Size questions per 1,000 views is the number to compare across products.

After the purchase, returns carry the same signal. Shopify records a standardized return reason on each returned line item, drawn from a library of reasons that Shopify suggests based on the product’s category (ReturnReasonDefinition). Group every size-related reason (too small, too large, and any length or fit reasons your categories use) into a single sizing bucket, and keep style, color, and quality reasons separate. If you pull returns through the API, key on each reason’s handle rather than its display name, because the handle stays the same across API versions and languages.

Two more signals belong in the same sheet. Orders that contain the same product in two sizes are bracketing, which is sizing doubt paid for upfront, and the cost of bracketing goes well beyond the refund itself. And exchanges for a different size of the same item are a sizing return even when no refund happens.

For scale, the NRF projected total retail returns at $890 billion for 2024. That figure spans all of retail and every return reason, so treat it as background for the size of the problem. The breakdown of wrong-size returns in online fashion covers the reasons behind the size share.

Segment product types

Sizing doubt is not spread evenly. Jeans raise different questions from blazers, and blazers different questions from swimwear. A catalog-wide average hides the product types where the cost concentrates.

Product typeWhat shoppers askSignals to watchFit information that answers it
Jeans and trousersRise, inseam, stretchSize switching, too-small returnsRise and inseam per size, stretch in plain words, model height and size
Blazers and outerwearShoulder width, sleeve length, room to layerShoulder questions, size exchangesShoulder and sleeve per size, a note on what fits underneath
Swimwear and bodysuitsCoverage, torso lengthCoverage questions, style returnsCoverage described plainly, torso length, a back view
KnitwearDrape, stretch, lengthToo-large returnsGarment length and width, fabric weight
DressesLength on the body, bust fitLength questionsLength from shoulder to hem, model height

The detailed guides on rise and length in denim and shoulder structure in blazers and outerwear go further for those two types.

Segment by size group as well. Petite, tall, and plus shoppers ask different questions from each other, and a product can fit its regular sizes well while failing its petite range. If your product data marks size groups consistently, segmenting becomes a filter instead of a spreadsheet job. Google’s merchant listing structured data supports a size, a size group (regular, petite, plus, tall, big, or maternity), and a size system, which is a reasonable scheme to copy even if you never use it for search.

With the signals joined by product type, you can put a rough ceiling on the conversion cost. Compare products with high fit signals against products of the same type and price band with low ones:

LineHow to calculateExample (illustrative)
A. Monthly views, high-signal productsFrom product view events8,000
B. View-to-cart, high-signal productsFrom your events4.0%
C. View-to-cart, low-signal products, same type and price bandFrom your events5.5%
D. Cart additions at stakeA multiplied by (C minus B)120
E. Cart-to-order rate for the typeFrom your funnel50%
F. Average order value for the typeFrom sales reports$90
G. Monthly revenue at stakeD multiplied by E multiplied by F$5,400

The example numbers are made up. Treat line G as a ceiling, not a cost. Low-signal products may also differ in color, photography, or popularity, so the gap includes more than sizing. Add the handling cost of size returns on the same products to see the full picture, then use a test to find out how much of the gap fit information actually closes.

Test fit information honestly

Honest fit information describes the garment as it is, including when that makes it less appealing to some shoppers. “Cut close through the hip; if you are between sizes, size up” will talk some people out of buying. That is the point. A shopper who would have returned the item is better served by not ordering it, and so is your margin.

That changes what a win looks like. A fit information test can succeed with flat view-to-cart if size returns fall, and it can fail with rising view-to-cart if the new copy oversells the fit and returns rise with it. Measure both sides:

MetricControlWith fit informationDirection you want
View-to-cart rateUp, or flat with fewer size returns
Size guide open rateDown, if the page now answers the question
Size questions per 1,000 viewsDown
Orders with the same item in two sizesDown
Size-related return rate, checked 30 to 60 days laterDown

A few rules keep the test clean:

  • Pick one product type with high fit signals and enough views to read.
  • Split, do not swap. Show the new fit information on half the products, or to half the sessions if your tools allow it. A before-and-after comparison across a season change measures the season.
  • Change one kind of information at a time. Measurements, a fit note, and model details tested together cannot tell you which one worked.
  • Write the fit copy from evidence. Use your own measurements, the questions support receives, and return notes, not a supplier’s generic chart. Why size charts fail covers the most common gaps between a chart and the garment.
  • Wait for returns. Decide the result only after orders from the test window have had time to come back.

For the rest of what shoppers do between viewing a product and adding it, the PDP hesitation metrics guide covers signals beyond sizing.

Limits of this measurement

  • Theme changes break custom events. Re-test the events after every theme update, or the numbers go quiet without warning.
  • Consent reduces counts. Custom pixels can be set to require consent, so in regions with a cookie banner, shoppers who decline may be missing, and the missing share can differ by market.
  • Return reasons are self-reported. Shoppers pick the first plausible option, or “too small” when they simply changed their mind. Read the notes on a sample before trusting the buckets.
  • Support questions skew. The shoppers who write in are more engaged, and they differ by region and language. Use the count to rank products against each other, knowing it misses the shoppers who never write.
  • Observational gaps are ceilings. Only a split test tells you what fit information is worth.

Start with one product type

Pick the product type with the most size returns. This week, add two custom events, size guide opened and size changed, and tag the last month of size questions for that type. Pull size-related returns and exchanges for the same products. In a month you will have enough to rank the products, estimate the ceiling, and choose where to test fit information first.