sample size4 min read

How much traffic do you need to A/B test an online store?

Some version of this question shows up every week on r/shopify, r/CRO, and r/ecommerce: "My store does about 10,000 visitors a month. Everyone says test everything, but my tests never reach significance. Am I doing something wrong?" No. The math was against you before you started.

by Sebastian Zijlstra · more essays

Nobody selling testing software has much reason to answer this honestly. The real answer: below a certain amount of traffic, an A/B test can't do its one job, telling a real improvement apart from random luck. The good news is you can find your store's cutoff in about two minutes. Two words carry the rest.

the two words that matter Significance asks: is this difference real, or did one version just get lucky this week? Power asks the flip side: if B really is better, will your test even notice? A low-power test is a smoke detector with a flat battery. The fire is real; the alarm stays quiet. Small stores nearly always run low-power tests, and that one fact explains everything below.

01Run the arithmetic before the test

How long a test runs depends on four things: your baseline conversion rate (the share of visitors who buy today), the lift you want to catch (a "10% lift" means a 2.5% rate climbing to about 2.75%, not adding ten percentage points), and how careful you want to be about two mistakes, calling a dud a winner and missing a real one (the usual defaults, 95% confidence and 80% power). Plug in a typical small store, 2.5% conversion and 10,000 visitors a month split between two versions, and here is what comes out.

Bar chart showing weeks until an A/B test can call a winner for a store with 10,000 visitors a month and 2.5% baseline conversion. A 50 percent lift takes 2.6 weeks, a 30 percent lift 6.8 weeks, a 20 percent lift 14.6 weeks, and a 10 percent lift 55.6 weeks, far past the 8-week practical limit.
Weeks to reach a verdict at 10,000 visitors/month, 2.5% baseline CVR, 95% confidence, 80% power. Computed with the standard two-proportion sample size formula.

To reliably catch a 10% lift, a good result for a typical tweak, you need about 128,000 visitors. At 10,000 a month that is 56 weeks, over a year frozen on one test. A 20% lift still needs about 15 weeks; only past 30% does the timeline drop under eight weeks, which is roughly the longest a test should run before cookies clear and seasons shift. Flip it around and it is starker: catching a 10% winner inside a single month needs about 140,000 visitors a month. Reliably spotting small wins is a big-store luxury, and it is on the box of no testing tool.

02Running it anyway backfires

At low traffic the temptation is to run it and hope. But when a weak test does scrape past "significant," it is usually luck that pushed it over, so the win it reports comes out inflated, often two or three times. That is the winner's curse: the +14% you would announce is probably a +5% in a costume, and shipping it turns a coin flip into false confidence that gets budget. If you have a borderline winner in front of you now, Reality Check shrinks it to an honest number before you present it.

but doesn't Bayesian fix this? No. Bayesian and frequentist are just two languages for the same data, and the data is where the information lives. Dynamic Yield and VWO use Bayesian because it reads like a business ("72% chance B wins") and can start from a skeptical assumption that curbs the winner's curse. What it cannot do is invent evidence a thin sample does not hold. A 66% "probability to be best" is a coin toss in a nicer outfit. What decides whether a small store can test is not the method. It is sample size.

03What actually works at low traffic

The line in that chart is not fixed. None of these need more visitors:

04Two minutes before you start

All of this only pays off if it is decided before the test runs. Lockbox sizes the test and locks your metric and stopping rule up front, so there is nothing to fudge later. Ship a winner, then log it in the Program Ledger against your real monthly numbers to see whether the wins reach revenue. Both are free and run in your browser. Two minutes of math beats three months of arguing.

the stack

Seven free tools for honest ecommerce experimentation: platform validation, pre-registration & sample size, survival analysis, winner deflation, integrity receipts, the program ledger, and subscription valuation. All of it runs in your browser. Explore the stack →

SZ

Sebastian Zijlstra

I build tools for ecommerce experimentation that hold up under scrutiny, and write about where A/B testing quietly goes wrong. Everything on this site runs in your browser, free. See the stack, or connect on LinkedIn.