When Does an A/B Test Make Sense — the Traffic Threshold

Three days ago you tested two homepage headlines. Today variant B is ahead of A — or so it looks. Pleased, you keep B and close the test. A week later, checking the numbers again, the gap has vanished, and A is even back in front. The test didn't get it wrong — you read it too early.
An A/B test is a strong tool, but it only works under two conditions: enough traffic, and the patience to wait for the result. Miss either one and the tool doesn't give you an answer — it just hands random noise back to you dressed as an official number. Here's why that happens, and what to do if your traffic is thin.
Why looking early kills the test itself
Checking the result every day and stopping the moment "B is winning" is a natural instinct. But the math behind statistical significance assumes a sample size fixed in advance. Per Evan Miller's widely cited piece on the subject, stopping a test the moment it looks "significant" after every new observation inflates a claimed 5% false-positive rate to an actual 26.1%. So a "winner at 95% confidence" declared this way is, in the real world, a coin flip you're calling a trend nearly a quarter of the time.
The author puts it bluntly: no web experiment should run without a sample size fixed ahead of time, and that commitment has to be kept with "near-religious discipline." The answer to "after how many visitors do we look" has to be written down before the test starts, not decided halfway through it.
What a traffic threshold actually means
"Enough traffic" isn't a vague phrase — it's a function of three specific things: your current conversion rate, the size of the difference you're hoping to detect, and how many visitors arrive per day. Per Optimizely's own explanation of statistical significance, "larger samples generally provide more reliable results" and "for website tests, more traffic means quicker, more accurate results." There's no universal number like "a thousand visitors is enough," because the answer depends on each site's own conversion rate.
Why a low-traffic "winner" is usually false
A site with 300 monthly visitors runs a week-long test on a few dozen clicks total. At that sample size, a statistically "significant" difference is usually just the natural wobble of small numbers — like flipping a coin twice and getting heads both times. If the same gap holds up across thousands of observations on a larger site, it's more likely to be real.
How to ask the question properly
Instead of "how many visitors do I need," ask "how many weeks will it take to close this test at my current conversion rate, my expected difference, and my daily traffic." The answer often comes out to several months — and for most small businesses, that isn't a realistic time budget.
What statistical significance says, and what it doesn't
Statistical significance doesn't tell you "B will win forever" — per Optimizely's definition, it measures "the likelihood that the difference in conversion rates between a given variation and the baseline is not due to random chance." That's a probability estimate for the period already observed, not a guarantee for the future. The classical method requires setting a minimum detectable effect and a sample size in advance, then waiting without peeking — skip that, and the reported significance level is simply wrong, because checking the result repeatedly always raises the false-positive rate.
A test closed early doesn't answer anything — it just gives randomness an official number.
The low-traffic alternative: sequential improvement
If your traffic isn't enough to support a test, that doesn't mean the site can't get better — sequential improvement takes the place of the split test. Apply changes grounded in known UX principles (a clear call to action, a short form, a fast load) and watch the trend over a longer stretch instead of hunting for "proof" through a split test. Heatmaps, session recordings and direct customer questions give a more reliable direction at low traffic than statistics do, because they answer "where exactly do people get stuck," not "which number won."
For a services site with 500 monthly visitors, for instance, a two-month headline test can be replaced by a day spent watching session recordings to see exactly which form field people stall on — and you can roll that fix out to every visitor at once, instead of artificially splitting anyone into a test group.
Other mistakes worth knowing
- Seasonality distorts the test. A test running during a big sale, a holiday, or a one-off ad spike measures that week's unusual behaviour, not an ordinary day. Don't act on a result without re-checking it in a quiet, typical period.
- The novelty effect looks like a win. Visitors react to a new variant simply because it looks different — that creates an artificial "winner" in the first few days, and the gap often disappears on its own a few weeks later. Not closing the test early is exactly what neutralises this.
- Parallel tests on the same page interfere with each other. Test the headline and the button colour at once, and a visitor sees both changes together — you'll never be able to tell which one moved the result. Run one test per page at a time.
- Expecting a big result from a small change. Assuming that swapping one word in a headline will sharply lift conversion isn't realistic — small edits usually produce small, hard-to-measure effects. If you want a large result, test a large change (the whole page structure, say), not a minor detail.
Four checks before you start a test
- The sample size is written down first — how many visitors or weeks the test runs before closing is decided before launch, not chosen once you see the result.
- Only one variable changes — if the headline, the button and the image all change at once, you will never know which one did the work.
- There is one success metric — whether it's the click, the sale, or both is settled before the test starts.
- No one looks until it closes — checking is tempting, but acting on the interim number is off-limits, which is exactly why the end date is fixed in advance.
If your traffic is enough, setting the test up correctly the first time is far cheaper than acting on a wrong conclusion for months. In our conversion A/B-testing service, we agree the sample size, the single variable and the stopping rule with you in writing before the test starts — so no one looks early and acts on it.
Photo by konat umut budak · Pexels