A/B Testing a Website: When It Works and When It Wastes Time

A/B testing has a reputation problem in both directions. Large sites treat it as the only valid form of evidence, which slows everything down. Small sites treat it as a growth ritual, run tests that could never have produced a signal, and act on noise. The useful question is not whether testing works, but whether it works at your traffic level, on the thing you are about to change.

The maths that decides whether you can test at all

Statistical significance needs conversions, not visitors, and the number is higher than most people expect.

Rough shape: to detect a 20 percent relative improvement on a page converting at 3 percent, you need somewhere in the region of a few thousand visitors per variant, and you need them inside a sensible window. Detecting a 5 percent improvement takes an order of magnitude more. So if your site sees 2,000 visits a month and generates 25 leads, an A/B test on the contact page would take the better part of a year to resolve, by which time your market, your ads, and your copy have all changed.

That does not mean low traffic sites cannot improve. It means their evidence comes from somewhere else: session recordings, five user interviews, the analytics you already have, and the reliable fixes that are known to work. Testing whether an accessible label helps is not a good use of a year.

Test big differences, not tints

The most common wasted test is a button colour. Even when it wins, the effect is small enough to be indistinguishable from seasonality.

Worthwhile tests change something structural: a different headline promise, a different first section, a form with three fields versus seven, a pricing page with a recommended plan versus one without, social proof placed high versus low. Big swings resolve faster because the effect size is larger, and they teach you something about your buyers rather than about your CSS.

If you are unsure what to test, look at where the drop off is steepest and test the thing directly above it. A test on a page nobody reaches is a very slow way to learn nothing, and the structure of the page above the fold is usually where the money is.

The mistakes that produce confident wrong answers

Peeking and stopping early. Watching a dashboard and declaring a winner the moment it crosses 95 percent is how most false positives are born. Decide the sample size and the end date before starting, then look when it is over.

Running for less than a full week. Weekday and weekend traffic behave differently. Anything shorter than seven days measures the day of the week, not the change. Two full weeks is a safer default.

Testing five things at once and calling it an A/B test. If the variant changes the headline, the image, and the button, a win tells you the bundle worked, not which part. That is fine if you only care about the outcome, and useless if you wanted to learn.

Measuring the wrong event. Optimising for clicks on a button that leads to a form nobody finishes will happily make your business worse while the test reports a win. Measure the outcome that pays: qualified leads, purchases, activations.

Ignoring segments. A change that helps mobile and hurts desktop can average out to “no difference” and get discarded, when the right answer was to ship it on the phone layout only.

Most changes should not be tested

This is the part that gets left out of testing advice. A large share of good design work is not a coin flip.

Fixing broken contrast, making a form usable on a phone, writing a headline that says what you sell, or making the site fast are not hypotheses. They are repairs, and testing them delays the fix while producing an answer everybody could predict. Ship the known good, then test the genuine unknowns: positioning, offer, proof, and price.

That is also why a redesign is a bad candidate for a clean test. Everything changes at once, so you learn whether the new site beats the old one overall, which is worth knowing but is a before and after measurement, not an experiment.

What to do when you cannot test

Run the cheaper evidence loops. Watch ten session recordings of people on the page in question and count where they hesitate. Ask five customers what nearly stopped them from buying, since a week of light research beats a quarter of guessing. Compare the copy against what your sales calls actually repeat, and write the page from that.

Then measure before and after honestly, with the same window, the same traffic sources, and an eye on the fact that a single good month is not proof.

We build measurement in from the start on the sites we ship, because the point of a redesign is a number that moved, not a screenshot.

Want help deciding what is worth testing on your site? hello@beconfidency.agency.

If you want the whole thing designed, built, and measured by one team, that is exactly what our web design service is for.

Next project

Have an ideaworth raising?