Christian Denehy
CRO
9 mins

How to run ecommerce A/B tests without fooling yourself

Testing starts with a decision worth making

An A/B test is useful when two credible approaches could produce different customer behaviour and the business has enough traffic to compare them. It is less useful when the answer is obvious, the experience is broken or nobody knows what decision the result will support.

Start with a customer problem rather than a preferred design. Analytics may show a drop between product views and add to basket, while research suggests customers are uncertain about delivery. That evidence creates a focused area to investigate.

Write the decision before the test. If the variation performs better, what will you change? If it performs worse or remains inconclusive, what happens next? This prevents teams from running experiments that create interesting charts but no action.

Do not test everything simply because software makes it possible. Each experiment consumes traffic, development effort and attention. Prioritise questions connected to meaningful journeys and commercial outcomes.

Write a causal hypothesis

A useful hypothesis explains the audience, proposed change, expected behaviour and reason. For example, showing delivery timing beside the purchase action may increase completed orders because customers can assess the full commitment before entering checkout.

This is stronger than saying a new design will improve conversion. It makes the assumption visible and helps the team choose supporting metrics. It also gives you something to learn when the result does not match the expectation.

Change as few meaningful variables as possible. If a variation introduces new copy, layout, imagery and pricing presentation together, you may discover that the bundle performs differently without knowing why. Larger redesign tests are sometimes necessary, but interpretation becomes harder.

Record the evidence that led to the hypothesis. Over time, this creates an experiment history that shows which customer problems recur and which assumptions have already been tested.

Choose one outcome and sensible guardrails

Select a primary metric connected to the decision. It might be completed purchases, add-to-basket rate or progression to checkout. Avoid declaring success based on whichever metric happens to look positive after the test.

Add guardrails for unintended effects. A variation may increase conversion while reducing average order value, increasing returns or causing more support contacts. Guardrails help the team see the broader commercial impact.

Define the audience before launch. New visitors, returning customers, mobile users and high-value segments may behave differently. Changing segment definitions after seeing results creates misleading stories from ordinary variation.

Check that tracking works in both experiences. Missing events, duplicate analytics or inconsistent consent behaviour can invalidate an experiment. Run test orders and review event volumes before trusting the dashboard.

Respect uncertainty and test duration

Agree the sample requirement and minimum duration in advance. Tests need enough observations to distinguish a real difference from random movement, and they should usually cover complete trading cycles rather than stopping after an unusually strong day.

Do not peek repeatedly and end the experiment as soon as one version moves ahead. Normal variance can create temporary winners. Following the agreed method protects the team from making confident decisions based on noise.

Watch for outside influences. Promotions, stock changes, acquisition campaigns and technical incidents can affect customer behaviour during the test. Record them and decide whether the result remains interpretable.

An inconclusive result is still information. It may show that the change is too subtle, the audience is too small or the assumed problem is not commercially important. Do not force a winner when the evidence does not support one.

Turn results into organisational learning

Review the result against the original hypothesis, not only the headline number. Ask whether the expected behaviour changed, whether guardrails remained healthy and what the outcome says about customer needs.

Document the setup, dates, audience, screenshots, metrics and conclusion. A searchable experiment record prevents repeated tests and helps new team members understand why the store works as it does.

Roll out winners carefully and confirm performance after release. Experiment code and permanent production code may behave differently, particularly across apps, markets and devices. A successful test is not complete until the change works reliably.

The aim is not to collect winning variations. It is to build a more accurate understanding of how customers shop. A disciplined sequence of honest tests produces better decisions than a high volume of experiments designed to prove existing opinions.

Read other insights