Case studies · free

Three experiments, taken apart.

Same five steps each time, ending with what the correct plan does to the money. Figures under a company name come from its published source; ours are labelled as ours.

Any team with a testing tool · ongoing Method

Checking the test every morning

What was done

A two week A/B test is launched. Somebody opens the dashboard each morning and stops the test the day it crosses 95% significance.

What was right

The instinct is right: nobody wants to leave a losing variant running. Watching the numbers is not the crime, stopping on them is.

What went wrong

Each look is another opportunity to cross the threshold by chance. The 5% false positive rate printed on the dashboard assumes exactly one look, at a sample size fixed in advance.

What should have happened

Fix the sample size before launch and look once, or use a sequential method built for repeated looks. Either is fine. Mixing the two is not.

What the right plan does to the economics

Our own simulation, 10,000 runs of an A/A test where both variants are identical: looking once gives false positives at the expected 5%. Checking daily across fourteen days pushes it to 27%. That means better than one shipped «win» in four is nothing at all, and the engineering time behind it was spent for no reason.

Our own simulation, code and method described in our article on peeking
Mid-size retailer · ongoing Design

The checkout test that could never have worked

What was done

A retailer with 40,000 orders a year tests a redesigned checkout, hoping for a few percent improvement, and runs it for four weeks.

What was right

The hypothesis was sound and based on real friction seen in session recordings. The change itself was probably an improvement.

What went wrong

At that traffic, four weeks cannot separate a 3% lift from noise. The test came back flat, the team concluded the redesign did nothing, and shelved work that was likely positive.

What should have happened

Compute the minimum detectable effect before building anything. If it exceeds the improvement you expect, do not run the test: ship on judgement, or test a bigger change.

What the right plan does to the economics

Our own arithmetic: 40,000 orders a year is roughly 3,300 a month, so about 1,650 per variant in a month. At a 2.5% baseline conversion rate that setup detects roughly a 12% relative change, not 3%. To see 3% would take about sixteen months. Knowing that in advance saves four weeks of traffic and, more importantly, stops you throwing away a change that worked.

Our own power calculation on a modelled retailer, arithmetic shown in full