Multivariate testing promises to tell you which combination wins and which element caused it. Both are true only if you have the traffic, and the traffic requirement grows faster than people expect.
Every term used above is defined in our experimentation glossary, and three experiments taken apart step by step sit in the case studies. Longer pieces with the workings attached are in the field notes.
Four elements with two versions each is sixteen combinations. Each needs its own sample, so the traffic requirement is roughly sixteen times a simple A/B on the same effect size.
Do this: Multiply your variations before agreeing to the test. If the result exceeds a month of traffic, run sequential A/B tests instead.
If you only want to know which four changes are individually good, four A/B tests are cheaper and clearer. Multivariate earns its cost when you suspect two changes behave differently together.
Do this: Write down the interaction you expect before starting. If you cannot name one, you do not need this design.
Send traffic and baseline rate. We reply with the smallest effect your test can actually detect. One email, no call.
With sixteen cells and normal traffic, the top cell is frequently top by chance. Shipping it and telling a story about why is how teams learn things that are not true.
Do this: Check whether the winning cell beats the second by more than the noise. Usually it does not.
Below roughly a hundred thousand relevant visitors a month, multivariate designs mostly produce confident noise. That is not a failure of the method, it is arithmetic.
Do this: Be honest about your traffic. Sequential A/B tests on bigger changes beat an underpowered multivariate every time.
This is teaching material and our own reading of standard practice, not advice for your specific site. Check anything important with your own specialist before you act on it.
Traffic, baseline rate and what you changed. We reply with the minimum detectable effect and whether the result meant anything. No cost, no call.