The appeal is obvious: one number everybody can name. The risk is equally obvious once somebody optimises it, which they will, because you asked them to.
Every term used above is defined in our experimentation glossary, and three experiments taken apart step by step sit in the case studies. Longer pieces with the workings attached are in the field notes.
Sign-ups, page views and downloads are things you receive. Weekly active teams, orders delivered on time, or reports actually read are things somebody got. Only the second kind survives being optimised.
Do this: Say your candidate metric out loud as a sentence about the customer. If it does not describe a benefit to them, pick again.
Any number attached to incentives gets optimised in the cheapest available way. Choose one where the cheapest way to move it is also the way you wanted.
Do this: Ask your team how they would move the metric with no regard for quality. If the answer is easy, the metric is wrong.
Send traffic and baseline rate. We reply with the smallest effect your test can actually detect. One email, no call.
One number alone always breaks something. Refunds, churn, support load and unsubscribes are the usual guardrails, and they exist to catch wins that were losses.
Do this: Publish two or three guardrails beside the main metric. A win that moves a guardrail the wrong way is not a win.
It is for deciding what to work on this quarter. It is not enough to run a business on, and treating it as such is how teams end up confidently wrong for a year.
Do this: Keep the north star for prioritisation. Keep a normal set of numbers for judging what happened.
This is teaching material and our own reading of standard practice, not advice for your specific site. Check anything important with your own specialist before you act on it.
Traffic, baseline rate and what you changed. We reply with the minimum detectable effect and whether the result meant anything. No cost, no call.