Glossary · 22 terms · free

The words that decide whether a test meant anything.

One plain sentence each, then the part that changes your result.

Statistical power

Testing

The chance your test detects a real effect of a given size, if one exists.

Most tests are run at power nobody calculated. A test with 30% power finds a real 5% lift in one run out of three, and the other two look like «no difference».

Minimum detectable effect

Testing

The smallest lift your test can reliably distinguish from noise, given traffic and duration.

Work it out before you start. If it comes to 18% and your realistic improvement is 4%, the test cannot answer the question at any duration.

Sample size

Testing

How many visitors per variant the test needs before you look at the result.

Fixed before the test starts, not adjusted when it looks promising. Adjusting it after seeing data is what breaks the maths.

Peeking

Testing

Checking a running test and stopping when it looks significant.

Every look is another chance to cross the line by accident. Checking daily on a two week test pushes the false positive rate far above the 5% your tool reports.

Sequential testing

Testing

Methods designed for looking at results as they arrive, with the threshold adjusted for it.

The honest way to peek. If your tool does not use one, the only safe rule is to look once, at the end.

Novelty effect

Testing

A short lift caused by the change being new rather than better.

It fades. Ending a test in the first days captures the novelty and none of the reality, which is why winners so often stop winning.

Segment slicing

Analysis

Splitting a flat result by device, source or country until something looks significant.

Twenty slices means roughly one false winner by chance. Decide your segments before the test, or treat what you find as a question rather than an answer.

Revenue per visitor

Analysis

Total revenue divided by all visitors, converters and not.

Conversion rate can rise while this falls, if the change pushes people towards cheaper items. It is the number that catches that.

Average order value

Analysis

Mean value of an order over a period.

Sensitive to a handful of large orders. A test that moves it by 8% has usually moved one customer, not the population.

Guardrail metric

Analysis

A number you watch to make sure a win did not break something else.

Returns, refunds, support tickets and unsubscribes. A checkout change that lifts orders and doubles refunds is a loss the primary metric cannot see.

Holdout group

Design

A slice of traffic deliberately excluded from a change, kept for comparison.

The only way to measure something that runs everywhere. Uncomfortable to leave money on the table, and it is the price of knowing.

Interaction effect

Design

Two changes running at once whose combination behaves differently from either alone.

It is why running four tests on one page at once produces four results and no knowledge.

Flicker

Implementation

The original version showing briefly before the test variant loads.

It biases the result against the variant on slow connections, and those visitors are usually the ones you most wanted to help.

Bucketing

Implementation

How visitors are assigned to variants and kept there across sessions.

If somebody sees version A on mobile and B on desktop, your data mixes two experiences into one row.

Sample ratio mismatch

Implementation

The split arriving unequal when it was set to 50/50.

A 52/48 split on large numbers means something is broken in assignment. Stop and fix it: the result is not trustworthy.

Loss recording

Practice

Writing down tests that produced no difference and keeping them.

Teams that drop losses repeat them every eighteen months. The archive of what did not work is worth more than the list of wins.

Test velocity

Practice

How many properly powered tests you complete per quarter.

Counting launched tests flatters you. Counting completed and powered ones tells you whether the programme exists.

Hypothesis

Practice

A statement of what you expect to change, for whom, and why.

Without the why, a win teaches you nothing you can carry into the next test.

Prioritisation model

Practice

A ranking of what to test next, scored rather than argued.

Any consistent model beats seniority. The value is in ending the meeting, not in the scores being exact.

Friction point

Qualitative

A step where people visibly hesitate, retry or leave.

Session recordings and form analytics find these far faster than testing does. Test the fix, not the discovery.

Micro conversion

Qualitative

A smaller step on the way to the real one: add to basket, start checkout, view sizing.

Useful when the real conversion is too rare to test on. Dangerous when you start optimising it instead of the sale.

Qualitative sample

Qualitative

The small number of people you watch or interview.

Five sessions find most usability problems and prove nothing about magnitude. Use it to generate hypotheses, never to size an effect.