Statistical testing: sample size, lift and full factorial designs

Most A/B tests I am shown were decided before the maths was done: a tool split the traffic, a dashboard turned green on day nine, and the team shipped. Then the lift quietly vanished. This chapter is about the discipline that stops that happening. I start from the lift a change has to produce to be worth shipping, and derive the sample size from it, with a table you can use for your own baselines. I then run a two by two by two full factorial on the Werkbank pricing page, six thousand visitors across headline, price anchor and proof, and show how the anchor hurts on its own and helps a lot next to proof, an interaction no one at a time test could ever see. Along the way: the two proportion z test, why the confidence interval on the lift matters more than the p value, how peeking at the dashboard inflates your false positives, and when a test is simply not worth running. This is part 20 of 21 of the Marketing Analytics series.

Everyone wants to test, few test well

A founder of Werkbank, the Rails SaaS for craft businesses I use as an example throughout this series, showed me a slide last spring. "New pricing page headline: plus 31 percent trial signups." The test had run nine days: 41 trials from about 1,050 visitors on the old page, 54 from roughly the same number on the new one. The team had already rewritten every headline on the site to match. Two months later the trial rate was exactly where it had been, and nobody could say why.

Nothing dishonest happened. The team had a tool that splits traffic and a dashboard that turns green, and they stopped when it turned green. What they did not have was a number written down before the start: how big a lift they needed to see, how many visitors it would take, and what they would do if the answer was "nothing". That is the whole of this chapter. The maths is not hard. The discipline is.

Where this sits in the series

This is chapter 20 of 21 and the first of part five, testing and big data. Part four closed with The customer loyalty journey: from segments to experiences, where I argued that a loyalty programme is a series of hypotheses about what a segment will do next. Testing is how you find out whether those hypotheses were right, and it is what turns every earlier chapter, from elasticity to media mix, from a model into a decision. The series ends next with Big data for marketers: what actually changes and what does not.

Lift first, then sample size

Start with what you actually care about. Not "is B different from A" but "is B better than A by enough to matter". The quantity is the lift:

lift=pB−pApA\text{lift} = \frac{p_B - p_A}{p_A}

Here pAp_A is the conversion rate of the control (the page you have today) and pBp_B is the conversion rate of the variant. Lift is relative: a move from 4.0 to 4.6 percent is a 15 percent lift, though the absolute gap is only 0.6 percentage points. Keep both in your head, because the sample size depends on the absolute gap while the business case is argued in relative terms.

The lift you need to detect is a business number, not a statistical one. For Werkbank: about 2,000 visitors a month reach the pricing page, 4 percent start a trial, roughly a quarter of trials become paying customers at 79 euro a month for around 22 months. A 15 percent lift in trials is three extra trials a month, less than one extra customer, worth somewhere near 1,400 euro of lifetime revenue per month of running the improved page. That is the smallest effect worth a redesign and a test; below it, I do not care whether it is "significant". This is the minimum detectable effect, and choosing it is the most important decision in the design of a test. I have written before about why statistical significance and error margins matter in growth marketing; here I want to go one step earlier, to the moment before the test exists.

The sample size equation

Two mistakes are possible when you read a test. You can declare a winner when there is none (a false positive, probability α\alpha), and you can miss a real winner (a false negative, probability β\beta). The conventions are α=0.05\alpha = 0.05 two sided and β=0.20\beta = 0.20, so a power of 80 percent. They are conventions, not laws.

For two conversion rates the sample size per variant is:

n=(z1−α/2 2 pˉ (1−pˉ)  +  z1−β pA(1−pA)+pB(1−pB))2(pB−pA)2n = \frac{\left( z_{1-\alpha/2}\,\sqrt{2\,\bar p\,(1-\bar p)} \;+\; z_{1-\beta}\,\sqrt{p_A(1-p_A) + p_B(1-p_B)} \right)^2}{(p_B - p_A)^2}

Reading the symbols: nn is the number of visitors you need in each arm. pAp_A is the baseline rate, pBp_B is the rate you hope to detect, so pB=pA(1+lift)p_B = p_A (1 + \text{lift}), and pˉ\bar p is their average. z1−α/2z_{1-\alpha/2} is the standard normal quantile for your false positive tolerance, 1.96 for a two sided 5 percent test. z1−βz_{1-\beta} is the quantile for your power, 0.84 for 80 percent. The denominator is the squared absolute difference, and that square is the reason small lifts on small baselines are so brutally expensive: halve the gap you want to detect and you need four times the visitors.

Plug in the Werkbank case. Baseline 4 percent, target 4.6 percent, gap 0.006. The first root gives 1.96×0.287=0.5621.96 \times 0.287 = 0.562, the second 0.84×0.287=0.2410.84 \times 0.287 = 0.241, the sum squared is 0.646, divided by 0.0062=0.0000360.006^2 = 0.000036 gives about 17,900 visitors per arm. Call it 18,000 each, 36,000 in total, a year and a half at 2,000 pricing page visitors a month. The nine day test with 1,050 per arm could only have reliably detected a lift of roughly 84 percent. It was never a test of a headline. It was a coin flip with a dashboard.

A sample size table you can actually use

Here is the equation evaluated for the baselines and lifts I meet most often. Visitors per variant, two sided 5 percent, 80 percent power. The numbers are exact from the formula and the scenarios around them are illustrative.

Baseline rate5% lift10% lift15% lift25% lift50% lift
1%637,000163,10074,20027,9007,750
2%315,20080,70036,70013,8003,830
4%154,30039,50017,9006,7501,860
8%73,90018,9008,5703,210880
15%36,3009,2604,1901,570420

Read it row by row. A 1 percent baseline is a checkout on a cold traffic landing page; a 5 percent lift there needs more than 600,000 visitors per arm, which for a small business means never. At 4 percent, a typical trial or lead form rate, the 15 percent cell is the Werkbank answer, 17,900 each. At 15 percent, an add to basket rate or a warm list's email click rate, a 25 percent lift is detectable with 1,570 per arm, a fortnight for most shops. The lesson: test where the baseline is high and the lever is big, further up the funnel and on bolder changes. Micro changes to low rate steps are not testable by SMEs, and no tool will change that.

Required sample size per variant by the lift you want to detect, 4 percent baseline, two sided 5 percent, 80 percent power.

A/B tests versus full factorial designs

An A/B test changes one thing. It is clean and slow, and its deeper problem is that a pricing page is not one thing. Headline, price anchor and proof all pull on the same visitor at the same time. Testing them one after another takes three times as long and tells you nothing about how they combine. I have made that argument in full in why testing one thing at a time is costing you money, so here I will only show the alternative.

A full factorial design tests every combination of every factor. Three factors at two levels each is a two by two by two, eight cells, and every visitor is randomised into one of them. The trick that makes this affordable is that each main effect is estimated from the whole sample: to measure the headline you compare the four cells with headline one against the four with headline two, 3,000 versus 3,000 in a 6,000 visitor test, exactly as in a simple A/B test on the headline alone. Three tests for the price of one, and the interactions come free.

The main effect of a factor is the average difference across all the other factors' levels:

MEproof=pˉproof−pˉno proof\text{ME}_{\text{proof}} = \bar p_{\text{proof}} - \bar p_{\text{no proof}}

where each pˉ\bar p is the conversion rate pooled over the four cells that share that level. An interaction is the difference of differences: does the anchor help more when proof is present than when it is absent?

INTanchor×proof=(pˉA,P−pˉN,P)−(pˉA,N−pˉN,N)\text{INT}_{\text{anchor} \times \text{proof}} = \left(\bar p_{\text{A,P}} - \bar p_{\text{N,P}}\right) - \left(\bar p_{\text{A,N}} - \bar p_{\text{N,N}}\right)

with A and N for anchor and no anchor, P and N in the second position for proof and no proof, each mean pooled over both headlines.

The Werkbank factorial: eight cells, 6,000 visitors

Werkbank ran the pricing page as a two by two by two over a quarter. Headline: the current feature headline ("Scheduling, quotes and invoices in one place") versus an outcome headline ("Get your evenings back"). Price anchor: the monthly price alone versus the price next to the crossed out cost of the three tools it replaces. Proof: nothing versus a strip of customer logos and one named quote from a joiner in Wels. The numbers below are illustrative, 750 visitors per cell.

CellHeadlineAnchorProofVisitorsTrialsRate
FNNFeatureNoneNone750304.00%
FNPFeatureNoneLogos750364.80%
FANFeatureAnchorNone750253.33%
FAPFeatureAnchorLogos750506.67%
ONNOutcomeNoneNone750344.53%
ONPOutcomeNoneLogos750405.33%
OANOutcomeAnchorNone750293.87%
OAPOutcomeAnchorLogos750547.20%

Line by line. FNN is the page as it was, 4.0 percent, the baseline we used for the sample size. FNP adds proof and moves to 4.8. FAN adds the anchor without proof and drops to 3.33: a crossed out competitor price with nothing behind it reads as a sales trick. FAP, anchor plus proof, jumps to 6.67, far more than the two single changes add up to. The outcome headline rows repeat the shape about half a point higher, ending with the best cell, OAP, at 7.20 percent.

Trial start rate for each of the eight cells. F and O are the feature and outcome headlines, N and A no anchor and anchor, N and P no proof and proof.

Now the main effects, each from 3,000 versus 3,000 visitors:

FactorLevel 1 rateLevel 2 rateDifferencezp value95% CI on the difference
Headline (feature vs outcome)4.70%5.23%+0.53 pts0.950.34minus 0.57 to +1.63 pts
Anchor (none vs anchor)4.67%5.27%+0.60 pts1.070.28minus 0.50 to +1.70 pts
Proof (none vs logos)3.93%6.00%+2.07 pts3.680.0002+0.97 to +3.16 pts

Proof is the only clear main effect: a 2.07 point gain, a relative lift of 52 percent, and the interval stays well clear of zero. The headline, the thing the team was so excited about, shows half a point with an interval from minus 0.57 to plus 1.63. That is not "the headline does nothing". It is "6,000 visitors cannot tell you what the headline does", a very different and more honest reading. The anchor's main effect is likewise a shrug at plus 0.6 points, and that is where the factorial earns its keep.

Reading the interaction

Look at the anchor on its own terms. Without proof, adding it moves the rate from 4.27 to 3.60 percent, minus 0.67 points. With proof, it moves the rate from 5.07 to 6.93, plus 1.87 points. The difference of those differences is 2.53 points with a standard error of about 1.12, so z=2.26z = 2.26 and p=0.024p = 0.024. The anchor is not a weak lever. It points in opposite directions depending on whether the page has earned the right to make a price comparison, and averaged over both conditions the halves cancel into a meaningless plus 0.6.

Had Werkbank tested the anchor alone on the existing page, they would have measured a drop, concluded that anchoring does not work for tradespeople and never tried it again. Had they tested proof alone they would have shipped it and left the anchor's contribution on the table. The interaction is the finding, and only a design that varies both factors at once can see it. My caveat: interactions are estimated with half the precision of main effects, because each side of the difference is a 1,500 versus 1,500 comparison. At p=0.024p = 0.024 this one is real enough to act on and not so certain that I would build a pricing philosophy on it. I would ship OAP and rerun anchor versus no anchor on the new page with proof, as a confirmation.

Reading results: the interval, not the verdict

The two proportion z test is the standard way to compare two rates:

z=p^B−p^Ap^ (1−p^)(1nA+1nB),p^=xA+xBnA+nBz = \frac{\hat p_B - \hat p_A}{\sqrt{\hat p\,(1-\hat p)\left(\dfrac{1}{n_A} + \dfrac{1}{n_B}\right)}}, \qquad \hat p = \frac{x_A + x_B}{n_A + n_B}

Here p^A\hat p_A and p^B\hat p_B are the observed rates, xAx_A and xBx_B the conversion counts, nAn_A and nBn_B the visitors per arm, and p^\hat p the pooled rate under the assumption that nothing differs. A zz beyond 1.96 either way is "significant at 5 percent", the least interesting thing the test can tell you.

The interesting thing is the confidence interval on the difference:

(p^B−p^A)  ±  z1−α/2 p^A(1−p^A)nA+p^B(1−p^B)nB(\hat p_B - \hat p_A) \;\pm\; z_{1-\alpha/2}\,\sqrt{\frac{\hat p_A(1-\hat p_A)}{n_A} + \frac{\hat p_B(1-\hat p_B)}{n_B}}

For proof that is 2.07 points plus or minus 1.10, so 0.97 to 3.16 points, or in relative terms a lift somewhere between 25 and 81 percent. That width is the true state of your knowledge. Report "plus 52 percent" and the finance director will plan on it; report "between 25 and 81 percent" and she will plan on the low end, which is correct. Every decision memo I write carries the interval, and where it includes zero I say so in the first line.

Running a test properly

The protocol is short and I have never seen a team regret following it.

The test protocol I use. The decision about whether to test at all is made before a single visitor is split.

Three details matter more than the diagram suggests. Randomise by visitor, not by day or week, because Tuesdays convert differently from Saturdays. Check the split on day one: if 750 per cell was the plan and one cell has 640, the assignment is broken and every downstream number is suspect. And pick one primary metric before you start: trials started, not trials plus pricing clicks plus scroll depth. Secondary metrics explain why; they never declare a winner.

Sequential peeking

The most common way honest teams fool themselves is by watching the dashboard. The statistic wanders as data comes in, and if you look every day for a month and stop the first day it crosses 1.96, you are not running one test at 5 percent, you are running thirty and taking the best. The false positive rate is several times the nominal 5 percent, and you can confirm it in ten minutes by simulating an A/A test where both arms share the same true rate:

import numpy as np
rng = np.random.default_rng(1)
p, n_day, days, hits = 0.04, 200, 30, 0
for _ in range(2000):
    a = rng.binomial(1, p, n_day * days).cumsum()
    b = rng.binomial(1, p, n_day * days).cumsum()
    for d in range(1, days + 1):
        n = d * n_day
        pa, pb = a[n-1] / n, b[n-1] / n
        pp = (pa + pb) / 2
        se = np.sqrt(pp * (1 - pp) * 2 / n)
        if se > 0 and abs(pb - pa) / se > 1.96:
            hits += 1
            break
print(hits / 2000)   # the share of A/A tests that "found a winner"

Run it and you will see a number a long way above 0.05. The cure is either to fix the sample size in advance and read once at the end, which is what I do for most SME tests, or to use a proper sequential method that spends the error budget across looks. The latter is fine if your tool implements one and you know which. The tool that turns green on day nine is not doing that.

When a test is not worth running

I decline tests more often than I run them. A test is not worth running when the required sample size exceeds what the page will see in about a quarter, because by then the season, the ad mix and the product have all changed under it. Not when the change is obviously right and reversible: fixing a broken form field needs no control arm. Not when the value of the information is smaller than the cost of delaying the change, as with small copy edits on low traffic pages. And not when the organisation has already decided what it will do regardless of the result. In that last case I say so and we spend the effort on something that can still change a decision.

The alternative to a test is not guessing. It is the earlier chapters of this series: a regression on historical data, a survey, a well read funnel. Tests are the gold standard for causation and also the most expensive instrument in the box. My guide to digital marketing testing methods walks through the less expensive instruments and when each one is enough.

Running it yourself

You need three things: a randomisation mechanism that assigns a visitor once and remembers it (a cookie or a hashed user id), an event that records the primary conversion against that assignment, and a way to count. Any tool works, and a spreadsheet works for the reading.

The sample size in a few lines; the z test is the formula above, one line each for p^\hat p and zz:

from math import sqrt
from statistics import NormalDist
N = NormalDist()

def sample_size(p_a, lift, alpha=0.05, power=0.8):
    p_b = p_a * (1 + lift); p_bar = (p_a + p_b) / 2
    z_a, z_b = N.inv_cdf(1 - alpha / 2), N.inv_cdf(power)
    num = z_a * sqrt(2 * p_bar * (1 - p_bar)) + z_b * sqrt(p_a * (1 - p_a) + p_b * (1 - p_b))
    return round(num ** 2 / (p_b - p_a) ** 2)

print(sample_size(0.04, 0.15))   # 17943

Before you trust a result, run four checks. The split: are the cell sizes within a few percent of what randomisation should give? The A/A check: do two arms of the control differ by less than the noise predicts? The segments: does the effect hold on mobile and desktop, new and returning, or is it driven by one burst of bot traffic on one day? The calendar: did a newsletter, a price change or a public holiday in Austria but not the UK hit the arms unequally? Only after those do I look at the interval.

How long does it take? Designing a factorial test with a client is an afternoon. Running it is the sample size divided by daily traffic, and the honest number is often longer than anyone likes. Reading it is an hour and a memo.

Pitfalls

Testing to the vanity precision. Teams ask "how many visitors do I need for the result to be accurate" and someone answers with a margin of error on a single rate. Wrong question. Precision on a rate is not power to detect a difference, and the difference you need is a business number you have to choose. If you have not chosen it, you have not designed a test.

Stopping on green. Covered above, and still the most common failure. If your tool lets you stop early without a sequential correction, decide the end date and hide the dashboard.

Too many cells. Four factors at three levels is 81 cells, and 6,000 visitors gives you 74 per cell. Main effects still pool, but the interactions become noise. Two or three factors at two levels is the sweet spot for SME traffic. Fractional factorial designs trade away the higher order interactions to save cells, a fair trade when you know those interactions are unlikely.

The novelty effect. A new headline does better in week one because returning visitors notice the change. Run long enough to cover full weekly cycles and look at new visitors separately. If the effect shrinks week on week, you have measured curiosity, not preference.

The wrong metric. Trial starts went up 52 percent; did paying customers? Proof may attract trials from people who were never going to buy. Track the primary metric to the money whenever the lag allows, and otherwise check that the downstream rates did not move the other way.

Mistaking a losing test for a wasted test. FAN, anchor without proof, lost. That is one of the most useful things Werkbank learned all year, because it explains a failed experiment from two years earlier and warns every future designer. A well designed test that returns "worse" has paid for itself.

How I do this for clients

The deliverable is a test programme, not a single test. It starts with the free workshop, where we list every change anyone has been wanting to make to the key page, put a lift and a sample size against each one, and cross out the ones the traffic cannot support. Most lists shrink by two thirds in that hour, and that is the point.

Then two weeks of real work on me, before any commitment. I need access to the analytics and the page, the historical baseline for the primary metric by day and by segment, and whoever owns the design of the variants. In those two weeks I write the test plan as a one page document: hypothesis with a number, primary metric, factors and levels, sample size, end date, the stopping rule and what we will do with each possible result. I set up the randomisation and the event, run an A/A check, and start the first factorial.

At the end of a test you get a decision memo with the lift, its interval and a risk statement in the first paragraph, the full cell table behind it, and a recommendation you can act on that week. Over a programme you get a running register of what was tested, what it did and what it taught, soon the most valuable document in your marketing folder. You own all of it, code, data and register.

This sits inside my data science work and, where the tests are on ad landing pages, inside performance marketing. Pricing is a fixed monthly retainer or, for growth work, a commission on the lift we actually measure; the pricing page has the shape of it. What I will not do is report a winner without an interval.

Questions to ask before your next test

  • What lift did we decide, before the test, was the smallest worth detecting, and where is that written down?
  • How many visitors per arm does that lift require at our baseline, and when will we reach it?
  • Which one metric decides the test, and who agreed to it in advance?
  • Is the assignment random by visitor, and did we check the split on day one?
  • Are we reading the result once at the planned end, or has someone been watching the dashboard?
  • What is the confidence interval on the lift, and does the decision still hold at its lower end?
  • Could the factors we are testing interact, and if so why are we testing them one at a time?
  • What will we do if the result is "no difference"?

This series is inspired by Mike Grigsby's Marketing Analytics (Kogan Page). The explanations, examples and numbers here are my own.

If you have a test running right now, or one you are about to start, send me the baseline rate, the monthly traffic on that page and the change you want to make. I will tell you honestly whether the traffic can support it, how long it would take and what lift it could actually detect. That conversation is what the free workshop is for, and it costs you nothing. Two weeks of real work on me follow before you commit to anything. Bring your data or your question and let us design a test worth running.

1%of every invoice goes to a UK charity you pick.

A donation, never sponsorship. You choose the cause at onboarding.

The story behind the pledge →

Stay ahead of your competition.

The latest innovative products and services, straight to your inbox before your competitors hear about them.

Get up to 5% off your first six months: 1% per topic you pick, the full 5% when you take everything. Limited offer · ends 31 December 2026.

New clients only. Terms apply.

* Up to 5% off your first six monthly invoices, new clients only. Full terms.

Questions about pricing, contracts or how we work together?

Read the FAQ