The statistics every marketer actually needs
Every marketing report is built on a handful of statistical ideas, and most of them are quietly misused. "Average order value" is usually the mean, inflated by a few corporate whales. Two campaigns are compared without anyone asking how many orders sat behind each number. A correlation between email opens and repeat orders becomes "email drives loyalty" by lunchtime. In this opening chapter I walk through the statistics you genuinely need, using orders from The Gift Bow, my Solidus demo hamper shop: the three averages and why the median saves you from the whales, variance and standard deviation, the normal curve and the 68/95/99.7 rule, the standard error and why bigger samples narrow it, confidence intervals and what they actually promise, a plain reading of a p value, and Pearson correlation with its causation trap. Every formula is written out and every symbol explained, with tables you can copy for your own shop. This is part 1 of 21 of the Marketing Analytics series.
The whale in the hamper basket
Every December the same conversation happens in The Gift Bow, my Solidus demo hamper shop at hampers.keferboeck.com. Someone opens the admin dashboard, sees "average order value £102" and starts planning next year's ad budget around a £100 basket. Then I pull the actual orders. Of the 1,000 orders in the run up to Christmas, roughly 900 sit between £30 and £150. A handful do not: a law firm in Linz ordered 40 hampers for its clients, a Vienna accountancy practice ordered 25, and two UK companies sent gift boxes to their whole staff. Twenty orders above £500 carry about a fifth of the revenue. Take them out and the "average" customer suddenly spends £83, not £102. The numbers in this article are illustrative, but the shape is real. I have seen it in every shop I have ever looked at.
That gap between £102 and £83 is not a rounding error. It decides how much you can pay for a click, whether the free delivery threshold sits at £60 or £90, and whether the corporate segment deserves its own landing page or its own salesperson. Statistics, the boring kind you may have slept through, is the toolkit that stops you building a strategy on a number that only exists in a spreadsheet cell.
Where this sits in the series
This is the first of 21 chapters and the opening of part one, which covers how analytics helps at all (chapters 1 to 3). It lays the foundations every later chapter leans on: averages, spread, the normal curve, confidence intervals and correlation. The next chapter, Consumer behaviour is the basis of marketing strategy, not the afterthought, moves from the numbers to the people who produce them. If you want a gentler ramp before this one, I wrote The math behind it all as a broader tour of the arithmetic growth marketing actually uses.
Three averages, and only one of them is honest
"Average" is not one thing. There are three common measures of the centre of a distribution and they answer different questions.
The mean is the one everybody calculates: add everything up, divide by the count.
Here (read "x bar") is the mean, is the number of orders and is the value of the th order. The mean has one lovely property and one nasty one. The lovely property: multiply it by and you get total revenue, which is why finance loves it. The nasty one: every single order pulls on it with its full weight, so a £2,400 corporate order pulls as hard as 48 orders of £50. In a shop with a long right tail the mean drifts towards the whales.
The median is the middle value when you sort all orders from smallest to largest. For The Gift Bow's 1,000 December orders the median is £68. Half the customers spent less than that, half spent more. The law firm's order could have been £24,000 instead of £2,400 and the median would not have moved a penny. That robustness is exactly why I use the median for "what does a typical customer spend".
The mode is the most frequent value. In a hamper shop with fixed price products it is usually just the price of the bestseller, here the £49 Classic hamper. The mode tells you what your shop is, in the eyes of most buyers: a £49 shop that occasionally sells a £2,400 order.
Mean £102, median £68, mode £49. When the three line up in that order, mean highest, you have a right skewed distribution, and almost every e commerce order value distribution is right skewed. When someone quotes you a single "average basket", ask which one. If they do not know, it was the mean, and it is too high. I wrote a whole article on this trap, Statistical thinking for growth: why averages mislead, because it costs SMEs real money every year.
How spread out is it: variance, standard deviation and the coefficient of variation
Two shops can share a mean of £102 and be completely different businesses. One sells everything at £95 to £110. The other sells £49 hampers and £2,400 corporate orders. You need a number for the spread.
Variance is the average squared distance from the mean:
is the sample variance, is how far each order sits from the mean, squaring makes every distance positive and punishes large distances more than small ones, and dividing by instead of is a small correction because we estimate the mean from the same sample. Variance is measured in pounds squared, which nobody can picture, so we take the square root:
is the standard deviation, back in pounds. For our 1,000 orders it is about £158. Yes, larger than the mean. That is what a long tail does. Here is the whole distribution as a histogram:
Illustrative order value histogram. Note the unequal bin widths: the last three bins are wide and still nearly empty, yet they hold a fifth of the revenue.
The coefficient of variation puts the spread in proportion to the mean:
For The Gift Bow . Anything above 1 tells me the mean is not a useful description of a typical order. A SaaS like Werkbank with plans at €29, €59 and €119 will have a CV around 0.5, and there the mean revenue per account is a perfectly good number.
Here is the summary I would put on the first page of any e commerce report:
| Measure | Value | What it tells you |
|---|---|---|
| Orders | 1,000 | The sample size, always first |
| Mean | £102 | Revenue per order; drives the total |
| Median | £68 | What a typical customer spends |
| Mode | £49 | The bestseller's price |
| Standard deviation | £158 | Spread; here larger than the mean |
| Coefficient of variation | 1.55 | Spread relative to the mean; above 1 means "do not trust the mean alone" |
| Minimum | £18 | A single jar of chutney |
| Maximum | £2,400 | The Linz law firm |
| Orders above £500 | 20 (2%) | About 20% of revenue |
Read it top to bottom. The mean and median disagree by a third, so the distribution is skewed. The standard deviation exceeds the mean, so the skew is severe. The last row names the culprit: 2% of the orders, 20% of the money. That single table already tells you The Gift Bow has two businesses inside it, a £68 gift shop and a corporate gifting service, and they should probably not share one ad budget. Chapter 13 on segmentation makes that separation formal.
The bell curve and the 68/95/99.7 rule
The normal distribution is the symmetric bell shape you were shown at school. It is described completely by two numbers, its mean and its standard deviation , and it has a rule of thumb worth memorising: about 68% of values lie within one standard deviation of the mean, about 95% within two, and about 99.7% within three.
The normal curve with its three bands. In a normal world, a value more than two standard deviations from the mean is a one in twenty event.
To place any single value on that curve you standardise it into a z score:
is the value you are looking at, the mean and the standard deviation of the population it came from. A z score of 2 means "two standard deviations above the mean", whatever the units. A Gift Bow customer who opened 15 newsletters in a year, when the mean is 6 and the standard deviation 4, has : unusually engaged, top few percent if opens were normal.
Now the uncomfortable bit. Are order values normal? Not remotely. If they were, with mean £102 and standard deviation £158, the 68% rule would say a sixth of orders are below minus £56. Nobody pays me to take a hamper away. The histogram above is skewed, bounded at zero and heavy on the right, and no amount of wishing turns it into a bell. So why does everyone still teach the normal distribution to marketers?
Because of the next section.
The standard error: why the mean of a skewed shop still behaves
Imagine you took 1,000 random December orders, computed the mean, then did it again with a fresh 1,000, and again, thousands of times. Each sample mean would be slightly different. The spread of those sample means is called the standard error, and it is remarkably well behaved:
is the standard deviation of the individual orders, the sample size. For The Gift Bow: . So although individual orders scatter by £158, the mean of 1,000 of them only scatters by about £5.
And here is the point that justifies the whole bell curve lesson: the distribution of those sample means is approximately normal, even though the orders themselves are not. That result, the central limit theorem, kicks in reliably once you have a few dozen observations and is the reason confidence intervals and significance tests work on ugly, skewed, real world marketing data. The individual customer is not normal. The average of many customers is.
The square root in the denominator has a business consequence I explain to clients every month. To halve your uncertainty you need four times the data. To cut it to a tenth you need a hundred times the data.
Diminishing returns. The first 400 orders buy you most of the precision you will ever get; the next 6,000 buy surprisingly little.
Confidence intervals: what they promise and what they do not
A confidence interval turns the standard error into a range you can put in a memo:
is the sample mean, the standard error, and is the multiplier for the confidence level you want: 1.96 for 95%, 1.64 for 90%, 2.58 for 99%. (For small samples, under about 30, replace with the slightly larger value from the t distribution; any statistics library does this for you.) For The Gift Bow's mean order value: , so £92 to £112.
What a 95% interval promises: if you repeated the whole exercise many times, 95% of the intervals you computed would contain the true mean. What it does not promise: that there is a 95% probability the true mean sits in this particular interval, and certainly not that 95% of orders fall inside it (they do not; the interval is about the mean, not about individual orders). Forgive an agency that fumbles this distinction, I do it myself when tired. Do not forgive one that never reports an interval at all.
Here is where the intervals earn their keep. Two Meta campaigns ran in November, "Family Christmas" aimed at gift buyers and "Treat Yourself" aimed at self buyers, plus the newsletter as the owned channel. The agency's report said "Treat Yourself" delivered a £7 higher basket and should get more budget.
| Channel | Orders | Mean | Median | Standard deviation | Standard error | 95% interval for the mean |
|---|---|---|---|---|---|---|
| Family Christmas (Meta) | 180 | £74 | £62 | £38 | £2.83 | £68.5 to £79.5 |
| Treat Yourself (Meta) | 95 | £81 | £59 | £61 | £6.26 | £68.7 to £93.3 |
| Newsletter (owned) | 240 | £69 | £64 | £30 | £1.94 | £65.2 to £72.8 |
Read it row by row. Family Christmas: 180 orders, a modest spread, so the interval is tight at about plus or minus £5.5. Treat Yourself: fewer orders and a bigger spread (a few self buyers bought several hampers at once), so the interval is more than twice as wide. The two intervals overlap almost entirely. The newsletter has the lowest mean but the narrowest interval, because it has the most orders and the calmest customers.
Note the medians too. Treat Yourself has the highest mean but the lowest median. Its "higher basket" is a handful of large orders, not a typical customer spending more.
To test the £7 difference properly, compute the standard error of the difference, , and divide: . That is a z score of about 1, which corresponds to a two sided p value of roughly 0.31.
A plain reading of that p value: if the two campaigns truly produced the same basket, you would see a gap of £7 or more about one time in three, just from the random mix of who happened to click. That is not evidence. It is a Tuesday. The p value is not "the probability the campaigns are the same" and it is not "the probability the agency is wrong"; it is the probability of data at least this extreme if there were no real difference. To detect a genuine £7 gap with these spreads at the usual 80% power, you would need roughly 800 orders per campaign; the order of magnitude is typical. I wrote more on this, including why 0.05 is a convention and not a law of nature, in The importance of statistical significance and error margins.
The honest decision memo says: "No detectable difference in basket size between the two campaigns. Allocate on cost per order instead, where the gap is larger and the sample is the same." That is a real decision, made with a number and a risk, and it beats moving budget on noise.
Correlation: sign, strength and the causation trap
So far each number described one variable. Marketing is about pairs: does spend move revenue, do email opens move repeat orders, do discounts move basket size. Covariance measures whether two variables move together, but its units are meaningless (pounds times opens). Pearson's correlation coefficient divides the units away:
and are the two measurements for customer , and their means. The numerator is the covariance (scaled), the denominator is the product of the two standard deviations (scaled the same way), so always lands between minus 1 and plus 1. Read two things separately: the sign tells you the direction, the size tells you the strength. And tells you the share of variation in one variable that a straight line through the other can account for.
For 640 Gift Bow customers who received the newsletter over twelve months:
| Pair of variables | r | r squared | Honest reading |
|---|---|---|---|
| Newsletter opens and repeat orders | 0.31 | 0.10 | Positive, modest. Opens go with reorders but account for a tenth of the variation |
| Discount depth and order value | minus 0.42 | 0.18 | Negative, moderate. Bigger discounts, smaller baskets |
| Pages per session and order value | 0.08 | 0.01 | Practically nothing, whatever the heatmap tool says |
| First order value and second order value | 0.56 | 0.31 | Strong for marketing data; big first buyers stay big |
The first row is the one everyone wants to act on: "opens drive repeat orders, so send more email". Slow down. Customers who love hampers open the newsletter and reorder; the love causes both. Customers who are about to place a corporate order open the email to find the phone number. Correlation cannot separate "email causes reorders" from "reorders cause email opens" from "a third thing causes both". The only clean answer is a holdout test, which is the job of chapter 20 on statistical testing. What the correlation does earn you is a hypothesis worth testing and a rough ceiling on how much it could matter.
Two warnings on correlation that bite constantly. Pearson's only sees straight lines: a relationship that rises and then falls (ad frequency against conversion is the classic one) can have near zero while being very real. And a single whale can manufacture a correlation from nothing: one customer with 30 opens and 6 orders will drag upward in a dataset of 640, which is the mean's weakness wearing a different hat. Always plot it before you quote it.
How to actually run this
You do not need a data warehouse. You need the order table and forty minutes.
Start in SQL against your shop database. For a Solidus store the completed orders live in one table:
select
count(*) as orders,
avg(total) as mean_value,
percentile_cont(0.5) within group (order by total) as median_value,
stddev_samp(total) as sd,
stddev_samp(total) / sqrt(count(*)) as standard_error,
min(total), max(total),
count(*) filter (where total > 500) as whales
from spree_orders
where state = 'complete'
and completed_at >= date '2025-11-01'
and completed_at < date '2026-01-01';
Then split it by channel (the UTM source you store on the order, or the campaign field), and you have the campaign table above. For the intervals and the correlation, a few lines of Python:
import numpy as np
from scipy import stats
def mean_ci(x, level=0.95):
x = np.asarray(x, dtype=float)
n = len(x)
se = x.std(ddof=1) / np.sqrt(n)
t = stats.t.ppf((1 + level) / 2, df=n - 1)
return x.mean(), x.mean() - t * se, x.mean() + t * se
mean_a, lo_a, hi_a = mean_ci(orders_family["total"])
mean_b, lo_b, hi_b = mean_ci(orders_treat["total"])
r, p = stats.pearsonr(customers["opens"], customers["repeat_orders"])
What data you need: one row per order with value, date, channel and customer id; one row per customer with the engagement measures you care about. Twelve months is enough for a first pass; less than ninety days and the seasonality will lie to you.
How long it takes: an afternoon to pull the numbers, a day to argue about the definitions (is a refunded order an order, does a corporate invoice count as one order or forty). The arguing is the valuable part.
Checks before you trust it. Plot the histogram before you compute anything, because your eyes will catch a data error (test orders, a currency mix up, duplicated rows) faster than any statistic. Compare mean with median; if they differ by more than 20%, report both. Check the sample size in every group before comparing groups. Look for the whales and decide, in writing, whether they belong in the analysis or in a separate one. And never report a mean without either a standard error or an interval next to it.
Pitfalls I keep seeing
Reporting the mean of a skewed variable as if it were typical. Average order value, average session duration, average customer lifetime value: all right skewed, all inflated by a tail. Report the median alongside, or model the logarithm of the value (chapter 4 does this for elasticities, where it also fixes the maths).
Comparing groups without their sample sizes. A campaign with 40 orders and a £95 basket is not beating one with 900 orders and an £80 basket; it is a coin toss with a good story. The interval width is the honest measure of how much a number is allowed to say.
Treating a p value as the probability that a claim is true. It is a statement about the data assuming no effect, not about the claim. Also, running the same comparison every day and stopping when it crosses 0.05 guarantees you will eventually "find" an effect that is not there. Decide the sample size first.
Confusing the spread of individual values with the spread of the mean. "Our customers spend £102 plus or minus £10" is a statement about the mean. Individual customers spend £18 to £2,400. If you set a stock level or a fraud threshold using the interval for the mean, you will be wrong often.
Dropping the whales silently. Excluding the 20 corporate orders is legitimate and often right, but it changes the mean by a fifth. Say so in the report, every time, and keep a second table with them in.
How I do this for clients
When a client asks me to "look at the numbers", the first deliverable is exactly the material in this chapter applied to their own data, not a model. It starts with the free workshop, two hours where we agree which decision the numbers should change: ad budget, delivery threshold, whether corporate deserves its own funnel. Then I take two weeks on my own account. I need read access to the order table (Solidus, Shopify, WooCommerce, or a CSV export if that is what exists), the marketing spend by channel and week, and whatever customer engagement data is already collected. I do not need a data warehouse, a tag manager rebuild or a new tool.
What comes back is a short document rather than a dashboard: the descriptive table above for the whole business and for each channel and segment, every mean with its interval and the median next to it, a page of correlations with scatter plots and a blunt note on which ones are worth testing, and a decision memo that says "do this, expect roughly that, here is the risk". Where the answer is "you do not have enough data to know", it says so, with the number of orders you would need. You own every query and every chart; they live in your repository, not mine.
If that first month is useful, it usually grows into ongoing data science work, or into performance marketing where I run the campaigns against the intervals rather than the averages. The commercial terms are on the pricing page: a fixed monthly retainer sized to the work, or a commission on measured growth where the tracking is clean enough to measure it honestly.
Questions to ask your own team or agency
- When you say "average basket", is that the mean or the median, and how far apart are the two?
- What is the standard deviation, and how many orders sit more than two of them above the mean?
- How many orders are in each group you are comparing, and what is the confidence interval on each mean?
- What would we need to see, and how many orders would it take, to be confident the difference is real?
- Did you plot the distribution before computing anything from it?
- Which correlations in this report have a plausible causal story, and which one would you actually test with a holdout?
- Were any orders excluded, and does the report show the numbers both with and without them?
This series is inspired by Mike Grigsby's Marketing Analytics (Kogan Page). The explanations, examples and numbers here are my own.