Five Ways to Prove a Campaign Actually Worked

Every owner has said it, and I have said it too. We ran the campaign, and sales went up, so the campaign worked. It feels airtight. It is one of the weakest claims in all of marketing, and once you see why, you cannot stop seeing it. In the weeks you were running that campaign, a dozen other things moved at the same time. The weather changed. A competitor put their prices up or ran out of stock. Payday landed. You got a bit of press. The season turned. Your own email list grew. Any one of those could have lifted sales all on its own, and you would still be standing there taking the credit for the ad. Proving a campaign worked is not about pointing at a number that went up. It is about ruling out the other reasons it might have gone up. That is a solvable problem, and you do not need a data science team to solve it. You need one of five experiment designs, picked to fit the situation you are actually in. Here they are, weakest to strongest, in plain English, with the one most owners already have the data to run and never do.

The sentence that fools everyone

Every owner has said it, and I have said it too. We ran the campaign, and sales went up, so the campaign worked. It feels airtight. It is one of the weakest claims in all of marketing, and once you see why, you cannot unsee it.

Here is the problem in one breath. In the weeks you were running that campaign, a dozen other things moved at the same time. The weather changed. A competitor put their prices up or ran out of stock. Payday landed. You got a bit of press. The season turned. Your own email list grew. Any one of those could have lifted sales all on its own, and you would still be standing there taking the credit for the ad.

Proving a campaign worked is not about pointing at a number that went up. It is about ruling out the other reasons it might have gone up. The whole discipline of causal inference, and the reason this article exists, is a set of tricks for doing exactly that without a laboratory. The taxonomy I am about to walk you through is the one laid out in Cutting Edge Marketing Analytics by Venkatesan, Farris and Wilcox, and it sorts the ways you can run a marketing experiment from the flimsiest to the most bulletproof. Once you know the five, you can look at any campaign and ask a much better question than did sales go up. You can ask which of these five did I actually run, and how much should I trust the answer.

And be clear about who this is for. You do not need a stats degree or a data team. Every design below can be run by an owner with a spreadsheet, an email tool and a bit of discipline. The hard part is never the maths. It is deciding, before you spend the money, how you are going to know whether it worked.

Why you cannot just trust the before and after

The reason all of this is hard has a name. It is called the counterfactual, and it is the thing you never get to see. To know for certain that your campaign worked, you would need to observe what your sales would have done over the exact same weeks if you had not run it. Same weather, same competitors, same everything, minus the campaign. That world does not exist. You only ever get to live in one version of events, the one where you did run it.

Everything below is a different way of building a stand in for that missing world. A believable guess at what would have happened anyway. The five designs differ only in how good that guess is, and how hard it is to get. Keep that single idea in your head and the rest falls into place.

It helps to name the villain. Statisticians call it a confounder, a thing that moves your sales and happens to coincide with your campaign, so you cannot tell the two apart. The season is a confounder. A competitor stockout is a confounder. Payday is a confounder. Every design here is a different tactic for cancelling confounders out. Compare a campaign period to the past and the confounders roam free. Compare two randomly split groups living through the same weeks and most confounders get trapped in both groups equally and quietly cancel. That is the whole trick, said once, before we get into the five ways of doing it.

Design one, after only

The weakest of the lot. You run the campaign, you look at sales during and after, and you judge it on the level. Sales are strong, so the campaign worked. There is no comparison at all here, no before, no held out group. You are comparing the result to a feeling about what good looks like.

When after only is your only option

This is where most gut driven marketing lives, and it is barely evidence. The only time it earns its keep is when you genuinely have nothing else. A brand new business with no trading history. A one off launch of a product with no predecessor. A channel you have never touched, in a market you have never sold to, where you do not even have a baseline to lean on. In those cases after only is not a choice, it is the floor. You take it because there is nothing under it.

How to run it without kidding yourself

If after only is all you have, at least run it honestly. Write down, before you launch, what number would count as good and what number would count as a flop. Guess the range you expect, in writing, in advance. This does not turn a feeling into proof, but it stops you from moving the goalposts after the fact, which is the classic sin of after only. Everyone declares victory when the target is drawn around the arrow that already landed. Pin the target first.

A worked example

You open a new coffee bar. No history, no list, no comparable site. You run an opening week promotion and take four thousand euros across the counter. Did the promotion work? You genuinely cannot say, because you have no idea what an opening week takes without a promotion. The most you can honestly write down is opening week took four thousand, which becomes the baseline for the next thing you try. After only is not proof. It is the first brick of a baseline you do not have yet.

Design two, before and after

A real step up, and the one most owners actually use without knowing its name. You measure sales for a stretch before the campaign, you run the campaign, you measure again, and you take the difference. Sales averaged one hundred a week before and one hundred and eighteen a week during, so the campaign added eighteen.

When before and after earns its place

This is better than after only because at least you have a baseline. It is genuinely useful in two situations. First, when the effect you are looking for is large and fast, a big obvious jump that dwarfs the background noise, so even a rough method catches it. Second, when the outside world is quiet and stable across your before and after windows, no season change, no competitor move, no trend. The more boring the period, the more you can trust before and after. The moment the world starts moving underneath you, this design starts lying.

Why the baseline betrays you

The trouble is the baseline is your own past, and the past is not a fair stand in for the present. Everything that would have changed between the before period and the after period on its own, the season, the trend, the competitor, is now baked silently into your result and wearing your campaign's name tag. If you launch your garden furniture push in April and sales climb, some of that climb is the campaign and some of it is simply that it stopped being February. Before and after cannot tell those two apart.

There is a free upgrade that helps a lot, and costs you nothing but a longer look back. Instead of comparing to last month, compare to the same period last year, and to the underlying trend. If the business was already growing five percent a month before you launched, then a five percent jump during the campaign is not a campaign effect at all, it is just Tuesday. Subtract the trend you were already on before you credit the campaign with anything.

Campaign effect≈(Sales during)−(Baseline×Expected growth)\text{Campaign effect} \approx (\text{Sales during}) - (\text{Baseline} \times \text{Expected growth})

A worked example

You sell garden furniture. March averaged one hundred thousand euros a week. You launch a push in April and April averages one hundred and eighteen thousand. Naive before and after says the campaign added eighteen thousand a week. But you check last year, and April always runs about twelve percent above March regardless of what you do, because it stops being winter. So your seasonal baseline for April is not one hundred thousand, it is about one hundred and twelve thousand. The honest campaign effect is closer to six thousand a week, not eighteen. Same numbers, a third of the story, and the two thirds you nearly claimed was just the spring arriving on schedule.

Design three, test and control, the one you are not running

Here is the first thing I really want you to take away, because it is the least expensive gold standard in marketing and almost nobody with a small business bothers with it.

Instead of comparing now to the past, you compare two groups living through the exact same now. You take your audience, you split it at random into two, and you show the campaign to one half, the treatment group, and nothing to the other half, the control group, the hold out. Then you compare the two over the same weeks. Because the split was random, both halves sat through the same weather, the same season, the same competitor, the same press. All of that cancels out. Whatever difference is left between treatment and control is the campaign, and almost nothing else. This is a randomised controlled experiment, the same logic medicine uses to test a drug, and you can run it from an email tool on a Tuesday.

Why random is the magic word

The word doing all the work here is random. It is tempting to build your control group by hand, the customers who did not open last time, the ones in the quiet region, the smaller accounts. Do not. The moment you choose the two groups on any rule, that rule becomes a confounder and you are back to square one. If your control group is the people who never open your emails, of course they buy less, they were less engaged before you started. Random assignment is the one thing that makes the two groups genuinely comparable, because on average everything else, engaged and unengaged, big and small, north and south, splits evenly between them. Let a coin decide, not your judgement.

The part that should sting

If you have an email list, a customer database, a loyalty scheme, an app, a retargeting audience, you already have everything you need to do this. The data is sitting there. The only reason you are not running test and control is that nobody ever told you to hold a slice of your own audience back on purpose and send them nothing. It feels wrong to deliberately not market to some of your customers. It is the single most valuable thing you can do, because those unmailed customers are the missing world made real. They are the counterfactual, walking around, spending money, showing you exactly what would have happened without the campaign.

You do not need the whole list held out. Ten percent is plenty for most SMEs. Keep it random, keep it untouched, and never look at it as lost revenue. Look at it as the only honest yardstick you own.

How to run one, step by step

Here is the whole thing, start to finish, in an order you can follow on your next campaign.

  1. Pick the outcome you actually care about before you start. Revenue per person is usually the right one, not opens or clicks. Opens do not pay wages.
  2. Take your eligible audience and split it at random. A hold out of ten to twenty percent is plenty. Tag the two groups so you can tell them apart later.
  3. Send the campaign to the treatment group only. Touch nothing in the control group for the whole test window, no email, no retargeting, no sneaky text.
  4. Wait long enough to capture the real buying cycle. If people take three weeks to decide, measuring after three days tells you nothing.
  5. Compare revenue per person across the two groups, then work out the lift and check it against the noise, which is the next section.
  6. Write the result down, campaign and control and lift, so it becomes a record you can compare the next campaign against.

The discipline that breaks people is step three. Someone always says it feels a shame not to mail the hold out too, and the second you cave, your yardstick is gone.

Putting a number on it, and clearing the noise

Once you have a treatment group and a control group, the sum is refreshingly simple. Lift is how much more the treatment group did than the control group, as a share of the control.

Lift=Treatment−ControlControl\text{Lift} = \frac{\text{Treatment} - \text{Control}}{\text{Control}}

A worked example. Over your test weeks the treatment group, the half you mailed, buys to the tune of one hundred and eighteen thousand euros. The control group, the matched half you held out and left alone, buys one hundred thousand euros. Same size groups, same weeks.

Lift=118,000−100,000100,000=18,000100,000=0.18\text{Lift} = \frac{118{,}000 - 100{,}000}{100{,}000} = \frac{18{,}000}{100{,}000} = 0.18

Eighteen percent lift, cleanly attributable to the campaign, because everything else hit both groups equally. That eighteen thousand is money the campaign made that would not have existed otherwise, and now you can set it against what the campaign cost and know, actually know, whether it paid.

From lift to a decision, not just a number

A lift on its own is a bragging number. The number that runs your business is what it earned against what it cost. If the campaign cost two thousand euros to run and generated eighteen thousand in extra revenue at, say, a forty percent margin, then it made about seven thousand two hundred in gross profit against two thousand of cost.

Return=(Incremental revenue×Margin)−CostCost\text{Return} = \frac{(\text{Incremental revenue} \times \text{Margin}) - \text{Cost}}{\text{Cost}}

Return=(18,000×0.40)−2,0002,000=5,2002,000=2.6\text{Return} = \frac{(18{,}000 \times 0.40) - 2{,}000}{2{,}000} = \frac{5{,}200}{2{,}000} = 2.6

Two hundred and sixty percent on the money you put in, and this time it is not a story, it is measured against a real hold out. The word incremental is the whole point. Test and control is the only design here that gives you truly incremental revenue, the sales that would not have happened anyway. Every other measure, especially the last click number in your ad dashboard, quietly counts sales you would have got for free.

Clearing the noise

One warning that matters more than the formula. A difference is not a result until it clears the noise. Sales wobble week to week for no reason at all. If your treatment group beat control by two percent, and your sales naturally bounce around by five or six percent from one week to the next anyway, you have not proved anything. You have caught a wobble and dressed it up as a win.

Before you trust a lift, you want the gap between the groups clearly bigger than the everyday jitter in the numbers, and enough people in each group that a few big spenders cannot swing it. A rough rule I use: if the lift is smaller than the week to week swing you normally see, treat it as unproven and keep running. If the gap is small and the groups are small, run it longer or hold your applause.

Design four, field experiments

Test and control on a list is the easy version because everyone is in the same inbox. A field experiment is the same idea let loose in the real world, across places and things you can physically separate. You run the campaign in Linz and Graz and deliberately not in Salzburg and Innsbruck, matched cities, then compare. You put the new packaging in half your stockists and the old in the other half. You switch a billboard on in one region and leave the neighbouring one dark.

When a field experiment is the right tool

Reach for this when the thing you want to test does not live in an inbox. Offline campaigns, radio, out of home, local sponsorship, packaging, in store display, brand advertising. None of that shows up cleanly in a click report, and last click attribution is worse than useless for it. A field experiment is how you put a real number on the spend your dashboard is blind to. That is the second thing worth sitting with. The channels hardest to measure with tracking are exactly the ones a field experiment can rescue, because it measures outcomes in the world rather than clicks in a dashboard.

How to run one

The logic is identical to design three, a treated group and a held out control living through the same period, so the strength is nearly as high. The craft is in the matching and the separation.

  1. Match your units before you start. Pick test and control regions that looked alike in the months before the campaign, similar size, similar trend, similar seasonality. Two cities that already move together are a fair pair.
  2. Use enough units. Two cities against two cities is thin, because one odd event in one city swings the whole result. More, smaller units beat two big ones.
  3. Watch for spillover. If your control town sits next to your test town and people commute across, the ad leaks into your control and shrinks the gap you are trying to measure. Separate them by geography or by time.
  4. Measure the same outcome in both, over the same window, and compute lift exactly as before.

A worked example

You want to know if a regional radio campaign moves the needle. You have eight towns you serve. You match them into four pairs by last year's sales, then flip a coin within each pair to decide which town gets radio and which stays quiet. After six weeks, the four radio towns are up nine percent on the four quiet towns, and the quiet towns confirm the season was no different for either. Nine percent lift on a channel your click dashboard would have scored as zero, because radio does not get clicked. That is the whole case for field experiments in one result.

Design five, natural experiments

The cleverest of the five, and the only one you do not run on purpose. A natural experiment is when the world hands you a treatment and control split for free, through some event you did not cause. A courier strike knocks out delivery to half your regions but not the other half. A rule change applies to one type of customer and not another. A stockout takes a product off the shelf in some stores for a fortnight. A pricing error goes live in one country overnight.

When to seize one

Whenever something splits your customers into affected and unaffected by pure accident, you have been given a free experiment, and if you were keeping clean records you can measure the effect after the fact. This is how you learn from things you would never be allowed to test deliberately, big price moves, outages, policy shifts, a supplier failure. Nobody signs off a deliberate test that doubles a price or cuts off delivery to half the country. The world runs those tests for you, for free, and all you have to do is notice and read the result.

How to read one

The trick is to treat the accident exactly like a planned test and control, then be honest about where it falls short.

  1. Define the affected group and the unaffected group by the event, not by anything you chose. The event did the randomising, or something close to it.
  2. Grab a clean before picture for both groups, so you can check they were tracking together before the event hit. If they were not alike before, the split was not as random as it looked.
  3. Compare the change in the affected group to the change in the unaffected group over the same window. The difference of those two differences is your effect.

Effect=(Affected after−Affected before)−(Unaffected after−Unaffected before)\text{Effect} = (\text{Affected after} - \text{Affected before}) - (\text{Unaffected after} - \text{Unaffected before})

That last formula has a name, difference in differences, and it is the workhorse of natural experiments. It nets out anything that hit both groups, so what remains is the bit that only hit the affected side.

A worked example

A courier strike stops next day delivery to your eastern regions for two weeks while the west runs normally. Sales in the east drop fourteen percent against the fortnight before. But the west also dropped four percent, because it was a quiet fortnight everywhere. So the strike did not cost you fourteen percent, it cost you the ten percent gap between east and west, the four percent was the general lull that hit both. Now you know what next day delivery is worth to those regions, a number you could never have got deliberately. The catch is you have to notice it happened and have the data to look back on, which is really an argument for keeping tidy records of who saw what and when, all the time, so that when the world runs an experiment on your behalf you are ready to read it.

The five ranked

Here is the whole menu on one screen, sorted by how strongly each design proves cause and effect against how hard it is to actually run. Strength is the thing you want. Effort is the thing you pay.

DesignHow strongly it proves causeHow hard to runUse it when
After onlyVery weakVery easyYou have no history and nothing to compare to
Before and afterWeakEasyYou need a rough read and effects are large
Test and controlStrongEasy to moderateYou have a list or database you can split at random
Field experimentStrongModerate to hardThe channel is offline or brand and lives in the real world
Natural experimentStrongOpportunisticThe world accidentally split your customers for you

Read down the strength column and one thing jumps out. The jump in proof happens between before and after and test and control, and the price of that jump is almost nothing. Splitting a list at random is barely more work than mailing all of it. That is the whole argument of this article compressed into one row of a table. The most affordable reliable design is the one nearly everyone skips.

There is a second way to read the same table, by which confounders each design controls. It explains the strength column rather than just asserting it.

DesignControls season and trendControls competitor and world eventsHow it does it
After onlyNoNoNothing, it is a level
Before and afterPartlyNoYour own past as a baseline
Test and controlYesYesRandom split, same weeks
Field experimentYesMostlyMatched places, same weeks
Natural experimentYesMostlyAn accident splits the groups

Which one can you run right now

You do not get to pick your favourite. You pick the strongest design your situation allows. This tree is how I decide in practice.

Work down from the top and stop at the first yes. Most SMEs with any kind of customer list land on the very first branch and should be running test and control, which is exactly the design they never reach for. If you cannot split a list, you drop to a field experiment. If you cannot stage anything, you keep your eyes open for a natural experiment the world hands you. And only when none of those are on the table do you fall back to before and after, and after only is the seat of last resort.

A worked mini case, all five in one business

Let me put it together with one shop, because seeing the ladder climbed in a single business makes it stick. Imagine a mid sized online and high street retailer, a few thousand names on the list, eight shops around the country. They want to know whether their spring campaign is worth repeating.

Year one, they do what everyone does. They run the campaign everywhere, sales rise, they call it a win. That proves nothing, because spring lifts them every year anyway.

Year two, I talk them into a hold out. Ninety percent of the list gets the spring emails, ten percent, chosen at random, gets nothing. The mailed group spends nine percent more per head than the hold out over the same six weeks. That nine percent is real, incremental and defensible, because the season hit both groups the same. For the first time they know the email half of the campaign pays.

Same year, they also run radio in four of their eight towns, matched by last year's takings against the four they leave quiet. The radio towns come in seven percent above the quiet ones. Now the offline half has a number too, one their click dashboard could never have given them.

Then the world chips in. A supplier fails and one product line is out of stock in three shops for a fortnight. They kept clean records, so afterwards they run the difference in differences against the five shops that stayed in stock, and learn what that line is worth on the shelf.

By year two this business is not guessing. Every part of the spring push has a defensible number attached, built from designs that each cost almost nothing over just spending the money. That is the whole point of the ladder. You climb as high as your situation lets you, and you refuse to pretend you climbed higher than you did.

Common mistakes that quietly ruin the answer

Even people who run a hold out manage to poison it. These are the ones I see most.

  1. Building the control group by hand instead of at random. The moment you pick the two groups by any rule, the rule becomes a confounder and the comparison is dead.
  2. Mailing the hold out anyway, because it felt a waste not to. One weak moment and your yardstick is gone.
  3. Measuring the wrong outcome. Opens and clicks are not revenue. A campaign can lift opens and move no money at all.
  4. Calling a wobble a win. If the lift is smaller than your normal week to week swing, you have proved nothing yet, however much you want to.
  5. Peeking early and stopping the moment it looks good. Stop the clock when you planned to, not when the number flatters you.
  6. Ignoring the cost side. A twenty percent lift that cost more than it earned is a failure with good PR.

None of these need a statistician to avoid. They need you to decide the rules before you look at the numbers, and then hold your nerve.

How to apply this on Monday

None of this is theory you file away. It changes one habit. Before your next campaign, ask a single question. What is my control group. If you can answer it, a held out slice of the list, a matched set of towns, a genuine baseline, you are about to learn something real. If you cannot answer it, you are about to spend money and then tell yourself a story about what it did.

Here is the whole article as a routine you can run every time.

  1. Before you launch, write down what you expect and what would count as a win.
  2. Decide your control group. If you have a list, hold ten percent back at random and leave it alone.
  3. Pick one outcome that pays wages, usually revenue per person, and ignore vanity numbers.
  4. Run the campaign, wait out the real buying cycle, compute lift, and check it clears your normal noise.
  5. Turn lift into return by setting incremental profit against cost, then write it all down so the next campaign has something honest to beat.

The owners who compound over the years are not the ones with the cleverest ads. They are the ones who, quietly and constantly, hold a bit back so they can tell what is working. It is the least glamorous habit in marketing and the one that pays for itself the fastest.

If you want a hand building this into how you run campaigns, wiring a proper hold out group into your email and ads and setting up the measurement so every push tells you the truth, that is precisely the kind of work I do. Have a look at how I approach performance marketing, or just book a call and bring the last campaign you were not quite sure about. We will find out whether it actually worked.

If your campaigns end with sales went up so it worked and you have never once held a control group back to check, let us fix that together. Book a call and we will design a simple test and control on your next push, so for the first time you will know what your marketing is really doing rather than guessing.

1%of every invoice goes to a UK charity you pick.

A donation, never sponsorship. You choose the cause at onboarding.

The story behind the pledge →

Stay ahead of your competition.

The latest innovative products and services, straight to your inbox before your competitors hear about them.

Get up to 5% off your first six months: 1% per topic you pick, the full 5% when you take everything. Limited offer · ends 31 December 2026.

New clients only. Terms apply.

* Up to 5% off your first six monthly invoices, new clients only. Full terms.

Questions about pricing, contracts or how we work together?

Read the FAQ