Why Testing One Thing at a Time Is Costing You Money

You run a clean test. You change the headline, hold everything else still, and the new headline wins. Next week you change the price, hold everything else still, and the lower price wins. So you ship the winning headline and the winning price together, because you tested them properly, one thing at a time, like every article on the internet told you to. And your conversion rate goes down. That is not bad luck, it is maths, and it happens far more often than anyone running these tests likes to admit. Testing one variable at a time has a blind spot built into the method itself, and the blind spot is exactly where a lot of the money hides. I want to show you what it is, why it happens, and what to do instead, with a worked example you can rebuild in a spreadsheet before lunch.

The test that picks the wrong winner

You run a clean A/B test. You change the headline, you leave everything else alone, and the new headline wins. Great. Next week you run another clean test. You change the price, everything else held still, and the lower price wins. Great again. So you ship the winning headline and the winning price together, because you tested them properly, one thing at a time, like every article on the internet told you to. And your conversion rate goes down.

That is not bad luck. That is maths. And it happens more often than anyone running these tests would like to admit, because testing one variable at a time has a blind spot built into the method itself, and the blind spot is exactly where a lot of the money hides. I want to show you what it is, why it happens, and what to do instead, with a worked example you can rebuild in a spreadsheet before lunch.

One at a time feels rigorous, and that is the trap

The reason everyone tests one thing at a time is that it feels like the careful, scientific thing to do. Change one variable, hold the rest constant, and you can point at the result and say the headline caused the lift. Clean cause, clean effect. It is the first rule of experiments you ever learn, and it is genuinely good advice for a lot of questions.

But hold that constant part up to the light. When you test the headline with the price held at its old value, you are only ever measuring what the new headline does at that one price. You learn nothing about how it behaves at a different price. You are assuming, without ever checking, that the headline's effect is the same regardless of what the price is doing next to it. Testing one variable at a time does not just fail to measure that assumption. It cannot measure it. The method is built to hold the very thing that would reveal the problem perfectly still.

That assumption has a name. It is the assumption that your changes are additive, that the effect of the headline plus the effect of the price is simply the sum of the two. And when it is wrong, it is wrong in a way that one at a time testing is structurally incapable of catching.

When two changes refuse to add up

Let me make it concrete with the smallest example that shows the problem. Two factors, two levels each. The headline is either the plain value message or the premium craftsmanship message. The price is either the full price or a discount. That gives four possible campaigns, and a full factorial test simply runs all four.

Here is what comes back, as conversion rate out of every hundred visitors.

CombinationFull priceDiscount
Value headline2.02.5
Premium headline2.62.4

Now watch what one at a time testing does with this. Start from the baseline in the top left, value headline at full price, 2.0.

Test the headline on its own, holding price at full. Premium beats value, 2.6 against 2.0. The premium headline wins, clearly.

Test the price on its own, holding the headline at value. Discount beats full price, 2.5 against 2.0. The discount wins, clearly.

So one at a time testing hands you two confident winners, the premium headline and the discount, and tells you to ship them together. That combination is the bottom right cell. It converts at 2.4. It is worse than three of the four options you had, including the baseline you started from. You ran two clean, rigorous tests and they marched you straight into nearly the worst square on the board.

The genuinely best campaign is the premium headline at full price, 2.6, sitting in the bottom left. One at a time testing can reach it in one of its two tests, then throws it away the moment it combines the winners, because it never suspected the winners might not get along.

The reason they fought, and the number that proves it

The story behind the numbers is not exotic. A premium craftsmanship message and a discount tag send opposite signals. The message says this is worth paying for, the price says we are knocking money off, and the visitor feels the mismatch even if they could not name it. The premium message works with a full price because the two agree. The discount works with the value message because those two agree as well. Cross them over and each undercuts the other.

This crossing over has a precise measure, the interaction effect. For a two by two grid it is the difference of the differences. Take how much the premium headline helps at the discount price, and subtract how much it helps at the full price.

Interaction=(μPD−μVD)−(μPF−μVF)\text{Interaction} = (\mu_{PD} - \mu_{VD}) - (\mu_{PF} - \mu_{VF})

Plug in the four cells, where P and V are premium and value, F and D are full and discount.

Interaction=(2.4−2.5)−(2.6−2.0)=(−0.1)−(0.6)=−0.7\text{Interaction} = (2.4 - 2.5) - (2.6 - 2.0) = (-0.1) - (0.6) = -0.7

That minus 0.7 is the whole story in one number. If the effects were additive, if the world were as tidy as one at a time testing assumes, this would come out at zero, and stacking the two winners would work exactly as promised. The further it sits from zero, the more the two factors are bending each other, and the more dangerous it is to test them apart. Here is the part worth sitting with. The interaction term is precisely the quantity that testing one variable at a time is designed never to show you. You can only see it by running the combinations together. The method that feels the most rigorous is blind to the one number that decides whether its own conclusion holds.

Three grids, three kinds of interaction

That minus 0.7 was negative, and negative is the dangerous one. But an interaction can land anywhere, and the sign is the whole plot. Let me put three grids next to each other, same layout, same baseline of 2.0 in the top left, and read the interaction off each.

Start with the tidy world, the one where one at a time testing is telling the truth.

Additive worldFull priceDiscount
Value headline2.02.4
Premium headline2.63.0

Interaction here is (3.0 minus 2.4) minus (2.6 minus 2.0), which is 0.6 minus 0.6, which is zero. The premium headline adds 0.6 whether the price is full or discounted, and the discount adds 0.4 whichever headline sits above it. Everything adds up. Stack the two winners, premium and discount, and you land on 3.0, which is genuinely the best cell. In a world like this, one at a time testing is fine and you should use it, because it is faster.

Now the friendly surprise, a positive interaction.

Synergy worldFull priceDiscount
Value headline2.02.2
Premium headline2.33.2

Interaction here is (3.2 minus 2.2) minus (2.3 minus 2.0), which is 1.0 minus 0.3, which is plus 0.7. The two winners do not just add, they amplify. Premium and discount together hit 3.2, which is more than the parts would predict. One at a time testing still lands on the right cell here, but it badly underestimates how good it is, so you would under invest in the winner. There is the next thing worth sitting with. Even when sequential testing picks the right combination, a positive interaction means it misreads the size of the prize, and budgets get set on the wrong number.

And the one we already met, the trap, a negative interaction.

Conflict worldFull priceDiscount
Value headline2.02.5
Premium headline2.62.4

Interaction here is (2.4 minus 2.5) minus (2.6 minus 2.0), which is minus 0.1 minus 0.6, which is minus 0.7. The winners fight, and stacking them lands you on 2.4, worse than where you began. Same method, three completely different verdicts, and the only thing that changed was one number you cannot see without running the combinations. That is the point. One at a time testing does not know which of these three worlds it is standing in, and it behaves identically in all three.

Reading a whole grid instead of two lines

Two factors at two levels is the toy version. The real lesson lands when you widen it, so let me borrow the shape of the message by price experiment laid out in Cutting Edge Marketing Analytics by Venkatesan, Farris and Wilcox. Three messages, three prices, every combination run at once. Nine campaigns, one grid.

Say the three messages are a value message, a premium message and an urgency message, and the three prices are low, medium and high. Here is the conversion rate for all nine cells, with the row and column averages worked out on the edges.

MessageLow priceMedium priceHigh priceMessage average
Value3.22.82.12.70
Premium2.43.03.42.93
Urgency3.12.92.52.83
Price average2.902.902.67

Now do what a one at a time thinker does with averages. The best message by its row average is premium, at 2.93. The best price by its column average is a tie between low and medium, at 2.90. So the averages tell you to pair the premium message with the low price. Go to that cell, premium message at low price. It converts at 2.4, the second worst square in the entire grid.

The genuinely best campaign is the premium message at the high price, converting at 3.4, and the averages actively point away from it. That is the second thing worth sitting with. When interactions are in play, the best row crossed with the best column can land you almost anywhere, because a marginal average smears together cells that behave completely differently. The premium message is not just good on average, it is good specifically when the price is high, and that pairing is invisible to anyone reading the margins instead of the middle. Only running the full grid shows you the cell that actually pays.

Splitting the grid into main effects and the bit that lies

There is a cleaner way to see why the margins mislead you, and it is worth two minutes because it is exactly how a proper model reads your test. Any cell in a factorial can be pulled apart into three pieces, the grand mean, the main effects, and the interaction.

Take our three by three grid again. The grand mean, the average of all nine cells, is 2.82. That is the baseline you would predict for a cell knowing nothing about which message or price it holds.

Now add what each factor does on its own, its main effect, measured as how far its average sits from the grand mean. The premium message averages 2.93, so its main effect is 2.93 minus 2.82, about plus 0.11. The high price averages 2.67, so its main effect is 2.67 minus 2.82, about minus 0.16.

An additive model, which is the world one at a time testing quietly believes in, predicts any cell by stacking these pieces.

μ^=grand mean+message effect+price effect\hat{\mu} = \text{grand mean} + \text{message effect} + \text{price effect}

For premium at the high price that gives

μ^=2.82+0.11−0.16=2.77\hat{\mu} = 2.82 + 0.11 - 0.16 = 2.77

The additive model says the best you can hope for from premium at a high price is about 2.77, a touch below average, which is exactly why the margins steer you away from it. But the real cell converts at 3.4. The gap between them is the interaction.

interaction=3.4−2.77=+0.63\text{interaction} = 3.4 - 2.77 = +0.63

Sit with those two numbers. The additive prediction is 2.77. Reality is 3.4. That plus 0.63 is not noise, it is a real, repeatable lift that lives only in the pairing, and it is the single richest cell in the grid. One at a time testing, at its very best, can find the additive prediction. The interaction residual is money it cannot reach by design, because reaching it means measuring the two factors together, which is the one thing the method refuses to do.

Adding a third factor, the cube you cannot see from the edges

Two factors give you a grid you can print. Three factors give you a cube, and a cube has a nasty habit the grid only hinted at. Let me add the hero image to the message and price, two images this time to keep it readable, and lay the cube out as two grids, one per image.

First image, a lifestyle photo.

Lifestyle imageFull priceDiscount
Value headline2.02.5
Premium headline2.62.4

Second image, a plain product shot.

Product imageFull priceDiscount
Value headline2.22.3
Premium headline3.12.7

The best cell in the whole cube is the premium headline, full price, product image, at 3.1, and notice it did not exist on the first grid at all. But here is the part that should make you slightly nervous. Remember the message by price interaction on the lifestyle image was minus 0.7, the winners fighting. Work it out again on the product image.

(2.7−2.3)−(3.1−2.2)=0.4−0.9=−0.5(2.7 - 2.3) - (3.1 - 2.2) = 0.4 - 0.9 = -0.5

Still negative, but not as fierce. The interaction itself changed when we swapped the image, from minus 0.7 to minus 0.5. That shift is a three way interaction, an interaction between interactions, and it means the image is quietly rewriting how message and price get along. So here is the next one to sit with. Interactions are not a one off correction you measure once and bank. They can depend on a third thing, and a fourth. This is not a reason to panic and vary twelve factors at once, it is the opposite. It is the reason to pick the two or three factors you most suspect are entangled and actually run them together, because the entanglement can run deeper than a single grid shows.

How fast the grid grows, and why that is fine

The obvious objection is cost. Surely running every combination explodes into an unaffordable number of tests. It grows quickly, yes, but not as scarily as the fear suggests, and the shape of the growth is worth knowing exactly.

For a full factorial design with kk factors, each at nn levels, the number of combinations is

N=nkN = n^k

Two factors at two levels is 22=42^2 = 4. Our three by three is 32=93^2 = 9. Add a third factor at three levels, say the image, and it is 33=273^3 = 27. A fourth and you are at 34=813^4 = 81. The exponent is the number of factors, so combinations grow fast in the number of things you vary and more gently in the number of levels each one takes.

Compare that with how many campaigns one at a time testing actually looks at. It runs the baseline, then changes each factor across its remaining levels one at a time, which is 1+k(n−1)1 + k(n-1) combinations. Here is the two side by side.

FactorsLevels eachFull factorialOne at a time
2243
2395
33277
43819

The gap is real and it grows. But here is the thing that changes how you plan a test. A full factorial does not usually need more traffic than a one at a time sweep. It needs the same visitors sliced across more cells. If you have the volume to run four clean sequential tests, you very often have the volume to run one factorial across the same weeks, because you were going to spend those visitors either way. What you are really deciding is whether to spend them learning about factors in isolation, which can lie to you, or learning about combinations, which cannot. And when volume genuinely is tight, you do not need the full grid. Fractional designs and sensible priors let you cover the important interactions without running all eighty one cells, which is exactly the sort of trade off worth planning deliberately rather than defaulting into.

One at a time against full factorial, in one picture

Here are the two approaches as the paths they actually walk.

The left path walks a line. It fixes one winner, then the next, then bolts them together and hopes they get along. The right path fills a grid. It never assumes the winners are friends, it just measures every pairing and reads off the best one. The line is faster to draw and easier to explain to a stakeholder. The grid is the one that does not hand you a losing campaign with a confident face.

A worked mini case, one week on a landing page

Let me put the whole idea on one page, literally, with a landing page you could imagine running this week. All the numbers here are made up to keep the arithmetic clean, so swap in your own, but the shape is the thing.

Say you are testing a quote request page. Two factors. The call to action copy, either Get my quote or Start now. And the form length, three levels, three fields, five fields or eight fields. That is a two by three factorial, six cells, and you run all six at once.

Here is what comes back, conversions per hundred visitors.

CTA / form3 fields5 fields8 fields
Get my quote4.13.82.9
Start now4.63.52.2

Now walk it the old way. Your live page today is Get my quote with the standard five field form, converting at 3.8, that is your baseline. You test the CTA first, holding the form at five fields. Get my quote at 3.8 beats Start now at 3.5, so you keep Get my quote. Then you test the form length holding the CTA at Get my quote, and three fields wins at 4.1. You ship Get my quote with three fields, 4.1, and you are pleased, because it beat your baseline.

But look at the grid. The best cell is Start now with three fields, at 4.6. You left half a point on the table, and half a point on a page that matters is a lot of quotes over a year. Why did the sequential test miss it? Because it tested the CTA at five fields, and at five fields Start now genuinely is worse. Start now only shines when the form is short, because a punchy, urgent button that promises a quick job and then asks for eight fields is writing a cheque the form cannot cash, whereas the same button over a three field form tells the truth. The winner of the CTA test depended entirely on the form length you happened to hold it at, and you held it at the wrong one.

The interaction spells it out. Compare the CTA effect at the short form and the long form.

(4.6−4.1)−(2.2−2.9)=0.5−(−0.7)=1.2(4.6 - 4.1) - (2.2 - 2.9) = 0.5 - (-0.7) = 1.2

A swing of 1.2 between the two ends of the form is enormous, far bigger than either factor on its own, and it sits entirely in how the two play together. Run them apart and it is invisible. Run them together and it is the headline result.

The common mistakes that quietly wreck a factorial test

Once people are sold on running combinations, they tend to trip over the same handful of things. Here are the ones I see most, and how to sidestep them.

Varying everything at once

The exponent in nkn^k is unforgiving. Six factors at three levels is seven hundred and twenty nine cells, and your traffic gets sliced so thin that every cell is noise. A factorial is a scalpel, not a fishing net. Pick the two or three factors you genuinely believe are entangled and leave the rest for later.

Levels that are not really different

Testing a price of 99 against 98 is not two levels, it is one level with a rounding error. If the levels do not differ enough to plausibly change behaviour, you are spending traffic to measure noise. Make each level a real, distinct choice.

Reading the margins instead of the cells

This is the original sin, the one this whole piece is about. The row and column averages are comforting and they lie whenever interactions are present. Always find the best cell, not the best row crossed with the best column.

Peeking and stopping early

A grid needs enough visitors per cell to trust the numbers, and cells are smaller than a single A/B test's two buckets, so they need patience. Deciding the winner the moment one cell edges ahead is how you crown a random wobble. Set the sample per cell up front and hold your nerve.

Not owning the tracking

You cannot read a nine cell grid if you cannot reliably tell which visitor saw which cell and what they did next. A rented dashboard that forgets the assignment turns your careful factorial into mush. Own the tracking end to end.

Assuming a flat row means the factor does not matter

This one is sneaky and it is a genuine surprise. A factor can have almost no main effect, its row average dead level with the grand mean, and still be one of the most important things on the page, because it only acts through interaction. If the image does nothing on average but flips which price wins, the image matters enormously, and a main effects only mindset throws it in the bin. Never read a main effect on its own and conclude a factor is dead.

What this means for how you actually test

You do not need to blow up how you work. You need three habits.

First, when two changes could plausibly talk to each other, test them together, not in sequence. Message and price talk to each other. Headline and hero image talk to each other. Offer and audience talk to each other loudly. Anything where one change alters the meaning of another is a candidate for a factorial, because that is precisely where a sequential test will mislead you. When two changes really are independent, a plain sequential test is fine, and quicker.

Second, be deliberate about how many factors and levels you take on at once, because nkn^k is unforgiving if you are careless. Three factors at three levels is a rich, readable experiment. Seven factors at four levels is a research project. Pick the two or three variables you genuinely suspect interact, give each a small number of meaningful levels, and resist the urge to vary everything.

Third, wire your testing to your own data so you can actually read a grid. Running a nine cell experiment and attributing outcomes correctly needs tracking you own and control, not a rented dashboard that forgets which visitor saw which combination. This is the same argument I keep making about owning the platform your growth runs on. If you want a hand designing a factorial test that fits your traffic, or turning a pile of past A/B results into a model that respects interactions, that is squarely the kind of work I do, and you can see the shape of it across my growth and analytics work. Bring your last three tests and we will find out whether your winners actually get along.

How to apply this on Monday morning

None of this earns its keep unless it survives contact with a real Monday. So here is the drill, start to finish.

  1. Write down the list of changes you were about to test one after another. Headline, price, image, CTA, form, whatever is in the queue.
  2. Draw a line between any two that could plausibly change one another's meaning. Message and price. CTA and form. Offer and audience. Those lines are your candidate factorials.
  3. Pick the two or three most entangled factors and give each a small number of genuinely different levels. Two or three levels each is plenty.
  4. Compute nkn^k and divide your expected traffic by it. If a cell would get too few visitors to trust, drop a factor or a level until it does not.
  5. Run all the cells at the same time, on tracking you own, so every visitor is cleanly assigned and every outcome is cleanly attributed.
  6. When the numbers are in, read the whole grid. Find the best cell. Then compute the interactions so you understand why it won, not just that it did.
  7. Confirm the winning cell with a quick follow up test against your old champion, so you know the lift is real before you bet the budget on it.

Here is the same drill as a decision you can hang on the wall.

The whole method fits on that one card. When two changes are strangers, test them in sequence and save the effort. When they might talk, put them in the same room and watch them together. If you want a second pair of eyes on which of your changes are strangers and which are entangled, that is exactly the sort of thing worth a quick call to sketch out before you burn a month of traffic on the wrong test.

The one idea to leave with

If you take a single thing from this, make it this. Testing one variable at a time quietly assumes your changes add up, and the whole reason interactions cost you money is that they do not. Your best headline and your best price, each a clear winner on its own, can combine into a campaign that loses to a pairing you would never have tried. The only way to see it is to run the combinations together and read the whole grid, not the edges. Test the things that talk to each other in the same experiment, and you will stop shipping confident losers and start finding the cell everyone else walked straight past.

If you have a drawer full of A/B results and a nagging feeling your winners do not add up the way the tests promised, let us design a full factorial experiment that fits your traffic and finds the combination that actually pays. Book a call and bring your last three tests.

1%of every invoice goes to a UK charity you pick.

A donation, never sponsorship. You choose the cause at onboarding.

The story behind the pledge →

Stay ahead of your competition.

The latest innovative products and services, straight to your inbox before your competitors hear about them.

Get up to 5% off your first six months: 1% per topic you pick, the full 5% when you take everything. Limited offer · ends 31 December 2026.

New clients only. Terms apply.

* Up to 5% off your first six monthly invoices, new clients only. Full terms.

Questions about pricing, contracts or how we work together?

Read the FAQ