The Long Tail of Keywords Is Draining Your Ad Budget

Open almost any Google Ads account that has been running for a year and you will find the same thing: a short head of ten or twenty keywords that get most of the traffic, and then a long, long tail of hundreds of keywords that each get a handful of clicks a month. And here is what nearly everyone does with that tail. They sort it by conversion rate, they pause the ones showing zero, and they pour budget into the ones showing a big number. It feels like diligent optimisation. It is actually statistical self harm, because a keyword with nine clicks and one conversion does not have an eleven percent conversion rate in any meaningful sense. It has almost no information at all, and you are making budget decisions off pure noise. I want to show you why the numbers on your long tail keywords are mostly random, why reacting to them week by week quietly drains your budget, and what to do instead, which is to stop judging cousin keywords on their own and start bidding them as families.

The tail is bigger than you think

Open almost any Google Ads account that has been running for a year and you will find the same shape. A short head of ten or twenty keywords that get most of the traffic, the ones everyone watches. And then a long, long tail of hundreds of keywords that each get a handful of clicks a month. Five clicks here, twelve there, three over there. Individually they look like rounding errors. Collectively they are often a third or more of your spend, and they are the part of the account nobody can actually manage, because there is never enough data on any single one of them to know what it is doing.

So people do the natural thing. They sort the tail by conversion rate. They pause the keywords showing zero. They pour budget into the ones showing a big fat number. It feels like diligent optimisation. Green means go, red means cut. It is actually statistical self harm, and I want to walk you through exactly why, because once you see the maths you will never trust a small sample conversion rate again.

The shape has a name. It is the same power law you see everywhere, a few winners doing most of the work and a vast crowd of stragglers each doing a little. In search that crowd is not a bug, it is where the good value intent hides, the oddly specific phrases a real buyer types when they are close to the till. You want the tail. What you do not want is to manage it one hair at a time, because the tail is precisely the region where the numbers are too small to mean anything.

Here is roughly what that shape looks like on a made up but painfully typical account. The exact figures do not matter, the pattern does.

BandShare of keywordsShare of clicksTypical clicks per keyword per month
Head3 percent55 percenthundreds
Body12 percent30 percentdozens
Tail85 percent15 percenta handful

Read the last row and let it sink in. Eighty five percent of your keywords, the overwhelming majority of the things you are asked to have an opinion about every week, live in a band where each one gets a handful of clicks a month. Fifteen percent of the clicks, spread so thin that no single keyword ever accumulates enough to be read honestly. That is not a reporting problem you can fix with a better dashboard. It is a data problem, and the only cure is to stop reading each row on its own.

A keyword with nine clicks knows nothing

Here is the uncomfortable truth. A keyword with nine clicks and one conversion does not have an eleven percent conversion rate. It has one conversion. The eleven percent is something you calculated, but it is built on so little data that it is almost pure chance whether that one conversion happened this month or not.

Let me make that precise, because precise is the only thing that changes minds. When you measure a proportion, a conversion rate, from a sample, the rough uncertainty around it is captured by the standard error.

SE≈p(1−p)nSE \approx \sqrt{\frac{p(1-p)}{n}}

Here pp is the conversion rate you observed and nn is the number of clicks. The important thing is the nn underneath. It sits inside a square root in the denominator, so the fewer clicks you have, the larger the uncertainty, and it grows fast as nn shrinks.

Take a long tail keyword with one conversion in twenty clicks. That is p=0.05p = 0.05 and n=20n = 20.

SE≈0.05×0.9520=0.047520=0.002375≈0.049SE \approx \sqrt{\frac{0.05 \times 0.95}{20}} = \sqrt{\frac{0.0475}{20}} = \sqrt{0.002375} \approx 0.049

So the standard error is about 4.9 percentage points, sitting under a measured rate of 5 percent. Read that again. The uncertainty is almost as large as the number itself. A rough confidence band would run from basically zero all the way up to around fifteen percent. That keyword is not telling you it converts at 5 percent. It is telling you it converts at somewhere between nothing and lots, and it has no idea which. You cannot bid on that. Nobody can bid on that.

Now take a head keyword with fifty conversions in a thousand clicks. Same measured rate, p=0.05p = 0.05, but n=1000n = 1000.

SE≈0.05×0.951000=0.04751000=0.0000475≈0.007SE \approx \sqrt{\frac{0.05 \times 0.95}{1000}} = \sqrt{\frac{0.0475}{1000}} = \sqrt{0.0000475} \approx 0.007

Now the standard error is about 0.7 of a percentage point. Same five percent on the label, but this time the band is roughly 3.6 to 6.4 percent. That is a number you can actually plan around. Same headline rate, wildly different trustworthiness, and the only thing that changed was the number of clicks underneath it.

That is the first thing worth sitting with. Two keywords can show you the identical conversion rate, and one of them is a fact while the other is a rumour. The dashboard prints them in the same font and the same colour, and gives you no hint which is which.

The standard error is a dimmer, not a switch

Here is a way to feel the maths without doing the maths every time. The standard error does not flip from useless to trustworthy at some magic click count. It fades in slowly, like a dimmer, because the clicks sit under a square root. To halve your uncertainty you do not need twice the clicks, you need four times the clicks. To cut it to a third you need nine times. That square root is the whole reason the tail is hopeless one keyword at a time.

Watch it move. Same five percent rate every row, only the clicks change.

ClicksConversions at 5 percentStandard errorRough band
2014.9 points0 to 15 percent
50about 2 or 33.1 points0 to 11 percent
10052.2 points1 to 9 percent
400201.1 points3 to 7 percent
1000500.7 points4 to 6 percent

Run your eye down the last column. At twenty clicks the honest band is so wide it contains almost every rate you could care about, from a disaster to a triumph. You have to climb to a few hundred clicks before the band tightens into something you would risk money on. Your tail keywords live at the top of this table and stay there forever. They never earn the bottom rows, because they never get the clicks. That is the aha that reframes everything: the problem is not that your tail keywords are bad, it is that they are permanently stuck in the fuzzy top of this table, and no patience moves them down.

Why chasing the tail actively costs you money

If small sample rates were just useless, that would be one thing. You would ignore them and move on. But they are worse than useless, because acting on them systematically drains budget, and here is the mechanism.

When you sort a long tail by conversion rate, the keywords at the top are not the best keywords. They are the luckiest keywords. With a handful of clicks each, some of them caught a conversion by chance and now show a flattering rate. The ones at the bottom, the zeros, are not the worst keywords. Many of them are perfectly decent keywords that simply have not caught their conversion yet, because they have only had eleven clicks.

So you pause the zeros and pile budget into the lucky spikes. And then next month the thing statisticians have been shouting about for a century happens: regression to the mean. The lucky keyword, given more clicks, drifts back down towards its true unremarkable rate, and now you are overpaying for it. The keyword you paused would have drifted up towards a perfectly fine rate, but you will never know, because you killed it. You are buying high and selling low, every single month, on noise you mistook for signal. That is the second thing worth sitting with, and it is the one that actually empties the account.

Put a number on the regret

Let me make the cost of chasing noise concrete, because regression to the mean sounds like a lecture until it is money.

Imagine you have twenty tail keywords that all share one true conversion rate, say five percent. Give each of them twenty clicks this month. By pure chance some will show zero, some will show one conversion, a couple will show two, and one lucky soul will show three, a fat fifteen percent. You do what the dashboard begs you to do. You pause the six that showed zero and you triple the bid on the fifteen percent star.

Next month, given fresh clicks, every one of them drifts back towards five percent, because five percent is what they always were. The star was never a fifteen percent keyword, it just had a good month, so now you are paying a tripled bid for a five percent return. The six you paused were never zero percent keywords, they were five percent keywords having a slow month, and you will never collect their conversions because they are switched off. You paid a premium for the fakes and you binned the perfectly good. Do that every month and you are not optimising, you are running a machine that systematically buys high and sells low. That is the aha most owners never see: the weekly ritual of pausing zeros and boosting spikes is not neutral, it has a direction, and the direction is down.

The fix is to stop judging keywords alone

The way out is not more data on each keyword. You will never get it. The long tail is long precisely because each keyword is rare, and no amount of patience turns three clicks a week into a reliable rate. The way out is to stop treating each keyword as its own little experiment and start borrowing strength from its neighbours.

This is the idea behind what the marketing analytics people call a keyword cloud, and the sparse data problem it solves is laid out nicely in Cutting Edge Marketing Analytics by Venkatesan, Farris and Wilcox. The move is simple to say. Most of your long tail keywords are not really separate. They are cousins, different phrasings of the same underlying intent, bought by the same kind of customer. Group the cousins into a family, pool their clicks and conversions, and judge the family, because the family has enough data to be judged even when no single member does.

Here is a family of cousins from that espresso shop I keep using, all of them variations on one intent, the commercial buyer.

KeywordClicksConversionsConversion rate on its own
commercial espresso machine2428.3 percent
espresso machine for cafe1600.0 percent
professional espresso machine2114.8 percent
restaurant espresso machine1200.0 percent
office espresso machine1516.7 percent
espresso machine for business1218.3 percent
Family pooled10055.0 percent

Look at the individual rates first, the way your account shows them. They run from a flattering 8.3 percent down to a scary two big zeros. If you managed these one by one you would pause "espresso machine for cafe" and "restaurant espresso machine" tonight, and you would crank the bid on "commercial espresso machine". Every one of those decisions is built on somewhere between twelve and twenty four clicks. Every one of them is noise. We already did the maths: at twenty odd clicks the standard error is about five points, so an 8.3 and a 0.0 are completely consistent with all six keywords having the exact same true rate. They probably do.

Now look at the bottom row. Pooled, the family has 100 clicks and 5 conversions, a rate of 5 percent. Its standard error is

SE≈0.05×0.95100=0.000475≈0.022SE \approx \sqrt{\frac{0.05 \times 0.95}{100}} = \sqrt{0.000475} \approx 0.022

about 2.2 points, against nearly 5 for the singles. Still not a laboratory measurement, but finally something you can bid on. The zeros were never zeros. They were members of a five percent family that had not had their turn yet. That is the third thing worth sitting with: the family knows what no single keyword can know.

A second family, so you believe it

One family could be a fluke. So here is a second, a different intent from the same shop, the home buyer hunting a machine for the kitchen. Same trick, different cousins.

KeywordClicksConversionsConversion rate on its own
home espresso machine1815.6 percent
best espresso machine for home2229.1 percent
small espresso machine kitchen1400.0 percent
espresso machine for beginners1915.3 percent
easy espresso machine home1100.0 percent
espresso machine under 5001616.3 percent
Family pooled10055.0 percent

Notice how the individual rates jump around exactly as before, from a flattering nine percent to a pair of zeros, all of it on eleven to twenty two clicks, all of it noise. And notice how the pooled family lands, once again, on a number you can hold: five percent on a hundred clicks, standard error about 2.2 points.

But here is the twist that matters for bidding. This home family converts at five percent just like the commercial family, yet a home buyer is worth far less to the shop than a cafe kitting out with commercial gear. Same conversion rate, different value, and that is the point where a lot of accounts go wrong. The rate tells you how often you win. It does not tell you what winning is worth. You need both, and you attach the value at the family level too.

From family to one sane bid

Once you trust the family rate you can bid it, and you bid it as one number across the whole family, not six separate guesses. Say your worked value model gives you an allowable acquisition cost of 200 euros for a commercial buyer, the most you will pay to win one and still make your margin. The most you can pay per click is the conversion rate times that allowable cost.

max bid=0.05×200=10 euros per click\text{max bid} = 0.05 \times 200 = 10 \text{ euros per click}

One sane bid, ten euros a click, applied to every cousin in the family, resting on 100 clicks of evidence instead of twelve. Here is the whole move in a picture.

Every loose keyword on the left flows into the family in the middle, the family gives you a rate you can trust, and the rate gives you a single bid you apply to all of them. No keyword gets paused for a run of bad luck. No keyword gets a fortune thrown at it for a run of good luck. The whole family rises and falls together on evidence that is actually big enough to mean something.

Now stack the two families side by side and you can see why one flat bid across the account is madness.

FamilyPooled rateAllowable cost per saleMax bid per click
Commercial buyer5 percent200 euros10 euros
Home buyer5 percent60 euros3 euros

Identical conversion rates, and yet the right bid on one family is more than three times the right bid on the other, purely because the customer behind it is worth more. If you had bid a single account wide number you would either be starving the commercial family or overpaying wildly on the home family. The family is the unit where rate and value both become knowable, which is exactly why it is the unit you should bid.

The formula behind every bid

Let me write the whole thing as one line you can keep, because every sane bid in a search account is really just this.

max bid per click=conversion rate×value per conversion×target margin\text{max bid per click} = \text{conversion rate} \times \text{value per conversion} \times \text{target margin}

Conversion rate you now get from the family, not the keyword. Value per conversion you get from your own numbers, ideally real revenue fed back rather than a flat conversion flag. Target margin is how much of that value you are willing to hand to the platform, say you keep thirty percent as headroom and bid to the remaining seventy. Everything else in performance marketing is plumbing to make these three numbers honest. Get them from a family with a hundred clicks under it and you are bidding on evidence. Get them from a keyword with nine clicks and you are bidding on a rumour.

A worked mini case, three months on one account

Let me put the whole method on its feet with a small story. Numbers are made up but the shape is one I see constantly.

An owner comes in spending, say, three thousand euros a month on search, with two hundred keywords. The head, twenty keywords, is fine and well managed. The tail, one hundred and eighty keywords, is where the bleeding is. It eats about a third of the budget, a thousand euros a month, and every Monday someone spends an hour pausing the zeros and boosting last week's spikes.

Month one, we do not touch bids. We just read the search terms report like humans and sort that tail into four intent families: commercial buyers, home buyers, spare parts and repairs, and pure browsers who will never buy. Four names, one hundred and eighty keywords sorted into them. Already the browsers family, pooled, shows a conversion rate near zero on a few hundred clicks, and now that is a fact rather than a guess, so we can cut it with a clear conscience. That one honest cut hands back real budget on day one.

Month two, we pool clicks and conversions per family, work out a pooled rate and a value per family, and set one bid per family from the formula. No keyword gets paused for a bad week. No keyword gets a fortune for a good week. The commercial family, worth the most, gets the confident bid it deserves. The home family gets a smaller bid that matches its smaller value. The weekly panic ritual stops, because there is nothing to panic about, the families move slowly and honestly.

Month three, a couple of individual keywords in the commercial family have quietly piled up a few hundred clicks each. Their own standard error is now small, so we graduate them out of the family and bid them on their own merits, a little higher or lower than the family as their real rate now justifies. Everything else stays pooled. The account has gone from one hundred and eighty coin flips a week to four honest bids plus a handful of graduates, and the thousand euros that was sloshing around on noise is now either bidding on evidence or switched off. Nobody added a data scientist. We just stopped pretending each hair on the tail was a data point.

Shrinking each keyword towards its family

There is a more refined version of this move, and it is worth knowing even if you never hand roll it, because it is exactly what good automated bidding does under the hood. Instead of either trusting a keyword's own wild rate or replacing it entirely with the family rate, you blend the two, leaning on the family when the keyword has little data and letting the keyword speak up as it gathers clicks.

adjusted rate=w×keyword rate+(1−w)×family rate\text{adjusted rate} = w \times \text{keyword rate} + (1 - w) \times \text{family rate}

The weight is where the cleverness sits.

w=nn+kw = \frac{n}{n + k}

Here nn is the keyword's own clicks and kk is a smoothing constant, roughly the number of clicks at which you start to half believe the keyword over its family. When the keyword has almost no clicks, ww is near zero and the adjusted rate is basically the family rate, which is the honest thing to do when you know nothing. As the keyword piles up clicks, ww climbs towards one and the keyword's own rate takes over, which is the honest thing to do once it has earned it.

Take our zero click orphan, "espresso machine for cafe", sixteen clicks, zero conversions, apparent rate zero percent. With a smoothing constant of, say, fifty, its weight is sixteen over sixty six, about 0.24. Its adjusted rate is 0.24 times zero plus 0.76 times five percent, which is about 3.8 percent. Not zero. Not the full family five either. A sober blend that says this keyword has shown me nothing alarming yet, so I will treat it as a slightly hesitant member of a five percent family. That is the number you bid, and it updates itself gently as clicks arrive. This is the whole idea of borrowing strength, written as one line of arithmetic, and it is the aha that ties the piece together: you never have to choose between the lying keyword and the blunt family, you can stand exactly as far between them as the data earns.

The common mistakes

I see the same handful of errors on nearly every account, so let me name them so you can catch yourself.

The first is grouping by string instead of by intent. A script that buckets keywords because they share the word "espresso" will happily throw a commercial buyer and a curious student into the same family, and now your pooled rate is a blend of two things that should never have been mixed. Families are about who is behind the query, not which words it contains. That is human work, and it is the work that matters most.

The second is pooling things with wildly different values. If you put your five percent home buyers and your five percent commercial buyers in one family because the rates match, you will bid a single number that is wrong for both. Rate can be pooled. Value must be respected. A family should be one intent and roughly one worth.

The third is never graduating anybody. Pooling is a way to survive a data shortage, not a permanent vow of poverty. When a keyword finally has the clicks to stand on its own, let it. An account that pools forever leaves money on the table just as surely as one that judges everything one click at a time.

The fourth, and the most expensive, is reacting weekly. The whole disease is the Monday ritual of nudging bids off last week's noise. Families move slowly precisely so you can leave them alone. Check them monthly, not daily. If you find yourself touching a bid because of what happened over the last twelve clicks, you have quietly walked back into the trap.

How to apply this on Monday morning

You do not need a project, you need an hour and a spreadsheet. Here is the whole thing as a short checklist.

Pull your search terms report for the last ninety days, not the last week, because you want enough clicks in each family to matter. Sort the tail, everything below the point where a single keyword has a few hundred clicks, into a small number of intent families. Aim for a handful, not fifty. Give each family a plain name that says who is behind it.

For each family, add up the clicks and the conversions and work out the pooled rate. Attach a value per conversion to each family from your own margins, not a guess. Run the one line bid formula, rate times value times the margin you are willing to give up, and set that single bid across every keyword in the family. Switch off any family whose pooled rate is honestly near zero on real click volume, because now you actually know.

Then leave it alone for a month. When you come back, look for individual keywords that have gathered enough clicks to graduate, promote those, and re read the families for any that have drifted. That is the loop. Ninety days of data in, a handful of honest bids out, reviewed monthly instead of poked daily.

Here is the loop as one picture.

Feeding it back to the machine

There is one more move that turns all of this from a monthly spreadsheet chore into something that compounds, and it is the one most owners skip. Feed real values back to the platform, not just conversion counts. Google and Meta will bid to a value you send them, and if you send them a value per family rather than a flat conversion flag, their automated bidding starts pooling data the same way you just did by hand, but continuously and at a scale no spreadsheet can match. Your families become the training signal, and the machine does the borrowing of strength for you.

This is exactly the kind of plumbing that pays for itself, and it is what my performance marketing work is built around. It also depends on owning your own customer data, on knowing the real margin behind each sale rather than a flat conversion flag, which is one more reason I keep pushing owners towards a shop and stack they control rather than a rented box that forgets where every customer came from. The families are only as honest as the values you can feed them, and you can only feed real values if you actually own the numbers.

The one idea to leave with

Stop asking what each little keyword converts at. With a handful of clicks the honest answer is that you do not know and cannot know, and every decision you make off that number is a coin flip you are paying for. Ask instead what the family of cousin keywords converts at, pool the clicks until the number means something, and bid the family with one sane figure. The long tail will stop draining your budget the moment you stop pretending each hair on it is a data point. If you want a hand turning your own tail into families that can actually be bid, that is squarely the kind of thing I do. Bring your search terms report and we will find the noise you have been paying for.

If your long tail keywords are being paused and boosted every week off nine clicks and a hunch, let us group them into families that actually carry enough data to bid on. Book a call and bring your search terms report, and I will show you how much of your spend is going on noise.

1%of every invoice goes to a UK charity you pick.

A donation, never sponsorship. You choose the cause at onboarding.

The story behind the pledge →

Stay ahead of your competition.

The latest innovative products and services, straight to your inbox before your competitors hear about them.

Get up to 5% off your first six months: 1% per topic you pick, the full 5% when you take everything. Limited offer · ends 31 December 2026.

New clients only. Terms apply.

* Up to 5% off your first six monthly invoices, new clients only. Full terms.

Questions about pricing, contracts or how we work together?

Read the FAQ