The First Data Science Month: What Actually Happens Before Any Model

Somebody sold you data science and you got a dashboard. I hear a version of this story most months. A business owner signed up for analytics, or AI, or growth intelligence, or whatever the agency called it that quarter, and six months later they have a screen full of charts, a monthly invoice, and no decision they make differently because of it. The dashboard updates. Nobody looks at it. The person who built it has moved on to the next client. So when I say the first month of working with me on data produces no model, no dashboard and no AI, that is not a warning, it is the point. The first month is where the questions get asked, the plumbing gets inspected, and the numbers that are quietly wrong get found before anything is built on top of them. This is what those four weeks actually contain, in the order they happen.

Somebody sold you data science and you got a dashboard

I hear a version of this story most months. A business owner signed up for analytics, or AI, or growth intelligence, or whatever the agency called it that quarter, and six months later they have a screen full of charts, a monthly invoice, and no decision they make differently because of it. The dashboard updates. Nobody looks at it. The person who built it has moved on to the next client.

So when I say the first month of working with me on data produces no model, no dashboard and no AI, that is not a warning. It is the point. The first month is where the questions get asked, the plumbing gets inspected, and the numbers that are quietly wrong get found before anything is built on top of them. Skip it and everything you build afterwards inherits the errors. I have rebuilt enough forecasting models on top of broken order data to know that the expensive mistakes are all made in the first fortnight, usually by not having a first fortnight.

This is what those four weeks actually contain, in the order they happen. I run the same shape for a shop doing eight hundred thousand a year and for a SaaS with two thousand accounts, because the shape is about decisions, not data volume. The data science page tells you what you get at the end. This is the bit before.

Week one starts with your diary, not your database

The first meeting has no laptop open. I ask the owner, and whoever else actually decides things, one question in several forms: what do you decide, and how often?

Weekly decisions in a typical shop: how much to spend on ads tomorrow and where, which products to push in the newsletter, whether to reorder stock, whether to discount what is not moving. Monthly: which channels to keep, which products to drop, whether to hire, whether to open a new market. Quarterly: budgets, pricing, the big bets. For a SaaS the list is different but the structure is the same. Which features to build, which plan to push, who to call before they churn.

Then the follow up, which is where the real work starts: what would you do differently if you knew X? An owner tells me she would double the Meta budget if she knew it was profitable. Fine. What does profitable mean, contribution margin after returns and after the ad cost, or revenue? Over what window, the first order or the first year? Attributed how? She has never been asked. Nobody has, because the agency that built the dashboard needed a metric to show going up and picked ROAS, and ROAS answers none of those questions.

By the end of that meeting we have a list. It is never longer than ten items and I make the owner rank it. The top three become the whole of month one. Everything else waits, however interesting it is, because the top three are the ones that move money.

For what it is worth, across a few dozen of these the top three are nearly always some version of the same questions. Which customers come back and what did they cost to acquire. Which channel is actually profitable after returns and after margin, not before. And what is going to run out, of stock or of cash or of customers, before I notice. Yours may differ. In my experience they do not differ much.

The tracking audit, or why your numbers disagree with each other

Week one, second half, I open the laptop. The first job is an audit of everything that claims to measure something. GA4, the ads platforms, the shop's own reports, the email tool, and any dashboards someone built before me.

I do this by reconciling. Pick one week. Count the orders in the shop database. Count them in GA4. Count them in the Meta Ads Manager, in Google Ads, in Klaviyo or whatever sends the newsletter. Add up the numbers the platforms claim. In every audit I have ever run, the platforms together claim more orders than the shop actually took. Sometimes half as many again. Each platform is honestly reporting what its own pixel saw, and its pixel saw a person who clicked a Google ad on Monday, opened a newsletter on Wednesday and came through a Meta retargeting ad on Friday, and all three take the credit.

That is the double counting. The missing counting is worse. Since consent banners became real in the DACH region, somewhere between a third and half of visitors decline tracking, and Google's consent mode fills the gap with modelled conversions that are, to put it kindly, a guess. A client's GA4 said their conversion rate had fallen by a third over eighteen months. It had not. Their consent rate had. The shop database, which does not need consent to count an order, showed conversions flat. They had cut their ad budget based on a chart that measured cookie acceptance.

This is the point where I explain the difference between client side and server side tracking, and I will keep it short because I wrote a whole guide on it. The short version: the browser is a hostile environment for measurement, and the only numbers you can fully trust are the ones your own server recorded. Everything else is a sample. Useful, but a sample, and month one is where we write down exactly how big the sample is so nobody mistakes it for the population again.

Then the event taxonomy. I list every event the site fires and what it is called. On a site of any age this list is archaeological. There will be a purchase event, a Purchase event, a transaction_complete event from a plugin that was removed in 2023 but whose tag is still in the container, and an add_to_cart that fires twice on the product page because two people added it in different years. I do not fix any of this in week one. I document it, because the fix has to wait until we know which events the top three questions actually need. Often that is four events, and the other thirty can be deleted.

And the UTMs. Somebody in every company has, at some point, tagged a campaign as "Newsletter", someone else as "newsletter", a third as "email" and a fourth as nothing at all. Google Ads auto tagging overrides the manual ones in some reports and not in others. I have a separate piece on getting UTMs right and in month one I do the unglamorous thing: export every campaign name from the last year, group the variants, and produce a mapping table. It is a spreadsheet. It is the most valuable spreadsheet the company will own for a while.

The data inventory, including the spreadsheets nobody admits to

Alongside the audit I build an inventory of every place data lives. The list is always longer than the owner expects.

The shop database is the obvious one, and it is the one I trust most, because it recorded every order whether or not anyone consented to anything. The payment provider has the truth about what money actually arrived, refunds included, which the shop database sometimes does not. The ad platforms have spend by day and by campaign, which nothing else has. The email tool has who opened what. The ERP, if there is one, has cost prices and stock, which is the difference between knowing revenue and knowing margin. The accounting system has the returns that came back through the door instead of through the shop.

And then, around day four, someone mentions the spreadsheet. There is always a spreadsheet. The purchasing manager keeps one with supplier lead times. The founder keeps one with the real ad budgets, because the platform numbers include a test account. Customer service keeps one of complaints by product. None of these are in any system, all of them are the answer to one of the top three questions, and none of them were mentioned in the first meeting because nobody thinks of a spreadsheet as data. It is data. Usually it is the best data in the building.

I write down, for each source, what it holds, who owns it, how I get access, how often it updates, and what its unit is. That last one matters more than it sounds. The shop counts orders. The payment provider counts transactions, and one order can be three transactions if it was partially refunded twice. The ad platform counts conversions, which are neither. If nobody writes this down, the first dashboard will add orders to transactions and call it sales.

What I need from you in week one

The owner's time in week one is about six hours. Two for the decisions conversation, one for the follow ups when I find things, and the rest spread over introductions to whoever owns each system. After that it drops to about an hour a week until the end of the month, when it goes back up for the review.

Access is read only everywhere. A read replica or a read only database user for the shop. Read access to the ad accounts, not admin. An export from the accounting system rather than a login to it. I never need write access to anything in month one and if someone offers it I say no, because the fastest way to lose the trust of a client's IT person is to be the contractor who could have changed something.

Two things I refuse to do, and I say so in the first meeting so there is no awkwardness later. I do not buy third party data to enrich your customer records. It is legally fragile under GDPR, it is usually wrong, and the improvement it delivers is smaller than fixing your own first party data, which is where month one goes anyway. And I do not fingerprint visitors to recover the ones who declined consent. Not because it does not work, it does, but because it is exactly the thing your consent banner promised you would not do, and the Datenschutzbehörde has started to agree.

Week two: the warehouse conversation, and the honest alternative

Around day eight the owner asks whether they need a data warehouse, because someone told them they do. The answer for most businesses under about ten million in revenue is: you need a warehouse in the sense of one place where the sources meet, and you do not need a Warehouse in the sense of a product with a sales team.

What most SMEs actually need is a Postgres database, separate from the shop, with a handful of tables: orders, order lines, customers, refunds, ad spend by day and campaign, email sends, and cost prices. Plus a few scheduled jobs that copy data in from each source every night. That is it. It runs on a server costing less than a business lunch a month. It holds years of data without noticing. And a Rails app or a plain SQL client can read it directly, which means the answers to the top three questions are a query away rather than a vendor away. I have written about the reasoning behind this setup before and I have not changed my mind.

BigQuery earns its place at a specific point, which is when you have GA4 raw event exports you want to join to orders, or your daily event volume is in the millions, or you have several people writing analysis at once and Postgres starts to feel it. If you are on Solidus that join is a known path and I documented it in the BigQuery and Solidus piece. Most clients do not reach that point in year one. Some never do. Starting on BigQuery because you might need it later is how people end up with a monthly Google Cloud bill and nobody who can explain it.

Week two is when the nightly jobs get written and the first copies land. It is boring and I like it, because at the end of it there is, for the first time, a single place where an order, its ad click, its email and its refund all sit in adjacent tables.

Identity stitching and where it stops

The question underneath most of the top three is: is this the same person? The same person who clicked the ad in March and bought in May. The same person who has a customer account and a guest checkout with a different email. The same person on the phone and the laptop.

Some of this is solvable. Email addresses can be normalised. A customer account with the same address as a guest order is probably the same household. A payment provider token can link two orders made with the same card. Doing this well recovers a surprising amount of repeat purchase behaviour that the shop's own reports miss, because the shop thinks every guest checkout is a new customer, and the actual repeat rate turns out to be higher than anyone believed.

Some of it is not solvable and should not be attempted. Matching a consented user to a non consented session is exactly what the consent was declining. Cross device matching without a login is fingerprinting by another name. I draw the line in writing in week two: here is what we join, here is the rule for each join, and here is what we deliberately leave unjoined. The document is short. It is also the thing I would hand to a data protection officer if one ever asked, and having it already written is worth more than the analysis it protects.

Definitions, or why every company disagrees with itself

The single most useful thing month one produces is a page of definitions. It sounds trivial. It has never once been trivial.

What is a customer? Someone who has placed an order, or someone with an account? Does a fully refunded order count? Does a B2B account with twelve users count as one or twelve? What is an order? Does a subscription renewal count as an order, and if so does the repeat rate become meaningless? What is a return? The item that came back, the refund that went out, or the exchange that was neither? What is revenue? With VAT or without, before discounts or after, before returns or after, on the order date or the shipping date or the payment date?

I ask these questions to the owner, the finance person and the marketing person separately, and I have never had three matching answers. Finance counts revenue net of VAT on the invoice date. Marketing counts it gross on the order date, because that is what the ad platform shows. The owner has a number in their head that is somewhere in between and is the one they use to decide things. When they all sit in the same meeting and look at the same month, they are looking at three different numbers and each thinks the other two are wrong.

Month one ends this. Not by picking the right definition, there is no right one, but by picking one per term and writing it down, with the reason. From then on the word revenue means one thing in this company. The arguments that used to happen about whether last month was good or bad simply stop, because there is no longer a way to have them.

A composite example, but one I have seen in several forms: a client believed their return rate was around nine percent, which is what the shop reported. The shop counted returns initiated through its portal. Returns that came back through customer service, and exchanges, and items refunded by the payment provider after a dispute, were in three other systems. The real rate was closer to nineteen. The entire ad budget had been set on margins that assumed nine. Nothing about the model that had been sold to them the year before was wrong, except its inputs, which were wrong by a factor of two.

Week three: the first answers, and the ones that surprise everyone

By week three the nightly jobs have run enough times to trust, the definitions exist, and I can start answering the top three. This is the first point where the owner sees anything, and I try to make sure the first thing they see is a number they did not expect.

Cohorts. Take every customer whose first order was in a given month and follow them. What share ordered again within ninety days, within a year? What did the group spend in total against what it cost to acquire? Laid out month by month this is the single chart that changes how owners think, because it turns "we did two million last year" into "the customers we acquired in March are worth this much and the ones from November are worth that much, and here is why". The March cohort came from search and reorders. The November cohort came from a discount campaign and never came back. Both looked identical in the monthly revenue chart.

Repeat rate, properly defined. After the identity work in week two, the repeat rate usually goes up, sometimes a lot, because guest checkouts stop looking like strangers. A client who thought fourteen percent of customers came back found it was twenty six. That changed what a new customer was worth, which changed what they could afford to pay for one, which changed the ad budget. Same data, same business, one join.

Margin by channel. This is the one that hurts. Take each acquisition channel, take the orders it brought, subtract cost of goods from the ERP, subtract returns using the honest return rate, subtract the ad spend, and look at what is left. The channel with the best ROAS is very often not the channel with the best contribution margin, because the discount codes that make a channel look good in the platform are invisible to the platform's own reporting. I have seen a marketplace channel that was the largest by revenue and negative by margin. Nobody had ever subtracted the fees.

These are not sophisticated analyses. They are arithmetic on clean, defined, joined data. That is why they were never done before. The sophistication was all being spent on the model, and the model did not have clean, defined, joined data to run on.

The moment the averages stop lying

Somewhere in week three the owner will ask why their average order value has been flat for two years while everything else changed. The answer is almost always that the average is made of two populations moving in opposite directions, and I usually pull out the piece I wrote on exactly this rather than repeat it.

But the practical version is this. Once the data is in one place, every number comes with a distribution instead of an average. Average order value becomes a histogram with two humps, the gift buyers and the stock up buyers, and the two humps respond to completely different marketing. Average delivery time becomes a long tail where eight percent of orders take more than a week and those eight percent are the ones leaving the reviews. Average customer value becomes a curve where the top tenth of customers are forty percent of the margin, and now the question of who to send the loyalty offer to has an answer.

None of this needs a model. It needs the data to be in a shape where you can ask for a distribution instead of a mean, and month one is where it gets into that shape.

The first dashboard is three numbers

At the end of week three I build a dashboard, and it has three numbers on it. Not thirty. Three.

Which three depends on the top three questions, but in a shop it is usually: repeat rate by cohort, contribution margin by channel this month, and days of stock or cash at current run rate. In a SaaS: net revenue retention, accounts showing the pre churn pattern this week, and payback period on the last quarter's acquisitions. Each number has a range on it, because I refuse to show a point estimate for something that is a sample, and the range is the honest answer to "how sure are we".

The reason it is three is that three numbers get looked at. I have watched thirty number dashboards being ignored in every company that had one, including ones I built early in my career before I learned better. An owner will look at three numbers every Monday morning for years. The thirty number version gets opened in the first week and then becomes a tab that is never closed and never read.

Everything else the data can answer is available as a question I can run, and in month two we decide which of those questions deserve to become the fourth number. Usually none of them do.

Week four: deciding what to build, and why forecasting is last

The final week is a decision meeting, and it is the first time in the month that anyone talks about building anything.

By now we have a list of what the data can answer today, what it could answer with a specific fix, and what it cannot answer without data the company does not have. The owner ranks the fixes the same way they ranked the questions. Usually the top of the list is unsexy: fix the four events that matter, delete the thirty that do not, get cost prices into the warehouse nightly instead of from a quarterly spreadsheet, tag campaigns consistently from now on. Each of those makes the three numbers more accurate, and accuracy compounds.

Then the owner asks about forecasting, because it is the thing they were originally sold, and I say: last. Not never, last. I wrote a full guide to forecasting and the first section of it is about data quality, because a forecast is a machine for extrapolating your definitions. If your return rate is wrong by a factor of two, the forecast will confidently predict a margin that does not exist, with a nice confidence interval around the wrong number. A forecast built in week one of an engagement is fiction with error bars. A forecast built in month three, on a warehouse with a year of clean history and definitions everyone agrees with, is a tool that budgets can be set from. The difference is not the algorithm. It is everything above this paragraph.

The same goes for the models people are actually excited about, the recommendation engines and the churn predictors and the dynamic pricing. They all need the same foundation and they all go wrong in the same way without it. Month one is not the delay before the interesting work. It is the part of the interesting work that decides whether the rest of it is real.

What you get at the end of month one, in writing

I finish the month with a document, because a month of work that lives in someone's head is not worth what was paid for it. It contains:

The decisions list, ranked, with the three we worked on and what we found. The tracking audit: every source, what it counts, where it double counts, where it misses, and by how much, with the consent rate written down so nobody mistakes GA4 for the truth again. The data inventory, spreadsheets included, with owners and access. The definitions page, one meaning per term, signed off by the owner. The join rules and the deliberate non joins. The nightly jobs, documented, running, and versioned in a repository the company owns. The three number dashboard, with ranges. And the ranked list of what to fix and build next, with rough effort against each.

Everything in it is yours. The warehouse is on your server, the code is in your repository, the document is in your portal. If we stop working together on day thirty one, you keep all of it and any competent developer can carry on from the document. That is not a generous gesture, it is how the concept and planning phase is meant to work for any project, and it is why I can offer the first month the way I do on the pricing page. The month is cheap for me to give because it produces something that is obviously either worth continuing or obviously not, and both of us can see which by the end of it.

The version of this that goes wrong

I should be honest about how this fails, because it does sometimes.

It fails when the owner does not have six hours in week one. Not because they are lazy, but because the business is on fire and the decisions conversation keeps getting pushed. Without it I am guessing at the questions, and a month spent answering questions I guessed is the dashboard nobody looks at, again.

It fails when access takes three weeks. A read only database user should take an afternoon. When it takes until day twenty because it has to go through an external agency that regards the database as theirs, month one becomes month two and the momentum goes.

And it fails when the definitions meeting produces a fight instead of a page, when finance and marketing would each rather keep their own number than agree on one. I can moderate that. I cannot resolve it, and a company that will not agree on what revenue means is not ready for a model that forecasts it.

None of those are common. All of them are visible by the end of week one, which is another argument for having one.

What it is not

It is not AI. Nothing in the first month uses a language model or a neural network or anything that would look impressive in a pitch. It is SQL, a few scheduled jobs, some spreadsheets that get retired, and a lot of conversations about what words mean.

It is not a dashboard project. A dashboard falls out of it, with three numbers, but the deliverable is the foundation the dashboard sits on, and that foundation is what makes month six's forecast worth reading.

And it is not the same as what you were sold last time, which is rather the point. If it were, you would already have the answers, and you would not be reading an article about why you do not.

If you want to know what your data can answer today, the first month is where we find out. Get in touch and we will start with your diary.

If you have a dashboard nobody looks at, or you are about to buy one, let's spend the first month on the questions instead. I will tell you what your data can answer today, what it cannot, and what it would take to close the gap.

1%of every invoice goes to a UK charity you pick.

A donation, never sponsorship. You choose the cause at onboarding.

The story behind the pledge →

Stay ahead of your competition.

The latest innovative products and services, straight to your inbox before your competitors hear about them.

Get up to 5% off your first six months: 1% per topic you pick, the full 5% when you take everything. Limited offer · ends 31 December 2026.

New clients only. Terms apply.

* Up to 5% off your first six monthly invoices, new clients only. Full terms.

Questions about pricing, contracts or how we work together?

Read the FAQ