Loyalty with structural equation modelling

Five hotel managers, one budget, two proposals: a new breakfast concept or a service programme. Both sides have a correlation with repeat intent and neither number settles it, because food, service, value and satisfaction are tangled together in a guest's head. Structural equation modelling untangles them. It combines the factor analysis from chapter 12, which turns survey questions into clean constructs like trust and perceived value, with regression among those constructs, so you can read how much each experience moves loyalty, directly and through everything in between. I walk through the Seeblick Hotels model step by step: the path diagram with standardised coefficients, how to check the constructs are real before you trust the arrows, the table of direct, indirect and total effects that ends the breakfast argument, what fit measures mean in plain words, and how I translate a standardised path into a money range with honest error bars. I also cover sample size rules, the small sample alternative PLS, and the six ways SEM gets misused. The numbers are illustrative; the reasoning is what I run for real clients. This is part 18 of 21 of the Marketing Analytics series.

The breakfast argument

Every spring the five general managers of Seeblick Hotels sit in the same room in Bad Ischl and argue about the same thing: where does next year's investment go. This year the food and beverage lead wants €180,000 for a new breakfast concept across all five houses, with regional produce, a live station and a longer service window. The operations lead wants €120,000 for reception staffing and a proper service training programme. The owner wants neither of them to win on charm. She wants to know which one brings guests back.

Both sides have a chart. Food scores correlate with repeat intent at 0.41. Service scores correlate with repeat intent at 0.47. Neither number settles anything, because food, service, value and satisfaction are all tangled up with each other in a guest's head. A guest who loved the breakfast also rates the staff kindly and finds the price fair. Correlations tell you that things move together. They do not tell you which one to push.

Structural equation modelling is the method I reach for when a client has a survey, a theory about how the pieces connect, and a real budget decision hanging on it. It lets you draw the theory as a diagram, fit it to the data, and read off how much each lever moves loyalty, directly and through everything in between. The numbers are illustrative, but the shape of the argument is the one I run with hotel groups, SaaS companies and shops alike.

Where this sits in the series

This is chapter 18 of 21 and belongs to part four, media and loyalty. The previous chapter, Loyalty: the three Rs, the spectrum, and designing earn and burn, was about what a loyalty programme is for and how to design its economics. This one is about the psychology underneath: which experiences actually create loyalty and by how much. The next chapter, The customer loyalty journey: from segments to experiences, takes these drivers and turns them into a journey design per segment. The survey factors we use here come straight from chapter 12, Principal components and factor analysis for marketers, where 24 guest questions collapsed into seven clean constructs.

Two models glued together

Structural equation modelling, SEM from here on, is not one technique. It is two familiar techniques bolted together and estimated at the same time.

The first is a measurement model. Loyalty, satisfaction, trust and perceived value are not things you can observe. Nobody has a loyalty gland. What you observe are answers to questions: "I would book here again", "I would recommend Seeblick to a friend", "next time I will book directly rather than through a platform". Each question is a noisy indicator of the thing you care about. The measurement model says each indicator is a latent construct plus its own error, exactly like the factor analysis in chapter 12. The difference is that in SEM you decide in advance which indicators belong to which construct, because you have a theory, and the model then tests whether the data agrees.

The second is a structural model: plain regression, except the variables being regressed on each other are the latent constructs rather than observed columns. Service quality drives satisfaction. Satisfaction drives repeat intent. Trust drives repeat intent too. Every arrow is a hypothesis and gets a coefficient.

Why not skip the latents and regress an average of the loyalty questions on an average of the service questions? Because averages carry measurement error, and measurement error in a predictor biases its coefficient towards zero. A guest who was tired when they filled in the form drags down every score a little. SEM estimates how much of each indicator is signal and how much is noise, and then runs the regressions on the signal. In practice this often makes drivers look 20 to 40 percent stronger than the naive regression suggests, and, more importantly, it stops the noisiest construct from looking like the weakest driver just because it was measured badly.

The path diagram is the language everybody uses. Ovals are latent constructs. Rectangles are the survey questions. A single headed arrow is a hypothesised causal path. A double headed arrow is a correlation you allow without claiming direction. Once you can read one, you can draw your own theory on a whiteboard before a single number exists, which is the most useful part of the whole exercise.

The maths, kept honest

The measurement model says every observed indicator is a loading times its construct plus an error:

xi=λi ξ+δix_{i} = \lambda_{i}\,\xi + \delta_{i}

Here xix_i is the standardised answer to question ii, ξ\xi (xi) is the latent construct that question is supposed to measure, λi\lambda_i (lambda) is the loading, meaning how strongly the question reflects the construct, and δi\delta_i (delta) is everything else in that answer: mood, wording, the fact that the guest misread the scale. A loading of 0.85 means the construct explains about 72 percent of the variance in that question, because the squared loading is the explained share.

The structural model is a set of regressions among constructs. For Seeblick I use two. Satisfaction is driven by service, food and value:

ηSat=γ1 ξService+γ2 ξFood+β1 ηValue+ζ1\eta_{\text{Sat}} = \gamma_{1}\,\xi_{\text{Service}} + \gamma_{2}\,\xi_{\text{Food}} + \beta_{1}\,\eta_{\text{Value}} + \zeta_{1}

And loyalty is driven by satisfaction, trust and value:

ηLoy=β2 ηSat+β3 ηTrust+β4 ηValue+ζ2\eta_{\text{Loy}} = \beta_{2}\,\eta_{\text{Sat}} + \beta_{3}\,\eta_{\text{Trust}} + \beta_{4}\,\eta_{\text{Value}} + \zeta_{2}

The Greek letters follow the convention most software uses. ξ\xi constructs are exogenous: nothing in the model explains them, they are the levers. η\eta (eta) constructs are endogenous: the model explains them. γ\gamma (gamma) is a path from a lever to an outcome, β\beta (beta) is a path from one outcome to another, and ζ\zeta (zeta) is the part of each outcome the model does not explain. When everything is standardised, each coefficient reads as "one standard deviation more of this gives so many standard deviations more of that, holding the other arrows fixed".

The third formula is the one the general managers actually need. The total effect of a lever on loyalty is its direct arrow plus the sum of every indirect route, where each route is the product of the coefficients along it:

TotalService→Loy=0⏟direct+γ1β2⏟via Sat+γ5β3⏟via Trust\text{Total}_{\text{Service} \to \text{Loy}} = \underbrace{0}_{\text{direct}} + \underbrace{\gamma_{1}\beta_{2}}_{\text{via Sat}} + \underbrace{\gamma_{5}\beta_{3}}_{\text{via Trust}}

Service has no direct arrow to loyalty in my model. It works entirely through satisfaction and through trust (γ5\gamma_5 is the service to trust path). That is the point, not a weakness: it tells you how the experience turns into the behaviour.

The Seeblick path model

Seeblick sent its post stay survey to every guest for fourteen months and got 1,140 complete responses across the five houses. Chapter 12 gave us seven factors. I used them as the seven constructs below, and drew the structure before I looked at a single coefficient. The theory is simple hotel common sense: rooms and food shape whether the price feels fair, service and food and fairness shape overall satisfaction, service builds trust, and satisfaction, trust and fairness together decide whether a guest comes back and books direct.

Path diagram of the Seeblick loyalty model with standardised coefficients. Ovals are latent constructs; the indicator questions are omitted for readability.

Read the arrows as levers. One standard deviation more perceived service quality gives 0.38 standard deviations more satisfaction and 0.44 more trust. Satisfaction is the strongest single arrow into loyalty at 0.52. Value has both a route through satisfaction and a small direct route of 0.14, which in hotels usually reflects guests who felt the price was fair and would rebook out of habit even on a middling stay.

Reading the measurement model first

Before anyone trusts the arrows, I check that the constructs are real. The table shows three of the seven constructs. Every loading is standardised, so it runs from 0 to 1.

ConstructIndicatorLoadingComposite reliabilityAVE
ServiceStaff were warm and attentive0.840.880.71
ServiceProblems were solved quickly0.86
ServiceReception knew who I was0.82
ValueThe price was fair for what I got0.810.850.66
ValueI would pay this rate again0.83
ValueNo surprises on the bill0.79
LoyaltyI will stay at Seeblick again0.880.900.75
LoyaltyI would recommend Seeblick0.87
LoyaltyNext time I will book direct0.85

How to read it. Loadings above 0.7 mean the question is a good measure of its construct; below 0.5 I drop the question or move it. Composite reliability is the SEM cousin of Cronbach's alpha and should sit above 0.7. AVE, the average variance extracted, is the mean squared loading, and above 0.5 means the construct explains more of its indicators than error does. "No surprises on the bill" is the weakest item here at 0.79, still comfortably fine. If the loyalty construct had come back with a loading of 0.4 on "next time I will book direct", that would have told me direct booking is a different animal from loyalty, and I would have had to model it separately. It did not, which is itself a useful finding for a hotel fighting platform commissions.

Direct, indirect and total effects

Now the part that ends the breakfast argument. For each lever I add up every route to loyalty.

DriverDirect effectIndirect effectTotal effectMain indirect route
Satisfaction0.520.000.52none, it is the mediator
Service0.000.320.32via Satisfaction 0.20, via Trust 0.12
Value0.140.150.29via Satisfaction 0.15
Trust0.270.000.27none
Food0.000.200.20via Satisfaction 0.11, via Value 0.09
Room0.000.100.10via Value then Satisfaction 0.05, via Value direct 0.05

Line by line. Satisfaction is the biggest number and the least useful one: you cannot buy satisfaction directly, you buy the things that create it. Service totals 0.32, none of it direct. The route through satisfaction is 0.38 times 0.52, which is 0.20, and the route through trust is 0.44 times 0.27, which is 0.12. Value totals 0.29: a direct 0.14 plus 0.29 times 0.52 through satisfaction. Food, the star of the €180,000 proposal, totals 0.20: 0.22 times 0.52 through satisfaction gives 0.11, and 0.31 times 0.14 plus 0.31 times 0.29 times 0.52 through value gives another 0.09. Rooms come last at 0.10, working only through the sense of fair value.

Two things I want you to notice. First, food's effect on loyalty is real but roughly 60 percent of service's, and half of it runs through perceived value rather than pleasure. Guests do not come back for the breakfast; they come back because the breakfast made the room rate feel fair. Second, none of this would have appeared in the simple correlations, which had food and service almost level. The correlation of 0.41 for food was borrowing strength from service, because the houses with the best kitchens also have the longest serving staff.

Which driver matters most

Driver importance as total standardised effect on loyalty. Satisfaction is left out because it is the mediator, not a lever.

Importance alone is still not the decision. The second half of the picture is current performance. On a five point scale, Seeblick's mean scores were food 4.3, room 4.1, trust 4.0, service 3.8 and value 3.6. Food is already the best rated thing the hotels do. Service is the third biggest driver and the second worst rated experience, and value is the second biggest driver and the worst rated. That combination, high importance and low performance, is where money earns the most. It is the same importance versus performance logic as the driver analysis in chapter 12, only now the importance numbers are causal paths rather than correlations, which makes them safer to spend against.

Does the model fit, in plain words

There are a dozen fit statistics and every paper reports a different subset. Four are enough for a business decision.

MeasureSeeblick valueWhat it says in plain wordsRule of thumb
Chi square over degrees of freedom2.1How far the model's implied correlations sit from the observed ones, per parameter you did not estimateBelow 3
CFI0.95How much better this model is than assuming nothing correlates with anythingAbove 0.90, ideally 0.95
RMSEA0.048Average misfit per degree of freedom, penalising complexityBelow 0.06
SRMR0.041Average gap between observed and predicted correlationsBelow 0.08

Chi square on its own is nearly useless with 1,140 responses, because with a big sample it flags every trivial misfit as significant. The other three are what I report to clients. A CFI of 0.95 and an RMSEA below 0.05 means the structure I drew on the whiteboard reproduces the survey correlations well enough that I am not obviously missing an arrow. It does not mean the arrows are true. Fit says the model is consistent with the data; it never says it is the only model that is.

The investment decision

Here is the memo I wrote for the owner, condensed. A realistic service programme, based on what comparable groups have achieved, lifts perceived service by about 0.3 standard deviations in a year. Multiply by the total effect: 0.3 times 0.32 gives roughly 0.10 standard deviations of loyalty intent. The same spend on breakfast plausibly lifts food perception by 0.4 standard deviations, because the current score is already high and improvements at the top of a scale are hard, giving 0.4 times 0.20, which is 0.08. Service wins, and it costs a third less.

Then I translate intent into behaviour, because nobody banks a standard deviation. Across Seeblick's booking history, a one standard deviation rise in the loyalty construct went with about seven percentage points more guests rebooking within eighteen months. So 0.10 standard deviations is roughly 0.7 points on a base repeat rate of 31 percent, on around 21,000 room nights a year, with average direct booking revenue of €210 per night and about €45 saved per night that moves from a platform to direct. That is a defensible range of €120,000 to €200,000 a year of repeat and direct revenue from a €120,000 programme, with the honest caveat that the link from intent to actual rebooking is the weakest joint in the chain and I put wide error bars on it.

The risk I named: the service effect is driven mostly by two of the five houses, where reception turnover is highest. If the programme does not fix turnover, the coefficient will not move. Breakfast is the safer, smaller win. I recommended service first, breakfast in year two, and a repeat survey after twelve months to see whether the arrow moved.

How to actually run it

You need a survey with at least three indicators per construct, five or seven point scales, and enough responses. The rule of thumb is ten responses per estimated parameter, or at minimum 200 in total, whichever is larger. Seeblick's model has roughly 50 free parameters, so 1,140 is comfortable. You also need a theory drawn before the estimation, otherwise you will fit the data rather than test an idea.

The workflow in R with lavaan, which is free and what I use most:

library(lavaan)

model <- '
  # measurement model
  Service =~ svc1 + svc2 + svc3
  Food    =~ food1 + food2 + food3
  Room    =~ room1 + room2 + room3
  Value   =~ val1 + val2 + val3
  Trust   =~ tru1 + tru2 + tru3
  Sat     =~ sat1 + sat2 + sat3
  Loyalty =~ loy1 + loy2 + loy3

  # structural model
  Value   ~ Food + Room
  Trust   ~ Service
  Sat     ~ Service + Food + Value
  Loyalty ~ Sat + Trust + Value
'

fit <- sem(model, data = survey, estimator = "MLR", std.lv = TRUE)
summary(fit, standardized = TRUE, fit.measures = TRUE)

The MLR estimator gives robust standard errors, which matters because survey scales are never truly normal. std.lv = TRUE fixes each construct's variance at one so the loadings are comparable.

From clean survey data to a defensible model is about a week of work: two days on the measurement model, one day fitting and checking the structural model, one day on the effects and the money translation, one day writing it up. The checks before I trust it: every loading above 0.6, every AVE above 0.5, the square root of each construct's AVE larger than its correlation with any other construct so the constructs are genuinely distinct, CFI above 0.90 and RMSEA below 0.06, and the same structure fitting acceptably in at least two of the five houses separately. If the model only works pooled, one house is doing all the talking.

When the sample is small: PLS

Covariance based SEM, which is what I have described, wants a few hundred responses and behaves badly below that. A Linz law firm with 85 client survey responses, or a Werkbank cohort of 140 trial users, cannot run it honestly. For them I use partial least squares SEM.

PLS SEM builds each construct as a weighted composite of its indicators and then runs ordinary regressions among the composites, iterating until the weights stop changing. It is happy with 100 responses, it does not assume normality, and it copes with constructs that are formed by their indicators rather than reflected in them, for example "onboarding effort" in Werkbank being made of time to first invoice, support tickets and the number of settings changed. The price is that it has no global fit statistic worth the name, its path coefficients are biased slightly upwards, and it is a prediction tool more than a theory test. I use it to rank levers when the sample is small, and I say so in the memo. I do not use it to publish a causal claim.

Pitfalls

Cross sectional causality. Every arrow in the Seeblick model comes from one survey at one moment. The arrow from satisfaction to loyalty could partly run the other way: guests who already feel loyal rate their stay kindly. SEM cannot rescue you from that; only time can. Two waves of survey, or a survey linked to actual rebooking twelve months later, turn a plausible model into a defensible one.

Fitting the model to the data. Software will happily print modification indices telling you which extra arrow would improve fit most. Adding them one at a time until CFI looks nice is the SEM equivalent of p hacking. If you change the structure after seeing the data, say so, and treat the result as a new hypothesis for the next survey, not a finding.

Too many constructs for the sample. Each latent construct adds loadings, error variances and paths. A twelve construct model on 250 responses will estimate, but its standard errors will be wide enough to drive a bus through and it will fit the noise. Fewer constructs, measured well, beat many constructs measured once each.

Everything from one questionnaire. When the same guest answers every question in the same mood on the same afternoon, part of every correlation is that mood. This is common method variance and it inflates all the paths a little. Mixing in behavioural data, actual rebooking or actual spend, as at least one construct is the cleanest fix.

Reading standardised paths as euros. A path of 0.32 is not 32 percent of anything. It is a standard deviation ratio. The translation to money needs a separate, honest link from the loyalty construct to observed behaviour, with its own uncertainty. That step is where most SEM decks go quiet, and it is the step the owner actually cares about.

Skipping the measurement model. If two of your constructs are really one, or one of your constructs is really two, every structural coefficient is wrong. Twenty minutes with the loadings and the AVE table saves a week of arguing about arrows that do not mean what you think.

How I do this for clients

An SEM engagement starts with the free workshop, where I want three things on the table: the survey instrument you have or want, the business decision hanging on it, and any behavioural data we can tie to respondents, even if it is only whether they bought again. If the survey does not yet exist, we design it in the workshop with three indicators per construct and a theory drawn on the whiteboard before anyone writes a question.

Then two weeks of real work on me, before any commitment. In that time I clean the responses, build and test the measurement model, fit the structural model, run it per segment or per location, and translate the total effects into a money range with error bars. You get the path diagram with coefficients, the effects table, the fit summary in plain words, and a two page decision memo: which lever, expected impact as a range, the main risk, and what to measure to know whether it worked. You also get the model code and the cleaned data, because you own everything I build.

If you want the drivers monitored rather than measured once, I set up a small dashboard that answers one question, "which experience is moving loyalty this quarter", fed by your ongoing survey and your booking or order data. That belongs with my data science work, and where the recommendation turns into campaigns or a retention programme, my performance marketing side picks it up, with a commission based arrangement available when the outcome is measurable. What it costs is on the pricing page in plain terms: a fixed scope for the model and memo, a monthly figure if you want it kept alive.

Questions to ask before you spend

  • Which survey questions measure which construct, and did we decide that before or after looking at the data?
  • What are the loadings and the AVE for each construct, and did any item get dropped or moved?
  • Are the paths standardised, and if so what does one standard deviation of service actually look like in our operation?
  • What is the total effect of each lever, not just the direct arrow, and which indirect route carries most of it?
  • Where does our current performance sit against each driver, so we are spending on high importance and low performance?
  • How was loyalty intent linked to observed behaviour, and how wide is the error bar on that link?
  • Does the same structure hold in each location or segment, or only in the pooled sample?
  • When do we survey again to see whether the arrow moved?

This series is inspired by Mike Grigsby's Marketing Analytics (Kogan Page). The explanations, examples and numbers here are my own.

If you have a customer survey and a budget argument nobody can settle, send me the questionnaire and a sample of the responses, anonymised is fine. I will tell you within a few days whether the constructs hold up and whether the sample is big enough for a proper model or needs the lighter PLS route. We start with a free workshop where we draw your theory on a whiteboard, then I put two weeks of work in before you commit to anything. You get the path diagram, the effects table and a decision memo with a number and a risk. Everything I build stays with you.

1%of every invoice goes to a UK charity you pick.

A donation, never sponsorship. You choose the cause at onboarding.

The story behind the pledge →

Stay ahead of your competition.

The latest innovative products and services, straight to your inbox before your competitors hear about them.

Get up to 5% off your first six months: 1% per topic you pick, the full 5% when you take everything. Limited offer · ends 31 December 2026.

New clients only. Terms apply.

* Up to 5% off your first six monthly invoices, new clients only. Full terms.

Questions about pricing, contracts or how we work together?

Read the FAQ