Measurement & statistics
Holdout designs for SEO testing
Intervention-vs-holdout as the lift-claim bar: no trend-line-only causality.
Budget season, and the person across the table has one question about your best slide. How do you know it was you? The slide says traffic rose after you shipped. True sentence. Proves nothing.
That sentence has funded more SEO roadmaps than any other. It is a trend line with a story attached. In a market where Google ships updates, seasons turn and three other teams deployed the same month, that story has too many authors.
So the bar for a lift claim is fixed. Intervention against holdout: you change one set of pages, leave a comparable set alone, and read the difference. No causality from a trend line alone, and no exceptions for exciting charts.
See also: what generates the candidates this test promotes →
Why a trend line cannot testify
Because it assumes the world held still while you worked, and the world never does.
The same weeks that held your refresh batch also held an algorithm tremor, a demand season, and a template change another team shipped without telling you. A trend line cannot separate your work from that weather. It was never built to.
A holdout can. The pages you left alone lived through the identical weather without the treatment, so whatever separates the two groups afterwards is your work.
The authorship problem is worst exactly when you most want the credit. Big pushes ship near seasonal peaks and algorithm cycles, because everyone plans around the same calendar. So the market co-authors your best quarters far more than your quiet ones.
Three designs that clear the bar
All three answer the same question with increasing sophistication: what would have happened without you?
The plain holdout treats a batch of pages, holds back a comparable batch, compares their paths. Difference-in-differences subtracts each group's own history first, so a gap that predates your work cannot masquerade as an effect. Synthetic control is for when no clean holdout exists: you build a stand-in from untreated pages, weighted to match how the treated group used to behave.
Pick the simplest one your situation allows. Let the comparison do the arguing.
Comparability is the craft. A holdout built from low-traffic stragglers flatters your treatment automatically, and you will not notice, because the numbers look wonderful. Matching on traffic, intent and template keeps it honest, and that matching is most of the design work.
The designs scale down gracefully. A modest site can treat twenty pages, hold twenty matched back, and learn something real. The logic does not care about size. It cares about comparability.
No cloaking, no split-serving
One boundary is absolute. You split by page or page group, and that is the only way you split.
You do not serve crawlers different content than people, and you do not show variant content to slices of the same URL's search audience. Those tactics contaminate your test and break the engines' rules at once, a rare twofer, and buy you nothing.
Holdout designs need no tricks. The units are pages, the treatment is real work, and the engines see exactly what your users see.
Testing clean protects the result later, too. A lift measured on honest pages survives every policy change and every new crawler generation, because there was never a trick inside it for the ecosystem to catch up with.
Patience is part of the design
The reading date gets fixed before the test starts, not after you have seen the numbers.
Early signals arrive in days, interim reads in weeks. The significance test then runs on the readout itself, a statistical check on whether the change is bigger than the normal week-to-week wobble, at the same α = 0.05 bar every other claim in your reporting meets. Read a holdout too early and you have a coin flip with a methodology section.
The design pays at the moment it hurts. Early numbers look wonderful, somebody wants to announce, and the only answer available is the reading date. It does not negotiate.
It defuses panic in the other direction too. An early wobble downward is not a failed test. It is an unread one, and the point of fixing the date in advance is that neither hope nor fear gets to move it.
The same bar everywhere else
Every kind of work Quattr runs meets it, which is what stops it being a statistics-page opinion.
Refresh pilots read against held-out pages. Internal-linking batches, the work that runs first because it proves fastest, scale only on lift a holdout confirmed. Even paid spend faces it, through brand-suppression holdouts that measure what a brand click buys.
The Quattr Method page shows that standard under all eight kinds of work. The article on correlation and causation in search analytics explains what happens before the test: co-movement generates your candidates, and this bar decides which become causes.
Between the two, nothing reaches a deck on charisma.
Ask it yourself
Design a holdout test for our refresh batch: which pages are treated, which are held out, and when do we read the result?
See also: the eight kinds of work, and the proof standard under all of them →