SEO reporting
Significance-gated SEO reporting: only real changes make the deck
Why 2 of your 5 headline movers should stay off the deck.
Last month your headline slide said the number was up. This month the same number is down by roughly as much. Somebody in the room has already done the subtraction while you were still talking.
Nothing was falsified. Your mover was real in the arithmetic sense and never real in the statistical one, and on the way back down it takes a piece of your credibility with it.
Pull five headline movers for a monthly deck and, in a typical month, a couple of them are variance wearing a trend costume. The fix is one rule, applied every month, without negotiation. No delta headlines your deck without first passing a statistical test, which is a check on whether the change is bigger than that metric's normal week-to-week wobble.
The rule, stated plainly
Any change that could drive a decision gets tested, at α = 0.05.
Movers that pass can headline. Movers that fail either stay off your deck or appear labelled as not yet significant, in a watch list, never as news.
Decision-driving is the trigger phrase, and it does real work. A number that decorates a slide does not need the test. Any delta that could move budget, headcount or roadmap gets it automatically. That is precisely where a wrong number turns expensive.
The α = 0.05 default sits in the open where anyone can argue with it, which is the point of stating it at all. It means one false alarm in twenty is the price of being allowed to speak.
Every reporting culture pays that price. Most never name it, and the ones that do get a far easier conversation in the room the first time a headline reverses.
See also: how the test itself works →
Fewer claims, each worth more
Because your first walk-back is expensive. And permanent.
Once an executive has caught a reported trend reversing itself, every later slide of yours gets read with one eyebrow up, and there is no run of good months that buys the eyebrow back down again. The compounding runs the other way too. A deck that has never had to retract a headline earns the default trust that makes the hard months survivable.
You are trading volume for weight. Three real movers land harder than five maybes, and your audience learns the difference faster than you would expect.
There is an internal effect nobody predicts. Your analysts stop pre-writing narratives for movers that might not survive the test, and that time flows into the movers that did. Half of the monthly writing hours were always the borderline cases. Testing first deletes that work before it starts.
Year over year is where it bites
Because seasonality hands everybody a plausible story in both directions.
Annual comparisons are the most abused numbers in search reporting. The same rule applies to them without softening. The year-over-year delta takes the test, and the result ships attached to the number rather than to a narrative written in advance.
The worked example below shows that comparison arriving with its statistical result already on it, rather than with a story bolted on afterwards.
The other chronic offender is the small segment, where a handful of clicks swings a percentage wildly. Small samples rarely reach significance. So the rule catches those by construction, and your deck stops carrying triple-digit percentage swings built on twelve clicks.
The workflow that does this: Year-over-year, with verdict →
Gating is not hiding
Failed movers do not vanish.
They sit in the appendix with their p-values, labelled as watch items, so nobody can mistake discipline for concealment. If one becomes real next month, it graduates to your headlines with its history attached. That is a better story than it would have been the first time.
Visibility is where the rule gets its authority. Everyone can see what you tested, what passed and what is still cooking. A silent filter would invite exactly the suspicion this exists to end.
Watch items earn unexpected respect over time. Executives learn that they leave that list with evidence attached, and your appendix turns from a graveyard into a preview.
Ask it yourself
Which of this month's movers are statistically significant, and which should stay off the deck?
What this does to your deck
It makes your headline slides survive a fact-check.
Scope and freshness on every slide are the rest of that armour, and the executive readout article covers them. Freshness here means the date the data was last complete, printed where a reader can see it. Testing first supplies the part that matters most: headlines that stay true after the meeting ends.
The significance-referee skill, a packaged analysis you run by name in your assistant, runs the test on demand.
One move makes the whole thing self-sustaining. Put the result on the slide next to the number it defends. Readers learn the grammar in a single meeting, and after that an untested number looks naked on the page.