Measurement & statistics
Real or noise: statistical significance for search metrics, in plain language
p-values and confidence intervals in plain language, before the panic escalates.
By the time the chart reached Tuesday's standup, three people had a theory. The metric had dipped on Friday. Nobody had checked whether the dip was bigger than that metric's ordinary week-to-week wobble, and by the following Tuesday it had reversed itself.
That afternoon of theories is this article's subject, and one question would have saved it. Real or noise? It is a statistical check on whether your change is bigger than the normal weekly wobble, and it runs before anyone escalates, so your hunt for an explanation starts only when there is something to explain.
The test compares the change to its weather
Every metric has natural weekly variance. Call it the weather.
The job is to ask whether this week's swing is bigger than the weather can explain. The check compares your change against that metric's own history and returns a p-value: the probability of seeing a swing this size by chance alone.
The default bar is α = 0.05. Below it, we call the change real. Above it, the answer is noise, and noise is a complete answer rather than a failed one.
Variance differs by metric and by site, so the bar is fitted to your own history rather than to a rule of thumb. A metric of yours that swings hard in ordinary weeks needs a bigger move to clear its bar than one that barely breathes. The significance-referee skill, a packaged analysis you run by name in your assistant, does this as a conversation.
What a p-value does not buy
It buys you discipline. It does not buy truth, and the distance between those two is where most misreadings live.
Your answer comes with a confidence interval, the honest range the true change plausibly sits in. A drop reported as somewhere between minus nine and plus four percent is a very different meeting from a drop reported as twelve.
Two readings keep it honest. Significant is not the same as important, because a tiny change can be statistically real and still not worth your slide. And a noise verdict does not mean nothing happened. It means your data cannot yet tell this apart from ordinary variation, which argues for patience while the evidence accumulates.
The bar itself is a convention and we state it as one. α = 0.05 means accepting a one-in-twenty false alarm rate. Stricter bars exist for costlier decisions, and picking one is a real choice rather than a formality. The habit matters more than the threshold: pick your bar before the number arrives, not after you see which way it went.
In practice the interval persuades where the label cannot. A stakeholder who resists being told something is noise will still concede that a range straddling zero is not a story yet.
Is this change real or noise?
Analysis plan, every step, resolved for you
- Routed to the governed workflow
significance-referee - Scope resolved before filtering
segment: Overall Market · Jul 20 to Jul 26, 2026 - 1 analysis queued
significance_check - Honesty gates armed
significance α=0.05 · scope echo · freshness
Clicks dipped week over week, but the change sits inside normal weekly variance.
- Verdict
- Noise, not signal
- p-value
- 0.31
- 95% CI
- −9% to +4%
Wait for more data before reorganizing anything.
Jul 20 to Jul 26, 2026 vs prior weekorganicsegment: Overall Market
⚠ Observationalas of Jul 28, 2026 (GSC lags 1 to 2 days)IllustrativeOpen in Quattr ↗
Noise, not signal. The deck stays quiet this week, and that's the discipline.
The honest answer is often wait
The demo above shows the output most tools will not sell you: not significant, p = 0.31, wait.
It is the least dramatic thing analytics produces and the most valuable, because every false alarm you skip is a day of your team's attention handed back.
Waiting is a decision, and this makes it defensible. When your boss asks why nobody is reacting, the answer is a p-value.
False alarms cost more than afternoons. Each one teaches your team to discount the dashboard, so by the quarter something real happens, the alarm has no audience left. This protects its credibility as much as your team's time.
Spikes get the same test as drops
It is symmetric on purpose. A jump in conversions, counted in the named goals your team defined in your analytics tool (Purchase, Demo request), gets the same test before anyone declares victory.
Celebrating variance is how your team learns to distrust its own reporting, and it happens more often than the reverse. Nobody volunteers to check a good month.
The example below shows a conversion spike getting exactly that treatment: the plan, the test, and an answer before the congratulations.
The symmetry also catches the pleasant lie. A good month that was really a calendar artifact, once claimed in a quarterly review, becomes a debt the next quarter collects with interest.
See also: why conversions always mean the goals your team named →
The workflow that does this: Conversion spike: real? →
What the rest of reporting leans on
Quattr runs this check automatically on any change that drives a decision, and no causal language survives without it.
That has consequences downstream. Only real changes make your deck, and the deeper statistics live in the correlation and holdout articles, where a holdout means changing some pages, leaving comparable ones alone, and reading the difference.
It also sets the tone for everything the assistant says next. A real result opens the where-does-it-concentrate investigation. A noise result closes the thread in one line, and it stays closed until the data reopens it.
One phrase arms the whole thing. Ask it about any number that surprises you.
Ask it yourself
Is that change significant, real or noise?
See also: holdout designs: change some pages, leave comparable ones alone →
Frequently asked
- Is 0.05 the right bar for us?
- It is a convention, not a law, and stricter bars make sense when a decision is expensive to get wrong. What matters more than the number is picking it before you see the result. A bar chosen after the fact is not a test, it is a justification.
- The answer came back noise, but I am sure something changed. Now what?
- Wait for more data, or look at a slice where the sample is thicker. Noise does not mean nothing happened. It means this data cannot yet tell your change apart from an ordinary week, and acting anyway means acting on a hunch while calling it a finding.
- Does a real result mean our work caused the change?
- No. It means the change is bigger than ordinary variation. What caused it is a separate question, and answering it properly needs a controlled design rather than a timeline. The test tells you something is worth explaining, not what the explanation is.