Card comparing ranking signal evidence types and their limits. 4 plain facts about ranking signals comparison
Image: Search Ranking Tactics

Rules

4 plain facts about ranking signals comparison

Ranking signals comparison of what each kind of evidence can carry, why a correlation study cannot support advice, and the comparison worth running yourself.

Most published claims about ranking factors come from one kind of study: take a few thousand items that currently rank, measure some attributes, and report which attributes track position. The results get written up as advice.

The method cannot support advice, and the gap between what it measures and what it is used to claim is the single biggest source of confident nonsense in this field. This page compares the kinds of evidence you can actually get, and what each one will and will not carry.

What to take away

  • A correlation study describes what already ranks. It does not show what would happen if you changed something.
  • The sample in these studies is selected on the outcome, which is the one sampling choice that makes causal reading impossible.
  • The only evidence that supports a decision about your own site is a change you made on your own site, measured against something that did not change.

What a ranking factor study actually measures

The procedure is roughly: collect items in ranked positions, measure attributes the tool can see, compute a correlation between each attribute and position. Everything that follows depends on that sample.

The sample contains items that ranked. It contains nothing about the items that did not, so you are looking at survivors and reasoning about the population. What those attributes are standing in for is set out in the overview of ranking signals.

Attributes arrive in clusters, so the study cannot separate them. An operation with long pages also has an editor, a design budget, an established audience and a trusted domain.

Correlate any one with position and you will find something, because you are measuring the cluster. That is a textbook confounder, and sample size does not fix it.

Direction is the third problem. Several of the attributes people measure are consequences of ranking rather than causes of it. Something visible collects links, mentions and engagement because it is visible. The correlation is real and the arrow points the other way, which is the ordinary meaning of the warning that correlation does not imply causation.

Finally, the number is an average over wildly different queries and audiences. A relationship that is positive in one segment and negative in another can average to nothing, and a relationship that holds only in one segment can average to something. A single figure across a whole corpus hides both.

Comparing what each kind of evidence can carry

EvidenceWhat it can honestly supportWhat it cannot supportMain weakness
One person's before and afterA hypothesis worth testingAnything generalNo control, and only the wins get published
Correlation study across ranked itemsA description of what currently ranksAny claim about the effect of a changeSample is selected on the outcome
Watching a competitor gainA prompt to look at what else changed that weekAttribution to any one thing they didYou see their output, never their inputs
A platform's own published guidanceWhat the platform says it wants and will penaliseThe weighting, or how the words map to codeWritten for a wide audience, deliberately general
A natural experiment on your propertyA reasonable inference, if the timing is cleanA precise effect sizeThe world changes at the same time you do
A controlled test on your own propertyA specific effect, on your site, for that periodGeneralisation to other sitesSlow, and needs enough traffic to see a signal

The order matters more than the rows: evidence near the bottom costs more and answers a narrower question. That is why it is worth more.

Evidence Strength and Limits

Evidence

Before and after
Hypothesis
Correlation study
What ranks now
Competitor gain
Prompt to look
Platform guidance
Stated intent
Natural experiment
Reasonable inference
Controlled test
Specific effect

Can support

Before and after
Anything general
Correlation study
Effect of change
Competitor gain
Attribution
Platform guidance
Weighting
Natural experiment
Precise effect
Controlled test
Generalisation

Cannot support

Before and after
No control
Correlation study
Selected sample
Competitor gain
Inputs unseen
Platform guidance
Deliberately general
Natural experiment
World changes too
Controlled test
Slow, needs traffic

Weakness

Before and after
Correlation study
Competitor gain
Platform guidance
Natural experiment
Controlled test

A study covering ten thousand domains tells you about the average of ten thousand situations, none of which is yours. Which part of the pipeline any of it could touch is a separate question, answered by the stage map of a ranking pipeline.

How a correlation study earns a place

Its one honest use: generating candidates. If an attribute tracks position across a large sample, that is a reason to put it on a list of things to test, and no more than that.

Treat the output as a queue of hypotheses, and the study becomes useful. Treat it as a set of instructions, and you will spend a quarter on the strongest confounder in the data.

There is a second honest use, which is describing a moment. A well-run study is a snapshot of what the visible results look like now. That is genuinely informative about the current shape of a results page, and it needs no causal claim at all.

Reading one that way is closer to reading a census than reading an experiment. The record of how these claims have changed is a useful check on how quickly the snapshots go stale.

The comparison you should actually run

Against the studies, set the one thing you control: a change you make deliberately, on a slice of your own property, with a comparable slice left alone. It answers a narrow question, and it answers it about you.

Knowing which stage a change could plausibly touch is worth an hour before you spend a month on one that could not have mattered.

Two habits make the comparison honest. Write your expectation before you look, so you cannot fit the story to the result. Decide in advance what result would make you abandon the idea.

A test you would explain away either way is not a test. Both habits come from formal practice; how recommendation research is designed shows them at a larger scale.

Common questions

Are ranking factor studies dishonest?

Usually not. Most are careful about their method in the methodology section and get misread in the summary. The failure is in the reading far more often than in the arithmetic.

Why do two studies disagree about the same attribute?

Different samples, different measurement tools, different corrections, different periods. Any of those can flip a weak correlation. Persistent disagreement across studies is a signal that the underlying effect is small or conditional, not that one team is wrong.

Can I compare my site against a competitor as a control?

Not really. A control has to be similar in every respect except the change, and another company differs in every respect at once. The nearest usable version is comparing two sections of your own site. The failure modes worth knowing before you try are worth reading first.

What about studies run by the platform itself?

They have access nobody else does and a strong interest in the conclusion. Read them for the mechanism they describe rather than the effect size they report, and note what the comparison group was.

More in Rules

Latest from Review Desk