Card outlining controlled ranking signal test design steps. A working review of ranking signals research
Image: Search Ranking Tactics

Costs

A working review of ranking signals research

Ranking signals research you can run: setting the question narrowly, building a split that cannot flatter you, and reading the result without fitting a story.

You cannot see inside a ranking system, but you can run an experiment against it. The system is a black box that takes your pages as input and returns positions and traffic. Change one input on part of your property, leave a comparable part alone, and compare the two.

That is the only method an outsider has that produces evidence about your own site rather than an average of other people's sites. It is slower and narrower than reading a study, and it is worth more.

What to take away

  • The unit you randomise is the thing you changed, usually a page or a template, not a visitor.
  • A control group is not optional. Before and after on the treated group alone measures the season, the news, and everything else that happened that month.
  • Decide the metric, the duration and the decision rule before you look at any data. Everything after that is arithmetic.

Set the question narrowly

A test answers one question about one property over one period. "Does adding a comparison table to product pages change organic entrances to those pages" is testable, but "Does quality matter" is not.

Testable vs Untestable Questions

Testable

Question
One property, one period
Change
One change
Settles
A number decides
Attribution
Which part earned it

Untestable

Question
Quality matters
Change
Three edits bundled
Settles
No number settles
Attribution
Bundle helped, unclear

If a number would not settle the question, you are not ready to run anything. Naming the input you test gets easier once vocabulary is settled, and the ranking signals overview sets out what counts as one.

Narrow means one change. Ship three edits together, and a positive result says the bundle helped. That is fine if you keep the bundle forever. It is useless if you need to know which part earned its place.

Bundles are reasonable when the changes are cheap and attribution does not matter. Be honest about which situation you are in.

Build the split so it cannot flatter you

The most common mistake is splitting on something related to the outcome. New pages against old pages, or one category against another, guarantees a difference that has nothing to do with your change.

Building a Non-Flattering Split

  1. Avoid splitting on outcome-related traits
  2. Assign by arbitrary stable hash
  3. Check groups tracked before change
  4. Start clock at recrawl, not deploy
  5. Account for external position shifts
  6. Extend duration for slow response

Assign pages to treated and control groups by something arbitrary and stable, like a hash of the identifier. Then check the two groups looked alike for several weeks before you touched anything.

If they did not track each other before, they will not track each other after. The general design is ordinary split testing with a slow and noisy response variable.

Two constraints make this harder than a usual on-site experiment: nothing takes effect until the change is discovered, so the clock starts at recrawl, not at deploy.

The response is also a position in someone else's system, which moves on its own the whole time. Discovery belongs to machinery you do not control, and how a feed gets assembled describes the stages behind it. Both constraints push the required duration up.

Decide these before you start

DecisionSet it toWhy it has to be beforehand
Primary metricOne number, chosen nowFive metrics give five chances to find a win by accident
Unit of analysisThe thing you changedCounting sessions when you changed pages overstates certainty
Minimum effect worth acting onA number you would act onWithout it, any result becomes interesting after the fact
DurationLong enough to cover a full weekly cycle and recrawlStopping when the line looks good is how noise becomes a finding
Stopping ruleWritten down, including the null caseA test you would keep running until it wins is not a test
What would make you abandon the ideaSomething specificIf nothing would, you are collecting support, not evidence

The row that gets skipped most often is the minimum effect. It is also the one that decides whether the test is possible at all. If the change you expect is a few percent and your weekly variation is larger than that, no length of run will separate them.

Pre-Test Decisions Checklist

  • Primary metricone number now
  • Unit of analysisthe changed thing
  • Minimum effect worth acting on
  • Durationfull weekly cycle and recrawl
  • Stopping rulewritten, includes null
  • Abandon criteriasomething specific

Work out the statistical power you have before spending a quarter on a question your traffic cannot answer. Being told on day one that you cannot detect it beats an ambiguous chart in month three.

Read the result honestly

Compare the change in the treated group against the change in the control group over the same window. Not treated before against treated after. That single discipline removes most of the ways a seasonal swing or a platform-wide shift gets reported as a win.

Reading Results Honestly

Do

Comparison
Treated vs control change
Null result
Effect not detectable
Positive result
Site, period, system version
Record
Size you could detect

Don't

Comparison
Treated before vs after
Null result
Effect is zero
Positive result
A universal law
Record
Only the win

A null result means the effect was too small to detect with this much data. It does not mean the effect is zero or the idea was stupid.

Record it anyway with the size you could have detected. The next person asking deserves to know what was ruled out and at what sensitivity.

A positive result depends on your site, your period, and the system version. It is not a law. Publishing it as one gives the field confident advice that later stops working, and no one can say when.

The wider argument about which kinds of evidence carry what weight sits in the comparison of evidence types.

Common questions

How long should a test run?

Long enough for the change to be discovered, plus enough full weekly cycles to average out the day of week pattern. In practice that is weeks, not days. The exact number falls out of your traffic volume and the effect size you set in advance.

Can I test on a small site?

You can run the test. You often cannot detect anything, because the noise on small numbers swamps a modest effect. On a small property, spend the effort on changes that are obviously right for a reader rather than on measuring them.

What if the platform changes mid-test?

Your control group is exposed to the same change, which is precisely why it is there. A platform-wide shift moves both groups and the difference between them survives. A shift that hits only one group means your split was not arbitrary after all.

Should I test the same thing more than once?

Yes, if the decision matters, a single positive result at a conventional threshold is weaker than people treat it. Repeat on a different section or another quarter, the cheapest way to learn if you measured an effect or a moment.

Worked cases of that idea behaving differently across different contexts appear in examples of recommendation behavior.

More in Costs

Latest from Review Desk