> "The first principle is that you must not fool yourself — and you are the easiest person to fool."
Prerequisites
- 1
Learning Objectives
- Explain what a ranking 'signal,' a ranking 'factor,' and a ranking 'system' each mean — and why 'there are 200 ranking factors' is folklore, not a checklist.
- List the ranking signals Google has actually confirmed, and state honestly how much weight each one carries (including the 'lightweight' ones).
- Distinguish signals that are strongly evidenced but unconfirmed — engagement, topical authority, freshness — from those that are confirmed and those that are debunked.
- Retire four persistent 'ranking factors' — meta keywords, keyword density, word count, and domain age — with the specific evidence that debunks each.
- Explain how machine-learned systems (RankBrain, BERT, MUM) changed ranking, and why not even Google's engineers can fully explain a single result.
- Apply a four-tier evidence habit to any SEO claim you meet, and read a ranking-correlation study without being fooled by it.
In This Chapter
- Overview
- Learning Paths
- 2.1 What a "ranking factor" even means — and why the count is a myth
- 2.2 What Google has actually confirmed
- 2.3 Strongly evidenced, but unconfirmed
- 2.4 Debunked "factors" that will not die
- 2.5 Machine learning changed the game — and hid the rulebook
- 2.6 How to reason under uncertainty: the evidence-tier habit
- 2.7 Reading a ranking-correlation study without being fooled
- 📈 The Strategy File
- Conclusion
- Key Terms
- Spaced Review
Chapter 2: The Ranking Algorithm — What We Know, Think We Know, and Are Guessing
"The first principle is that you must not fool yourself — and you are the easiest person to fool." — Richard Feynman, "Cargo Cult Science" (1974)
Overview
Here is the question this chapter exists to answer honestly: when you sit down to make a page rank better, which of the hundred things the internet insists you must do actually matter — and how would you ever know?
Open any list of "SEO ranking factors" and you will find a confident, numbered inventory: 200 items, each described as though Google published it, many with a helpful little weight next to it. Almost none of that is true. Google runs a ranking system it does not publish, changes it thousands of times a year, and has publicly confirmed only a small handful of the signals inside it. Into that vacuum the industry poured two decades of folklore, half-remembered conference talks, vendor marketing, and outright invention — and then sold it back to you as expertise. The single most valuable skill you can build in this entire book is the one that lets you walk into that noise and tell the three categories apart: what we actually know, what we have good reason to believe, and what we are frankly guessing.
That skill is not cynicism. A good SEO is not someone who believes nothing; that person is as useless as the one who believes everything. A good SEO is calibrated — able to say "this is confirmed, so I'll build on it," "this is a strong correlation but unproven, so I'll act on it while holding it loosely," and "this is someone's guess dressed up as a law, so I'll ignore it." Chapter 1 gave you the pipeline — the machine that turns a page into a result. This chapter goes straight at the last and most contested stage of that pipeline, ranking, and does something the rest of the field mostly refuses to do: it tells you the truth about how little is known, and then teaches you to work well anyway.
We will bust the "200 factors" myth properly, not with a slogan but with an autopsy. We will lay out the
signals Google has genuinely confirmed — and be equally clear that "confirmed" does not mean "powerful."
We will look hard at the famous contested ones, like whether Google uses your clicks. We will retire four
zombie "factors" that will not die. And we will meet the reason nobody — not even Google — can hand you the
rulebook: the machine-learned systems at the heart of modern ranking. By the end you will own the book's
signature move, the ⚖️ Evidence Check, and be able to run it on any claim you ever meet again.
In this chapter, you will learn to:
- Define a ranking signal, a ranking factor, and a ranking system, and explain why counting "factors" is a category error.
- Name the signals Google has actually confirmed — and rank them honestly by how much they seem to matter.
- Tell strongly-evidenced-but-unconfirmed signals (engagement, topical authority, freshness) from the rest.
- Debunk meta keywords, keyword density, word count, and domain age with the evidence that kills each.
- Explain what RankBrain, BERT, and MUM changed, and why a machine-learned ranking has no readable rulebook.
- Apply the four evidence tiers to any claim, and read a correlation study without being deceived by it.
Learning Paths
This chapter is the honesty engine of the whole book — everyone should read it, because the
⚖️ Evidence Checkhabit built here is used in every chapter that follows. 🏪 Local Business and 📝 Content Creator: focus on §2.1 (the myth), §2.2 (what's confirmed — especially relevance), and §2.4 (stop wasting effort on debunked "factors"). 🛒 E-Commerce: weight §2.2 and §2.7 (you will be sold a lot of correlation studies about product pages). 🔧 Developer: §2.5 (machine-learned ranking) is your section — it explains why there is no config file for rankings. 📊 Strategist: internalize §2.6 whole; the evidence-tier discipline is the thing that will make your recommendations trustworthy when everyone else's are guesses.
2.1 What a "ranking factor" even means — and why the count is a myth
Before we can sort claims into true and false, we have to fix the vocabulary, because most of the confusion in this corner of SEO is a vocabulary problem wearing a technical costume. Three words get used interchangeably and shouldn't be: signal, factor, and system.
A ranking signal is a piece of information Google can measure about a page, a query, or the context of a search, and use to help decide the order of results. The words on your page are a signal. The links pointing at it are a signal. The searcher's location is a signal. The language of the page is a signal. A signal is just an input — an observable fact the ranking machinery is allowed to look at.
A ranking factor is the looser, industry word for "a signal, considered as something you might act on." It is not wrong, exactly, but it smuggles in a dangerous assumption: that each factor is a lever you can pull independently, and that ranking is the sum of how many levers you have pulled. Hold that suspicion; we will dismantle it in a moment.
A ranking system is a component of Google's software that uses signals to do a job. PageRank (which we build properly in Chapter 22) is a system that turns the link signal into a measure of importance. RankBrain is a system that helps interpret what a query means. The "helpful content" system (Chapter 6) is a system that tries to reward content written for people over content written for search engines. Google itself increasingly talks about ranking systems, plural — a whole roster of them — not a single "algorithm."
Here is the distinction that matters: signals feed systems, and systems produce the ranking. There is no tidy one-to-one line from "a factor" to "a slot in the results." A single signal (say, links) is consumed by several systems; a single system (say, RankBrain) blends many signals at once. Asking "how many ranking factors are there?" is a little like asking "how many ingredients are in cooking?" — the number is not the point, and any specific total is a sign the person answering has misunderstood the question.
Now let's do the autopsy Chapter 1 promised on the most famous number in the field.
🚫 SEO Myth: "There are exactly 200 ranking factors, and here is the complete list." The "200" traces back to Google's own communication from many years ago, when the company would say, in round terms, that Search used "more than 200 signals." That was a rhetorical, illustrative figure — a way of saying "a lot, and more than you'd think" — never a precise, stable count you could enumerate. The SEO industry did three things to it, each worse than the last. First, it hardened the fuzzy "more than 200 signals" into a firm "200 ranking factors." Second, it published lists purporting to name all 200 — lists in which a handful of confirmed items sit shoulder to shoulder with dozens of guesses, debunked ideas, and inventions, all formatted identically so you can't tell which is which. Third, it kept selling those lists long after Google had said the real number is far larger, constantly changing, and not a meaningful target. The truth is less tidy and more useful: ranking is not a checklist you complete; it is a contest you win by being a better answer than the other candidates, judged by systems that weigh many signals in ways no list can capture. When you next see "the 200 ranking factors," you are not looking at Google's algorithm. You are looking at someone's content-marketing asset.
The deeper problem with the checklist mindset is not that the count is wrong; it is that it points you at the wrong activity. Checklist-thinking says: complete the items, earn the ranking. Contest-thinking says: there are ten other pages fighting for this query — why would Google prefer mine? A page can satisfy every box on any "200 factors" list and still sit at position 11, because a competitor simply answers the question better. Throughout this book, whenever you feel the pull toward completing a checklist, drag yourself back to the contest: who else wants this ranking, and what would make us the more deserving result?
One more piece of vocabulary earns its place here, because it quietly dissolves a lot of bad arguments. Some signals are query-dependent signals — their relevance changes depending on what was searched. Freshness is the classic example: for the query "election results" or "iPhone release date," how recently a page was published or updated matters enormously; for "how to tie a bowline knot," it matters almost not at all. Location is query-dependent too — decisive for "coffee shop near me," irrelevant for "who wrote Hamlet." This is why the eternal SEO debate "is freshness a ranking factor — yes or no?" is malformed. The honest answer is "for some queries, strongly; for others, not at all," and anyone who answers it with a flat yes or no has already told you they don't understand how ranking works.
🔎 How Search Sees It Picture the ranking machinery not as a scorecard with 200 rows you tick off, but as a courtroom that convenes fresh for every single query. The "witnesses" are the signals — the page's words, its links, its speed, the searcher's location and language. The "judges" are the systems, several of them machine-learned. They hear the case for this specific query, weigh the witnesses differently depending on what was asked, and hand down an ordering. The next query convenes a different courtroom with different weights. Nothing about this process resembles filling in a form — which is exactly why the form-shaped "factor lists" mislead you. You are not completing a document. You are making a case.
2.2 What Google has actually confirmed
Let's stand on solid ground for a while. Strip away the folklore and there is a real, if short, list of things Google has stated, on the record, that its ranking systems use. The cleanest starting point is Google's own public "How Search Works" material, which frames ranking around five broad ideas: the meaning of your query, the relevance of a page to it, the quality of that page, the usability of the page, and the context and settings of the search (your location, language, some history). That framing is not a leak or a guess; it is Google describing its own goals in its own words, and it is a good skeleton to hang the confirmed specifics on.
Two of those specifics deserve proper definitions, because they are the load-bearing beams of everything else.
Relevance is how well a page actually addresses what the searcher wants. Note the phrasing: not "how many times the page contains the query words," but how well it satisfies the need behind them. Relevance is, by a wide margin, the most important thing ranking rewards — and it is inseparable from intent, the concept so central that it gets its own chapter next. A page can be fast, secure, mobile-friendly, and richly linked, and still lose to a page that simply matches what the searcher meant better. Keep that hierarchy straight: relevance first, everything else in support of it.
Authority is the rough idea that some pages and sites are more trustworthy, more established, and more widely vouched-for than others, and that Google tries to prefer them, all else equal. Historically the signal most associated with authority is links — the founding insight that a link is a kind of vote, which we develop fully in Chapter 22. "Authority" is a genuinely useful concept and a genuinely slippery one: it is not a single number Google publishes (beware the third-party "authority" scores we dissect in Chapters 22 and 30), and it is earned slowly. But that Google prefers more authoritative sources, especially where trust matters, is not in serious dispute.
With those defined, here are the signals Google has confirmed, stated at honest strength:
| Confirmed signal | What Google has said | Honest weight |
|---|---|---|
| Relevance / content match | The page must actually address the query and its intent | The single biggest lever — everything else is secondary |
| Links | Links are used as a signal of importance/authority (PageRank lineage) | Strong, but not the raw 1998 formula — see Ch 22 |
| Mobile-friendliness / mobile-first indexing | Google predominantly indexes and ranks the mobile version of a page | Real; a floor to clear more than a lever to pull — see Ch 17 |
| HTTPS | Announced (2014) as a "lightweight" signal, affecting well under 1% of queries | Confirmed but tiny; do it because it's correct, not to rank |
| Core Web Vitals / page experience | Part of ranking; a tiebreaker between pages of similar relevance | Real but light — Google says great content matters more (Ch 16) |
| Language & locale | Results are matched to the searcher's language and location | Decisive for local/multilingual queries; query-dependent |
| Freshness (for some queries) | Some queries "deserve freshness"; recency is weighed then | Query-dependent — powerful for news, irrelevant for evergreen |
Read that table twice, because it contains two lessons that most SEO advice gets backwards.
The first lesson: "confirmed" is not the same as "powerful." Look at HTTPS. Google confirmed it as a ranking signal, which means it is one of the few items on any list you can state as fact — and in the same breath Google called it lightweight and said it affected a tiny slice of queries. A whole cottage industry treats "it's a confirmed ranking factor!" as if it settled the question of importance. It doesn't. The confirmed signals include several — HTTPS, Core Web Vitals — that are real, worth doing, and nearly weightless as ranking levers on their own. Confirmation tells you a signal exists; it does not tell you it will move your rankings. Those are two different claims, and conflating them is one of the most common ways smart people waste months.
The second lesson: most of the confirmed signals are things you should do anyway, for reasons that have nothing to do with ranking. A secure site (HTTPS), a fast site that works on a phone (Core Web Vitals, mobile-friendliness), a page in the reader's language — these are baseline decencies of a modern website. Their ranking value is a bonus on top of the real reason to do them: they serve the human being. That alignment is not a coincidence, and it is the through-line of this whole book. The confirmed levers are the ones where Google's interest and your reader's interest are the same interest.
⚖️ Evidence Check Claim: "Core Web Vitals is a ranking factor, so improving your speed score will lift your rankings." Let's sort it precisely, because it is a perfect specimen. — Confirmed by Google: that Core Web Vitals — the page-experience metrics we cover in Chapter 16 (Largest Contentful Paint, LCP; Interaction to Next Paint, INP; Cumulative Layout Shift, CLS) — is part of ranking. Yes. Google has stated this on the record. — The honest caveat: Google has been equally explicit that it is a tiebreaker — a factor that can help decide between pages of similar relevance and quality, not one that promotes a weak page over a strong one. Google's own guidance says, in effect, that great page experience does not override having great, relevant content. So the claim's first half is true and its second half ("improving your score will lift your rankings") is not reliably true: if you are losing on relevance, a perfect speed score fixes nothing. The move is to treat Core Web Vitals as a floor to clear for your users' sake, then spend your real energy on relevance and quality. Confirmed ≠ decisive.
There are other genuinely confirmed items we will meet in their home chapters and won't belabor here: that Google ignores the meta keywords tag (we bury it in §2.4), that intrusive interstitials can hurt (Chapter 17), that structured data earns eligibility for rich results but is not itself a ranking boost (Chapter 18), and that the quality framework Google summarizes as E-E-A-T — Experience, Expertise, Authoritativeness, and Trustworthiness — shapes what its systems are trying to reward, without there being any single "E-E-A-T score" inside the machine (Chapter 5 is its proper home). Notice how much smaller this confirmed list is than "200." That gap between the short list of things we know and the long list of things people claim is the whole subject of this chapter.
2.3 Strongly evidenced, but unconfirmed
Between "Google confirmed it" and "someone made it up" lies the most interesting territory in SEO: signals that Google has not clearly confirmed as ranking factors, but that the weight of evidence — reproducible observation, large-scale correlation, and the plain logic of how a search engine would have to work — strongly suggests are real. This is where a calibrated practitioner earns their keep, because you have to act on these without the comfort of certainty. Three examples carry the lesson.
Engagement and click signals are the great contested question of modern SEO, and the honest history of it is a masterclass in reading evidence. For years, Google's public spokespeople denied that Google uses click-through rate, dwell time, or bounce rate directly as ranking signals — the reasoning being that raw clicks are noisy and trivially gameable, so using them naively would be a disaster. Many SEOs took those denials as final. Then reality got complicated. During the United States Department of Justice (DOJ) antitrust case against Google (tried in 2023), sworn testimony and disclosed documents revealed that Google does operate systems that use click and interaction data — one of them named NavBoost — as part of ranking. In 2024, a large trove of internal Google Search API (Application Programming Interface) documentation leaked and was widely analyzed; it referenced numerous features, including ones that appear to involve clicks and user interaction. So which is it — does Google use engagement or not?
⚖️ Evidence Check Claim: "Click-through rate is a ranking factor — get more people to click your result and it will rank higher." This is the single most argued-about claim in SEO, so let's be careful. — What the public statements said: for years, Google representatives stated that clicks are not used directly as a ranking signal, because they are too noisy and too easy to manipulate. — What later came to light: testimony in the 2023 DOJ antitrust trial and the 2024 documentation leak indicate that Google does use click/interaction data inside specific systems (such as NavBoost). This is strong evidence — court testimony is public record — that engagement plays some role. — The honest synthesis: engagement signals almost certainly matter, but not in the naive form the claim states. There is no simple "raise your CTR, rank higher" dial. The data is aggregated across many users, denoised, and used within particular systems — and Google is very good at ignoring manipulation (bots clicking your result do nothing). So the tier here is strong evidence, unconfirmed in the specific form marketers sell. The move is to earn clicks and satisfaction honestly — a compelling, accurate title and a page that delivers on it (Chapter 9) — and to distrust anyone selling a "CTR manipulation" service. Note, too, the meta-lesson: a company's public statements and the full truth are not always the same thing. Weigh testimony and documents over press-friendly denials.
Topical authority — the idea that Google recognizes a site as a credible, comprehensive source on a subject and rewards its pages accordingly — is strongly evidenced and central enough that Chapter 4 owns it. For our purposes here, place it firmly in this middle tier: Google has spoken about wanting to identify sites that are good sources for a topic, and practitioners reliably observe that broad, deep coverage of a subject lifts a site's pages across that subject. Is "topical authority" a single named factor with a knob? No. Is covering a topic thoroughly a good bet backed by real evidence? Yes. Both are true.
Freshness we met in §2.1 as the model query-dependent signal, and it belongs here with a sharpened point. Google has effectively confirmed that some queries deserve freshness and that recency is weighed for them — so freshness is more confirmed than engagement. But the folklore version — "update your pages constantly, Google loves freshness" — is false as a general rule. Re-saving a page to change its date without improving it is a classic cargo-cult move: it mimics the appearance of the thing Google rewards (genuinely updated, more accurate content) without the substance, and it fools no one. Freshness matters where the world has changed, not where your timestamp has.
There is a decision rule hiding in this whole section, and it is one of the most practical ideas in the book:
🔎 How Search Sees It When a signal is strongly evidenced but unconfirmed, ask a second question before you act: is this also just good for the reader, regardless of whether it's a ranking factor? Earning genuine engagement, covering a topic thoroughly, keeping information actually current — every one of these is good for the human being on the page whether or not it moves a ranking. That makes them no-regret moves: you win if the signal is real, and you still win (a better site) if it isn't. Contrast that with chasing an unconfirmed factor that is only about gaming a ranking — there, if you're wrong, you've simply wasted the effort and possibly harmed the page. The safest way to act under uncertainty is to prefer moves that pay off on both sides of the bet.
2.4 Debunked "factors" that will not die
Some "ranking factors" are not uncertain at all — they are settled, and settled in the direction of "no." Yet they persist, taught in courses and repeated in audits, because folklore is durable and because a plausible-sounding factor is easy to sell. Here are four zombies, each with the specific evidence that killed it. Learn to recognize them, because the effort you don't spend on debunked factors is effort freed for the things that work.
Meta keywords. There is an HTML tag, <meta name="keywords">, into which you can stuff a list of terms
you'd like to rank for. It does nothing. Google stated plainly, back in 2009, that it does not use the
meta keywords tag in web ranking — and it has never reversed that. The tag survives on countless sites and
in countless "SEO plugins" purely as a fossil. Worse than useless, it is a small gift to competitors, who
can read your source code and see exactly which terms you're targeting. Fill it in if a form makes you;
expect nothing from it.
Keyword density. This is the belief that there is an ideal percentage of times your target keyword should appear on the page — 2%, or 3%, or whatever a particular course insists on — and that hitting it helps you rank. There is no such number, Google has repeatedly said there is no such number, and tools that score your "keyword density" are measuring a quantity the ranking systems do not care about in that form. Modern language understanding (§2.5) reads meaning, not term frequency ratios. The only real phenomenon nearby is the opposite one: keyword stuffing — cramming a term unnaturally — is a spam signal that can hurt you. So the honest guidance, which we develop in Chapter 9, is to write naturally about the subject and never think about density at all.
Word count. "Longer content ranks better; aim for 2,000 words" is perhaps the most expensive myth in this list, because it directly wastes writers' time. Google's search advocates have stated flatly, more than once, that word count is not a ranking factor. Where does the myth come from? From a real correlation mistaken for a cause — the exact trap §2.7 is built to defuse. Longer pages often rank well because thorough answers tend to be longer, and thoroughness (relevance, comprehensiveness) is what actually gets rewarded. Padding a 600-word answer to 2,000 words does not add thoroughness; it adds fluff a reader has to wade through, which makes the page worse. Length is a side effect of covering a topic well, never a target.
Domain age. "Google trusts older domains, so an aged domain ranks better" is a myth with a small, lucrative black market attached (people buy expired domains partly on this belief). Google's John Mueller has stated directly that domain age does not help ranking. The correlation that fuels the myth is real and entirely explicable: older domains have simply had more time to earn links, build content, and accumulate authority — it is the links and content doing the work, not the birthday. A ten-year-old domain with nothing on it ranks like what it is: nothing.
Notice the shape these last two share. In both cases a true correlation ("long pages rank," "old domains rank") gets misread as a causal factor ("add words," "age your domain"), when a third thing — thoroughness, or accumulated authority — is causing both. Hold that pattern; it is the master key to §2.7.
🚫 SEO Myth: "Every page needs at least 1,500–2,000 words to rank." No. Word count is not a ranking factor; Google has said so explicitly. What ranks is the page that best and most completely satisfies the searcher's intent — which is sometimes a 300-word direct answer and sometimes a 4,000-word guide, decided by the query, not by a target. The right length for a page is "exactly as long as it takes to answer the question well, and not one sentence longer." Chapter 3 (intent) and Chapter 9 (writing) turn that principle into method. Any tool or course that hands you a universal word count is selling you a cargo cult.
🛠️ Try It on Your Site Right now, open one of your own pages, right-click, and choose "View Page Source." Search (Ctrl/Cmd-F) the raw code for
keywords. If you find a<meta name="keywords" content="...">tag stuffed with terms, you've found a fossil doing nothing for you and quietly showing competitors your targets — a candidate for removal. While you're there, search the source forviewport(a sign the page declares itself mobile-ready) and note whether the address bar showshttps://. In sixty seconds you've checked one debunked "factor" and two genuinely confirmed ones. That is the evidence habit in miniature: look at what's actually there, not at what a checklist told you to assume.
To make the whole sort visible in one place, here is the map this chapter builds — the book's recurring confirmed / likely / debunked table, in miniature:
THE RANKING-CLAIM MAP [what this chapter establishes]
CONFIRMED (build on it) LIKELY (act, hold loosely) DEBUNKED (stop)
───────────────────────── ───────────────────────── ─────────────────────
Relevance / intent match Engagement / click signals Meta keywords tag
Links (as a signal) Topical authority Keyword density %
Mobile-first indexing Freshness (general chasing) Word-count targets
HTTPS (lightweight) "Comprehensiveness helps" Domain age
Core Web Vitals (tiebreaker) "Post daily to rank"
Language / locale Third-party "authority" as
Freshness (query-dependent) a thing Google reads
└─ confirmed ≠ powerful └─ evidence, not proof └─ folklore, sometimes
(several are light) (prefer no-regret moves) harmful
Pin this somewhere. Ninety percent of the bad advice you will encounter is a debunked item smuggled into the confirmed column, or a confirmed-but-lightweight item (HTTPS, Core Web Vitals) sold as a heavyweight.
2.5 Machine learning changed the game — and hid the rulebook
Everything so far might leave the impression that ranking is a set of rules — weight relevance this much, links that much, mobile-friendliness a touch — that Google could publish if it chose to. That picture was roughly right around 2010. It is badly wrong today, and understanding why is what finally explains the uncertainty this whole chapter keeps insisting on. Modern ranking is not a rulebook a person wrote. Large parts of it are machine-learned: systems trained on enormous amounts of data to recognize patterns, whose resulting behavior no human hand-coded and no human can fully read back out. Three named systems mark the shift.
RankBrain, introduced in 2015, was Google's first major use of machine learning in ranking. Its job is to help understand queries, especially the roughly one-in-seven searches Google has never seen before (Chapter 1). Faced with a novel or ambiguous query, RankBrain helps connect it to concepts and results it has seen, rather than matching raw strings. At its launch Google told reporters that RankBrain had quickly become one of the most important of its many ranking signals — a striking thing to say about a system whose decisions emerge from training rather than from written rules.
BERT — Bidirectional Encoder Representations from Transformers — arrived in Search in 2019 and was a leap in natural-language processing (NLP), the field concerned with getting software to understand human language. BERT reads a query by considering each word in the context of the words around it, which lets it grasp the way small words change meaning — the difference "to" makes in "travel to Canada" versus "travel from Canada," or the role of "no" and "stand" in a phrase. Google said at launch that BERT would affect around one in ten English-language queries. Its significance for us is conceptual and large: with BERT and its successors, Google stopped treating your page as a bag of keywords and started reading it for meaning — which is the technical reason keyword density (§2.4) is dead and covering a topic in natural language (Chapter 4) is the game.
MUM — the Multitask Unified Model, announced in 2021 — is a still more powerful model that Google described as many times more capable than BERT, able to work across languages and even across formats (text, images). Here honesty requires a firmer hand than the marketing: Google has been comparatively vague about exactly where and how much MUM is used in live ranking, deploying it in particular features rather than describing it as a general ranking dial. So we file MUM as announced and real, with an unconfirmed and probably narrow role in ranking — a Tier-1 fact about its existence, a Tier-2-or-worse guess about its day-to-day ranking weight. Notice that we can be precise about our imprecision. That is the skill.
Why does any of this mean Google itself can't fully explain a ranking? Because a machine-learned system does not store its knowledge as readable rules. It stores it as millions of numerical weights, tuned by training, that collectively produce good outputs without any one of them corresponding to a statement a person could read. Add to that the fact that modern ranking is a system of systems — RankBrain and BERT and the helpful-content system and the spam systems and PageRank's descendants and more, each contributing — and you get behavior that is emergent: it arises from the interaction of many trained components, no one of which is the answer. Google's engineers can tell you how the systems are built, what they are trained to do, and which broadly matter. What no one can hand you is a sentence of the form "page A beats page B for this query because of factors X, Y, and Z, weighted thus." That sentence does not exist inside the machine either.
HOW A RANKING IS ACTUALLY PRODUCED [schematic — not to scale]
SIGNALS (measurable inputs) SYSTEMS (learned + rule-based) RESULT
┌───────────────────┐ ┌────────────────────────┐
│ words / meaning │──┐ │ query understanding │
│ links │ │ │ (RankBrain, BERT, …) │──┐
│ language / location │ ├───────▶├────────────────────────┤ │ ┌────────────┐
│ freshness │ │ │ core ranking systems │ ├────▶│ page → #1 │
│ page experience │ │ ├────────────────────────┤ │ │ page → #2 │
│ …many more… │──┘ │ helpful content / spam /│──┘ │ page → #3 │
└───────────────────┘ │ reviews / … │ │ … │
└────────────────────────┘ └────────────┘
No single signal owns a slot in the output. Learned systems combine signals
in ways that are not readable as a rulebook — not by you, and not by Google.
This is not a counsel of despair; it is a counsel of humility with direction. You cannot reverse-engineer the rulebook because there isn't one. But you can do the thing the whole system is trained to reward — be the genuinely most relevant, most trustworthy, most useful answer — and trust that a system built to find that will, over time and imperfectly, find you. That is not resignation. It is the only strategy that is robust to a machine you cannot read.
🔗 Connection The language-understanding thread begun here (BERT, meaning over strings) is developed into entities, the Knowledge Graph, and semantic search in Chapter 4. Machine-learned quality judgment, and how it moves in the big "core updates," is Chapter 6. And the frontier of machine learning in search — AI Overviews and generative answers, where the uncertainty is greatest of all — is Chapter 36, which §2.6 introduces next.
2.6 How to reason under uncertainty: the evidence-tier habit
We have now seen confirmed signals, strongly-evidenced ones, debunked ones, and a machine no one can fully read. The obvious question is: so how do I decide what to do? The answer is a habit, and building it is the real deliverable of this chapter — the thing you will use in all thirty-eight chapters that follow. We call it the evidence check, and it is nothing more than the discipline of sorting every SEO claim you meet onto a ladder before you act on it.
THE EVIDENCE LADDER — how much weight a claim can bear [schematic]
STRONGEST ┌──────────────────────────────────────────────────┐
▲ │ 1. CONFIRMED BY GOOGLE │ public docs, on-record
│ │ "HTTPS is a lightweight ranking signal." │ statements, testimony
│ ├──────────────────────────────────────────────────┤
│ │ 2. STRONG CORRELATION / EVIDENCE, NOT PROVEN │ large studies,
│ │ "More referring domains ~ higher rankings." │ reproducible tests
│ ├──────────────────────────────────────────────────┤
│ │ 3. PRACTITIONER EXPERIENCE │ "I've seen this
│ │ "This fix moved the page — for me, that time." │ work" — real, anecdotal
│ ├──────────────────────────────────────────────────┤
▼ │ 4. SPECULATION / SALES PITCH │ "Google's secret rule
WEAKEST │ "The hidden 2% keyword-density factor." │ the gurus won't share"
└──────────────────────────────────────────────────┘
The rule: never let a claim carry more weight than its rung allows.
Each rung is legitimate in its place; the error is always a claim pretending to a higher rung than it has earned. Confirmed facts (rung 1) you build on. Strong correlations (rung 2) you act on while holding them loosely and watching for the reversal §2.7 describes. Practitioner experience (rung 3) — including your own — is genuinely valuable and genuinely limited: "it worked on my site" is one data point, confounded by everything else you changed and by the fact that you remember your wins. Speculation (rung 4) is not worthless as a hypothesis, but it is worthless as a basis for spending a client's money, and most of what is sold as SEO secret sauce lives here.
When any claim lands in front of you — "X is a ranking factor," "you must do Y," "this update was about Z" — run three fast questions:
- Which rung is this on? Is it something Google confirmed, a correlation, someone's anecdote, or a guess? Half the time the honest answer reveals the claim can't bear the weight being put on it.
- What would prove it wrong? A claim that no observation could ever falsify ("Google rewards quality, and if your quality page didn't rank, it wasn't quality") is not knowledge; it's an unfalsifiable belief. Real claims make risky predictions you could check.
- Who benefits if I believe it? The vendor whose study finds that links matter sells a link tool. The course that insists on 2,000-word posts sells a writing course. Follow the incentive; it won't tell you the claim is false, but it tells you how hard to squint.
That is the entire method, and its power is that it works on claims about a system nobody can read. You do not need to know Google's algorithm to know that "the hidden 2% density factor" is a rung-4 pitch and "HTTPS is lightweight" is a rung-1 fact. Calibration, not omniscience, is the goal.
Now let me show you the uncertainty at its sharpest, with a scenario that will run through the rest of the book — because reasoning under uncertainty is not an abstraction here; it is the defining condition of modern search.
📄 Read the SERP
text FIGURE 2.1 — "The AI Overview that took 30% of the clicks" [constructed teaching example] THE QUERY / PAGE "how long does a furnace last" — an informational query, and a genuinely excellent explainer page that has ranked #1 for it for three years. WHAT'S THERE Above the #1 result now sits an AI Overview: an AI-written summary that answers the question ("typically 15–20 years, depending on…") directly on the results page, with a few cited links folded into it. The blue links, including our #1, are pushed down. WHAT IT SHOWS The page still ranks #1 — it did everything the confirmed signals reward — and yet its clicks have fallen sharply (in this constructed case, by about 30%), because many searchers now get their answer without clicking anything. Ranking well was necessary and no longer sufficient. WHAT IT DOESN'T It does NOT tell us this is permanent, universal, or precisely measurable; AI Overviews appear unevenly, change constantly, and hit informational queries far harder than transactional ones. The "30%" is illustrative, not a measured law. THE MOVE Don't panic-rewrite a page that's already winning the ranking contest. Diversify what the page (and the business) depend on — depth and originality a summary can't replace, plus traffic from email, brand, and direct — and read Chapters 13 and 36 before acting. THE LESSON The algorithm's frontier is genuinely uncertain, and "rank #1" is not the finish line it used to be. Build for the uncertainty, not against a single snapshot of it.
This is the AI Overview — Google's AI-generated answer atop the results page — and this scenario, "the AI Overview that took 30% of the clicks," is one of the running examples of the book. We introduce it here, in the honesty chapter, on purpose: it is the single clearest case of a place where a diligent, well-optimized site can do everything the confirmed evidence says to do and still watch the ground move under it. We are not going to pretend to resolve that here. Its full treatment — how AI Overviews work, who still clicks, how to be cited by them, and why traffic diversification is now insurance rather than a nicety — is Chapter 36, and the content response to it is Chapter 13. For now, hold it as the emblem of this chapter's whole posture: be honest about what you don't know, build the robust things anyway, and never bet the business on a single snapshot of a system that changes every day.
🔄 Check Your Understanding A consultant tells Rivertown's owner: "Bounce rate is a Google ranking factor — I ran the numbers, sites with lower bounce rates rank higher, so we need to cut yours." Run the three questions on this claim.
Answer
(1) Which rung? "Sites with lower bounce rates rank higher" is a correlation (rung 2) presented as a confirmed causal factor (rung 1) — the first red flag. Google has, in fact, repeatedly said it does not use Google Analytics bounce rate as a ranking signal. (2) What would falsify it? The consultant hasn't said — and the correlation is easily explained the other way around: good, relevant pages both rank well and keep visitors, so relevance causes both, and "cutting bounce rate" cosmetically (autoplay, pop-ups) could rank you nothing and annoy users. (3) Who benefits? The consultant selling a "bounce-rate optimization" engagement. Verdict: a rung-2 correlation wearing a rung-1 costume. Improve the page's actual relevance and usefulness (which will, incidentally, lower bounce as a side effect); don't chase the metric.
2.7 Reading a ranking-correlation study without being fooled
The most persuasive-looking evidence in SEO comes in the form of large correlation studies: a tool vendor crawls hundreds of thousands of search results, measures some trait of the top-ranking pages (their number of links, their word count, their speed, their use of a keyword in the title), and publishes a chart showing that higher-ranking pages tend to have more of it. These studies are genuinely useful — and they are the single most common way smart people get fooled, because they invite a leap the data does not support. The leap has a name, and un-learning it is a superpower.
Correlation vs. causation is the distinction between two things moving together (correlation) and one thing causing the other (causation). "Ice cream sales and drowning deaths rise together" is a true correlation; ice cream does not cause drowning — summer heat causes both. Nearly every ranking-correlation study reports a correlation and then, explicitly or by insinuation, invites you to read it as causation: "top pages have more links, so build more links." Sometimes the causal story is even roughly right (links do matter). But the study did not establish it, and three specific traps explain why.
The first trap is the confounder — a third variable causing both things you measured, exactly like the summer heat. Long pages rank well (§2.4) not because length causes ranking but because thoroughness causes both length and ranking. Any correlation can be a confounder in disguise, and the studies rarely rule them out.
The second trap, and the one that catches nearly everyone, is reverse causation — the arrow pointing the opposite way from how you read it. A study finds that #1-ranked pages have far more links than page-two pages, and you conclude "links cause ranking." But consider: a page that already ranks #1 is seen by vastly more people, so it earns more links precisely because it ranks well. Ranking causes links at least as much as links cause ranking. The same reversal haunts almost every engagement and "authority" correlation in SEO: high rankings produce the traffic, the citations, the brand searches, and the engagement that studies then present as the cause of high rankings. Whenever you read "top pages have more X," ask immediately: could ranking be causing X, rather than the other way around? Astonishingly often, it could.
The third trap is sampling and selling — who was studied, and who is paying. What queries were in the sample (commercial? informational? English-only?), how big and representative was it, and — never skip this — what does the publisher sell? A study by a link-tool vendor that concludes "links are the top ranking factor" is not necessarily wrong, but its incentive is not neutral, and it belongs on rung 2 with an extra pinch of salt.
📄 Read the Report
text FIGURE 2.2 — "Top pages have 3x the referring domains" [constructed teaching example] THE QUERY / PAGE A vendor's study of 100,000 search results, charting referring domains (distinct sites that link to a page) against ranking position. WHAT'S THERE A clean downward curve: position-1 pages have, on average, roughly three times as many referring domains as position-10 pages. Headline: "Backlinks remain the #1 ranking factor." (All figures constructed for teaching.) WHAT IT SHOWS A real, unsurprising CORRELATION: higher-ranked pages tend to have more links. This is consistent with links mattering — which, from other evidence, they do (Ch 22). WHAT IT DOESN'T It does NOT show that links CAUSED the rankings. Reverse causation is wide open (ranking #1 earns links), confounders abound (better sites have both more links and better content), the sample and query mix are unstated, and the publisher sells a link tool. "The #1 factor" is an unearned leap from a correlation. THE MOVE Treat it as rung-2 support for a thesis you already hold on better grounds (links matter), not as proof of a weight or a mandate to buy links. Let it inform, not command. THE LESSON A correlation study is a hypothesis generator, never a verdict. Read the caveats the headline omitted, and always ask which way the arrow points.
None of this means correlation studies are worthless — they are one of the few sources of large-scale evidence we have, and a well-run one can genuinely sharpen your priors. It means you read them like a scientist and not like a customer: as rung-2 evidence that suggests and constrains, never as rung-1 proof that settles. Here is the compact checklist to keep beside any study you meet:
- Correlation or causation? Almost always the former; the headline almost always implies the latter.
- Could a confounder explain it? Is a third thing (quality, authority, thoroughness) driving both?
- Could the arrow point backward? Could ranking be causing the trait, not the trait causing ranking?
- What's the sample? Which queries, how many, how representative, whose data?
- Who published it, and what do they sell? Follow the incentive; then weigh accordingly.
- What would I actually do differently? If the honest answer is "nothing I wasn't already doing for good reasons," the study is interesting, not actionable — and that's fine.
⚖️ Evidence Check Claim: "This study proves that pages using the keyword in the H1 heading rank higher, so we must put the exact keyword in every H1." Sort it. Rung 2 at best — a correlation from a study, not a Google confirmation. Confounder? Yes: pages written by people who understand the topic tend both to name it in the heading and to be more relevant. Reverse causation? Less so here, but the correlation is weak evidence for a mandate. What Google actually says: headings help it understand structure (Chapter 9), and naming the topic in your H1 is sensible — but as clear communication, not as a keyword-placement ritual with a causal guarantee. The move: write an accurate, descriptive H1 because it helps readers and machines understand the page, and drop the word "must." Same action, honest reason, no false precision.
📈 The Strategy File
Rivertown Home Services — the family-owned, second-generation HVAC, plumbing, and electrical company with five locations (Rivertown headquarters plus Cedar Hills, Northgate, Westbrook, and Millhaven), roughly 75 employees, 35 trucks, about \$16M in revenue, a WordPress site, and around 8,000 mostly-branded organic visits a month — has a very Chapter 2 problem. One of its owners, Tony Delgado, has spent the past month asking around, and everyone has an opinion. His nephew took an online course. A cold-calling agency sent a "free audit." The expiring "SEO guy" left behind a list. Tony now has a page of proposed "fixes" and no way to tell the gold from the folklore. That is precisely the skill this chapter built, so before Rivertown spends a dollar, we run every item through the evidence check and sort it into confirmed / likely / folklore. We are not doing any of these yet — sequencing and execution come later — we are deciding which deserve to be on the list at all.
FIGURE 2.3 — "Rivertown's proposed fixes, sorted by evidence" [the Strategy File]
PROPOSED FIX (as Tony heard it) VERDICT WHY / WHERE WE'LL ACTUALLY DO IT
───────────────────────────────────────── ───────── ──────────────────────────────────
"Move every page to HTTPS." CONFIRMED Real (lightweight) signal + basic
security. One-time, do it. (Ch 16/14)
"Fix the slow mobile site / Core Web CONFIRMED Confirmed page-experience signal
Vitals." (light) (tiebreaker) AND vital for panicked
"no heat" mobile searchers. (Ch 16/17)
"Make location & service pages actually CONFIRMED Relevance/intent is the #1 lever;
match what people search; kill the (BIG) also fixes the thin/near-duplicate
near-duplicate city copy." pages. This is the main event. (Ch 3/9/25)
"Earn real links; add internal links to CONFIRMED Links are a confirmed signal; internal
the buried pages." (how-much linking flows authority. (Ch 15/22/23)
uncertain)
"Cover heating/cooling/plumbing/electrical LIKELY Topical authority is strongly evidenced,
thoroughly as topics, not keywords." not a named knob. Good bet. (Ch 4/8)
"Keep genuinely time-sensitive info current LIKELY Freshness is query-dependent; update
(rebates, codes) — not just re-dating." (nuanced) where the WORLD changed, not the clock.
"Add a meta keywords tag listing every FOLKLORE Google ignored it since 2009; also leaks
city + service." targets to competitors. Delete.
"Hit 2% keyword density for 'furnace repair' FOLKLORE No such factor. Write naturally. Stuffing
on every page." is a SPAM risk, not a win. (Ch 9)
"Every page needs 2,000+ words." FOLKLORE Word count isn't a factor; padding thin
pages makes them worse. Length follows intent.
"Our domain's old, so we're set" / FOLKLORE Domain age doesn't help; it's the accrued
"buy an aged domain." links/content that do. (Ch 22)
"Publish a blog post every single day." FOLKLORE Velocity isn't a ranking factor; quality and
intent-match are. (Ch 8)
"Buy 100 backlinks for $99 / submit to 500 FOLKLORE Link schemes are detectable and punishable —
directories." (this is what the old guy did) + RISK exactly the inherited liability. (Ch 24/26)
Two honest notes on what this sort does and does not settle. It does not tell us how much any confirmed fix will move Rivertown's rankings — that depends on the competitors in Cedar Hills and Northgate we don't control, and it's the difference between "confirmed lever" and "guaranteed result" that this book will never blur. And it does not sequence the work; the prioritized punch list is assembled much later, in the full audit (Chapter 38). What it does settle is enormous and immediate: it stops Rivertown from spending its first, scarcest dollars on the bottom half of that table — the meta keywords, the density targets, the word padding, the bought links — which is exactly where the previous "SEO guy" spent them, and exactly why Rivertown inherited a spammy-link liability instead of rankings. The evidence filter is now the first gate every future Rivertown recommendation passes through. That gate is this chapter's contribution to the Strategy File. (All Rivertown figures are a constructed teaching example.)
Conclusion
We set out to answer an uncomfortable question — which of the hundred things you're told to do actually matter, and how would you know? — and we answered it not with a longer list but with a sharper mind. Google runs a machine-learned system of systems it does not publish and changes constantly; it has confirmed only a short list of signals, several of which (HTTPS, Core Web Vitals) are real but lightweight; the biggest confirmed lever by far is relevance; a fascinating middle tier (engagement, topical authority, freshness) is strongly evidenced but unproven; and a stubborn set of "factors" (meta keywords, keyword density, word count, domain age) is simply debunked. Most important, we saw why certainty is impossible — a learned system has no readable rulebook, not even for its own makers — and we turned that from a source of anxiety into a working method: the four-tier evidence check, the three questions, the correlation-study checklist, and the no-regret preference for moves that serve readers whether or not they move rankings.
That method is the real inheritance of this chapter. Every remaining chapter carries an ⚖️ Evidence Check,
and now you know what it is for: not to hedge, but to keep you calibrated in a field that punishes both the
credulous and the cynical. You will meet confident numbers, secret factors, and can't-miss tactics for the
rest of your career. You now have the one tool that makes you safe among them — the willingness, in Feynman's
words, not to fool yourself.
We also planted the emblem of everything we can't yet know: the AI Overview quietly taking a third of a winning page's clicks. Hold it. The book returns to it honestly, and never pretends the frontier is settled.
Next, we turn from how ranking is judged to what it is judging for — and it turns out the single biggest confirmed lever, relevance, rests entirely on a concept most SEOs still get wrong. In Chapter 3, we take on search intent: the recognition that Google ranks pages that satisfy what a searcher wants, not pages that merely contain the words they typed. If Chapter 2 taught you to reason honestly, Chapter 3 gives that honesty its most important single target.
→ Continue to Chapter 3: Search Intent.
Key Terms
- Ranking factor — the loose industry term for a signal considered as something to act on; useful but misleading when it implies ranking is a checklist of independent levers rather than a contest.
- Ranking signal — a measurable piece of information about a page, query, or context that Google's systems may use to order results (e.g., the page's words, its links, the searcher's location).
- Relevance — how well a page satisfies the need behind a query (not merely whether it contains the query words); the single most important thing ranking rewards.
- Authority — the rough notion that some pages and sites are more trustworthy and widely vouched-for than others; historically associated most with links, and not a single number Google publishes.
- Query-dependent signal — a signal whose relevance changes with the query, such as freshness (decisive for news, irrelevant for evergreen topics) or location (decisive for "near me," irrelevant for facts).
- RankBrain — Google's first major machine-learning ranking system (2015), which helps interpret the meaning of queries, especially novel or ambiguous ones.
- BERT — Bidirectional Encoder Representations from Transformers (in Search from 2019); a natural-language model that reads words in context, moving ranking from string-matching toward meaning.
- MUM — the Multitask Unified Model (2021); a more powerful, multimodal, multilingual model whose exact role in live ranking Google has kept vague (announced and real; ranking weight unconfirmed).
- Correlation vs. causation — the distinction between two things moving together and one causing the other; the trap behind most ranking-correlation studies (watch for confounders and reverse causation).
Spaced Review
Retrieval practice. Try each before revealing the answer. (Mixing this chapter with Chapter 1.)
- Explain the difference between a ranking signal, a ranking factor, and a ranking system, and use the difference to say why "there are 200 ranking factors" is a myth.
- Google confirmed that HTTPS is a ranking signal. Why is it a mistake to treat "it's a confirmed ranking factor" as meaning "improving it will lift my rankings"?
- A vendor study shows that #1-ranked pages have three times the backlinks of page-two pages. Give two distinct reasons this does not prove that building links will move your rankings.
- (From Chapter 1.) Name the five stages of the pipeline a page passes through to become a result — and say which stage this chapter's "ranking factors" all belong to.
- (From Chapter 1.) A page is indexed but ranks #11. Why is deleting it and starting over usually the wrong move — and which stage is actually the one to work on?