> "Search traffic answers a question the reader asked. Discover traffic answers a question the reader didn't
Prerequisites
- 12
- 13
Learning Objectives
- Explain how a site actually gets into Google News today — automated inclusion, news sitemaps, Top Stories — and debunk the myth that you submit or pay your way in.
- Describe Google Discover as a query-less, interest-driven surface, and explain why there is no keyword to target and no lever that guarantees entry.
- Distinguish breaking, analysis, and evergreen content by freshness need, traffic shape, and decay — and apply the content-audit discipline to a publisher's archive at scale.
- Explain what NewsArticle schema, publisher and author E-E-A-T signals, and Web Stories do and do not accomplish.
- Set up a paywall that Google can index without cloaking, using flexible sampling and paywall structured data.
- Manage content syndication with canonical attribution honestly — knowing a cross-domain canonical is a hint, not a guarantee.
- Make the diversification argument for publishers: why depending on a single algorithmic traffic source is the most fragile position in SEO.
In This Chapter
- Overview
- Learning Paths
- 34.1 Google News and the news ecosystem
- 34.2 Google Discover: the feed that giveth and taketh
- 34.3 Breaking, analysis, and evergreen: the publisher's portfolio
- 34.4 NewsArticle schema, author E-E-A-T, and Web Stories
- 34.5 Paywalls and flexible sampling
- 34.6 Syndication and canonical attribution
- 34.7 The fragility of one traffic source
- 📈 The Strategy File
- Conclusion
- Key Terms
- Spaced Review
Chapter 34: Publisher and Media SEO — News, Evergreen Content, and the Discover Feed
"Search traffic answers a question the reader asked. Discover traffic answers a question the reader didn't know they had — which is why it arrives like a gift and leaves like a thief." — a working principle of this book (constructed epigraph, in the narrator's voice)
Overview
Here is the question that keeps publishers awake, and it is a sharper version of a fear every site in this book shares: what happens to a business when almost all of its audience arrives through a door that one company owns and can move, narrow, or wall off at any time — without warning, and with nothing you can point to as "broken"? Rivertown Home Services depends on Google, but it also has a phone that rings, five storefronts, and trucks with its name on them. A digital publisher often has none of that. It has articles, and it has traffic, and a frightening share of that traffic comes from two Google surfaces — search results and a feed most people outside the industry have never heard of, called Discover. When Google adjusts either one, a publisher can lose half its audience in a day and spend the following month with no error to fix, because nothing is broken. The rules simply changed.
This chapter is about that world: news, media, and content publishing, where the stakes of everything the book has taught are turned up to their maximum. It is also, deliberately, the chapter where the book's sixth theme — SEO is a long game, and traffic diversification is insurance — stops being an abstraction and becomes a survival strategy. Publishers are the canaries. What breaks them first, and how they survive, is a preview of pressures every site will eventually feel.
We will use a new running example, because Rivertown is not a publisher and pretending otherwise would teach you the wrong reflexes. Meet WellPath, a constructed mid-size health-and-wellness publisher: a large archive of health explainers and timely health news, credentialed contributors, a medical-review process, and a business that lives on Search and Discover traffic. (WellPath is a health site, which makes it a Your Money or Your Life — YMYL — publisher; we will touch that lightly here and hand the deep treatment to Chapter 35, its true home.) Everything about WellPath is a labeled teaching construction, including every number.
In this chapter, you will learn to:
- Get into Google News the way it actually works now — and stop paying for "News inclusion" services that do nothing.
- Understand Google Discover as a surface with no query, no keyword to target, and no guaranteed way in.
- Balance a content portfolio of breaking, analysis, and evergreen — and prune a decayed archive the way Chapter 12 taught, at publisher scale.
- Mark up articles with NewsArticle schema and build the author and publisher E-E-A-T that news demands.
- Build a paywall Google can read without treating it as cloaking.
- Syndicate content without handing your rankings to a bigger site.
- Make the honest, urgent case for diversifying away from a single algorithmic traffic source.
Learning Paths
📝 Content Creator: this is your chapter — every section is aimed at you, and §34.2 (Discover), §34.3 (the breaking/evergreen portfolio), and §34.7 (why you cannot live on Google alone) are the ones to internalize. 📊 Strategist: weight §34.3 and §34.7 — the content-portfolio economics and the diversification imperative are the strategic core, and they apply far beyond publishing. 🔧 Developer: §34.4 (NewsArticle schema, Web Stories, news sitemaps) and §34.5 (paywall structured data without cloaking) are your build tickets. 🏪 Local Business and 🛒 E-Commerce: you are not publishers, but read §34.2 and §34.3 anyway and then the Strategy File — the evergreen depth and compelling title-and-image thinking that publishers live by is exactly what your neglected blog or content-marketing program should borrow, and knowing what not to copy (the news treadmill) is half the lesson.
34.1 Google News and the news ecosystem
Start by separating two things people constantly conflate. Google News is a specific product — the News tab in Search, the standalone news.google.com site, and the Google News mobile app — a curated, personalized stream of journalism. But most of the search traffic that news content actually earns does not come from that product at all; it comes from the Top Stories carousel that appears inside ordinary Google Search results for queries Google judges newsworthy, and from Discover (§34.2). So when a publisher says "we need to get into Google News," what they usually mean, and should mean, is "we need to be eligible for the news-related surfaces Google shows," which is a broader and more useful goal.
THE THREE PUBLISHER DOORS [schematic — not to scale]
GOOGLE SEARCH (a QUERY) GOOGLE DISCOVER (NO query)
┌───────────────────────────────┐ ┌────────────────────────────────┐
│ the ten organic links │ │ a personalized feed of cards │
│ ▸ TOP STORIES carousel ◀──┐ │ │ ▸ driven by inferred INTEREST │
│ (fresh, newsy queries) │ │ │ ▸ no search terms at all │
│ ▸ "News" tab / Google News │ │ │ ▸ mobile only (app / Android) │
└─────────────────────────────┼──┘ └────────────────────────────────┘
│
one publisher, three doors ┘ — and Google controls the hinges on all three.
News SEO is mostly doors #1 and #2; Discover is door #3 and behaves nothing like the others.
The single most important fact about Google News in the modern era is one that saves publishers real money: you do not submit your site to be included, and you cannot pay to be included. For years there was a manual application process, and a cottage industry grew up around "get your site into Google News" services. That process is gone. Since around 2019–2020, Google automatically considers any site that publishes news content and meets its content policies; inclusion is algorithmic, not an application. The Publisher Center still exists, but its job is presentation and organization — claiming your publication, setting its name and logo, grouping your sections — not gatekeeping whether you appear at all. If someone is selling you "guaranteed Google News inclusion," they are selling you something Google gives away and does not let anyone guarantee.
🚫 SEO Myth: "You have to submit your site to Google News (or pay a service) to appear in it." This is folklore left over from a system Google retired. Today, eligibility for Google News and Top Stories is automatic for sites that publish genuine news content and follow Google's news content policies — no submission, no fee, no "News inclusion package." What the Publisher Center actually does is let you manage how your publication is presented (its title, logo, and sections). Anyone charging you to "get into Google News" is charging for a door that is already open. Spend the money on reporting and on the technical eligibility basics below instead.
What does make a site eligible and competitive for news surfaces? A cluster of things, none of them a trick:
- Genuine, original news content that follows Google's news policies (no scaled spam, no dangerous or deceptive content, transparent authorship and dates).
- Clear publication dates and datelines, so Google can tell how fresh a story is — freshness is the currency of news.
- A stable, crawlable URL structure and the technical foundation from Part III; a story Google cannot crawl in time is a story that misses the news window entirely.
- A news sitemap for very fresh content (below), on top of your normal XML sitemap.
- Author and publisher trust signals — bylines, author pages, an about/masthead, a corrections policy — the publisher form of E-E-A-T we deepen in §34.4.
A word on the news sitemap, because it is genuinely news-specific and it is the one piece of technical
plumbing this chapter adds to what you already know. Your ordinary XML sitemap (owned by Chapter 14) is
a discovery aid for your whole site. A news sitemap is a specialized, additional sitemap that lists only
articles published in the last 48 hours, with a <news:publication> block giving the publication name
and the article's publication date. Its entire purpose is speed: it tells Google "these specific URLs are
fresh news, right now," so the crawler prioritizes them during the short window when a breaking story is
worth ranking. Articles roll off the news sitemap after two days. It does not replace your regular sitemap;
it sits alongside it, and it only matters if you actually publish time-sensitive news.
🔎 How Search Sees It There is no switch in Google's systems labeled "this is a news publisher." Google infers newsworthiness from a combination of signals: content that looks like journalism (timely, dated, reported), a site with a history of publishing such content, freshness signals (a news sitemap, fast crawling, clear timestamps), and enough authority and trust that Google is willing to surface you for queries where showing bad information does real harm. For queries Google judges to deserve freshness — a developing event, a "what happened" search — it assembles a Top Stories set from eligible, trusted, recent sources. For the same topic a week later, when the query no longer deserves freshness, those same news articles fade and evergreen explainers take over the results. The surface you can win depends less on your page than on what Google thinks the query wants right now — which is Chapter 3's intent lesson, wearing a press badge.
That last point deserves a name you will hear in publishing circles: QDF, "query deserves freshness." It is old Google terminology, describing the long-acknowledged behavior that for some queries — breaking events, trending topics — Google boosts fresher results, while for stable queries ("how does a heat pump work") freshness barely matters. Google has confirmed, in general terms, that freshness is a query-dependent signal: it helps for some searches and is irrelevant, or even counterproductive, for others. Do not treat "freshness" as a universal good to chase on every page; treat it as something specific queries reward and most do not.
What news SEO can do is put timely, well-reported content in front of a large audience during the brief window a story is hot, through Top Stories and the News product. What it cannot do is manufacture that window for content that isn't genuinely newsy, keep a story ranking after its moment passes, or substitute for the authority and trust that decide which trusted sources Google is willing to feature. News is a speed-and-trust game played in a window that closes fast.
34.2 Google Discover: the feed that giveth and taketh
Now the strangest, most seductive, and most dangerous surface in all of SEO. Google Discover is a personalized content feed — a scrolling stack of article cards — that appears in the Google app and on the home screen of many Android phones, below the search box, before anyone searches for anything. That last clause is the whole point, and it breaks the mental model this entire book has been building. Every other surface we have studied answers a query: a person types words, and you try to be the best match. Discover has no query. Google decides, based on a user's inferred interests, activity, and location, what to show them — and if it decides your article fits, it puts your card in front of them unbidden.
Sit with the implication, because it is disorienting: there is no keyword to target in Discover. There is no query to match, no SERP to read, no position to track. You cannot "optimize for a Discover keyword" because there isn't one. Everything you learned about intent and keyword research aims at a searcher who told you what they want. In Discover, nobody told you anything. Google is guessing at interest, and your job is to be the kind of content it is willing to guess with.
So how do you become eligible? Google's own guidance is unusually humble here, and honesty compels us to match it. You do not opt in to Discover, and there is no guaranteed way in. Eligibility is essentially: be indexed in Google Search (Discover draws from the same index), follow Discover's content policies (which are stricter than Search's), and — the part you can actually influence — publish content people find genuinely interesting, with a few presentation ingredients that Google has said help:
- Compelling, accurate titles that capture the essence of the content without being clickbait. This is a real tension: Discover rewards curiosity but explicitly penalizes exaggerated or withheld-information headlines ("You won't believe what happened next").
- High-quality, large images. Google has said that large images — roughly 1,200 pixels wide or more —
can improve a page's presence in Discover, and that you enable them by allowing large image previews (the
max-image-preview:largerobots meta setting). A great image is not optional garnish in Discover; it is much of the card. - Content that demonstrates expertise and trustworthiness — the E-E-A-T signals from Chapter 5, which matter more in Discover, not less, because Google is putting content in front of people who didn't ask for it and must be more careful about what it endorses.
- Timeliness and interest, without requiring hard news. Discover loves "interesting to a human right now" — a strong evergreen explainer can surface months after publication if it matches a rising interest.
⚖️ Evidence Check Claim: "Add a 1,200px image and
max-image-preview:largeand you'll get Discover traffic." Sort it. — Confirmed by Google: large images and enabling large previews are documented as things that help a page's appearance and eligibility in Discover; the guidance is real and worth following. — The honest limit: these are necessary-ish enablers, not sufficient causes. Google states plainly there is no guaranteed way to appear in Discover; eligibility is not entry. Countless well-illustrated, correctly-configured pages never get a single Discover impression, and the ones that do are chosen by an interest-matching system Google does not expose. So: do the image and preview work because it removes a known blocker — but anyone promising Discover traffic because you did it is selling correlation as control. — Speculation to reject: any specific "Discover ranking factors" list with weights. There is no position and no published factor set; treat detailed Discover "algorithms" you read online as guesswork.
Here is why Discover is called "the feed that giveth and taketh." When it works, it is astonishing: a single article can earn more traffic in a day from Discover than from a month of Search, because Google pushed it to hundreds of thousands of feeds. When it stops, it stops without notice or explanation — the same article goes cold, and the next twenty like it never catch at all. Publishers who reorganize their whole operation around Discover — chasing big images and curiosity-gap headlines — build a business on a surface with no dials, no query, and no accountability. That is the trap this chapter keeps circling back to.
📄 Read the Report
text FIGURE 34.1 — "The Discover report: a different shape of traffic" [constructed teaching example] THE QUERY / PAGE Search Console's "Discover" performance report for WellPath, last six months. WHAT'S THERE A jagged line. Most days sit near ~4,000 clicks; then four sharp spikes to 40,000–90,000 clicks on single days, each tied to ONE article, each falling back to baseline within ~48 hours. There is NO "queries" view — Discover has no search terms, only pages, impressions, clicks, and CTR. WHAT IT SHOWS Discover traffic is real and can be huge, but it is bursty and article-driven, not a steady base. One well-matched, well-illustrated piece can out-earn a month of Search — once. WHAT IT DOESN'T It cannot tell you WHY those four pieces hit while 400 others didn't; there is no keyword to target and no lever to pull to reproduce a spike on demand. Correlation, never control. THE MOVE Treat Discover as UPSIDE, never as baseline. Bank each spike — capture emails, add internal links to evergreen pages, retarget the reader — but forecast and budget on Search plus owned channels, which you can actually influence. THE LESSON You do not "rank" in Discover; you become eligible and hope to be matched. Build for Search, be delighted by Discover, and never let it become the floor you stand on.🛠️ Try It on Your Site Open Google Search Console. If a Discover report appears in the left sidebar, your site has received Discover traffic in the last 16 months — open it and look at the shape: is it a steady stream or a few giant spikes? (If there's no Discover report at all, you've simply never been surfaced there, which is the norm for most sites.) Then check one thing you can control: view any article's page source and search for
max-image-preview. If it saysmax-image-preview:large, you've enabled large Discover images; if it'sstandardor missing, you may be capping your own image previews. That is the rare Discover lever that is actually in your hands.
What Discover can do is deliver enormous, top-of-funnel exposure to content that matches a rising interest — reach you could never buy or target. What it cannot do is be relied upon: it offers no query to optimize for, no guaranteed entry, no stable baseline, and no explanation when it giveth or when it taketh. Use it as a windfall; never as a foundation.
34.3 Breaking, analysis, and evergreen: the publisher's portfolio
Every publisher is really running two businesses with opposite economics, and a smart one runs a third in between. Understanding the three — and the very different way each behaves in search — is the strategic heart of publisher SEO.
Breaking news answers "what just happened." Its value is almost entirely in the first hours; its traffic is a sharp spike followed by a fast decay, because tomorrow the story is old and QDF has moved on. Evergreen content — and this is a term this chapter owns — is content whose value and relevance persist over time: a guide to "how a heat pump works," an explainer on "what causes migraines," a definitive "best practices" piece that is as useful in two years as today, provided it is kept current. Its traffic builds slowly and compounds, the way an asset does. Between them sits analysis and explainer content — "what does it mean," "why this matters" — which rides a current topic but with depth that outlives the initial spike, and which, at its best, ages into evergreen.
THREE CONTENT TYPES, THREE TRAFFIC SHAPES [schematic — not to scale]
BREAKING ANALYSIS / EXPLAINER EVERGREEN
traffic traffic traffic
│ █ │ ▄ │ ▄▄▄▄▄▄▄
│ █ │ ▄█▄ │ ▄▄▄▄▄▄
│ █▄ │▄███▄▄ │ ▄▄▄▄
│ ██▄▄ │██████▄▄▄▄ │▄▄
└──────── time └──────────── time └──────────────── time
spike, dies in medium spike, slow slow build, COMPOUNDS
hours–days decay over weeks for years (if updated)
| Breaking news | Analysis / explainer | Evergreen | |
|---|---|---|---|
| Answers | "what just happened" | "what does it mean / why" | "how / what is X" (timeless) |
| Freshness need | extreme (minutes–hours) | high (days) | low (refresh ~yearly) |
| Traffic shape | huge spike, fast decay | medium spike, slow decay | slow build, compounds |
| Typical half-life | hours to days | weeks to months | years, if maintained |
| Main SEO lever | speed, news sitemap, QDF | depth, angle, timeliness | comprehensiveness + updating (Ch 12) |
| The failure mode | a dead archive at scale | never updated, decays | decays silently if abandoned |
The trap that catches most publishers is treating everything as breaking. A newsroom optimized purely for speed produces an enormous archive of once-hot articles that are now permanently cold — thousands of pages that earned their spike and will never earn another click. This is exactly the situation the book's running content-audit anchor describes: a site with a huge archive, most of it drawing zero traffic, where the counter-intuitive cure is not "publish more" but prune, update, and consolidate so the site's overall quality signal rises and the pages that can still work get room to breathe. Chapter 12 is the home of that discipline (the keep / update / merge / delete audit, historical optimization, and cannibalization). Here we simply note its publisher-scale version, which is severe.
Picture WellPath's archive: 9,000 published articles accumulated over a decade. Suppose an audit finds that roughly 6,000 of them are decayed breaking-news and time-bound pieces that now earn essentially nothing, while a core of maybe 800 evergreen health explainers drive the majority of the site's durable Search traffic. The move is not to delete the 6,000 reflexively — some are historical record, some still earn a trickle, some can be merged into evergreen hubs — but to stop letting a mountain of dead content define the site's quality to Google's systems. The Helpful Content and core-update machinery (Chapter 6) assesses quality partly at the site level: an archive that is 70% abandoned, thin, or outdated is a site telling Google, at scale, that it produces low-value pages. Pruning and updating is not tidying; it is repairing a site-wide signal, and for a large publisher it can matter more than any single new article.
🔗 Connection The audit mechanics — how to score pages by traffic × rankings × conversions × quality, and the four decisions (keep / update / merge / delete), plus historical optimization (updating an existing page often beats publishing a new one) and keyword cannibalization (two pages fighting for one query) — are owned by Chapter 12 (Content Updating, Pruning, and the Lifecycle of Published Content). This section is that chapter applied at publisher scale; the freshness signal it leans on is defined there and in Chapter 6 (Google Updates). Do not re-audit from scratch — bring Chapter 12's method to the archive.
The strategic payoff is a portfolio, not a monoculture. Breaking news buys you relevance, brand, and the occasional Discover or Top Stories windfall — but it is a treadmill: stop running and the traffic stops. Evergreen content is slower and less glamorous, but it is the compounding asset that still earns while you sleep and does not evaporate when the news cycle turns. The healthiest publishers run both deliberately: they use breaking coverage to build audience and authority, and they convert the durable interest into evergreen explainers that keep paying out for years. Theme 6 — the long game — is not a slogan for a publisher; it is the difference between a business and a hamster wheel.
🔄 Check Your Understanding A news site's editor is proud that the team publishes 40 breaking-news posts a day and has 200,000 indexed URLs. Traffic, however, is flat and a recent core update hurt them. Name the likely problem in the archive, and the two moves (from Chapter 12) that address it — and say why "publish even more" is the wrong response.
Answer
The likely problem is a vast archive of decayed, thin, once-breaking pages dragging down the site-level quality signal that core and Helpful Content systems assess — most of those 200,000 URLs earn nothing and collectively say "low value at scale." The two moves: (1) prune/consolidate — noindex, merge, or remove the genuinely dead thin pages; (2) update (historical optimization) the evergreen and still-relevant pieces so they regain freshness and completeness. "Publish more" adds to the pile that is already the problem; the cure is subtraction and improvement, not addition. (This is the content-audit anchor at scale.)
34.4 NewsArticle schema, author E-E-A-T, and Web Stories
Three publisher-specific tools help Google understand and present your content. None of them is a ranking trick, and it matters to say so up front: two of the three are about being understood and being eligible for richer presentation, and the third is a content format. Let's take them in turn.
NewsArticle schema. Structured data — the general mechanism of describing your page's meaning to
machines in a format like JSON-LD (JavaScript Object Notation for Linked Data) — is owned by Chapter 18.
What this chapter adds is the news-specific type. NewsArticle schema is the Schema.org type for a news
article (a subtype of the broader Article) that lets you state, explicitly and machine-readably, the
things a news surface needs to know: the headline, the datePublished and dateModified, the author
(as a person or organization, ideally linked to an author entity), the publisher, and the article images.
Marking this up does not push you up the rankings — like all schema, it is not a direct ranking factor —
but it removes ambiguity about when your story was published and updated and who stands behind it, which
is exactly the metadata news and Top Stories eligibility lean on. For a publisher, getting datePublished
and dateModified honest and correct is not cosmetic: misrepresenting freshness (stamping an old article
with today's date to fake newness) is a trust violation Google actively discourages.
{
"@context": "https://schema.org",
"@type": "NewsArticle",
"headline": "New State Rules Change How Clinics Report Wait Times",
"datePublished": "2026-05-14T08:00:00-05:00",
"dateModified": "2026-05-14T15:30:00-05:00",
"author": { "@type": "Person", "name": "Dr. Lena Ortiz", "url": "https://wellpath.example/authors/lena-ortiz" },
"publisher": { "@type": "Organization", "name": "WellPath", "logo": { "@type": "ImageObject", "url": "https://wellpath.example/logo.png" } }
}
Read that as a label, not code you must write: it simply tells Google, in a format it trusts, "this is a news article with this headline, published and last updated at these exact times, written by this specific author, published by this organization." Your content management system (CMS) or a plugin usually generates it; your job is to make sure the dates are truthful and the author links to a real author page.
That author link is the doorway to the second tool, which is the most important for a publisher and the one no markup can fake: author and publisher E-E-A-T. E-E-A-T — Experience, Expertise, Authoritativeness, and Trustworthiness — is Chapter 5's framework, and it is not an SEO tactic bolted onto news; it is the substance of what makes journalism trustworthy, made legible. For publishers specifically, it lives in a recognizable set of signals:
- Real, named bylines linked to substantive author pages — a bio, credentials, areas of coverage, and (for expert content) qualifications. A story attributed to "Admin" or "Staff" is a story with no one standing behind it.
- An author entity Google can connect across the web — the same author, consistently identified, linked
from their profiles and other bylines (the
sameAsidea from Chapter 4). Expertise is easier for Google to credit when it recognizes the person, not just the page. - Organizational transparency — a masthead or "about" page, editorial standards, an ownership disclosure, a corrections policy, and clear contact information. These are the institutional cousins of a personal résumé.
- For expert or YMYL topics, a documented review process — content reviewed by a credentialed expert, with the reviewer named and dated. WellPath's health articles carrying a "Medically reviewed by [named clinician], [date]" line is the canonical example.
🔎 How Search Sees It There is no "author authority score" you can inspect or inject, and Google has been explicit that E-E-A-T is not a single number in the algorithm. What Google's systems actually do is triangulate: they read the byline and author markup, follow the author page and its links, and cross-reference what the wider web says about that person and that publication — citations, other bylines, references, reputation. An author who is a genuine, recognized expert accumulates corroborating signals across the web that a fabricated persona cannot. This is why you cannot fake author E-E-A-T at scale, and why the honest move — real experts, real bylines, real review — is also the only durable one. The markup helps Google find the signals; it cannot manufacture them.
This distinction — markup surfaces a trust signal, it does not create one — is the through-line of every trust topic in the book, and it reaches its highest stakes exactly where content can affect a person's health, money, or safety.
🔗 Connection Author and organizational E-E-A-T for the highest-stakes topics — health, finance, legal — is developed fully in Chapter 5 (E-E-A-T) and, for YMYL specifically, in Chapter 35 (YMYL SEO), which is WellPath's home chapter. Because WellPath is a health publisher, everything here about review processes and credentialed authors is a preview; Chapter 35 is where the YMYL bar, the Medic update, and the ethics of never faking authority get their full treatment. This chapter stays on the publisher mechanics.
The third tool is a format, not a signal: Web Stories. A Web Story is a full-screen, tappable, mobile-first visual story — a sequence of image- or video-led cards with short text overlays, built on an open-web framework — that can appear in Discover, in Search, and in a dedicated Stories surface. It is, lineage-wise, the open-web successor to what used to be called AMP Stories. Web Stories shine for visual, snackable, mobile content — a photo-led explainer, a "5 signs of X" walkthrough, a visual recap — and they are one of the few formats explicitly surfaced in Discover. Their limits are real too: they are production-intensive, they suit visual topics far better than dense reporting, and they are a supplement to your article strategy, never a replacement for it.
🚫 SEO Myth: "Web Stories are just AMP Stories rebranded, and AMP is dead, so don't bother." Two half-truths mangled into a bad conclusion. It's true that Google's AMP (Accelerated Mobile Pages) lost its privileged status — the June 2021 Page Experience update ended AMP's role as the requirement for the Top Stories carousel, and many publishers subsequently dropped AMP (that story is Case Study 1). But Web Stories are not a dead technology; they are an actively supported content format that appears in Discover and Search, and they stand on their own regardless of AMP's diminished role. The right read is nuanced: you no longer need AMP to compete in Top Stories, and Web Stories remain a legitimate — if niche and effortful — way to reach the visual, mobile audience. Don't build Web Stories because you think they're a ranking cheat (they aren't); build them where a visual format genuinely serves the content.
What these three tools can do is help Google correctly understand who published what, when, and by whom, make you eligible for richer news and visual presentation, and — through real E-E-A-T — earn the trust that news surfaces require. What they cannot do is substitute for the underlying substance: schema cannot make a page rank, a masthead cannot manufacture expertise, and a Web Story cannot rescue thin content. They make genuine quality legible; they do not create it.
34.5 Paywalls and flexible sampling
Publishers who charge for content face a problem the rest of the web does not: Google has to be able to read your content to index and rank it, but if you show all of it to Google you are also showing it to any reader who knows how to look — and if you hide it from Google, you cannot rank at all. For years this felt like a forced choice between visibility and revenue. It isn't, and the resolution is a documented, sanctioned system that every subscription publisher should understand.
First, the rule you must not break. Showing Googlebot different content than you show human visitors — the full article to the crawler, a paywall to the reader — is, by default, cloaking, a violation of Google's spam policies (the anti-deception ethos this whole book argues for). A publisher who serves Google the whole piece while gating humans, with no declaration, is cloaking, even with good intentions. So the naïve "just let Googlebot through" approach is not a clever loophole; it is a guidelines violation waiting to be caught.
The sanctioned resolution is flexible sampling, Google's approach (introduced in 2017, replacing an older policy called "First Click Free") that lets you gate content and remain indexable without cloaking — by declaring the paywall to Google with structured data. There are two flavors:
FLEXIBLE SAMPLING: TWO HONEST WAYS TO GATE CONTENT [schematic — not to scale]
METERING LEAD-IN
reader gets N free articles per period, reader gets a free EXCERPT of every article,
then the paywall appears then the paywall appears
│ │
└─────────────┬────────────────────────────┘
▼
Googlebot is shown the FULL article, BUT the page carries
PAYWALL STRUCTURED DATA that marks the gated section:
"isAccessibleForFree": false
+ the paywalled body identified via a CreativeWork / hasPart block
│
▼
Google can index and rank the full content AND knows it is gated —
so serving the paywall to humans is NOT treated as cloaking. Declared, not hidden.
Metering gives each reader a set number of free articles in a period before the wall goes up (the
familiar "you have 2 free articles left this month"). Lead-in shows everyone a free excerpt — the first
few paragraphs — then gates the rest. Either way, the key is the paywall structured data: you mark the
page with isAccessibleForFree: false and identify the paywalled portion (using a CreativeWork with a
hasPart that flags the gated section). That declaration is what turns a potential cloaking violation into a
legitimate, Google-blessed arrangement: you are not hiding that the content is gated, you are telling
Google, in its own structured-data language, exactly what is free and what is paid. Google can then index
the full article, rank it on its merits, and show searchers a result they may have to subscribe to read —
which Google is fine with, as long as the gating was declared.
🔗 Connection The structured-data mechanics — JSON-LD syntax, the
CreativeWork/hasPartpattern, and testing markup with the Rich Results Test — belong to Chapter 18 (Structured Data and Schema Markup). The cloaking rule this section leans on is part of Google's spam policies, first met in Chapter 6 and threaded through the book's white-hat ethos. This section is where those two ideas — declare your markup, don't deceive the crawler — meet the subscription business model.
There is an honest strategic tension flexible sampling does not dissolve, and the book won't pretend it does. The more you gate, the less a reader (and, in a softer sense, Google's quality assessment of the experience) gets from your page; the more you give away, the weaker your subscription incentive. Metering and lead-in are levers on that trade-off, not solutions to it. And a hard paywall that gives searchers a frustrating dead end can, over time, dampen the engagement signals and reader satisfaction that Google's systems ultimately chase. The technical setup is settled; the business judgment — how much to sample, for whom — is yours, and it is genuinely hard.
⚖️ Evidence Check Claim: "Paywall structured data will hurt (or help) my rankings." Sort it. — Confirmed by Google: paywall/subscription structured data is the supported, correct way to have gated content indexed without it being treated as cloaking. Using it protects you; omitting it while gating Googlebot differently risks a cloaking violation. — Not a ranking lever: the structured data itself is not a ranking boost or penalty — it's a declaration that keeps you compliant and indexable. Your gated article still competes on its content's merits like any other page. — Professional experience / open question: whether a heavy paywall indirectly dampens rankings via worse engagement is plausible and debated, not proven; treat it as a UX-and-business consideration, not a confirmed algorithmic factor. The confident part: declare the paywall correctly. The uncertain part: how aggressively to gate.
What flexible sampling can do is let a subscription publisher be fully indexable and rankable while still charging for content, with no cloaking risk. What it cannot do is resolve the core tension between giving content away for reach and withholding it for revenue — that is a strategic decision no markup makes for you.
34.6 Syndication and canonical attribution
Publishers rarely publish in isolation. A story gets picked up by a wire service; a strong piece is republished by a larger partner; a network shares content across its member sites. This is content syndication — the practice of republishing the same content on other domains, by agreement — and it is both a reach opportunity and a quietly dangerous SEO problem, because it creates the same content on multiple domains and forces Google to choose which one to rank.
Here is the failure mode that stings: you write an original, excellent article. A partner with far more domain authority republishes it verbatim, by arrangement. A searcher looks for the topic, Google sees two (or ten) near-identical pages, and it ranks the high-authority republisher's copy — not yours. You did the work; someone bigger got the traffic. Google is not penalizing you (there is no duplicate-content penalty, as Chapter 31 established); it is simply consolidating and selecting, and authority often tips the selection away from the original author. Syndication done carelessly is a machine for handing your best work to your biggest competitor.
The defense is canonical attribution — using signals, chiefly the cross-domain rel="canonical" link
(and its alternatives), to tell Google which URL is the original that should receive the ranking credit. And
here the book must be scrupulously honest, because the industry oversells this:
⚖️ Evidence Check Claim: "A cross-domain canonical guarantees the original ranks and the syndicated copy won't." Sort it. — The reality:
rel="canonical"is a hint, not a directive (Chapter 14). Google usually respects a clear canonical, but across different domains it is more likely to override the hint when other signals — especially the republisher's authority — point elsewhere. There is no guarantee. — Practitioner experience: partners frequently don't implement the canonical you ask for, or implement it wrong, precisely because they want the traffic too. The attribution you negotiated on paper often doesn't exist in the HTML. — Google's own guidance has shifted toward bluntness: rather than relying on cross-domain canonicals to protect syndicated originals, Google has advised that if you don't want a syndicated copy competing in Search, the more reliable move is for the republisher tonoindextheir copy. A canonical is a polite request; anoindexis a closed door. So: ask for the canonical, but know it is a hint you don't control, and negotiatenoindexor non-compete terms when the ranking genuinely matters to you.
Your realistic options, from strongest to weakest control:
| Attribution option | What you ask the republisher to do | Reliability | The catch |
|---|---|---|---|
noindex the syndicated copy |
Keep their republished version out of Google entirely | Strongest — if honored | The partner may refuse; they often want the search traffic too |
Cross-domain rel="canonical" |
Point their copy's canonical at your original URL | A hint, sometimes ignored | Google may still rank their copy if their authority is higher |
| A prominent link back to the original | Editorial credit + a followed link near the top | Weak on its own | Doesn't stop their copy from outranking yours |
| Syndicate a variation, or delay | Publish a differentiated or later version | Reduces direct competition | More work; not always possible under the deal |
| Publish first, get indexed first | You publish and are crawled before partners | Helps establish origin | Not decisive against a much stronger domain |
The strategic reading is simple: syndication trades control for reach. If reach is the goal — brand
exposure, referral traffic, relationships — syndicate freely and don't fret about the ranking. If search
rankings for that content are the goal, protect the original with the strongest attribution the deal
allows (ideally the partner's noindex), publish and get indexed first, and consider syndicating a
differentiated version rather than a verbatim copy. What you must not do is assume a cross-domain canonical
has quietly protected you; verify it exists, and understand it is a hint that a stronger domain can override.
🔗 Connection The
rel="canonical"tag, how Google treats it as a hint, and duplicate-content selection are owned by Chapter 14 (Technical SEO Fundamentals); the "duplicate content is not a penalty, it's selection and consolidation" principle was developed for catalogs in Chapter 31 (E-Commerce SEO). Syndication is the cross-domain version of the same mechanism — with the twist that you don't control the other domain.
What canonical attribution can do is improve the odds that Google credits the original when you syndicate. What it cannot do is guarantee it across domains, or force a republisher's authority to defer to yours. Reach and ranking-control are, in syndication, partly at odds — decide which you're buying before you sign the deal.
34.7 The fragility of one traffic source
Everything in this chapter converges on a single, uncomfortable truth, and it is the reason publishers are the test case for the whole book's sixth theme. A publisher whose traffic comes almost entirely from Google Search and Discover has built its business on ground it does not own and cannot control. A core update, a Helpful Content adjustment, a Discover mood swing, or an AI Overview that answers the question before the click — any one of these can cut the audience in half overnight, and the cruelest part, as we saw in Chapter 6, is that a broad quality reassessment often leaves nothing specific to fix. You cannot file a bug report against an algorithm's opinion of your site.
This is not hypothetical, and it is not only a publisher problem — it is just worst for publishers, because they have the least outside the funnel. A local business has a physical location, a phone, repeat customers, and word of mouth. A store has an email list and a brand. A pure content publisher that neglected everything but Google has articles and a cliff edge. When the ground moves, the diversified site wobbles and the monoculture falls.
So the durable response — the strategic conclusion of this chapter — is diversification: deliberately building traffic and relationships that do not depend on Google's next decision.
- Owned audience: email and newsletters. The single most valuable asset a publisher can build, because it is a direct line to readers that no algorithm mediates. A subscriber who opens your newsletter is yours in a way a Discover impression never is. Every Discover spike (§34.2) should be spent, in part, converting anonymous readers into subscribers.
- Direct and brand traffic. People who type your name, bookmark you, or open your app are people who found you without Google's permission. Brand is the moat; a publication readers seek out by name is resilient to ranking changes in a way a publication readers only stumble onto is not.
- Community and membership. Forums, comment communities, paid memberships, events — relationships that give readers a reason to return that has nothing to do with a search result.
- Multiple discovery channels. Social platforms, podcasts, syndication (§34.6, for reach), video — no single one a substitute for Google, but collectively a portfolio rather than a single point of failure.
🔗 Connection The systematic treatment of traffic diversification — as a named strategy and the answer to the zero-click, AI-Overview future — is owned by Chapter 36 (AI Search, SGE, and AI Overviews), where the book's "AI Overview that took 30% of the clicks" anchor pays off. Surviving specific algorithm updates is Chapter 6. This section makes the publisher-specific case; Chapter 36 makes it general and takes it into the AI era. The two are one argument: never let a single algorithmic channel become your only floor.
If all of this reads as an indictment of Google, it isn't — and seeing why sharpens the strategy rather than softening it.
🔎 How Search Sees It It's worth naming why Google itself would tell you to diversify. Google's systems are built to reward the best result, and "best" is a moving, contested, machine-learned judgment reassessed constantly. Google does not owe any publisher its previous traffic, and it changes its systems hundreds of times a year in pursuit of better results for users, not stable traffic for sites. That is not hostility; it is the job. Understanding that Google optimizes for searchers, not for your revenue, is what makes diversification obvious rather than paranoid: you are insuring against a system that was never designed to keep you safe, only to keep its users well-served.
The content-audit anchor closes the loop here. The publisher that pruned its dead archive, invested in compounding evergreen, built real author E-E-A-T, and converted spikes into an owned email audience is resilient: when an update hits, it has depth Google still values and channels Google doesn't touch. The publisher that chased Discover with thin, curiosity-gap content and lived entirely on the feed is fragile: when the feed turns, it has nothing. Same industry, same algorithms, opposite outcomes — and the difference is every principle in this chapter, applied or ignored.
What diversification can do is turn a catastrophic single-point-of-failure into a survivable wobble, and buy a publisher the time and independence to recover from any one channel's bad quarter. What it cannot do is replace the reach of Google at its best, or be built overnight — an email list and a brand are themselves long-game assets, which is why the time to start is before the algorithm moves, not after.
📈 The Strategy File
A candid note first, in the spirit of Chapter 31's: Rivertown Home Services is not a publisher, and most of this chapter does not apply to it. Marisa and Tony Delgado — running the HVAC, plumbing, and electrical company their father Ray founded in 1984 — are not going to compete in Google News, and they should not build Web Stories at scale, chase Discover traffic as a strategy, or install a paywall. Trying to make a home-services company behave like a media outlet would be exactly the wrong lesson. So this is a brief: what Rivertown's neglected blog can legitimately borrow from publisher and Discover thinking — and, just as important, what it should ignore.
FIGURE 34.2 — "What Rivertown's blog borrows from publishers — the brief" [the Strategy File]
WHAT TO BORROW (yes) HOW IT APPLIES TO RIVERTOWN
─ Evergreen depth over the treadmill Rivertown's wins are EVERGREEN, not breaking: "why is my furnace
blowing cold air," "how to relight a pilot light," "AC vs. heat
pump." These compound for years (theme 6) — the opposite of news.
Invest here; never try to be "timely."
─ Compelling, HONEST titles + strong A blog post titled "Furnace Blowing Cold Air? 6 Causes and How to
images Fix Them" with a clear, real photo beats "HVAC Services." Borrow
the publisher's craft of an inviting, accurate headline and a good
image — WITHOUT the clickbait Discover punishes.
─ Author E-E-A-T, small-business scale Bylines from named, licensed Rivertown technicians with a short
bio ("Tony Delgado, master electrician, 18 years") — the §34.4
author-trust idea sized for a local business (deepened in Ch 5).
─ The content-audit discipline The neglected blog gets the keep/update/merge/delete pass — but
that audit is CHAPTER 12's Rivertown task; here we just note the
blog should be pruned and the evergreen pieces kept fresh.
WHAT TO IGNORE (no) WHY
─ Google News / Top Stories / news Rivertown publishes no news; it is not a news publisher. No news
sitemap sitemap, no QDF chasing, no newsroom cadence.
─ Discover as a strategy Nice if a how-to post ever surfaces there, but never a plan —
there's no keyword to target and no reliable entry (§34.2).
─ Paywalls / syndication Rivertown wants its content SEEN by local searchers, not gated or
republished. The whole monetization model is service calls.
THE HONEST READ Borrow the publisher's EVERGREEN + QUALITY + author-trust mindset;
skip the entire news/breaking/Discover machine. And take the
diversification lesson to heart at local scale — Rivertown's phone,
reviews, and repeat customers ARE its traffic diversification.
(All Rivertown details are a constructed teaching example.)
Your Strategy-File task for this chapter: look at your own site's (or a client's) content and sort a sample of it into the three buckets from §34.3 — breaking, analysis, evergreen. Most non-publishers will find they have almost no true "breaking" content and that their durable traffic comes from a handful of evergreen pages. Write one sentence naming your best evergreen asset and one naming a page that decayed and needs the Chapter 12 treatment. Then answer the harder question, the one that matters most: if Google cut your organic traffic in half next month, what other channel would keep the business alive? If the honest answer is "nothing," you have just found your most important project — and it is not an SEO project at all.
Conclusion
Publisher and media SEO is the whole book at maximum intensity. It takes intent (Chapter 3), content quality and the audit discipline (Chapters 9, 12), E-E-A-T (Chapter 5), structured data (Chapter 18), and technical foundations (Part III), and it plays them for the highest stakes, because a publisher has the least to fall back on when Google changes its mind. We separated the three doors — Search and Top Stories, Google News, and the query-less Discover feed — and saw that News inclusion is automatic (not something you buy) and that Discover offers no keyword to target and no guaranteed way in. We distinguished breaking, analysis, and evergreen content by their opposite traffic economics, and brought Chapter 12's pruning-and-updating discipline to a bloated archive at publisher scale. We covered the tools — NewsArticle schema, real author and publisher E-E-A-T, and Web Stories — and were careful that none of them is a ranking trick. We built a paywall Google can read without cloaking, via flexible sampling and paywall structured data. We faced the syndication trap honestly, where a cross-domain canonical is a hint a stronger domain can override. And we ended where the theme demanded: with the fragility of one traffic source, and the diversification that is a publisher's only real insurance.
We were honest about the limits throughout. You cannot guarantee Discover traffic, cannot force a republisher's authority to defer to yours, and cannot fix a core-update reassessment with a checklist. What you can do is build compounding evergreen depth, earn genuine author trust, keep your archive clean, and own an audience Google doesn't mediate — so that when the algorithm moves, and it will, you wobble instead of fall.
Next, we go deeper into the highest-stakes corner this chapter kept deferring. Chapter 35 takes on YMYL SEO — health, finance, and legal, where wrong information harms people and Google raises its bar the highest — and it is WellPath's true home. Everything we previewed here about credentialed authors, review processes, and never faking authority gets its full, careful treatment there.
→ Continue to Chapter 35: YMYL SEO.
Key Terms
- Google News — Google's dedicated news product (the News tab in Search, news.google.com, and the News app); today, inclusion is automatic for sites that publish genuine news content and follow Google's news policies — you do not submit or pay to be included.
- Google Discover — a personalized, query-less content feed on mobile (the Google app and many Android home screens) that surfaces content based on inferred user interest rather than a search; there is no keyword to target and no guaranteed way in.
- Web Stories — a full-screen, tappable, mobile-first visual story format built on an open-web framework (the successor to AMP Stories) that can appear in Discover and Search; a supported content format, not a ranking trick.
- NewsArticle schema — the Schema.org structured-data type for a news article (a subtype of
Article) that states the headline, publish/modified dates, author, and publisher; it aids understanding and eligibility for news surfaces but is not a direct ranking factor. - Flexible sampling — Google's approach (introduced 2017, replacing "First Click Free") that lets a paywalled publisher stay indexable without cloaking, using metering (N free articles) or lead-in (a free excerpt) plus paywall structured data that declares the gated content to Google.
- Content syndication — republishing the same content on other domains by agreement (wire services, partner networks); a reach opportunity that creates cross-domain duplicate content and forces Google to choose which copy to rank.
- Canonical attribution — using
rel="canonical"(and stronger measures like the republisher'snoindex) to tell Google which URL is the original that should receive ranking credit for syndicated content; a hint across domains, not a guarantee. - Evergreen content — content whose value and relevance persist over time (and with periodic updating), producing traffic that builds slowly and compounds — the opposite economics of breaking news, and the durable asset in a publisher's portfolio.
Spaced Review
Retrieval practice. Try each before revealing the answer. (This set mixes Chapter 34 with the content audit and pruning of Chapter 12, AI content from Chapter 13, and E-E-A-T from Chapter 5.)
- Explain why there is "no keyword to target" in Google Discover, and what that means for how you should treat Discover traffic in a forecast or budget.
- A publisher believes it must "submit and pay to get into Google News." Correct the myth, and name two things that actually make a site eligible and competitive for news surfaces.
- (From Chapter 12.) A news site has 200,000 mostly-decayed archived URLs and flat traffic after a core update. Which two content-audit moves address the site-level quality problem, and why is "publish more" the wrong response?
- (From Chapter 5.) WellPath is a health (YMYL) publisher. Name three author/publisher E-E-A-T signals it should demonstrate — and explain why author E-E-A-T cannot be faked at scale.
- (From Chapter 13.) A publisher plans to mass-produce breaking-news posts with AI to feed Discover. Using the Helpful Content lesson, explain the risk to the whole site, not just those posts.