Chapter 33 — Further Reading

Annotated and grouped by the book's three evidence tiers (see Chapter 2 and the style bible). Tier 1 is what you can stand behind; Tier 2 is real but attributed and specifics-unverified; Tier 3 is constructed and illustrative. Start with the Google Search Central documentation — for this chapter, it is unusually authoritative and directly on point.

Tier 1 — Verified / canonical (read these first)

  • Google Search Central — Spam Policies for Google web search ("Doorway pages" and "Scaled content abuse"). The primary source for this chapter. Google names doorway pages and large-scale content abuse as violations and describes what they are. This is the definition your quality bar defends; read Google's own words, not a third party's paraphrase.
  • Google Search Central — "Large site owner's guide to managing crawl budget." Google's own explanation of crawl budget as crawl capacity limit + crawl demand, who actually needs to manage it (very large and large-fast-changing sites), and the recommended techniques. The backbone of §33.4–33.5.
  • Google Search Central — "Verifying Googlebot and other Google crawlers." The official method for confirming a request is really from Googlebot (reverse/forward DNS lookup; published IP ranges). Read before you trust a single log line.
  • Google Search Central — "How Google Search crawls pages" / crawling documentation, and the Crawl Stats report help. How crawling works and how to read the free Crawl stats report — the small-site substitute for raw log analysis.
  • Google Search Central — "Multi-regional and multilingual sites" and general guidance on faceted/duplicate URLs. Useful background for keeping a large site's URL space clean (pairs with Chapter 14 and Chapter 31).
  • Google Search Central Blog — Matt Cutts / Google Webmaster Central posts on the BMW.de removal (2006). The contemporaneous public record of the doorway-page delisting in Case Study 1. Primary-source confirmation that doorway enforcement reaches the largest brands.
  • Google Search Central — the Panda / Helpful Content system explanations (see also Chapter 6). Background for the site-level quality assessment that Case Study 2 turns on.

Tier 2 — Attributed, specifics unverified (useful, read critically)

  • The SEO community's analyses of eBay's May 2014 organic decline (e.g., contemporaneous write-ups and third-party visibility indices such as Searchmetrics). The source of the causal story in Case Study 2 — informative, but analysis rather than Google confirmation, and reported magnitudes vary by source. Treat as Tier 2 and label it so.
  • Practitioner guides to programmatic SEO (industry blogs and conference talks). Good for pattern and tooling ideas — how sites structure datasets, templates, and internal-link modules at scale. Read them for the how, and apply this chapter's quality bar to the whether: much programmatic-SEO advice online underweights the doorway line.
  • Large-scale internal-linking and crawl-optimization write-ups from enterprise SEO teams and tool vendors. Useful, concrete tactics; keep in mind that vendor content is also marketing, and that "we saw X% lift" claims are correlational and unverified (Chapter 2).
  • Log-file-analysis tooling documentation (e.g., Screaming Frog Log File Analyser, Botify, OnCrawl, and similar). For the mechanics of parsing and segmenting logs at scale. Chapter 30 covers the professional toolset generally; this is its crawl-analysis corner.

Tier 3 — Illustrative / constructed (this book's own scaffolding)

  • Rivertown Home Services — the constructed running project; the service×city matrix designed in this chapter's Strategy File. All figures and local details are illustrative teaching examples.
  • Summit Gear Co. — the constructed e-commerce example from Chapter 31, referenced here for the faceted-navigation and index-bloat parallels.
  • Figures 33.1–33.3 — constructed teaching examples (the two-draft comparison, the crawl-distribution report, and the service×city system design). Numbers are round and illustrative, never presented as real data.

Suggested order

  1. Google's Spam Policies — Doorway pages & Scaled content abuse (Tier 1). Anchor the quality bar in Google's own definition before anything else.
  2. Google's "Large site owner's guide to managing crawl budget" (Tier 1). Get the capacity/demand model and the honest scoping straight from the source.
  3. Google's "Verifying Googlebot" + the Crawl Stats help doc (Tier 1). The practical tools of §33.4–33.5.
  4. The BMW.de (2006) record (Tier 1) and then the eBay 2014 analyses (Tier 2). Read the clean enforcement case first, then the contested enterprise-scale one — and notice the tier shift between them.
  5. A reputable programmatic-SEO guide and a log-analysis tool's docs (Tier 2), applying this chapter's quality bar as your filter throughout.