40 min read

> *"hreflang is the only corner of SEO where a page can be flawless on its own and still be wrong — because

Prerequisites

  • 14
  • 15

Learning Objectives

  • Decide honestly whether a site actually needs international SEO — and recognize the far more common case where it would add risk and cost with no benefit.
  • Compare the three URL structures for multi-country and multi-language sites (ccTLD, subdomain, subdirectory) and choose one against real constraints rather than folklore.
  • Explain precisely what hreflang does and does not do, and lay out a correct, self-referential, reciprocal hreflang set that obeys the return-tag rule.
  • Distinguish localization from translation, and say why machine translation with no local adaptation is a ranking and trust liability.
  • Diagnose the common hreflang failure modes and validate an implementation with free and standard tools.
  • Describe how Google now infers country targeting after the retirement of Search Console's manual country-target setting.
  • Explain why serving the same content across languages or regions is not a 'duplicate content penalty,' and how hreflang and self-referential canonicals keep the wrong version from showing.

Chapter 20: International SEO — hreflang, ccTLDs, and Ranking in Multiple Languages and Countries

"hreflang is the only corner of SEO where a page can be flawless on its own and still be wrong — because correctness does not live in any single page. It lives in the relationships between them." — a working principle of international SEO [constructed for this book]

Overview

Here is the question this chapter is really about, and it is not the one most people expect: should you be doing international SEO at all?

Almost every other chapter in this book assumes you want to do the thing it teaches. This one does not, because the honest answer for most sites is no — and the most valuable skill an advanced technical SEO can have here is the judgment to say so out loud. International SEO is the machinery for showing the right language and country version of your content to the right searcher: country-code domains, hreflang annotations, localized pages, geotargeting signals. It is genuinely useful for the sites that need it. It is also the single most error-prone corner of technical SEO, the place where competent teams ship broken implementations that quietly bleed traffic for months, and the place where an ambitious "let's go global" project most often produces a fragile, half-localized site that ranks worse than the single clean one it replaced.

So we are going to do something a little unusual. We will spend the first section talking you out of this chapter unless you truly need it — and giving you the specific test for whether you do. Then, for the sites that pass that test, we will teach the machinery properly and honestly: the three URL structures and their real trade-offs; hreflang and the reciprocal "return-tag" rule that trips up nearly everyone; the difference between translating your words and actually localizing your business; the failure modes and how to validate against them; what changed when Google quietly retired the manual geotargeting setting most blog posts still tell you to use; and why "duplicate content across languages" is a myth that scares people out of doing the right thing.

This is an advanced chapter, and it builds directly on two others. You need the crawl-and-index fundamentals from Chapter 14robots.txt, sitemaps, canonical tags, and the difference between crawled and indexed — because half of all hreflang disasters are really canonical-tag disasters wearing a costume. And you need the site-architecture thinking from Chapter 15, because your international structure is an architecture decision, made once, expensive to reverse.

In this chapter, you will learn to:

  • Run the honest "do I even need this?" test before touching a single tag.
  • Choose between a country-code domain, a subdomain, and a subdirectory against your real constraints.
  • Write a correct hreflang set — self-referential, reciprocal, with a sensible x-default.
  • Tell localization from translation, and know why the difference decides whether the pages rank.
  • Find and fix the failure modes that make hreflang silently stop working.
  • Understand how Google now decides which country a page is for, and stop chasing a setting that no longer exists.

Learning Paths

🔧 Developer and 📊 Strategist: this is your chapter — you own the URL-structure decision (§20.2), the hreflang implementation and its failure modes (§20.3, §20.5), and the validation loop. Read all of it. 🛒 E-Commerce: you are the most common legitimate user of this machinery (same product, many countries and currencies), so weight §20.2, §20.3, and §20.7 (cross-locale duplicate content) heavily. 📝 Content Creator: if you publish in more than one language, §20.3 and §20.4 (localization vs. translation) matter; if you publish in one, you can skim. 🏪 Local Business: read §20.1 and the Strategy File, then almost certainly close the chapter and go back to Chapters 15 and 25 — the honest answer for a single-market local business is usually "you don't need any of this," and knowing that is worth the price of admission.


20.1 Do you even need it? (most sites don't)

Let's start by defining the thing precisely, so we can then decide whether it applies to you.

International SEO is the practice of structuring and optimizing a website so that search engines show the right language version and the right country version of your content to each searcher, and so that each version can rank in the market it is meant for. Notice that there are two independent dimensions hiding in that sentence, and confusing them is the first mistake people make:

  • Language targeting — serving Spanish content to Spanish speakers and English content to English speakers, regardless of country. A person in Miami, Mexico City, and Madrid might all want the Spanish version.
  • Country targeting — serving a country-specific experience (prices in the local currency, local shipping, local phone numbers, local law) to searchers in that country, possibly in the same language. The classic case is English: a store might want a different English page for the United States (dollars, "sneakers") than for the United Kingdom (pounds, "trainers").

You can need one, the other, both, or — and this is the point of this section — neither. Most sites need neither. A plumber in one city, a blog written in one language for one audience, a software product sold in one market: for all of these, "international SEO" is not an opportunity, it is a complexity tax with no return. Every additional language and country you add multiplies your page count, multiplies the surface area for technical bugs, splits your authority and your team's attention, and creates pages you now have to keep current forever. If nobody is on the other side of that effort, you have made your site worse to serve an audience that does not exist.

DO YOU NEED INTERNATIONAL SEO?                          [decision sketch — not to scale]

  Do you have real customers — or clear, evidenced demand — in more than
  one COUNTRY, or who speak more than one LANGUAGE?
        │                                        │
        NO                                       YES
        │                                        │
        ▼                                        ▼
  You do not need this chapter's           Can you genuinely LOCALIZE and then
  machinery. One market, one language,     MAINTAIN a full version for each —
  one clean, focused site. Put the         translate it well, keep it updated,
  effort into intent, content, and         and actually serve those customers?
  local SEO (Ch 3, 7–9, 15, 25).                │                        │
                                                 NO                       YES
                                                 │                        │
                                                 ▼                        ▼
                                          STOP. A half-built,       Proceed — pick a URL
                                          stale, machine-           structure (§20.2) and
                                          translated locale is      implement hreflang
                                          worse than not having     (§20.3) carefully.
                                          it at all.

Walk that tree slowly, because the second branch is where good intentions go to die. Plenty of businesses have demand in another market — and still should not build for it yet, because they cannot localize and maintain it. Localization (which we treat fully in §20.4) is not a one-time translation job; it is an ongoing commitment to a second (or fifth) full version of your site, kept accurate, kept current, and backed by the ability to actually serve the customers it attracts. A Spanish landing page that generates leads you cannot answer in Spanish is not a win. A German store page that ranks and then ships nothing to Germany is not a win. If you cannot staff the far side of the effort, the honest move is to wait.

🚫 SEO Myth: "Adding more languages and countries will grow my traffic." It sounds obvious — more markets, more searchers, more clicks. In practice, launching languages and countries you cannot properly localize and maintain is one of the more reliable ways to lose organic traffic. Here is the mechanism: thin or machine-translated locale pages drag down your site's overall quality signal (Chapter 6's Helpful Content thinking applies to every language you publish); a botched hreflang setup can cause Google to show the wrong version and get it bounced; and spreading a fixed amount of authority and editorial effort across five sites instead of one makes each of them weaker. Growth comes from demand you can serve, localized well. It does not come from the existence of more URLs. More pages is not more traffic; more pages that deserve to rank in markets you actually serve is more traffic.

There is an honest middle case worth naming, because it catches a lot of businesses: a single-country company whose local market is multilingual. A U.S. company with a large Spanish-speaking customer base does not need country targeting at all — no country-code domains, no geotargeting — but it might genuinely benefit from a Spanish language version of its key pages, served on the same site, for the same country. That is the lightest possible slice of this chapter (language hreflang on one domain, and real localization), and it is a completely different, far cheaper proposition than "going global." We will return to exactly this case in the Strategy File, because it is the only version of international SEO that touches Rivertown.

⚖️ Evidence Check Claim: "hreflang and international SEO will help me rank in new countries." Sort it honestly. — What's true (and confirmed by Google): hreflang helps Google serve the correct existing version of your page to the right user — it is a documented mechanism for language/region selection. — What's overstated: hreflang is not a ranking booster. It does not lift you into results you would not otherwise reach; it only swaps among your own alternate versions once one of them is already competing. Ranking in a new country still depends on the ordinary things — relevance, quality, authority, local links — done in that market's language and context. — The honest summary: international SEO's technical layer prevents the wrong version from showing and makes the right one eligible. The winning still has to be earned, per market, the ordinary way. Anyone selling hreflang as a growth lever has skipped the sentence that matters.

The rest of this chapter is for the sites that walked the decision tree honestly and landed on "yes, and we can maintain it." If that is not you, you have already gotten this chapter's most important lesson — and you are free to spend the saved effort where it will actually move the needle.


20.2 URL structures: ccTLD vs. subdomain vs. subdirectory

If you do need international SEO, the first real decision — and the most expensive one to get wrong, because reversing it is a full site migration (Chapter 21) — is where each version lives. There are three standard answers, and the debate over them generates more heat than almost any topic in technical SEO.

The three options are:

  • Country-code top-level domain (ccTLD) — a separate domain per country, ending in that country's two-letter code: example.fr, example.de, example.co.uk. A ccTLD is a top-level domain assigned to a specific country or territory.
  • Subdomain — a prefix on your main domain, one per language or country: fr.example.com, de.example.com. A subdomain is a distinct host under your registered domain, treated by search engines as related to, but partly separate from, the main site.
  • Subdirectory (also called a subfolder) — a path on your single main domain: example.com/fr/, example.com/de/. A subdirectory keeps every version on one domain, separated only by the URL path.

Here is the comparison that actually drives the decision:

Structure Example Geo signal to Google Where authority sits Cost & effort Strongest when
ccTLD example.fr Strongest — the domain itself says "France" Split — each domain earns authority separately, from scratch Highest — buy, secure, and run many domains You have a real, standalone country presence and the budget to build authority in each
Subdomain fr.example.com Moderate — relies on other signals Somewhat separate from the root domain Medium Infrastructure or organizational reasons push you to separate hosts
Subdirectory example.com/fr/ Weakest by itself — leans entirely on hreflang and content signals Consolidated — one domain, all your authority in one pot Lowest — one domain, one setup, one certificate You are expanding by language and want every version to benefit from the whole domain's authority

Read the two middle columns together, because they are in tension, and the tension is the decision.

A ccTLD gives the clearest possible geographic signal. example.de is unmistakably for Germany, to both users and Google, and users in Germany often trust a .de address more. But that clarity comes at a steep price: authority does not flow between separate domains. Every ccTLD is, to Google, its own website that must earn its own links and its own reputation from zero. If your .com has spent a decade accumulating authority, your brand-new .fr inherits little of it. You have traded one strong site for a family of weak ones — which can be the right call for a large organization with genuine per-country operations and the resources to build each up, and a slow, costly disaster for anyone smaller.

A subdirectory does the opposite. Because example.com/fr/ lives on the same domain as everything else, it draws on the whole domain's accumulated authority from day one, and you maintain one domain, one security certificate, one analytics setup. The cost is that the URL carries no inherent geographic signal — a path segment /fr/ means nothing to Google on its own — so a subdirectory setup leans entirely on hreflang and on the content itself to communicate targeting. For language-based expansion, most practitioners default to subdirectories precisely because authority consolidation is such a powerful, compounding advantage (Chapter 22 explains why authority is so hard-won that you never want to split it without a strong reason).

The subdomain sits in between and is the option chosen most often for non-SEO reasons — a separate technical stack, a different team, a content-management constraint — rather than because it is optimal for search.

⚖️ Evidence Check Claim: "Subdirectories rank better than subdomains — it's proven." This is one of the longest-running arguments in SEO, and the honest picture is messier than either camp admits. — What Google says (Tier 1): Google has stated it can crawl, index, and rank all three structures, and that it does not have a blanket preference — the choice should be driven by what you can build and maintain. Google treats a subdomain as capable of being understood as part of your overall site. — What practitioners observe (Tier 2): many SEOs report that moving content from a subdomain into a subdirectory correlated with improved performance, and infer that authority consolidates more readily within one host. This is real-world observation and directionally believable — but it is correlational, confounded by the fact that such moves usually come bundled with other improvements, and it is not a Google- confirmed ranking rule. — The honest takeaway: the strongest defensible reason to prefer a subdirectory is authority consolidation, which is a first-principles argument (Chapter 22), not a magic ranking bonus. Choose on maintainability and where you want your authority to sit — not on a folklore claim that one structure is secretly favored.

🔗 Connection Your international URL structure is a site-architecture decision (Chapter 15) with migration stakes (Chapter 21). Because switching structures later means redirecting every URL — a ccTLD-to-subdirectory move is a full domain migration — this is a "measure twice, cut once" choice. Get the structure decision reviewed before you build, not after.

A last, easy-to-miss point about ccTLDs: a ccTLD targets a country, not a language. example.ca (Canada) cannot, by its domain alone, distinguish English-Canadian from French-Canadian visitors — Canada has both. Country and language are different axes, and a domain can only speak to one of them. That is a large part of why hreflang, which can address language and region together, exists at all. Which brings us to it.


20.3 hreflang: the syntax and the return-tag rule

hreflang is an annotation that tells Google which language and (optionally) which regional version of a page to show to which searchers. The name comes from the HTML attribute hreflang ("the language of the referenced document"). It is the central mechanism of international SEO, and it is where most implementations break — not because the syntax is hard, but because the rules that connect pages to each other are unforgiving.

Start with what a single annotation looks like. There are three places you can declare hreflang, and you pick one (never mix them for the same set):

  1. HTML <link> tags in the <head> of each page — the most common method.
  2. HTTP headers — used for non-HTML files such as PDFs, where there is no <head>.
  3. XML sitemap annotations — declared once per URL in the sitemap, which keeps the markup out of the page and is often easier to manage at scale.

Here is the HTML method for a small English/Spanish set with a global fallback:

<!-- Placed in the <head> of the U.S. English page: https://example.com/us/ -->
<link rel="alternate" hreflang="en-us" href="https://example.com/us/" />
<link rel="alternate" hreflang="en-gb" href="https://example.com/uk/" />
<link rel="alternate" hreflang="es-mx" href="https://example.com/mx/" />
<link rel="alternate" hreflang="x-default" href="https://example.com/" />

Three things about that block decide whether it works, and each is a rule people break:

The value is a language code, optionally plus a region code. The language uses the ISO 639-1 standard (en, es, fr, de); the optional region uses the ISO 3166-1 Alpha-2 standard (US, GB, MX), joined with a hyphen: en-us, en-gb, es-mx. You may specify language alone (hreflang="es" for "Spanish speakers anywhere"), but you may not specify a region alone — there is no such thing as hreflang="us" meaning "everyone in the U.S. regardless of language," because a page has to be in some language. Case does not matter to Google, but the codes must be valid. And note the trap already lurking: the United Kingdom's region code is GB, not UKen-uk is a silent error, one of the most common in the wild.

One of the tags must point at the page itself. This is the self-referential requirement: the U.S. page's set includes a tag for en-us pointing to the U.S. page. A set that lists every version except the current one is incomplete, and Google may not process it correctly. Every page names every version, itself included.

x-default names the fallback. The special value x-default marks the page to show when none of your language/region versions is a good match for the user — typically a country/language selector, or your primary global homepage. It is optional but strongly recommended: it is your answer to "a searcher we didn't plan for just arrived; where do we send them?"

Now the rule that breaks more hreflang implementations than everything else combined.

🔎 How Search Sees It Google treats hreflang as a bidirectional, confirmatory signal. It does not simply believe a page's claim that "the German version is over there." It checks that the German version claims this page back. If page A declares that page B is its German alternate, but page B does not declare A as its English alternate, Google sees an unconfirmed, one-sided claim — and it may ignore the pairing entirely. Think of it as a handshake: both hands have to close. This is why hreflang correctness cannot be verified one page at a time; the correctness lives in the relationships, and a single page that forgets to reach back can break the connection for the whole set.

That handshake requirement is the return tag rule, and it is worth stating as a rule you can recite:

The return-tag rule. Every version in an hreflang set must reference every other version and itself, and every reference must be returned by the page it points to. If page A points to page B, page B must point back to page A. A missing return tag means Google may discard that pairing.

THE RETURN-TAG RULE                                       [schematic — not to scale]

     English (US) ──names──▶ English (GB)
          ▲    ▲                 │   │
          │    └─────returns──────┘   │
          │                           │
          └──────────returns──────────┘      ...and each page ALSO names ITSELF
                                             (the self-referential tag).

   If A names B, B must name A. Every page in the set names every page in the
   set, itself included. One missing "return" and Google may throw out that pair.

The practical consequence is that hreflang is an all-or-nothing set, not a per-page tweak. Adding a new language means editing every existing page to add the new version — and adding a return tag on the new pages for every old one. Miss one, and you have a broken pair. This is exactly why the XML sitemap method is so popular at scale: you declare the whole set once, per URL, in one managed file, instead of hand-maintaining matching tags across dozens of templates. The same relationships, expressed in a sitemap, look like this:

<!-- One entry per URL in an XML sitemap; every URL in the set gets its own such block -->
<url>
  <loc>https://example.com/us/</loc>
  <xhtml:link rel="alternate" hreflang="en-us" href="https://example.com/us/" />
  <xhtml:link rel="alternate" hreflang="en-gb" href="https://example.com/uk/" />
  <xhtml:link rel="alternate" hreflang="es-mx" href="https://example.com/mx/" />
  <xhtml:link rel="alternate" hreflang="x-default" href="https://example.com/" />
</url>

Whichever method you choose, one discipline is non-negotiable: use absolute, canonical, 200 OK URLs in every href. Point at https:// (not http://), at the final URL (not one that redirects), at a page that actually returns successfully (not a 404). hreflang that points at redirects, broken pages, or non-canonical URLs is hreflang Google cannot trust — a theme we will make concrete in §20.5.

🛠️ Try It on Your Site If your site has more than one language or region version, run it through a free hreflang checker or a crawl of a few key URLs (Screaming Frog's free tier crawls up to 500 URLs and reports hreflang relationships; several web-based validators check a single URL). Ask three questions of the result: (1) does every page reference itself? (2) does every reference get returned? (3) are all the language and region codes valid (watch for en-uk)? If you do not run an international site, do the thought experiment instead: pick any global brand, view the page source, and find its hreflang block — you will start seeing them everywhere once you know the shape.

What hreflang can do: get the right existing version of your page in front of the right searcher, so a British visitor sees pounds and a Mexican visitor sees Spanish, reducing the "wrong version" mismatch that sends people bouncing back to the results. What it cannot do: make you rank where you otherwise wouldn't, fix thin or badly localized content, or substitute for the ordinary work of being relevant and authoritative in each market. It is a routing signal, not a ranking signal — a crucial distinction we will keep returning to.


20.4 Localization vs. translation

Suppose you have chosen a structure and you understand hreflang. You still have not answered the question that actually decides whether your international pages rank: what goes on them? And here the industry's default answer — "translate the existing pages" — is quietly, expensively wrong.

Localization is the adaptation of content to a specific locale's language and its culture, conventions, and commercial reality — not only the words, but the currency, units, date formats, imagery, examples, idiom, tone, legal and regulatory details, payment methods, and the actual terms local people use to search. Translation is a subset of localization: converting the words from one language to another. You can translate without localizing, and the result usually reads as foreign, converts poorly, and ranks worse than you hoped.

The gap between the two is easiest to feel through search terms, because it is not merely a matter of dialect — it is a matter of what people type:

SAME LANGUAGE, DIFFERENT MARKETS: WHAT PEOPLE ACTUALLY SEARCH   [illustrative]

  Concept          United States (en-US)      United Kingdom (en-GB)
  ────────────     ───────────────────────    ───────────────────────
  athletic shoe    "sneakers"                 "trainers"
  mobile phone     "cell phone"               "mobile phone"
  vacation         "vacation"                 "holiday"
  apartment        "apartment"                "flat"
  trunk (of a car) "trunk"                    "boot"

  A "translation" from English to English changes nothing.
  LOCALIZATION rewrites the page around the words that market's searchers use —
  which means fresh keyword research per locale (Chapter 7), not find-and-replace.

That table is within a single language. Across languages the gap is wider still, and it swallows literal translations whole: a keyword translated word-for-word is frequently not the phrase native speakers search, because search behavior is cultural, not lexical. This is why real localization begins with keyword research in the target market (the Chapter 7 toolkit, re-run per locale) and search-intent analysis for that market (Chapter 3) — because the intent behind a "translated" query can differ, and the format that satisfies it can differ too. You are not translating a page; you are building the right page for a different set of humans and then making sure it speaks their language.

🚫 SEO Myth: "Just run the site through Google Translate and publish it — instant new markets." This is the most common and most damaging shortcut in international SEO, and it fails on three fronts. First, quality: raw machine translation, published without expert review, reads as unnatural to native speakers and undermines trust — and trust is the center of the quality framework (E-E-A-T, Chapter 5). Second, Google's own guidance: Google's documentation specifically discourages publishing untranslated or raw machine-translated content without human review, treating text created purely for search engines with no added value as exactly the kind of low-effort content its systems target (Chapter 6, Chapter 13). Third, relevance: a literal translation misses the local search terms entirely (see the table above), so even a grammatically perfect machine translation targets phrases nobody in that market types. Machine translation is a fine starting draft for a skilled human localizer to adapt — the same human-in-the-loop posture Chapter 13 argues for AI content generally — but publishing its raw output is publishing pages that deserve not to rank, and usually don't.

There is a cost dimension here that the decision tree in §20.1 was pointing at. Localization done properly is expensive and ongoing: every time you update the source page — a new price, a new policy, a new product — every localized version drifts out of date until someone updates it too. A site with five real locales is, in maintenance terms, five sites. This is not a reason to avoid international SEO; it is a reason to enter it with clear eyes, localize the pages that carry real commercial weight, and resist the temptation to auto-generate the long tail.

🔄 Check Your Understanding A team translates its 400-page English site into German using a high-quality machine-translation service, adds correct hreflang, and waits. Six months later the German pages get almost no traffic. Name two distinct reasons rooted in this section — not in hreflang — that could explain it.

Answer (1) No local keyword research. The pages were translated, not localized, so they target literal translations of English keywords rather than the phrases German searchers actually type — they may be optimized for terms with little or no search demand. (2) Quality / trust. Unreviewed machine-translated text reads as unnatural to native speakers and can be treated as low-value, thin content, so it neither earns engagement nor clears the quality bar to rank. (Correct hreflang only ensures the right version is shown; it cannot make a badly localized page deserve to rank.)


20.5 The hreflang failure modes (and validation)

hreflang is famous for being broken in the wild. It is worth understanding why it fails so reliably: it is a distributed, reciprocal system maintained by hand (or by a plugin) across many templates, where a single inconsistency on one page can silently disable a relationship — and where nothing turns red on the page itself. The page looks fine. The set is broken. Here are the failure modes that account for the overwhelming majority of cases:

Failure mode What it looks like Why it breaks The fix
Missing return tag Page A names B; B does not name A Google discards the one-sided pairing Make every version reference every version
No self-reference A page omits the tag pointing to itself The set is incomplete/ambiguous Add the self-referential tag to every page
Invalid codes en-uk, fr_FR (underscore), hreflang="Spanish" Not valid ISO codes — the annotation is ignored ISO 639-1 language + ISO 3166-1 region, hyphen-joined
Canonical conflict The fr page has a canonical pointing to the en page Tells Google the fr page is a duplicate to drop Each version's canonical points to itself
Wrong URLs in href Points to a redirect, an http:// URL, or a 404 Google can't trust or follow the reference Absolute, canonical, https, 200 OK URLs only
Language mismatch Page tagged es-mx is actually written in English The signal contradicts the page's real language Tag the language the page is genuinely in

Two of those deserve special emphasis because they are both common and fatal — they don't degrade the setup, they destroy it.

The canonical conflict is the deadliest, and it is why this chapter has Chapter 14 as a prerequisite. A canonical tag (Chapter 14) tells Google "this is the master URL for this content; index this one." A very common mistake is to build a French page that (a) carries hreflang annotations declaring itself the French alternate, while (b) carrying a canonical tag pointing at the English page — usually because a plugin was configured to canonicalize everything to the "main" language. Read those two signals together the way Google does: the hreflang says "I am a distinct French version, show me to French users," and the canonical says "I am a duplicate of the English page, don't index me, index that instead." They contradict each other, and the canonical usually wins — so the French page is dropped from the index and never shows to anyone. The rule is absolute: in an hreflang set, every version's canonical must be self-referential (the French page canonicalizes to the French page). hreflang points across versions; canonical points at itself. Never cross the streams.

The missing return tag is the most frequent, precisely because of the all-or-nothing property from §20.3. It usually appears when someone adds a new locale and updates the new pages but forgets to go back and add the corresponding return tags to the existing ones — or when two different teams manage two different locales and their tag sets drift apart.

📄 Read the Report

text FIGURE 20.2 — "An hreflang set that quietly fell apart" [constructed teaching example] THE QUERY / PAGE A crawl-tool audit of a retailer's /uk/ and /us/ versions of one product page. WHAT'S THERE /us/ lists hreflang for en-us, en-gb, and x-default. /uk/ lists only en-gb and omits en-us — and has no self-referential tag. On top of that, BOTH /uk/ and /us/ carry a canonical tag pointing to /us/. WHAT IT SHOWS Two independent, fatal problems. (1) The return tag from /uk/ back to /us/ is missing, so Google may ignore the pairing and stop swapping the UK page in for British users. (2) The canonical on /uk/ points to /us/ — telling Google the UK page is a duplicate to be dropped, which would remove it from the index entirely. WHAT IT DOESN'T The crawl cannot show you Google's live decision — hreflang is a signal, not a guarantee, and Google may already be doing something partial. It shows only that the annotations are self-contradictory and cannot be relied on. THE MOVE Give /uk/ the full, self-referential, reciprocal set (en-us, en-gb, self, x-default); change its canonical to point to /uk/. Then re-crawl to confirm before waiting on Google to re-process. THE LESSON hreflang correctness lives in the RELATIONSHIPS between pages. A page can look fine on its own and still break the whole set — and a stray canonical can undo all of it.

That figure points at the central validation problem: because the errors are relational, you cannot find them by eyeballing one page. You need a tool that crawls the set and cross-checks the references. The standard approach:

  • Crawl and cross-check. A crawler such as Screaming Frog (free up to 500 URLs; Chapter 30 covers the professional toolset) fetches your pages, reads their hreflang, and reports missing return tags, missing self-references, invalid codes, and hrefs that don't resolve. This is the workhorse of hreflang validation.
  • Inspect individual URLs with Google's URL Inspection tool (Chapter 27) to see how Google fetched and understood a given page, and to confirm the page is indexed and canonicalized as you intend.
  • Watch the field. After launch, use Search Console's Performance report filtered by country to confirm that traffic for each market is landing on the intended version, not the wrong one.

A candid note about tooling, because it is a place folklore has gone stale: Google used to provide a dedicated "International Targeting" report inside Search Console that listed hreflang errors like "no return tags." Google retired that report (the change rolled out in 2022). A great deal of older advice still tells you to "check the International Targeting report for hreflang errors" — and that report no longer exists. Validation today lives in third-party crawlers and the URL Inspection tool, not in a dedicated Search Console hreflang dashboard. This is a small but perfect example of the book's third theme: SEO advice rots, and an instruction that was correct in 2019 can quietly become impossible to follow. Always confirm that the tool a tip names still exists.


20.6 Geotargeting and Search Console settings

Geotargeting is telling Google which country a site — or a section of it — is meant for, so that its pages are more likely to be shown to searchers in that country. This is the country axis from §20.1 (distinct from the language axis that hreflang handles), and how you set it has changed in a way most guides have not caught up with.

For years, the standard advice was: if you use a generic domain (a .com, .org, or other gTLDgeneric top-level domain, one not tied to any country) and you want it associated with a specific country, open Search Console and set a country target in the International Targeting report. That lever let you say "treat this .com as targeting Canada," and it was the canonical way to geotarget a gTLD.

That setting is gone. When Google retired the International Targeting report (§20.5), it removed the manual country-target control along with it, on the stated reasoning that the setting was little used and that Google had gotten good enough at inferring country relevance from other signals that the manual override was no longer needed. So the honest, current answer to "how do I set my target country in Search Console?" is: you don't — that control no longer exists.

What determines country targeting now, then? Google infers it from a combination of signals:

HOW GOOGLE INFERS WHICH COUNTRY A PAGE IS FOR             [schematic — not to scale]

  STRONG   ccTLD (example.de) ─────────────▶ unambiguous: this domain is for Germany
     │     hreflang region tags (en-gb) ───▶ this version is for British English searchers
     │     On-page & business signals ─────▶ local currency, address, phone (NAP), local language
     │     Links from within the country ──▶ who references you, and from where
  WEAK     Server / IP location ───────────▶ historically minor; with global CDNs, now largely moot

Two of those deserve a comment. On-page and business signals — prices in the local currency, a local address and phone number, the local language, mention of local regions — are doing more of the geotargeting work than people realize, and they are things you localize anyway (§20.4). And server location, once cited as a geotargeting signal, is today essentially irrelevant: with content delivery networks (Chapter 16) serving every site from everywhere, where your server physically sits tells Google almost nothing about who your content is for, and Google has said as much. If you read a guide insisting you must "host in-country to rank in-country," you are reading a guide that stopped being true a long time ago.

⚖️ Evidence Check Claim: "Set your target country in Google Search Console to rank in that country." Sort it. — Status: obsolete. The Search Console country-target setting was removed (with the International Targeting report, in 2022). Following this advice is now impossible — the button is gone. — What replaced it (Tier 1): Google infers country relevance from ccTLD, hreflang, on-page/business signals, and links — not from a manual switch. — The lesson (theme 3): this is a live example of SEO folklore outliving its facts. A confident, still-widely-repeated instruction points at a control that no longer exists. Trust primary documentation over the age-indeterminate blog post, and re-verify any "go here and click this" advice before you rely on it.

🔗 Connection The tools that now carry the geotargeting-verification job — the Performance report filtered by country and URL Inspection — are covered in Chapter 27 (Google Search Console). After an international launch, the country filter is how you confirm each market is finding the version you built for it.

The practical upshot is freeing: with the manual lever gone, geotargeting is no longer a setting you toggle — it is a consequence of doing the rest of this chapter well. Pick a structure that signals geography (a ccTLD) or that leans on hreflang (a subdirectory), annotate correctly, and genuinely localize, and you have already told Google which country each page is for. There is no switch to find because the signals are the switch.


20.7 Duplicate content across locales

The last worry that keeps people from doing international SEO correctly is a ghost: the fear that publishing similar content across languages or regions will trigger a "duplicate content penalty." It is worth killing this ghost directly, because the fear pushes people toward genuinely bad fixes — like canonicalizing every locale to one master page, which (per §20.5) deletes those locales from the index.

First, the foundational fact, which belongs to Chapter 14 and which we are only applying here: Google does not levy a duplicate content penalty in the punitive sense people imagine. Duplicate or near-duplicate pages are filtered — Google picks one to show and suppresses the others — not penalized. That distinction matters enormously for international SEO.

Now apply it across the two axes:

  • Different languages are not duplicates. A page in German and its counterpart in French are, to Google, entirely different content — different words, different everything. There is no duplicate-content issue between languages at all. The fear is simply misplaced.
  • Same language across regions is the real (and manageable) case. An en-US page and an en-GB page with near-identical English text can look duplicative — this is where the concern has a grain of truth. Left unmanaged, Google might filter one out or show the "wrong" one to a given searcher. This is precisely the job hreflang exists to do: it tells Google "these are regional variants of the same content; show the U.S. one to U.S. searchers and the U.K. one to U.K. searchers," so the variants don't compete with or cannibalize each other and the right one reaches each market.

🚫 SEO Myth: "Translating my content will cause a duplicate-content penalty." No. Translation produces different content in a different language; it is not duplication, and there is no penalty. The only place the concern is even partly real is same-language regional variants (US/UK/AU English), and even there the remedy is not to hide or merge them — it is to (1) apply correct hreflang so Google shows the right regional version to each audience, (2) keep each version's canonical self-referential so none is dropped as a duplicate, and (3) actually localize — different currency, spelling, examples, local terms — so the variants genuinely differ where it counts. Do those three things and "duplicate content across locales" is a non-issue. The penalty was never real; the routing problem is, and hreflang solves it.

There is a genuine quality dimension underneath the myth, though, and it is worth stating honestly so you don't over-correct into complacency. If your "regional variants" are identical — the same English page copied to /us/, /uk/, and /au/ with nothing changed but a flag in the corner — you have created thin, redundant pages that add no value, and Google's systems may well decide only one deserves to be shown. The fix is the same as the honest fix for everything else in this chapter: make each version genuinely serve its market. If you cannot make the U.K. version meaningfully different from the U.S. one — different prices, terms, examples, spelling — then perhaps you do not need a separate U.K. version at all, and a single English page with a broad hreflang="en" (or just a normal single page) is the cleaner answer. The technical machinery is not a substitute for having something locally worth showing.

🔗 Connection The mechanics of duplicate content, canonical tags, and how Google chooses a representative URL from a set of similar pages are established in Chapter 14 (§14.3). This section only applies that framework to the multi-locale case; if the canonical-vs-hreflang interaction in §20.5 felt fast, Chapter 14 is where the underlying canonical concept is taught in full.


📈 The Strategy File

Every chapter so far has added a component to Rivertown Home Services' strategy. This one, honestly, mostly removes one — and that is exactly the lesson. The advanced skill in international SEO is recognizing when it does not apply and saying so in writing, before someone talks the client into an expensive mistake.

Recall the frozen facts. Rivertown is a family firm founded in 1984 by Ray Delgado and now run by his second-generation children, Marisa and Tony Delgado. It is an HVAC, plumbing, and electrical company serving one metropolitan area through five branches — Rivertown, Cedar Hills, Northgate, Westbrook, and Millhaven — with about 75 employees and 35 trucks. Its customers have one defining constraint: they must be physically reachable by a service truck. That single fact settles the international question.

FIGURE 20.1 — "Rivertown and international SEO: a one-line answer, and the one exception"   [the Strategy File]
  THE SITUATION      One U.S. metro, five branches, an English-speaking market. Every customer must be
                     within driving distance of a truck. ~8,000 organic visits/mo, mostly branded.
  DOES IT APPLY?     No. There is no second country, and no second-country customer Rivertown could serve
                     even if it ranked there. ccTLDs, geotargeting, cross-country hreflang: none of it
                     applies. For the international MACHINERY, Rivertown should skip this chapter entirely.
  THE ONE EXCEPTION  IF a meaningful share of the LOCAL market preferred Spanish, the lightest correct
                     move would be Spanish versions of the top service pages on the SAME domain
                     (rivertownhome.example/es/), with es/en hreflang + self-referential canonicals —
                     genuinely localized and reviewed by a bilingual, licensed technician, NOT machine-
                     translated. That is a LANGUAGE decision inside one country, not "going global."
  WHAT WOULD HURT    Buying ccTLDs, standing up country sites, or auto-translating into many languages
                     would split Rivertown's already-small authority across domains, add fragile hreflang
                     the team would maintain badly, and generate leads for markets no truck can reach.
                     Complexity with zero addressable demand is SEO you are doing AGAINST yourself.
  THE LESSON         The advanced call is knowing this chapter is not for you — yet — and documenting why,
                     so the effort goes to local SEO (Ch 25) instead, where Rivertown's real prize is.

Notice how the one exception threads back into decisions you already understand. A Spanish-language slice would live in a subdirectory (/es/) — chosen for authority consolidation exactly as §20.2 argued, since Rivertown has no reason to split its small authority across domains — and it would use language hreflang (es / en) with self-referential canonicals, the §20.3 and §20.5 discipline applied at its smallest honest scale. It would demand real localization (§20.4): a licensed technician who speaks Spanish reviewing the terms customers actually use, not a plugin's raw output. And it would only be worth doing at all if Marisa and Tony had evidence — from their call logs, their neighborhoods, their existing customers — that the demand is real. Absent that evidence, the correct deliverable for the Strategy File is a single documented sentence: "International SEO does not apply to Rivertown; a Spanish-language subdirectory is a possible future option contingent on demonstrated local demand, and nothing else in Chapter 20 is relevant."

Your Strategy-File task for this chapter (on your own site, using Appendix C's worksheet): run the §20.1 decision tree honestly. Do you have real, evidenced demand in another country or language, and the ability to localize and maintain a version for it? If not, write one sentence recording that international SEO does not apply to you and why — and move on, richer by the effort you did not waste. If so, sketch which of the three structures fits your case and why, and note the one thing (usually: the ability to genuinely localize) you would need in place before you build.


Conclusion

We opened this chapter by trying to talk you out of it, and that framing was not a rhetorical trick — it is the single most important judgment in international SEO. Most sites do not need this machinery, and building it without the demand and the maintenance capacity to back it produces a fragile, half-localized site that ranks worse than the focused one it replaced. The advanced skill is the honest triage of §20.1.

For the sites that genuinely do need it, we drew the real map. The URL-structure decision (ccTLD vs. subdomain vs. subdirectory) is an architecture-and-authority choice, made once and expensive to reverse: a ccTLD gives the clearest geographic signal but splits your authority; a subdirectory consolidates authority but leans entirely on hreflang. hreflang itself is a routing signal — it swaps the right existing version in for the right searcher, and it does not boost rankings — governed by the unforgiving return-tag rule: every version references every version and itself, every reference is returned, and a single self-referential canonical pointed at the wrong place can delete a locale from the index. Localization, not translation, is what actually earns the ranking, and raw machine translation is a trust and relevance liability. The failure modes are relational and invisible on any single page, so validation means crawling the whole set. Geotargeting is no longer a Search Console switch — that setting was retired — but a consequence of your structure, your annotations, and your genuine local signals. And cross-locale duplicate content is a myth in the punitive sense: different languages aren't duplicates, and same-language regional variants are exactly what hreflang and self-referential canonicals are for.

Two themes carried this chapter. Theme 4 — technical SEO is the foundation — because hreflang and URL structure are pure plumbing that either lets your international content rank or silently sabotages it. And theme 3 — evidence over folklore — because this is a topic where the folklore has rotted: the geotargeting setting people still cite is gone, the "duplicate content penalty" they fear was never real, and the "just translate it" shortcut fails on facts Google has published. The professional posture is to verify that the tool still exists and the fear is still founded before acting on either.

Next, we take on the most dangerous operation in all of SEO — the one where getting the redirects wrong can erase years of authority overnight. If choosing an international structure late means a migration, Chapter 21 is where you learn to run that migration without losing your traffic.

→ Continue to Chapter 21: Site Migrations.


Key Terms

  • International SEO — the practice of structuring and optimizing a site so search engines show the right language and country version of its content to each searcher, and so each version can rank in its intended market.
  • hreflang — an annotation (in HTML <link> tags, HTTP headers, or an XML sitemap) that tells Google which language and optional region version of a page to show to which searchers; a routing signal, not a ranking booster.
  • ccTLD (country-code top-level domain) — a top-level domain assigned to a specific country or territory (.fr, .de, .co.uk); the strongest geographic signal, but it splits site authority across separate domains.
  • subdirectory (subfolder) — an international structure that keeps every version on one domain, separated by URL path (example.com/fr/); consolidates authority but relies on hreflang for the geographic signal.
  • subdomain — an international structure using a host prefix on the main domain (fr.example.com); treated as partly separate from the root site, often chosen for infrastructure rather than SEO reasons.
  • localization — adapting content to a locale's language and its culture, currency, units, conventions, and the actual terms local people search; the superset of which translation is only the words.
  • geotargeting — signaling which country a site or section is meant for; now inferred by Google from ccTLD, hreflang, on-page/business signals, and links rather than a manual Search Console setting.
  • x-default — the special hreflang value marking the fallback page to show when no language/region version is a good match for a searcher (often a selector or the global homepage).
  • return tag — the reciprocal hreflang reference: every version must reference every other version and itself, and every reference must be returned by the page it points to, or Google may ignore the pairing.

Spaced Review

Retrieval practice. Try each before revealing the answer. (This set revisits Chapter 14 per the schedule.)

  1. In one sentence each, state what hreflang can do and what it cannot do.
  2. State the return-tag rule, and name the single most common way an hreflang set breaks.
  3. (Chapter 14) A team wants to keep a staging page out of Google's index. Why is a Disallow line in robots.txt the wrong tool for that, and what is the right one?
  4. (Chapter 14) What is a canonical tag for — and, applying it to this chapter, why must each language version's canonical point to itself rather than to the "main" language version?
  5. (Chapter 14) Name two ways an XML sitemap helps Google, and one thing it does not do.
Answers 1. **Can:** get the correct existing language/region version of your page shown to the right searcher (so the British visitor sees the U.K. page, the Mexican visitor the Spanish one), reducing wrong-version bounces. **Cannot:** make you rank where you otherwise wouldn't — it is a routing signal among your own alternates, not a ranking boost, and it cannot rescue thin or badly localized content. 2. **Return-tag rule:** every version in an `hreflang` set must reference every other version *and itself*, and every reference must be *returned* by the page it points to (if A names B, B must name A). The most common break is a **missing return tag** — usually when a new locale is added and the existing pages aren't updated to point back. 3. `robots.txt` `Disallow` blocks *crawling*, not *indexing* — a blocked page can still be indexed (e.g., from links) but now Google can't even read the `noindex` you might want it to obey, so it can stick in the index with no snippet. The right tool is a **`noindex`** directive (meta robots or `X-Robots-Tag`) on a page Google is *allowed* to crawl, so it can read and obey the instruction. 4. A **canonical tag** names the master URL for a piece of content, telling Google which version to index when several are similar. Each language version must canonical to *itself* because it is *not* a duplicate of the other languages — it is the master of its own content; pointing its canonical at the "main" language would tell Google to drop it as a duplicate, removing that locale from the index entirely (the fatal §20.5 canonical conflict). 5. An XML sitemap helps Google by (a) **listing URLs** you want discovered (aiding discovery, especially for pages weakly linked internally) and (b) providing metadata such as last-modified dates (and, relevantly here, `hreflang` annotations). What it does **not** do: **force** indexing — a URL in a sitemap is a suggestion for discovery, not a guarantee Google will crawl, keep, or rank it.