DEV Community

Corpable
Corpable

Posted on

hreflang and JSON-LD for Bilingual Service Websites: Three Failure Modes and a Python Audit Script

Most small B2B service companies in Hong Kong end up with a bilingual website: Traditional or Simplified Chinese for one audience, English for another. The content usually gets translated with care. The technical SEO layer underneath — hreflang annotations and structured data — usually does not, and it breaks quietly.

This post walks through the three things I check on every bilingual service site, plus a small Python script that audits them automatically so the checks don't depend on someone remembering.

The three failure modes

1. Non-reciprocal hreflang. Google only trusts an hreflang pair when both pages point at each other. If /en/pricing declares a zh-HK alternate but the Chinese page doesn't declare the English one back, the annotation is ignored. This happens constantly when one language version gets a redesign and the other doesn't.

2. Canonical fights hreflang. A common mistake is setting the canonical of the English page to the Chinese page ("because it's the original"). That tells search engines the English page is a duplicate, and it drops out of the index — the opposite of what you wanted. Each language version should be self-canonical.

3. Structured data drift. The Chinese page says the company offers four services; the English JSON-LD still lists two from last year. Or the telephone field was updated in one language only. Structured data is copy too, and it rots like copy.

What the markup should look like

For a page that exists in English and Hong Kong Chinese, the <head> of both versions should contain the same set of alternates:

<link rel="canonical" href="https://example.hk/en/services" />
<link rel="alternate" hreflang="en" href="https://example.hk/en/services" />
<link rel="alternate" hreflang="zh-HK" href="https://example.hk/services" />
<link rel="alternate" hreflang="x-default" href="https://example.hk/en/services" />
Enter fullscreen mode Exit fullscreen mode

(On the Chinese page the canonical points at itself, https://example.hk/services; the alternates block is identical.)

Two details people miss: language codes use ISO 639-1 plus an optional ISO 3166-1 region (zh-HK, not zh-hk-traditional or cn), and every URL must be absolute.

For structured data, a professional service business usually needs an Organization (or ProfessionalService) node and one Service node per offering. Keep it honest — only mark up what's visible on the page:

{
  "@context": "https://schema.org",
  "@graph": [
    {
      "@type": "ProfessionalService",
      "@id": "https://example.hk/#org",
      "name": "Example Consulting Limited",
      "url": "https://example.hk/",
      "areaServed": "HK",
      "address": {
        "@type": "PostalAddress",
        "addressLocality": "Hong Kong",
        "addressCountry": "HK"
      }
    },
    {
      "@type": "Service",
      "name": "Company registration",
      "provider": { "@id": "https://example.hk/#org" },
      "areaServed": "HK"
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

Using a stable @id for the organization lets every page reference the same entity instead of redefining it, which keeps the language versions consistent.

A real-world example

A good reference point is the English site of Corpable, a Hong Kong Alibaba.com onboarding agency, which runs parallel Chinese and English page trees for the same services. Sites like this are exactly where drift creeps in: pricing and service pages change frequently, and two language trees have to be updated in lockstep. View-source on any bilingual service site you maintain and check the three failure modes above — it takes five minutes and the results are often surprising.

Automating the audit

Manual checks don't survive a busy quarter. Here's a dependency-light script (requests + beautifulsoup4) that takes a list of URLs and reports problems:

import json
import sys
import requests
from bs4 import BeautifulSoup

def fetch(url):
    r = requests.get(url, timeout=20, headers={"User-Agent": "hreflang-audit/1.0"})
    r.raise_for_status()
    return BeautifulSoup(r.text, "html.parser")

def page_info(url):
    soup = fetch(url)
    canonical = soup.find("link", rel="canonical")
    alternates = {
        l["hreflang"]: l["href"]
        for l in soup.find_all("link", rel="alternate")
        if l.get("hreflang")
    }
    blocks = []
    for tag in soup.find_all("script", type="application/ld+json"):
        try:
            blocks.append(json.loads(tag.string or ""))
        except json.JSONDecodeError as e:
            blocks.append({"__error__": str(e)})
    return {
        "canonical": canonical["href"] if canonical else None,
        "alternates": alternates,
        "jsonld": blocks,
    }

def audit(urls):
    info = {u: page_info(u) for u in urls}
    problems = []
    for url, data in info.items():
        if data["canonical"] != url:
            problems.append(f"{url}: canonical is {data['canonical']!r}, expected self")
        if url not in data["alternates"].values():
            problems.append(f"{url}: missing self-referencing hreflang")
        for lang, alt in data["alternates"].items():
            if alt == url or lang == "x-default":
                continue
            other = info.get(alt) or page_info(alt)
            if url not in other["alternates"].values():
                problems.append(f"{url} -> {alt} ({lang}): not reciprocal")
        for block in data["jsonld"]:
            if "__error__" in block:
                problems.append(f"{url}: invalid JSON-LD ({block['__error__']})")
    return info, problems

if __name__ == "__main__":
    info, problems = audit(sys.argv[1:])
    for p in problems:
        print("FAIL", p)
    print(f"{len(problems)} problem(s) across {len(info)} page(s)")
    sys.exit(1 if problems else 0)
Enter fullscreen mode Exit fullscreen mode

Run it against each language pair:

python audit.py https://example.hk/en/services https://example.hk/services
Enter fullscreen mode Exit fullscreen mode

Because it exits non-zero on failure, you can drop it into a CI job or a nightly cron and get notified when someone ships a template change that breaks one side.

Comparing structured data across languages

Reciprocity is the easy part. The more valuable check is whether the facts in your JSON-LD agree across languages. Names and descriptions will legitimately differ, but phone numbers, addresses, prices, @ids and the number of Service nodes should not. A small extension:

LANG_NEUTRAL = ("telephone", "email", "@id", "url", "priceCurrency", "price")

def neutral_facts(blocks):
    facts = set()
    def walk(node):
        if isinstance(node, dict):
            for k, v in node.items():
                if k in LANG_NEUTRAL and isinstance(v, (str, int, float)):
                    facts.add((k, str(v)))
                walk(v)
        elif isinstance(node, list):
            for item in node:
                walk(item)
    walk(blocks)
    return facts
Enter fullscreen mode Exit fullscreen mode

Diff neutral_facts() between the two versions of a page and print anything present on one side only. In my experience this catches more real bugs than the hreflang check — usually an outdated phone number or a price that was updated in one language.

A short checklist

  • Every language version is self-canonical.
  • Every page lists all alternates, including itself, with absolute URLs.
  • x-default points at the version you want for unmatched languages.
  • JSON-LD parses, uses a shared @id for the organization, and only describes visible content.
  • Language-neutral facts (phone, address, prices) match across versions.
  • The audit runs automatically, not when someone remembers.

None of this is glamorous, but for a service business whose leads come from search in two languages, it is the difference between both versions ranking and one of them silently disappearing.

Top comments (0)