DEV Community

Cover image for Building a Sync Pipeline for Multilingual Help Centers (So They Don't Rot in Three Languages)
Diogo Heleno
Diogo Heleno

Posted on Originally published at m21global.com

Building a Sync Pipeline for Multilingual Help Centers (So They Don't Rot in Three Languages)

Translating your help center is the easy part. Keeping it in sync with a product that ships weekly is where most teams quietly give up and end up with three French articles from 2022 sitting next to an interface that's been redesigned twice since then.

The original article on help center localisation makes a good case for why this matters and how to think about vendors, glossaries and tiers of review. This post is about the other half: what the actual pipeline looks like if you're the engineer responsible for making translated content show up correctly, on time, without someone manually copy-pasting HTML into a spreadsheet every Friday.

The core problem is a sync problem, not a translation problem

Translation vendors sell you words. Your job is making sure those words land in the right place, in the right format, without breaking internal links, code placeholders or embedded screenshots, every time a source article changes.

Treat it like you'd treat any other content pipeline with an external dependency:

  • Source of truth (your docs platform or CMS)
  • Export/trigger mechanism
  • External processing (translation, in this case)
  • Import and validation
  • Publish with rollback

If you wouldn't ship a feature without tests and a rollback plan, don't ship translated content that way either.

Don't export plain text. Ever.

The single most common failure mode is flattening HTML into plain text for translation and then trying to reconstruct formatting afterward. You lose:

  • {{variable}} placeholders used for personalization
  • Internal anchor links (<a href="#step-3">)
  • Code blocks that shouldn't be translated at all
  • alt text on images, which often should be translated but gets forgotten

The fix is translating structured formats, not raw strings. If you're on Zendesk Guide or Intercom, use their native translation fields and connectors rather than exporting rendered HTML. If you're building this in-house, something like this keeps structure intact:

import re

PLACEHOLDER_PATTERN = re.compile(r"(\{\{.*?\}\}|<code>.*?</code>)", re.DOTALL)

def extract_protected_segments(html):
    segments = PLACEHOLDER_PATTERN.findall(html)
    masked = PLACEHOLDER_PATTERN.sub("__PROTECTED_{}__", html)
    return masked, segments

def restore_segments(translated_html, segments):
    for i, segment in enumerate(segments):
        translated_html = translated_html.replace(f"__PROTECTED_{i}__", segment, 1)
    return translated_html
Enter fullscreen mode Exit fullscreen mode

It's a crude example, but the principle holds for any pipeline: mask what shouldn't be translated, send the rest through, restore afterward, and validate the output against the original DOM structure before publishing.

Build a diffing step before re-translation

The expensive mistake is re-sending an entire article for translation every time you fix a typo. You want a diff-based trigger:

git diff --name-only HEAD~1 HEAD -- docs/en/*.md
Enter fullscreen mode Exit fullscreen mode

If you're not versioning docs in git (you should be, even if the CMS is the real source of truth), at minimum store a content hash per article per locale:

{
  "article_id": "guide-sso-setup",
  "source_hash": "a1b2c3",
  "translations": {
    "de": { "hash_translated_from": "a1b2c3", "status": "current" },
    "fr": { "hash_translated_from": "9f8e7d", "status": "stale" }
  }
}
Enter fullscreen mode Exit fullscreen mode

This lets you flag stale translations automatically instead of discovering them when a support ticket references an outdated screenshot. A simple cron job comparing hashes and opening a ticket in your translation queue does 90% of the work here.

Translation memory is a database problem, not a vendor promise

A shared translation memory (TM) only works if it's actually queryable and enforced, not just something your vendor says they maintain. If you have any engineering control over the pipeline, keep your own terminology store alongside theirs:

CREATE TABLE glossary_terms (
    term_en TEXT PRIMARY KEY,
    term_de TEXT,
    term_fr TEXT,
    context TEXT, -- 'ui_button', 'feature_name', 'plan_tier'
    forbidden_synonyms TEXT[]
);
Enter fullscreen mode Exit fullscreen mode

Run a lint step against translated content before publishing: scan for forbidden synonyms, flag UI term mismatches between the interface strings (your i18n JSON/PO files) and the help center copy. This is the kind of check that catches "Workspace" in the product UI versus "Espace de travail" in one French article and "Environnement" in another, before a customer notices.

Screenshots are a release-blocking dependency

If an article embeds a screenshot of the UI, that screenshot is now coupled to your product's release cycle in that locale. Decide this at the pipeline level, not article by article:

  • If the product UI isn't localized in a given language yet, don't localize the screenshot. Caption it instead. Trying to half-translate visuals creates more inconsistency than leaving them in English with a note.
  • Tag every screenshot-containing article with the UI version it was captured against, so your stale-content job (above) also fires when the UI changes, not just when the source text changes.
{ "article_id": "guide-sso-setup", "ui_version_at_capture": "2.14.0" }
Enter fullscreen mode Exit fullscreen mode
  • For video, subtitle tracks (WebVTT/SRT) are far easier to version and diff than dubbed audio. Store them as separate files per locale and diff them the same way you diff text.

Validation before publish, not after complaints

Before anything goes live, run automated checks:

  • HTML/Markdown structure matches the source (same number of headings, links, code blocks)
  • All internal links resolve
  • No leftover placeholder artifacts (__PROTECTED_0__, raw {{var}} not re-rendered)
  • Glossary lint passes

A basic version of this can run in CI if your docs are in git:

- name: Validate translated docs
  run: |
    python scripts/validate_structure.py docs/de/*.md
    python scripts/glossary_lint.py docs/de/*.md
Enter fullscreen mode Exit fullscreen mode

This catches the embarrassing stuff (broken links, stray placeholders) before a customer does, and it's the difference between a translation workflow and a translation system.

The takeaway

Localizing a knowledge base is a content ops problem with a translation step in the middle, not a translation problem with some ops around it. If you build the sync, diffing and validation layer properly, the actual translation (human, AI-assisted, or hybrid) becomes a replaceable component rather than the thing your entire support quality depends on.

If you're evaluating vendors or tiers of human review for this kind of work, the source article covers that side well. This is the engineering checklist to pair with it.

Top comments (0)