DEV Community

Cover image for Building a Terminology Pipeline for Multilingual Regulatory Documents (SFDR Case Study)
Diogo Heleno
Diogo Heleno

Posted on Originally published at m21global.com

Building a Terminology Pipeline for Multilingual Regulatory Documents (SFDR Case Study)

Regulatory disclosure translation is usually framed as a legal or linguistic problem. It's also an engineering problem, and a fairly interesting one once you look at the constraints.

Take SFDR (Sustainable Finance Disclosure Regulation) documentation for asset managers. A fund classified as Article 8 or Article 9 has to publish pre-contractual disclosures, periodic reports and website disclosures in every market where it's sold. If a fund is distributed in Portugal, Spain and Germany, you need Portuguese, Spanish and German versions, all referencing the same fixed set of legal terms, all updated in sync whenever the underlying regulation changes.

The source article on the M21Global blog, SFDR and Taxonomy fund translation, covers this from the compliance and linguistic-services angle: which documents need translation, why terms like "Principal Adverse Impacts" or "Do No Significant Harm" can't be paraphrased, and why second-linguist review matters. That's the right framing for compliance teams and translation vendors.

This article is for the people who have to build and maintain the system that keeps those documents in sync. If you're a developer supporting a compliance, legal, or investor-relations team that ships multilingual regulatory content, here's what that pipeline actually looks like.

The core problem is drift, not translation quality

The risk isn't bad grammar. It's semantic drift between document versions over time. A glossary term gets translated correctly in the 2023 prospectus and then translated slightly differently in the 2024 periodic report because a different translator or vendor touched it. Nobody catches it until an auditor or regulator cross-references both documents.

This is the same class of problem as API contract drift or schema versioning. The fix looks similar too: single source of truth, validation, and automated checks before publication.

Treat your glossary as structured data, not a spreadsheet

Most asset managers keep regulatory glossaries in a spreadsheet that lives in someone's inbox. That doesn't scale past two languages or two document types. Model it as data instead:

{
  "term_id": "PAI",
  "en": "Principal Adverse Impacts",
  "definition_ref": "SFDR Art. 4",
  "translations": {
    "pt": "Principais Impactos Adversos",
    "es": "Principales Incidencias Adversas",
    "de": "Wichtigste nachteilige Auswirkungen"
  },
  "locked": true,
  "last_reviewed": "2024-11-02"
}
Enter fullscreen mode Exit fullscreen mode

Once terms live in a structured, versioned format, you can:

  • Feed them into a translation memory (TMX) or termbase (TBX) that CAT tools like memoQ, Trados or Phrase can consume directly
  • Validate translated documents against the termbase programmatically
  • Diff glossary versions over time and flag every document that references a changed term

A basic consistency checker

You don't need a full localization platform to catch obvious drift. A simple script that extracts text from the translated document and checks locked terms against the termbase catches a large share of real incidents:

import re
import json

def load_termbase(path):
    with open(path) as f:
        return json.load(f)

def check_consistency(doc_text, termbase, lang):
    issues = []
    for term in termbase:
        expected = term["translations"].get(lang)
        if not expected:
            continue
        # crude check: does the expected translation appear at all?
        if expected.lower() not in doc_text.lower():
            issues.append({
                "term_id": term["term_id"],
                "expected": expected,
                "lang": lang
            })
    return issues

termbase = load_termbase("sfdr_termbase.json")
with open("periodic_report_es.txt") as f:
    doc = f.read()

results = check_consistency(doc, termbase, "es")
for r in results:
    print(f"Missing or inconsistent term: {r['term_id']} -> expected '{r['expected']}'")
Enter fullscreen mode Exit fullscreen mode

This won't replace a human reviewer, but it's a useful CI-style gate before a document goes to legal sign-off. Run it as part of your document build pipeline, the same way you'd lint code before merge.

Version control for regulatory text, not just source code

SFDR technical screening criteria get updated. When that happens, every language version of every affected document needs the same update, at the same time. Git works fine for this if your documents are stored as Markdown or XML rather than locked Word files:

  • One repo per fund or per document family
  • Branches per language, or a single branch with locale-tagged files (prospectus.en.md, prospectus.pt.md)
  • Tags for regulatory milestones (sfdr-rts-2024-update)
  • CI job that flags any file whose source (English) version changed but whose translated counterparts weren't touched afterward
# .github/workflows/translation-sync-check.yml
name: translation-sync-check
on: [pull_request]
jobs:
  check-stale-translations:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Check for stale translations
        run: python scripts/check_stale_translations.py
Enter fullscreen mode Exit fullscreen mode

The check script compares the git commit hash or last-modified date of the English source file against each locale file, and fails the pipeline if a locale file is older than a threshold after the source changed.

Where this fits with human review

None of this replaces the second-linguist review the source article rightly emphasizes for Article 9 disclosures distributed to institutional investors. Automated checks catch mechanical drift; they don't catch a translator misunderstanding "sustainable investment" as defined under SFDR Article 2(17) versus a colloquial reading of the phrase.

What automation buys you is fewer surprises before the document reaches the reviewer, and a paper trail showing when a term was locked, who approved it, and which documents were checked against it. That paper trail matters as much to an auditor as the translation itself.

Practical takeaways

  • Model your regulatory glossary as versioned, structured data, not a spreadsheet
  • Build a lightweight consistency checker and run it in CI before documents go to legal or compliance review
  • Store multilingual regulatory documents in a system that supports diffing and tagging, so a regulatory update propagates visibly across all languages
  • Automation reduces the volume of drift a human reviewer has to catch; it doesn't replace that reviewer

If you're setting up this kind of pipeline for the first time, start with the glossary. Everything else, TM, CI checks, version tagging, depends on having that term list in a format a script can read.

Top comments (0)