DEV Community

Akmal Urunboev
Akmal Urunboev

Posted on

Terminology drift is the quiet tax on a large content corpus

Translation bugs that crash are easy. You find them, you fix them, they stay fixed.

The expensive problem is slower: the same concept, translated three different ways in three parts of your product, by three people, over two years. Nothing is wrong. Everything parses. And the reader cannot tell whether the three words mean three things.

This is terminology drift, and on a large corpus it will cost you more than any individual mistake.

Why it happens

None of the causes are anyone being careless.

Multiple translators, inevitably. Any corpus large enough to matter is touched by more than one person, usually across years. Each makes a defensible choice. The choices differ.

Translation memory rewards local consistency, not global. A TM suggests what was used in similar strings. It has no concept of "this is the same domain concept as that thing over there, phrased completely differently."

The source drifts first. This is the one people miss. If your English says "account" in one place, "profile" in another and "user" in a third for the same object, you have already lost. Every target language will faithfully reproduce your inconsistency, multiplied by the number of locales.

Nobody owns it. Terminology sits between engineering, content and translation, which means it belongs to no one's checklist.

Fix the source first

The cheapest terminology work happens before anything is translated.

Write down the nouns your product uses for its own concepts, pick one word per concept, and enforce it in English. Do this and you have improved every locale simultaneously, at the cost of one document.

Keep the list small. Twenty to fifty terms covers the concepts that actually recur. A list of four hundred is a list nobody reads.

The lesson: a target-language glossary built on an inconsistent source just encodes the inconsistency more durably.

Then build a real termbase

For each term, a target-language glossary needs more than a word:

  • the source term, and the one approved target equivalent
  • a definition, in plain language, so a translator knows which sense is meant
  • forbidden alternatives — the near-synonyms you have decided against, which is the field that does the most work
  • part of speech and any grammatical notes the target language needs
  • a note on whether it is a brand term that should not be translated at all

That forbidden-alternatives field is worth the effort on its own. "Use X" is advice. "Use X, never Y or Z" is checkable, and it is what turns a glossary from a document into a test.

Enforce it mechanically

A glossary nobody checks is a glossary nobody follows.

Run a check over your translated resources: for every source string containing a glossary term, confirm the target contains the approved equivalent and none of the forbidden ones. Fail the build, or at minimum raise it in review.

Two caveats from experience.

Inflection. In a language that declines nouns, the approved term will legitimately appear in half a dozen forms. Naive substring matching produces false failures, your team learns to ignore the check, and the check is now worse than nothing. Match on stems, or maintain the inflected forms explicitly for the languages that need it.

Legitimate exceptions. Sometimes the approved term genuinely does not fit a sentence. Give the check an explicit opt-out that has to be written down. The goal is that every deviation is a decision someone made on purpose, not that there are no deviations.

Measure it

Terminology consistency is one of the few translation-quality properties you can actually count: of all the places a glossary term should appear, what share use the approved equivalent?

Track it per locale. It tells you where to spend review budget, and it tells you whether a new translator is drifting before the complaints arrive.

The retrofit

If you are reading this with a large corpus already translated inconsistently, you do not need to fix it all.

Rank by exposure. A term on the first screen everyone sees matters more than one buried three levels deep. Fix the top of that list, freeze the glossary so new content is correct from now on, and let the long tail be corrected opportunistically when those strings are touched for other reasons.

The goal is not a perfectly consistent corpus. It is a corpus that stops getting worse, with the most-seen parts cleaned up first.


The mistake I made was treating terminology as a translation problem. It is a source-side product decision that translation merely propagates — and the work to fix it is almost all in English, before a single word gets sent anywhere.

Top comments (0)