I ship in ten languages. I speak two of them.
For a long time my quality process was "the translator seemed competent," which is not a process, it is a hope. Here is the system I use now. None of it requires me to read the target language.
Separate the translator from the reviewer
This is the single highest-value change, and it is organisational rather than technical.
Nobody is a reliable reviewer of their own work, and that is doubly true in translation, where the translator has already made a hundred interpretive decisions they have stopped being able to see. One person translates. A different native speaker reviews. They should not be the same person and ideally should not be in the same room.
If budget only allows one native speaker per language, spend them on review and use a lower-cost first pass for the draft. Reviewing a flawed translation is faster than producing a good one, and the output is better than one unreviewed expert pass.
Automate everything that doesn't need a human
A surprising share of translation defects are mechanical, and mechanical defects can be caught by a script that does not speak the language either. I run all of these in CI against the translated resource files:
- Placeholder integrity. Every placeholder in the source exists in the target, with the same name, and no extras. This is the single most common serious bug, and it crashes at runtime rather than reading oddly.
- ICU syntax validity. Plural and select blocks parse, and the categories present are the ones that locale actually requires. A translation missing a required plural category will fall back silently and look fine until the number is right.
- Untranslated strings. Target identical to source, across a whole file, usually means a pipeline failure rather than a word that happens to be the same.
- Length budget. Flag anything substantially longer than source. It will not throw, it will just overflow a button somewhere you aren't looking.
- Terminology adherence. If you have a glossary — and you should — check that the agreed term was actually used. This one catches real meaning drift, not just mechanics.
- Punctuation and spacing conventions. Some locales take a space before a colon, some use different quotation marks, some use a different decimal separator. These are rule-checkable.
None of this tells you whether the translation is good. All of it tells you whether it is broken, and it removes the broken ones from the human reviewer's plate so their attention goes where it is actually needed.
Review in context, never in a spreadsheet
Translators are usually handed a column of strings with no indication of where any of them appear. Then we are surprised when a word that means "free of charge" shows up where we meant "unoccupied."
Give the reviewer the actual screens. A build, a staging link, or at minimum a screenshot attached to each string. The defects that only show up in context — wrong register, wrong sense of an ambiguous word, a button label that is grammatically fine and situationally absurd — are exactly the ones that automation cannot catch and that users notice first.
If you do nothing else from this post, do this one. It is the highest ratio of quality gained to effort spent of anything I have tried.
Give the review a rubric
"Does this look right?" produces a thumbs up. Ask four separate questions and you get usable answers:
- Accuracy. Does it say what the source says?
- Terminology. Does it use the agreed terms from the glossary?
- Register. Is the formality level right and consistent? Does it address the reader the way we intend to?
- Fluency. Would a native speaker write it this way, or does it read as translated?
Scoring those separately matters because they have different fixes and different severities. A fluency complaint is a polish item. An accuracy or terminology failure is a bug.
Sample, don't boil the ocean
Full review of every string on every release does not survive contact with a shipping schedule.
I review everything once at launch for a locale, then sample afterwards — weighted toward strings that are new, strings that changed, and strings on screens that matter most. If a sample comes back with defects above your tolerance, widen it for that locale. If it comes back clean repeatedly, narrow it.
The point is that you are measuring translation quality as an ongoing property, not certifying it once and assuming it holds.
Back-translation is weaker than it looks
A common suggestion is to have the target translated back into English and compare. I use it rarely and trust it less than I used to.
It catches gross errors — a flipped negation, a dropped clause. It is poor at everything subtler, because a fluent back-translator will unconsciously repair an awkward target, and an awkward back-translator will make a perfectly good target look wrong. You end up investigating noise.
Use it as a spot check on your highest-stakes strings. Do not build a quality process on it.
The thing I wish someone had told me earlier: you cannot eyeball your way out of not speaking the language, and you do not have to. Almost everything above is either a script or an organisational decision. The part that genuinely requires a native speaker is smaller than it first appears — but it is not zero, and the mistake is pretending it is.
Top comments (0)