We run a small text-rewriting tool. English feedback was decent. Then we turned on the other twelve languages and support tickets changed shape completely.
Nobody complained that "the text is still robotic". They complained about things like "why is there a space before my question mark", "my boss name is spelled wrong in Hebrew", and "this is insulting in German".
Here are the six failures worth writing down. All of them were invisible to our English test suite.
1. "Shorter is better" is an English habit
Our English pass was tuned to cut filler and shorten sentences. Applied to German it shredded perfectly natural compound nouns. Applied to Japanese it did the opposite — the rewrite came out longer than the source because our "simplify" rule kept splitting clauses that Japanese keeps joined with commas.
A 100-word English paragraph is roughly 130 words in German and about 70 characters in Chinese. There is no single length target. If your rewrite rules reference word or character counts, they are English rules wearing a neutral costume.
2. Punctuation is a fingerprint, and it differs per script
The em dash obsession that AI detectors flag is an English signal. In French, the giveaway is different — and so is the correct typography:
| Language | Correct | What our tool shipped |
|---|---|---|
| French |
Bonjour ! (narrow NBSP before ! ? ; :) |
Bonjour! |
| German | „Anführungszeichen" |
"Anführungszeichen" |
| Spanish |
¿Qué? / ¡Vamos! (opening marks) |
Qué? / Vamos!
|
| Chinese |
,。 full-width |
, . half-width |
Half-width punctuation in Chinese text is the single loudest "a foreigner or a machine wrote this" signal there is. We shipped it for a week.
3. Your "AI-tell" word list does not translate
Every language has its own overused vocabulary, and it is almost never a translation of the English list. Machine-translating "delve / tapestry / bustling" into Spanish produces words nobody overuses. Meanwhile the actual Spanish tells — en el mundo actual, es importante destacar — were not on our list at all.
You have to build the list per language from real text, not translate the English one. There is no shortcut here.
4. Formality is a required field, not a style option
English has one "you". Many languages do not:
- French:
tuvsvous - German:
duvsSie - Spanish:
túvsusted, andvosin Argentina - Chinese:
你vs您
Defaulting to the informal form is not "casual", it is rude — the German one in particular. If the input does not tell you the register, the honest move is to ask or stay formal. Guessing friendly is a bug, not a feature.
5. RTL is not "same text, right aligned"
Arabic and Hebrew need the direction handled in the layout, not just in the CSS direction property. Mixed content is where it breaks: an English product name inside an Arabic sentence, a phone number next to a parenthesis. They reorder visually, so a name that reads correctly in the database renders wrong on screen and in a quoted email.
If you cannot test with a native reader, at minimum wrap Latin runs and numbers with proper isolates and check the rendered string, not the source string.
6. Numbers, dates, and units
1,000 means one thousand in the US and one in Germany. 3/4/2026 is two different dates depending on the country. Temperature units, paper sizes, and currency placement all differ.
This is the boring one, and it is also the one that produces refund requests, because a wrong date in a contract line is not a stylistic problem.
The lint we ended up writing
We stopped trying to be clever in the rewriter and moved all of this into a per-language checker that the rewriter must pass:
CHECKS = {
"fr": [("missing_nbsp_before_punct", lambda s: bool(
__import__("re").search(r"[A-Za-z\u00C0-\u017F][!?;:]", s)))],
"zh": [("halfwidth_punct", lambda s: bool(
__import__("re").search(r"[\u4e00-\u9fff],[\u4e00-\u9fff]", s)))],
"de": [("straight_quotes", lambda s: '"' in s)],
"es": [("missing_opening_mark", lambda s: ("?" in s and "¿" not in s)
or ("!" in s and "¡" not in s))],
}
def lint(text, lang):
return [name for name, fn in CHECKS.get(lang, []) if fn(text)]
Then a human who actually speaks the language reads the output. Not a translator — a reader. Every language we shipped without that step came back with a ticket.
The lesson is not "AI text rewriting is hard in other languages". It is that an English-only test suite will tell you everything is fine right up until a customer quotes you.
The tool we build this into is ShipCopy, if you want to see the multilingual side in action: launchcraft.io — 10 rewrites free, no card.
Top comments (0)