DEV Community

24hTrack
24hTrack

Posted on AI-assisted

Translating carrier scan text: all-or-nothing, or don't bother

Every cross-border tracking UI eventually hits the same wall: a scan line arrives in a language your users cannot read, and you have to decide what to render.

邮件离开处理中心【广州】
已妥投【US California Olivehurst 95961】
包裹已签收,如有疑问请联系收件人
Enter fullscreen mode Exit fullscreen mode

The naive fix is a phrase dictionary plus str.replace. That is the right idea and the wrong implementation, and the failure modes are specific enough to be worth writing down. Everything below came out of shipping this on a live corpus of Chinese-language scan lines.

1. A half-translated sentence is worse than an untranslated one

This is the rule that decides the whole design.

包裹已签收,如有疑问请联系收件人 contains a phrase you probably have in your dictionary (已签收 = delivered, signed for) and a clause you probably do not (if you have questions, contact the recipient). Replace the part you know and the user gets:

Delivered (signed for),如有疑问请联系收件人

That reads as broken software, not as foreign text. Worse, it hides which half you actually understood — the untouched half might be an instruction the user needs to act on.

So: translate all of it, or hand back the original byte for byte.

function gloss(text) {
  let out = text
  for (const [re, en] of PHRASES) out = out.replace(re, ` ${en} `)
  if (out === text) return text            // recognised nothing: unchanged
  if (PROSE_RE.test(strip(out))) return text  // real prose survived: give up
  return tidy(out)
}
Enter fullscreen mode Exit fullscreen mode

2. You need a boundary between "place name" and "sentence"

The all-or-nothing check needs to answer: is what's left a proper noun, or a clause I failed to read?

Chinese does not put spaces between words, so you cannot count tokens. What works is a character-run threshold. Measured against real place names and real clauses:

  • 东莞市 (Dongguan City) — 3 characters
  • 深圳 (Shenzhen) — 2
  • 如有疑问请联系收件人 — 10

Four is the boundary. A run of 4+ Han characters is a clause; shorter is a name.

const PROSE_RE = /[぀-ヿ一-鿿가-힯]{4,}/
Enter fullscreen mode Exit fullscreen mode

And strip bracketed payloads before testing, because carriers put the location inside 【…】 and you deliberately never translate those:

const strip = s => s.replace(/【[^】]*】/g, '')
Enter fullscreen mode Exit fullscreen mode

3. Ordering is load-bearing, and getting it wrong is silent

Dictionary entries overlap. 邮件离开处理中心 contains 离开处理中心 contains 离开. Put the short one first and the long one can never match:

邮件到达处理中心【SZ】
  → 邮件 Arrived at sorting centre 【SZ】   // orphan characters, prose gate trips, whole row reverts
Enter fullscreen mode Exit fullscreen mode

When we put the bare forms first by accident, it made thousands of rows worse — and nothing threw. The output was just slightly wrong on a page nobody was diffing.

Two habits that catch this:

  • Longest-prefix first, enforced by a test that asserts entry order, not just entry presence.
  • Diff old vs new over your whole live corpus before shipping, and count three things: rows changed, rows emptied, rows that lost more than half their length. "Rows emptied" should be zero, always.

4. Never gloss a bare noun

We had a 处理中心 → "sorting centre" entry. It reached inside a facility name:

【深圳邮区中心邮件处理中心】  →  【深圳邮区中心邮件 sorting centre】
Enter fullscreen mode Exit fullscreen mode

Every entry should be a verb plus its object, or a pattern that consumes the brackets:

[/离开(【[^】]{0,40}】)处理中心/g, 'Departed $1 sorting centre'],
[/到达(【[^】]{0,40}】)处理中心/g, 'Arrived at $1 sorting centre'],
Enter fullscreen mode Exit fullscreen mode

5. Pad your replacements, then collapse

No spaces between words means a replacement welds itself to whatever follows:

邮件离开处理中心Dongguan  →  Departed sorting centreDongguan
Enter fullscreen mode Exit fullscreen mode

Pad both sides on insert, then collapse runs and trim before punctuation:

out.replace(/\s*,\s*/g, ', ')
   .replace(/\s{2,}/g, ' ')
   .replace(/\s+([,.;:)\]】])/g, '$1')
   .trim()
Enter fullscreen mode Exit fullscreen mode

Note the full-width comma , — leave it in and you get Chinese punctuation inside English sentences.

6. Do this on the read path only

The glossed string is for display. The stored value stays exactly as the carrier sent it.

That is not tidiness. Anything derived from status text — a delivered_at stamp, a downgrade guard, a "stop polling this parcel" rule — must key off the carrier's own words. Translate on write and you have quietly made your business logic depend on the current contents of a phrase table. Change the table, change the meaning of historical rows.

One function, called from every render surface. No new call sites.

The short version

  • All-or-nothing: fully translated, or byte-identical.
  • 4+ CJK characters in a row means a clause, not a name — bail out.
  • Longest phrase first, and test the ordering.
  • Verb + object, never a bare noun.
  • Pad on insert, collapse after.
  • Display layer only; never the write path.

If you want to see the output rather than the code: 24hTrack is a free package tracker for 3,200+ carriers — paste any tracking number, the carrier is detected automatically, no sign-up — and the recognised Chinese scan lines are rendered in English on the timeline. There is a plain-English walkthrough of the phrases themselves here.

Written with AI assistance; the measurements and the failure cases are from our own production corpus.

Top comments (0)