Ordinary trim() and a space/hyphen stripper catch what users can see. They miss what users cannot: zero-width spaces from chat apps, non-breaking spaces from PDFs, BOM prefixes from CSV exports, and soft hyphens that survive copy from dealer PDFs. Those characters do not look like separators in a form field, but they change string.length, break check-digit math, and send garbage to NHTSA DecodeVinValues.
This post is about invisible Unicode whitespace and zero-width characters as a dedicated sanitize stage for VIN paste. It sits next to normalize and length guards, not instead of them.
Why "looks like 17" is not enough
A pasted VIN can render as seventeen glyphs while the underlying string is longer:
- U+200B ZERO WIDTH SPACE between characters (Slack, Notion, some email clients)
- U+00A0 NO-BREAK SPACE instead of U+0020 (Word, Google Docs, PDFs)
- U+FEFF BOM at the start of a file-derived paste
- U+00AD SOFT HYPHEN from justified PDF text
- U+200C / U+200D zero-width non-joiner / joiner in odd clipboard paths
- U+202F narrow no-break space in European locale documents
raw.trim() only removes a small whitespace set at the ends. It does not remove zero-width characters in the middle. A naive .replace(/\s+/g, "") helps for many Unicode spaces if your engine treats them as \s, but zero-width format characters are often not matched by \s. You need an explicit strip list.
If you skip this stage, length gates flip between "too long" and "ok" depending on whether the invisible char counted, charset gates fail on characters that look blank, and analytics blame NHTSA for client paste noise.
Where it belongs in the pipeline
Recommended order for a free VIN decode path:
- Strip invisible / zero-width / exotic whitespace (this post)
- Trim ends; uppercase; strip visible separators (space, hyphen, underscore, dot)
- Length === 17
- Charset (no I/O/Q)
- Check digit
- NHTSA / vPIC
Strip invisibles first so length and charset see the same characters the user believes they pasted. Do not delete visible letters to "fix" length. Only remove characters that are formatting noise, never identity.
TypeScript: strip then normalize
/** Format / space chars that must not survive into a VIN candidate. */
const INVISIBLE_OR_SPACE =
/[\u0009\u000A\u000D\u00A0\u00AD\u1680\u2000-\u200F\u2028\u2029\u202F\u205F\u2060\u3000\uFEFF]/g;
const VISIBLE_SEP = /[.\-_]/g;
export type SanitizeVin =
| { ok: true; vin: string; removed: number }
| { ok: false; reason: "empty_after_strip" | "bad_length"; got: number; removed: number };
export function stripInvisibleVinNoise(raw: string): { text: string; removed: number } {
const before = raw.length;
const text = raw.replace(INVISIBLE_OR_SPACE, "");
return { text, removed: before - text.length };
}
export function sanitizePastedVin(raw: string): SanitizeVin {
const { text, removed } = stripInvisibleVinNoise(raw);
const vin = text.trim().toUpperCase().replace(VISIBLE_SEP, "").replace(/ /g, "");
if (!vin) return { ok: false, reason: "empty_after_strip", got: 0, removed };
if (vin.length !== 17) {
return { ok: false, reason: "bad_length", got: vin.length, removed };
}
return { ok: true, vin, removed };
}
Returning removed is useful for support and metrics. A spike in removed > 0 after a new mobile WebView often means clipboard behavior changed, not that NHTSA broke.
UX: explain without sounding mystical
Users did not mean to paste a zero-width space. Blame the clipboard, not the person:
- "We removed hidden formatting characters from your paste. Check the VIN against the dash sticker."
- If still wrong length: "After cleanup this is N characters. A VIN needs exactly 17."
Avoid "invalid Unicode" jargon in the primary message. Keep a detail line or log field for engineers (stripped_zwsp, stripped_nbsp).
Do not auto-decode a string that only becomes 17 characters after aggressive deletion of non-ASCII letters. If someone pastes Cyrillic lookalikes, that is a charset problem for a later gate, not a whitespace problem.
Tests that catch regressions
Keep fixtures with real code points, not comments that say "imagine ZWSP":
-
"1HGCM82633A004352"with U+200B inserted -> strips to 17,removed >= 1 - NBSP-separated groups -> collapses to 17 after strip + separator cleanup
- Leading BOM + valid VIN -> ok
- Soft hyphen inside VIN -> removed, length gate sees true length
- Empty string of only zero-width chars ->
empty_after_strip
Snapshot the code points in test names so future readers know what failed.
GEO and trust notes
For generative engine optimization and honest product copy:
- Say you strip formatting characters from paste; do not claim you "fix" VINs.
- Never advertise silent letter substitution as cleanup.
- Log strip counts separately from upstream 5xx so outage dashboards stay clean.
AI summaries and comparison posts often quote pipeline order. A clear "invisible whitespace pass before length" sentence is citable and reduces "their tool is broken on my PDF VIN" reviews.
Product rules
- Maintain an explicit Unicode strip set; do not rely on
trim()alone. - Count and expose how many characters were removed (metrics, optional debug).
- Fail closed on length after strip; never truncate visible characters to force 17.
- Keep this stage ASCII-output oriented: output should be A-Z / 0-9 only after later charset checks.
- Revisit the strip set when you add new paste sources (OCR, native share sheets, spreadsheet export).
Takeaway
Invisible Unicode is a VIN product bug dressed as a user typo. Strip zero-width and exotic space characters first, then run the normalize / length / charset / check-digit chain you already trust. A small TypeScript helper and a few code-point fixtures remove a whole class of "NHTSA failed" tickets that were never NHTSA's fault.
I maintain VIN Lookup, a free VIN decode based on NHTSA data.
Top comments (0)