A basic VIN check-digit function looks finished after length, charset, and mod-11. Production traffic then proves otherwise: position 9 can be X, letter maps are not A=1 through Z=9, and users paste banned I/O/Q. This post hardens the usual transliterate-weight-mod-11 flow in TypeScript without lying to the user.
Why edges matter more than the happy path
Marketplace forms, OCR, and CSV imports feed messy strings into the same validator. Digit-only check digits reject about one in eleven valid North American VINs. Auto-mapping I/O/Q can invent a "valid" VIN for a different vehicle. Soft and hard failures need different copy, and that starts with these edges.
Edge 1: X as a legal check digit
The check digit is sum % 11. Values 0 through 9 become "0" through "9". Value 10 becomes "X". That is the standard for North American vehicles under the usual ISO / 49 CFR rules.
Common bugs: a regex that forces position 9 to be \d; a UI mask that only allows digits there; tests that never include remainder 10; error copy that says "check digit must be a number."
Fix the contract first: position 9 is one of 0-9 or X. Then fix the branch:
const rem = sum % 11;
const expected = rem === 10 ? "X" : String(rem);
Add one fixture where the expected digit is X and one where the actual digit is X but the expected digit is not. Both directions matter. X elsewhere in the VIN is a normal letter with transliteration value 7; only position 9 uses X as the stand-in for remainder 10.
Edge 2: transliteration mistakes
Digits map to themselves. Letters use a fixed table with gaps. The table is not "A=1, B=2, ... and skip I/O/Q." In particular: S maps to 2 (not 1); P maps to 7; R maps to 9; I, O, and Q have no legal values.
People derive the table with arithmetic, then encode the wrong values for S through Z. The validator still runs. Some random strings still pass. Real VINs with those letters fail, or pass for the wrong reason if another bug cancels out.
Hard-code the map. Do not compute it:
const TRANSLIT: Record<string, number> = {
A: 1, B: 2, C: 3, D: 4, E: 5, F: 6, G: 7, H: 8,
J: 1, K: 2, L: 3, M: 4, N: 5, P: 7, R: 9,
S: 2, T: 3, U: 4, V: 5, W: 6, X: 7, Y: 8, Z: 9,
};
Unit-test the table itself. Assert TRANSLIT.S === 2 && TRANSLIT.P === 7 && TRANSLIT.R === 9, and assert that I, O, and Q are undefined so a later change cannot "helpfully" add them.
Edge 3: invalid characters
The VIN alphabet is A-H, J-N, P, R-Z, and 0-9. No I, O, or Q. Users still enter letter O for digit 0 (and the reverse), I/Q from OCR, lowercase from PDFs, hyphens or spaces from printouts, and zero-width junk from web copy.
A policy that works:
- Trim, uppercase, and strip common separators and zero-width characters.
- If length is not 17 after that, fail with a length reason.
- If any remaining character is outside the alphabet, fail and name I/O/Q when they appear.
- Do not auto-map O to 0 or I to 1. That hides the mistake and can invent a VIN the plate never had.
const VIN_RE = /^[A-HJ-NPR-Z0-9]{17}$/;
function normalizeVin(raw: string): string {
return raw
.trim()
.toUpperCase()
.replace(/[\s-]/g, "")
.replace(/[\u200B-\u200D\uFEFF]/g, "");
}
Return structured reasons (length, illegal_char, check_digit) so the UI can say "remove I/O/Q" instead of a generic "invalid VIN."
Soft fail vs hard fail
Not every check-digit miss means "reject forever."
- Hard fail for North American flows where the check digit is required.
- Soft warn when users may enter non-NA VINs that do not use the same rule.
- Skip the algorithm for pre-1981 / non-17-character identifiers.
Copy should match the mode. "Likely typo in character 9" beats "fake VIN" on a soft path. "Cannot decode until the VIN uses only legal characters" beats silently correcting O to 0.
A compact TypeScript shape
type VinEdgeResult =
| { ok: true; vin: string }
| {
ok: false;
vin: string;
reason: "length" | "illegal_char" | "check_digit";
detail?: string;
};
export function validateVinEdges(raw: string): VinEdgeResult {
const vin = normalizeVin(raw);
if (vin.length !== 17) {
return { ok: false, vin, reason: "length", detail: `got ${vin.length}` };
}
if (!VIN_RE.test(vin)) {
const bad = [...new Set([...vin].filter((ch) => !/[A-HJ-NPR-Z0-9]/.test(ch)))];
return { ok: false, vin, reason: "illegal_char", detail: bad.join(",") };
}
const weights = [8, 7, 6, 5, 4, 3, 2, 10, 0, 9, 8, 7, 6, 5, 4, 3, 2];
let sum = 0;
for (let i = 0; i < 17; i++) {
const ch = vin[i];
sum += (/\d/.test(ch) ? Number(ch) : TRANSLIT[ch]) * weights[i];
}
const rem = sum % 11;
const expected = rem === 10 ? "X" : String(rem);
if (vin[8] !== expected) {
return {
ok: false,
vin,
reason: "check_digit",
detail: `expected ${expected}, got ${vin[8]}`,
};
}
return { ok: true, vin };
}
Keep decode behind this gate. Edge validation is cheap and specific; decode should not be your first typo detector.
Test matrix worth keeping
Minimum cases beyond one known-good VIN: expected check digit X; actual X when expected is a digit; illegal I/O/Q in different positions; transposed adjacent serial characters; hyphenated or spaced input that normalizes cleanly; lowercase input; a VIN containing S, P, or R that only passes with the correct table.
Disclosure
I maintain VIN Lookup, a free VIN decode that validates format and check-digit edges before showing manufacturer attributes. Use the patterns above in your own forms; use a proper history report and inspection when purchase money is on the line.
Ship the boring rules
Accept X. Hard-code transliteration. Reject I, O, and Q out loud. Normalize separators without inventing characters. Those four habits remove most haunted VIN-field tickets, and they keep free decode honest: a passing check digit means internal consistency, not proof that the car on the lot matches the ad.
Top comments (0)