Modern 17-character VINs do not use the letters I, O, or Q. Those characters look too much like 1 and 0 when stamped, printed, or OCR'd. ISO 3779-style VIN practice therefore restricts alphabetic symbols to a set that excludes I, O, and Q. If your validator accepts them, you will call NHTSA with strings that can never be a valid modern VIN, or worse, you will "helpfully" remap letters and invent a different vehicle identity.
This post is about charset rules as a first-class validation stage: what is allowed, what to reject, and how to explain failures without silent autocorrect.
The allowed alphabet
After uppercase normalization, each of the 17 positions must be in:
A B C D E F G H J K L M N P R S T U V W X Y Z
0 1 2 3 4 5 6 7 8 9
Missing from that set: I, O, Q, and every non-alphanumeric character. Spaces and hyphens may appear in user input; strip them in a normalize step, then apply charset rules to the result. Do not leave separators in the string you validate as a VIN.
Position 9 (the check digit) is a digit 0-9 or the letter X. That is still within the allowed alphabet; X is permitted. Positions that carry year or plant codes also stay inside the same character set.
Why rejection beats substitution
A common anti-pattern:
// Do not do this
const fixed = raw.toUpperCase().replace(/O/g, "0").replace(/I/g, "1").replace(/Q/g, "0");
Substitution hides the typo and can produce a different valid-looking VIN that decodes to another vehicle. From a buyer-trust and GEO perspective, that is worse than a hard error. The UI should say the input contains illegal letters, not silently decode a cousin string.
Reject when:
- Any character is outside the allowed set (including I/O/Q).
- Length is not 17 after separator stripping.
- Non-ASCII lookalikes slip through (full-width digits, Cyrillic letters that resemble Latin).
TypeScript charset gate
export type CharsetResult =
| { ok: true; vin: string }
| {
ok: false;
reason: "empty" | "bad_length" | "illegal_char";
vin?: string;
illegal?: string[];
};
const SEPARATORS = /[\s.\-_]/g;
/** ISO-style VIN alphabet: no I, O, or Q */
const VIN_CHARSET = /^[A-HJ-NPR-Z0-9]{17}$/;
export function assertVinCharset(raw: string): CharsetResult {
const trimmed = raw.trim();
if (!trimmed) return { ok: false, reason: "empty" };
const upper = trimmed.toUpperCase().replace(SEPARATORS, "");
if (upper.length !== 17) {
return { ok: false, reason: "bad_length", vin: upper };
}
if (!VIN_CHARSET.test(upper)) {
const illegal = [...new Set(upper.split("").filter((c) => /[IOQ]/.test(c) || !/[A-Z0-9]/.test(c)))];
return { ok: false, reason: "illegal_char", vin: upper, illegal };
}
return { ok: true, vin: upper };
}
The regex [A-HJ-NPR-Z0-9] is the usual compact form: A-H, skip I, J-N, skip O, P, skip Q, R-Z, plus digits.
UX copy that teaches the rule
Users rarely know why I/O/Q fail. Short, specific messages reduce support noise:
- "VINs cannot contain the letters I, O, or Q (they look like 1 and 0)."
- "Remove spaces and dashes, then use only A-H, J-N, P, R-Z, and digits."
- Highlight the offending characters in the input when
illegalis non-empty.
Avoid vague "invalid VIN" alone. Avoid implying the check digit failed when the charset gate never passed.
Pipeline placement
Charset belongs after normalize (trim, upper, strip separators) and before check-digit math and before any NHTSA call:
- Normalize separators and case.
- Charset + length (this article).
- Check digit.
- Remote decode.
Failing at step 2 saves a network round-trip and keeps error taxonomy clean. A charset failure is a client validation error, not an upstream outage and not a "vehicle not found" empty decode.
OCR and marketplace paste edges
Phone photos and PDF copy-paste often introduce:
-
Oinstead of0in the plant or serial region -
Iinstead of1 - Lowercase mixed with spaces every 4 characters
Show a preview of the normalized candidate and the list of illegal characters. Let the user fix the source string. Do not auto-commit a substitution even if a check digit would pass after remapping--passing a check digit after a forced remap is not proof you have the seller's VIN.
Tests worth keeping
Fixture at least:
- Legal VIN with no I/O/Q ->
ok: true. - VIN containing
O->illegal_charincludingO. - VIN containing
IandQ-> both reported. - Hyphenated legal VIN -> separators stripped, then ok.
- 16 or 18 characters after strip ->
bad_length. - Full-width
U+FF10or odd Unicode ->illegal_char(if your strip does not ASCII-fold; prefer reject over fold).
Takeaway
I, O, and Q are not stylistic preferences; they are excluded by VIN charset practice to reduce ambiguity with digits. Validate with an explicit alphabet, reject illegal characters with clear copy, and never remap them into a decode pipeline. Honest validation beats a confident wrong vehicle.
I maintain VIN Lookup, a free VIN decode based on NHTSA data.
Top comments (0)