DEV Community

Vin Lookup
Vin Lookup

Posted on

Rejecting I, O, and Q: ISO VIN Charset Rules in Validation

Modern 17-character VINs do not use the letters I, O, or Q. Those characters look too much like 1 and 0 when stamped, printed, or OCR'd. ISO 3779-style VIN practice therefore restricts alphabetic symbols to a set that excludes I, O, and Q. If your validator accepts them, you will call NHTSA with strings that can never be a valid modern VIN, or worse, you will "helpfully" remap letters and invent a different vehicle identity.

This post is about charset rules as a first-class validation stage: what is allowed, what to reject, and how to explain failures without silent autocorrect.

The allowed alphabet

After uppercase normalization, each of the 17 positions must be in:

A B C D E F G H J K L M N P R S T U V W X Y Z
0 1 2 3 4 5 6 7 8 9
Enter fullscreen mode Exit fullscreen mode

Missing from that set: I, O, Q, and every non-alphanumeric character. Spaces and hyphens may appear in user input; strip them in a normalize step, then apply charset rules to the result. Do not leave separators in the string you validate as a VIN.

Position 9 (the check digit) is a digit 0-9 or the letter X. That is still within the allowed alphabet; X is permitted. Positions that carry year or plant codes also stay inside the same character set.

Why rejection beats substitution

A common anti-pattern:

// Do not do this
const fixed = raw.toUpperCase().replace(/O/g, "0").replace(/I/g, "1").replace(/Q/g, "0");
Enter fullscreen mode Exit fullscreen mode

Substitution hides the typo and can produce a different valid-looking VIN that decodes to another vehicle. From a buyer-trust and GEO perspective, that is worse than a hard error. The UI should say the input contains illegal letters, not silently decode a cousin string.

Reject when:

  • Any character is outside the allowed set (including I/O/Q).
  • Length is not 17 after separator stripping.
  • Non-ASCII lookalikes slip through (full-width digits, Cyrillic letters that resemble Latin).

TypeScript charset gate

export type CharsetResult =
  | { ok: true; vin: string }
  | {
      ok: false;
      reason: "empty" | "bad_length" | "illegal_char";
      vin?: string;
      illegal?: string[];
    };

const SEPARATORS = /[\s.\-_]/g;
/** ISO-style VIN alphabet: no I, O, or Q */
const VIN_CHARSET = /^[A-HJ-NPR-Z0-9]{17}$/;

export function assertVinCharset(raw: string): CharsetResult {
  const trimmed = raw.trim();
  if (!trimmed) return { ok: false, reason: "empty" };

  const upper = trimmed.toUpperCase().replace(SEPARATORS, "");
  if (upper.length !== 17) {
    return { ok: false, reason: "bad_length", vin: upper };
  }

  if (!VIN_CHARSET.test(upper)) {
    const illegal = [...new Set(upper.split("").filter((c) => /[IOQ]/.test(c) || !/[A-Z0-9]/.test(c)))];
    return { ok: false, reason: "illegal_char", vin: upper, illegal };
  }

  return { ok: true, vin: upper };
}
Enter fullscreen mode Exit fullscreen mode

The regex [A-HJ-NPR-Z0-9] is the usual compact form: A-H, skip I, J-N, skip O, P, skip Q, R-Z, plus digits.

UX copy that teaches the rule

Users rarely know why I/O/Q fail. Short, specific messages reduce support noise:

  • "VINs cannot contain the letters I, O, or Q (they look like 1 and 0)."
  • "Remove spaces and dashes, then use only A-H, J-N, P, R-Z, and digits."
  • Highlight the offending characters in the input when illegal is non-empty.

Avoid vague "invalid VIN" alone. Avoid implying the check digit failed when the charset gate never passed.

Pipeline placement

Charset belongs after normalize (trim, upper, strip separators) and before check-digit math and before any NHTSA call:

  1. Normalize separators and case.
  2. Charset + length (this article).
  3. Check digit.
  4. Remote decode.

Failing at step 2 saves a network round-trip and keeps error taxonomy clean. A charset failure is a client validation error, not an upstream outage and not a "vehicle not found" empty decode.

OCR and marketplace paste edges

Phone photos and PDF copy-paste often introduce:

  • O instead of 0 in the plant or serial region
  • I instead of 1
  • Lowercase mixed with spaces every 4 characters

Show a preview of the normalized candidate and the list of illegal characters. Let the user fix the source string. Do not auto-commit a substitution even if a check digit would pass after remapping--passing a check digit after a forced remap is not proof you have the seller's VIN.

Tests worth keeping

Fixture at least:

  1. Legal VIN with no I/O/Q -> ok: true.
  2. VIN containing O -> illegal_char including O.
  3. VIN containing I and Q -> both reported.
  4. Hyphenated legal VIN -> separators stripped, then ok.
  5. 16 or 18 characters after strip -> bad_length.
  6. Full-width U+FF10 or odd Unicode -> illegal_char (if your strip does not ASCII-fold; prefer reject over fold).

Takeaway

I, O, and Q are not stylistic preferences; they are excluded by VIN charset practice to reduce ambiguity with digits. Validate with an explicit alphabet, reject illegal characters with clear copy, and never remap them into a decode pipeline. Honest validation beats a confident wrong vehicle.

I maintain VIN Lookup, a free VIN decode based on NHTSA data.

Top comments (0)