DEV Community

zhihu wu
zhihu wu

Posted on

UUIDs That Look Identical Aren't: The Case, Braces, and Hyphen Trap

The bug that makes no sense

A row exists. The API returns its ID. You paste that ID into a query and get… nothing. You look at both strings side by side. They are the same. Except they are not, because "the same" for a UUID string means byte-for-byte identical, and four sneaky formatting differences break that all the time.

1. Case

.NET's Guid.ToString() returns lowercase. SQL Server hands back uppercase. Postgres normalizes to lowercase, and an API gateway somewhere in the middle may have done whatever it wanted. In every case-sensitive comparison — a JS ===, a SQL WHERE id = '...', a Redis key — "ABC..." and "abc..." are different values. Same UUID, failed lookup.

2. Hyphens

Canonical form is 8-4-4-4-12, 36 characters. But UUIDs travel through places that strip punctuation: URL slugs, log parsers, spreadsheet columns, that one time someone hit sed 's/-//g'. The 32-char hex form is a legitimate representation of the same value and an instant mismatch against a stored canonical string.

3. Braces and parentheses

C# and PowerShell are happy to hand you {550e8400-e29b-41d4-a716-446655440000} or (550e8400-...). If that string is what went into your DB or your cache key, your canonical-form lookups will never hit.

4. Hex vs bytes

MySQL's UUID_TO_BIN() stores 16 raw bytes for indexing efficiency — correct and fast. It also means a plain string comparison against that column returns nothing at all, and BIN_TO_UUID() has to be in both your SELECT and your WHERE, or you are comparing apples to BLOBs.

The fix: normalize at the boundary

The rule I now follow: parse on input, format on output, never compare raw strings from two different sources.

  • Normalize to lowercase, hyphen-free hex before any equality check or cache key. One toLowerCase().replaceAll('-', '') kills cases 1–3.
  • Parse, don't regex, when you can: UUID.fromString() in Java, uuid.UUID() in Python, uuid.Parse() in Go all accept the mixed forms and reject genuinely malformed input — a validation bonus you don't get from a loose pattern.
  • Validate with the version/variant bits in mind if you care about v4 vs v7, since a regex that only checks the 8-4-4-4-12 shape happily accepts a v7 ID where you expected v4 (and vice versa).
  • Log the raw form when debugging. Half of these bugs are invisible because whatever printed the ID for you already normalized it.

When I need a handful of IDs in a known-good canonical form to test a normalization path — or to check what my parser does with the uppercase, braced, hyphen-less variants — I generate them in a free in-browser UUID generator that runs entirely on the client, no upload, no signup. It does bulk output up to 100 at a time and the uppercase/lowercase toggle is handy for exactly the comparison bug above.

None of this is exotic. It is just the reason two strings that "look identical" quietly fail the only test that matters.

Top comments (0)