DEV Community

Cover image for Picking the Right Version (So Your Database Doesn't Hate You)
Yasir Jafri
Yasir Jafri

Posted on Originally published at yasir323.hashnode.dev

Picking the Right Version (So Your Database Doesn't Hate You)

I used to type uuid.uuid4() for years without giving it a second thought. It works, it's unique, ship it. Then one afternoon I was staring at a Postgres table where inserts had slowed to a crawl, the index was bloated way past what the row count /Users/yasir/Downloads/popular-ids-primary-keys-system-design-v2.mdjustified, and after an embarrassing amount of digging the culprit turned out to be that one line I'd copy-pasted into every project since 2021.

That's when I stopped thinking of UUIDs as one interchangeable thing you sprinkle into a schema. They're not. Each version makes a different trade between randomness, sortability, and how much it quietly tells other people about where and when it was created. Get the wrong one and nothing breaks today. It just costs you later, usually right when your write volume finally gets interesting.

So here's the full rundown, v1 through v8: what each one actually does under the hood, where I'd use it, and where I've seen people get burned. I generated every example below myself, and I'm showing the code so you can run it and get the same thing.

Why bother with UUIDs at all?

Quick detour before the versions, because it colors everything after it.

Auto-increment integers are cheap and they index beautifully, but they leak information. A sequential order ID hands a competitor your daily volume for free. Worse, they fall apart in distributed systems: two services can't both hand out the next integer without talking to each other first, and that coordination becomes a bottleneck the moment you scale past one database.

UUIDs sidestep that. Any service, anywhere, can mint an ID with odds of collision so low they're not worth worrying about, and nobody has to ask permission first. You pay for that in size (16 bytes instead of 4 or 8), and depending which version you pick, you might pay again in index performance. Which brings us to the versions.

UUIDv1: Timestamp + MAC Address

This one glues together a 60-bit timestamp (100-nanosecond ticks since October 15, 1582, I know, I had the same reaction) with the MAC address of whatever machine generated it, plus a clock sequence to keep things from colliding if the clock ever runs backward.

You mostly find it in older Java and .NET codebases these days, from back when it was the obvious default before better options showed up.

Here's one, generated with Python's standard library:

>>> import uuid
>>> uuid.uuid1()
UUID('f3df5f45-b007-11f1-a01d-2b2681fd5c7b')
Enter fullscreen mode Exit fullscreen mode

To actually see where those characters came from, you need the field layout. A UUID isn't one opaque blob, it's carved into named fields, and the hyphens in the string mark exactly where one field ends and the next begins:

f3df5f45 - b007 - 11f1 - a01d - 2b2681fd5c7b
time_low   mid    hi+ver  seq       node
(32 bit) (16 bit) (16 bit)(16 bit) (48 bit)
Enter fullscreen mode Exit fullscreen mode

The 60-bit timestamp gets split across three of those fields. The bottom 32 bits land in time_low untouched, the next 16 go into time_mid, and the top 12 bits have to share space with the 4-bit version marker inside time_hi_and_version:

>>> timestamp = 140086611627958085          # raw 100-ns ticks since 1582, from the OS clock
>>> time_low  = timestamp & 0xFFFFFFFF       # bottom 32 bits
>>> hex(time_low)
'0xf3df5f45'                                 # matches the first group

>>> time_mid  = (timestamp >> 32) & 0xFFFF   # next 16 bits
>>> hex(time_mid)
'0xb007'                                     # matches the second group

>>> top_bits  = (timestamp >> 48) & 0x0FFF   # remaining top 12 bits
>>> version   = 1
>>> time_hi_and_version = (version << 12) | top_bits
>>> hex(time_hi_and_version)
'0x11f1'                                     # leading '1' is the version,
                                              # '1f1' is the rest of the timestamp
Enter fullscreen mode Exit fullscreen mode

And it runs backward cleanly too. Feed those pieces back in and you get the original number:

>>> (0x1f1 << 48) | (0xb007 << 32) | 0xf3df5f45
140086611627958085
Enter fullscreen mode Exit fullscreen mode

The node field is the boring part, in a good way. No splitting, no reassembly. 2b2681fd5c7b is just the machine's 48-bit MAC address, read straight as an integer and printed as 12 hex digits. If the OS doesn't hand over a real MAC (plenty of containers and VMs won't), the spec has you fall back to 48 random bits with a flag set so it's marked as fake. Either way, the encoding itself doesn't change.

What I like about v1 is that it's roughly ordered by creation time and needs zero coordination between nodes. What I don't like: it hands out the generating machine's MAC address to anyone who has the ID, which is a genuinely bad look if those IDs are ever public, and the clock-sequence handling is easy to get subtly wrong if you're implementing it yourself. Honestly, at this point I'd only touch v1 if I inherited a system that already leaned on it.

UUIDv2: DCE Security

I'll be upfront: I've never used this one in production, and I doubt many people reading this have either. It's worth knowing it exists, mostly so you're not caught off guard when someone mentions it in an interview.

DCE Security is what you get when you take the v1 layout and overwrite time_low with a local identifier, usually a POSIX UID or GID, while borrowing part of the clock sequence byte to record a "domain" (0 for person, 1 for group, 2 for org). You're trading away most of the timestamp precision so the ID can carry an identity instead.

It shows up occasionally in older DCE/DFS enterprise systems, and that's about it. Python's uuid module doesn't even bother shipping a uuid2() function, which tells you roughly how much demand there is for it. Since nothing generates one for you, I built one by hand following the DCE 1.1 spec:

>>> import uuid
>>> base = uuid.uuid1()                      # borrow a real timestamp + node
>>> local_id = 1000                          # example POSIX UID
>>> domain = 0                                # 0 = 'person'
>>> time_low = local_id
>>> time_mid = (base.time >> 12) & 0xFFFF
>>> time_hi_version = ((base.time >> 28) & 0x0FFF) | (2 << 12)   # version 2
>>> clock_seq_hi_reserved = 0x80 | (base.clock_seq >> 8)
>>> clock_seq_low = domain
>>> fields = (time_low, time_mid, time_hi_version, clock_seq_hi_reserved, clock_seq_low, base.node)
>>> uuid.UUID(fields=fields)
UUID('000003e8-c637-207f-bd00-65e042fcfd6d')
Enter fullscreen mode Exit fullscreen mode

Look at the very first group: 000003e8. That's 1000 in hex, our fake POSIX UID, sitting right there in plain sight. That's the whole point of v2. Unlike the other versions, part of the identifier is meant to be readable, not opaque.

The upside is exactly that embedded identity, if a system genuinely needs it without a separate lookup table. The downside list is longer: you lose most of your timestamp resolution, there's essentially no library or database support so you're always rolling your own like I just did, and you inherit v1's MAC-leakage problem plus a UID/GID leak on top. I'd file this one under "know it for the trivia, skip it for the architecture."

UUIDv3 and UUIDv5: Name-Based, and Deterministic on Purpose

These two get forgotten a lot, which is a shame because they solve a real problem. Instead of randomness or a clock, they hash a namespace plus a name (v3 with MD5, v5 with SHA-1) into a UUID. Same inputs, same output, every single time, on every machine, with no coordination needed.

Where this earns its keep: deduplication across systems that don't share a primary key. Say three different upstream systems all send you the same customer, identified only by email. Instead of maintaining a lookup table to map "customer@example.com" to some canonical ID, you hash namespace + email and every service arrives at the identical UUID independently.

Run it twice and watch it not change:

>>> import uuid
>>> uuid.uuid3(uuid.NAMESPACE_DNS, "example.com")
UUID('9073926b-929f-31c2-abc9-fad77ae3e8eb')
>>> uuid.uuid3(uuid.NAMESPACE_DNS, "example.com")   # same inputs again
UUID('9073926b-929f-31c2-abc9-fad77ae3e8eb')         # identical output

>>> uuid.uuid5(uuid.NAMESPACE_DNS, "example.com")
UUID('cfbff0d1-9375-5685-968c-48ce8b15ae17')
>>> uuid.uuid5(uuid.NAMESPACE_DNS, "example.com")   # same inputs again
UUID('cfbff0d1-9375-5685-968c-48ce8b15ae17')         # identical output
Enter fullscreen mode Exit fullscreen mode

Look at the third group in each: 31c2 for the MD5 version, 5685 for SHA-1. The leading 3 and 5 are the version markers; everything else about the process is identical, they just hash differently.

Deterministic generation is genuinely handy for idempotency keys and cache keys, and you don't need to store anything to remember the mapping, you can just recompute it. The catch is that predictability cuts both ways: if someone knows your namespace and can guess the name you're hashing, they can predict the UUID too. Don't use these anywhere security matters, like session tokens or password reset links. And pick your namespaces carefully up front, because a sloppy one can quietly collide two unrelated entities into the same ID.

UUIDv4: Pure Random

This is the one everyone reaches for by default, and honestly, for most cases, that instinct is fine. It's 122 bits pulled from a cryptographically secure random source, with 6 bits fixed to mark the version and variant.

>>> import uuid
>>> uuid.uuid4()
UUID('cb2387dc-b97f-4eac-a0ec-2dd03d2526f8')
>>> uuid.uuid4()
UUID('d014cafb-3ce1-4668-87bb-c9a63cc187ab')
Enter fullscreen mode Exit fullscreen mode

Two calls, two values that share absolutely nothing. No common prefix, no pattern to spot. That's exactly the design goal.

It's simple, it leaks nothing about the machine or time it was created, and the odds of a collision are low enough that you'd need something like 2.7 quadrillion UUIDs before you'd hit a 50% chance of a repeat. Every language and database understands it.

The catch, and it's a real one at scale, is that random values are brutal for B-tree indexes. Every insert can land anywhere in the tree, which means page splits, fragmentation, and cache misses; on a busy Postgres table that adds up to a measurable performance tax. You also lose ordering entirely, so you end up adding a created_at column anyway just to know what happened when, which you'd probably want regardless, but it's an extra column carrying weight the ID itself could have carried for free.

This is the exact lesson my team learned the hard way once: unique and index-friendly are not the same promise, even though it's easy to assume they are.

UUIDv6: Reordered Timestamp

Think of this as v1 with the awkward parts fixed. Same underlying data, timestamp and node, but the fields are rearranged before formatting so the string actually sorts the way you'd expect from plain binary or lexicographic comparison, which v1 never managed.

>>> import uuid6
>>> uuid6.uuid6()
UUID('1f1b007f-6119-6ddf-8f5f-36c939af5891')
Enter fullscreen mode Exit fullscreen mode

Compare it to the v1 example from earlier. The information is the same, but v6 moves the high bits of the timestamp to the front instead of burying them in the last 12 bits of the third group like v1 does. That's why 1f1b007f-6119-6... sorts correctly and f3df5f45-b007-11f1... doesn't: v1's most significant timestamp bits are stuck at the tail end of that group, not leading the string where sorting needs them.

It's a solid upgrade path if you're already on v1 and want better sort behavior without redesigning anything. It still carries the option to embed MAC or node info though, so the same privacy caveats apply if you're not careful. And honestly, if you're building something new rather than migrating something old, v7 is the better target anyway.

UUIDv7: Unix Timestamp + Random (the one I'd actually reach for now)

This is where things get genuinely good. A 48-bit millisecond-precision Unix timestamp up front, random bits after it. You get time-ordered IDs without handing out any machine identity, and, this is the part that matters most in practice, without the indexing penalty v4 carries.

>>> import uuid6
>>> uuid6.uuid7()
UUID('01a09eaa-7591-7147-8ebe-bd610c0eda97')
>>> uuid6.uuid7()
UUID('01a09eaa-7592-7789-aaca-2f2005b422fb')
Enter fullscreen mode Exit fullscreen mode

Notice 01a09eaa-759 is shared between both calls. That's the millisecond timestamp prefix, and it's identical because both calls happened within the same couple of milliseconds. Everything after that point is random. That shared, steadily incrementing prefix is exactly what makes v7 so much friendlier to an index: new rows cluster at the end of the tree instead of landing randomly all over it.

For a high-write relational table (event logs, orders, chat messages, anything where rows pile up fast) this gives you sequential-ID-style insert locality while still letting any service mint IDs independently. Postgres 17 added native support for it, which tells you where the ecosystem is heading.

It's not free of trade-offs. Because the timestamp leads the ID, sequential values are somewhat guessable, so don't use v7 anywhere you need true unguessability, like password reset tokens or API secrets. And since it's a newer standard, formalized in RFC 9562 back in 2024, tooling support is excellent but not universal quite yet.

Still, for anything sitting on Postgres with real write volume, this is genuinely what I default to now instead of v4.

UUIDv8: Custom, Your Rules

RFC 9562 leaves v8 deliberately open. It only pins down the version nibble and the variant bits; everything else in those 122 bits is yours to define however your system needs.

>>> import uuid6
>>> uuid6.uuid8()
UUID('01a09eaa-7592-8206-8d60-31408822aac6')
Enter fullscreen mode Exit fullscreen mode

The uuid6 package's implementation keeps a timestamp-like prefix and fills the rest randomly, but that's just one convention someone picked, not a rule. You could just as easily put a 32-bit tenant ID right after the version nibble if you're building a multi-tenant system and want routing baked into the identifier itself, no lookup required.

That flexibility is the whole appeal, and also the whole risk. Nothing about v8 is standardized beyond the version marker, so you're documenting and maintaining the layout yourself. Get the design wrong early and fixing it later means a migration, not a config change, so it's worth sitting with the decision for more than five minutes before you commit.

The Practical Decision Table

Version Sortable Leaks Info Deterministic Best For
v1 Roughly MAC address No Legacy systems only
v2 Barely MAC + UID/GID No Essentially none, historical/niche
v3/v5 No No Yes Deduplication, idempotency keys
v4 No No No General-purpose, low-write-volume tables
v6 Yes MAC (optional) No Upgrading existing v1 systems
v7 Yes No No High-write primary keys, event logs, new systems
v8 Depends Depends Depends Custom routing/sharding schemes

Where I'd land, if you're starting fresh

If I'm building something new today, it mostly comes down to v4 versus v7. Pick v7 when the ID is a primary key on a table that gets hit with heavy writes and index performance actually matters. Pick v4 when you just need a safe, unpredictable identifier and write volume isn't the concern. And keep v5 tucked away for deduplication, it quietly solves a class of bugs that people otherwise reinvent, badly, with lookup tables.

None of this is exotic. It's just the kind of decision that's painless to skip past when you're shipping fast, and mildly painful to unwind once it's baked into a schema six months deep. Worth the five minutes up front.

Top comments (0)