DEV Community

Cover image for Designing Ticketmaster's Seat Hold: Why the Obvious Lock Breaks at 50K Requests per Second
Branden Floris
Branden Floris

Posted on

Designing Ticketmaster's Seat Hold: Why the Obvious Lock Breaks at 50K Requests per Second

A stadium has 20,000 seats. Sales open at 10:00 AM. At 10:00:01, 500,000 people hit refresh.

Your job is simple to state and hard to build. Never sell the same seat twice. Hold a seat for the buyer while they pay. Don't fall over.

This is one of the best system design interview questions because the naive answer works on a whiteboard and fails in production. Let's build it three times and break it twice.

Ticketmaster is also one of the full case studies in Grokking the System Design Interview, if you want to compare this walkthrough with a complete design.

The requirements

Functional:

  • Users see a seat map for an event.
  • Users pick seats and get a hold for 10 minutes while they check out.
  • If they pay in time, the seats are sold. If not, the seats go back on sale.

Non-functional:

  • No double selling. This is the one rule you can't bend.
  • Survive the spike. Popular on-sales see tens of thousands of requests per second in the first minute.
  • Be fair. Bots and fast refreshers shouldn't win every time.
  • The seat map can be a few seconds stale. The hold itself can't be.

Notice the split. Browsing can be eventually consistent. Holding and selling need strong consistency. That split drives everything below.

The seat lifecycle

Every seat moves through a small state machine.

Seat state machine

A seat is AVAILABLE, HELD, or SOLD. A hold expires on its own if the buyer walks away. The whole design is about making the arrows in this picture safe under heavy concurrency.

Attempt 1: Check, then update

The first answer most people give:

-- Step 1: is it free?
SELECT status FROM seats WHERE seat_id = 'A12';

-- Step 2: if yes, take it
UPDATE seats SET status = 'HELD', held_by = 'user_42'
WHERE seat_id = 'A12';
Enter fullscreen mode Exit fullscreen mode

This has a classic race. Two users run step 1 at the same moment. Both see AVAILABLE. Both run step 2. Both think they own seat A12.

It's like two people checking that a parking spot is empty, then both pulling in.

The fix: make it one atomic step

Fold the check into the write. Let the database decide who wins.

UPDATE seats
SET status = 'HELD',
    held_by = 'user_42',
    hold_expires_at = now() + interval '10 minutes'
WHERE seat_id = 'A12'
  AND (status = 'AVAILABLE'
       OR (status = 'HELD' AND hold_expires_at < now()));
Enter fullscreen mode Exit fullscreen mode

If the update touches one row, you got the seat. If it touches zero, someone beat you. No race. As a bonus, an expired hold is reclaimed in the same statement, so you don't need a cleanup job for correctness.

This is correct. So why isn't it the final answer?

Where it breaks: hot rows

At 10:00 AM, nobody wants a random seat. Everyone wants the same few hundred seats near the stage.

Thousands of conditional updates pile onto the same rows. Each one waits for a row lock. Only one can win, and every loser still holds a database connection while it waits. Your connection pool drains. Requests time out. Now even the users who want cheap seats in the back can't get in.

The database is correct, but it's doing a lot of expensive work just to say "no."

Attempt 2: A fast gate in Redis

Move the hold off the main database and into Redis. One command does the atomic check and the expiry.

SET hold:event_77:A12 user_42 NX EX 600
Enter fullscreen mode Exit fullscreen mode

NX means "only set if it doesn't exist." EX 600 means "delete it in 10 minutes." If the command returns OK, you hold the seat. If it returns nil, someone else does. When the timer runs out, Redis removes the key and the seat is free again.

Redis handles this in memory, in microseconds, without the row-lock pileup. The losers get a fast "no" instead of a slow one.

Holding several seats at once

Families buy four seats together. You want all four or none. Holding them one at a time can leave you with seats A12 and A13 but not A14.

Use a small Lua script. Redis runs it atomically, so no other command can sneak in between the checks.

-- KEYS = seat keys, ARGV[1] = user, ARGV[2] = ttl
for i, key in ipairs(KEYS) do
  if redis.call('EXISTS', key) == 1 then
    return 0
  end
end
for i, key in ipairs(KEYS) do
  redis.call('SET', key, ARGV[1], 'EX', ARGV[2])
end
return 1
Enter fullscreen mode Exit fullscreen mode

In a Redis Cluster, all keys in one script must live in the same hash slot. Use a hash tag like hold:{event_77}:A12 so every seat for one event lands on the same shard.

Where it breaks: Redis is not your ledger

Redis replication is asynchronous. If the primary dies right after granting a hold, the replica that takes over may never have seen it. A second user can now hold the same seat.

So Redis can't be the final word on who owns a seat.

The fix is to treat Redis as a fast gate, not the source of truth. Before charging the card, the booking service does one more conditional write to the database:

UPDATE seats SET status = 'SOLD', sold_to = 'user_42'
WHERE seat_id = 'A12' AND status != 'SOLD';
Enter fullscreen mode Exit fullscreen mode

Redis absorbs the stampede. The database makes the final call, but now it sees only a trickle of real purchases instead of a flood of attempts. If a rare failover creates a duplicate hold, the second buyer gets a clean "sorry, that seat just sold" before any money moves. If the charge then fails, the booking service flips the seat back to available.

This "fast gate in front, durable ledger behind" idea shows up in many systems, from rate limiters to inventory. System Design Patterns covers it alongside other reusable patterns.

Attempt 3: Don't let everyone in at once

Attempt 2 is correct and fast. There's still a problem. 500,000 people are fighting over 20,000 seats. Most of them will lose. Letting them all hammer the hold service makes the experience feel like a lottery and burns capacity on requests that can't succeed.

Real ticketing systems add a virtual waiting room in front of everything.

Seat hold architecture

Here's how it works:

  1. Everyone who arrives before the sale gets a random position in a queue. Randomizing early arrivals stops the "refresh fastest" race.
  2. The waiting room admits users at a controlled rate. Each admitted user gets a signed, short-lived access token.
  3. Only requests with a valid token reach the seat map and hold service.
  4. The admit rate is tuned to what checkout can handle. If roughly 2,000 people are checking out at once and checkout takes about 4 minutes, you admit around 500 people per minute.
  5. When seats run out, the waiting room tells everyone still in line. No one wastes 20 minutes waiting for nothing.

Think of a popular restaurant. The host doesn't let 300 people rush the kitchen. They hand out pagers and seat people as tables open.

The waiting room is also your best defense against bots. You can add a challenge at the gate and rate-limit tokens per account and per payment card, long before bots reach the scarce resource.

What about the seat map?

Showing live seat status to 500,000 people is a read-heavy problem. It can be a few seconds stale.

  • Build the map from a cache that refreshes every second or two.
  • Push updates to admitted users over WebSockets or server-sent events if you want it to feel live.
  • Accept that a user may click a seat that was just taken. The hold call will return a fast "no," and the UI picks the next best seat.

Don't try to make the map strongly consistent. That's the mistake that turns a read problem into a write problem.

Key Takeaways

  • Separate what must be strongly consistent (holds and sales) from what can be stale (the seat map).
  • "Check, then update" has a race. Fold the check into one atomic conditional write.
  • Atomic database writes are correct but collapse under hot-row contention.
  • Redis SET NX EX gives fast, self-expiring holds. Use a Lua script for multi-seat holds.
  • Redis is a fast gate, not the ledger. Confirm the sale with a conditional database write.
  • A virtual waiting room turns an uncontrolled stampede into a steady, fair flow.

FAQ

Why not use a distributed lock library instead of Redis keys?
A hold is already a lock with a timeout. SET NX EX is the simplest version of that. A heavier locking protocol adds latency without fixing the core issue, which is that the final sale still needs a durable check.

Why not just scale the database?
Adding replicas helps reads, not contended writes. The contention is on a few hundred specific rows. More hardware doesn't change the fact that only one buyer can win each row.

What happens if the hold service crashes?
Holds expire on their own. Users who lost their hold see their seats go back on sale and can try again. Nothing is sold without the final database write, so no seat is ever sold twice.

How do you stop one user from holding 200 seats?
Cap seats per hold and holds per account at the waiting-room token level. Enforce the same limit again in the hold service.

Is a waiting room overkill for an interview?
Mention it once the core hold logic is solid. It shows you're thinking about the whole user experience and load shape, not just the database.

How would you handle seat selection for general admission, where there are no seat numbers? Tell me in the comments.

Top comments (0)