DEV Community

Cover image for Address Space Exhaustion in Production
Francisco Perez
Francisco Perez

Posted on Originally published at uncorreotemporal.com

Address Space Exhaustion in Production

I went looking for why none of our registered users were using the product. I found something worse: for most of them, the product had simply stopped working. Address space exhaustion had been quietly breaking mailbox creation for weeks, and not one of our monitoring tools said so. One of them actively told us the opposite.

This is the writeup of that outage. The bug itself is almost embarrassingly simple — a combinatorial space that filled up because nothing ever left it. The interesting part is the twenty minutes I spent believing my instruments over my arithmetic, and what that says about how these failures hide.

A generator with 36,000 possible outputs

We hand out disposable email addresses. They're meant to be readable and quotable over a phone call, so they aren't random hashes — they're adjective-noun-NN@domain. The generator was about as simple as it looks:

_ADJECTIVES = ["mango", "coral", "verde", "plata", ...]  # 20 entries
_NOUNS       = ["panda", "tigre", "koala", "llama", ...]  # 20 entries

def generate_friendly_address(domain: str) -> str:
    """
    Genera una direccion amigable del tipo 'mango-panda-42@dominio.com'.    Hay 20x20x90 = 36 000 combinaciones posibles.
    Si hay colision en DB, el router reintenta (max 5 veces).
    """
    adjective = random.choice(_ADJECTIVES)
    noun = random.choice(_NOUNS)
    number = random.randint(10, 99)
    return f"{adjective}-{noun}-{number}@{domain}"```
{% endraw %}
Note the docstring. Past me did the multiplication correctly — 20 x 20 x 90 = 36,000 — wrote it down, and shipped it anyway. The number was never hidden. It was sitting in the source, in Spanish, three lines above the code that would eventually take the service down.

The creation endpoint handled collisions the obvious way: generate a candidate, check whether it exists, and if it does, try again up to five times.
{% raw %}


```python
for attempt in range(_MAX_ADDRESS_RETRIES):   # 5
    address = generate_friendly_address(settings.domain)
    existing = await db.execute(
        select(Mailbox.id).where(Mailbox.address == address)
    )
    if existing.scalar_one_or_none() is None:
        # ... insert and break
        break

if mailbox is None:
    raise HTTPException(status_code=503, detail="No se pudo generar una direccion unica.")
Enter fullscreen mode Exit fullscreen mode

Five attempts is plenty when the space is 5% full. It is not plenty at 93%.

Why address space exhaustion is invisible until it isn't

Here is the shape of this class of bug, and why it gives you no warning.

The probability that a single request fails is the probability that all five candidates collide: occupancy^5. At 50% occupancy that's 3%. At 80%, still only 33%. At 90% it's 59%. At 93.3% — where we actually were, with 33,590 of 36,000 addresses taken — it's:

0.933^5 = 0.707  ->  ~70% of requests should fail
Enter fullscreen mode Exit fullscreen mode

That curve is the whole story. For months the failure rate is a rounding error, then it goes vertical. There is no gentle degradation to alert on, no slow climb anyone notices in a weekly review. The system is fine, fine, fine, and then it is 70% broken.

And when I measured it live against production — just firing ten requests at the endpoint — nine of ten came back 503.

Lie #1: the analytics said 18 errors

Our frontend emits a funnel_error event when a step fails. Querying it for the whole history returned 18 errors of this type, spread over six weeks. Eighteen is a number you look at and move on from. Eighteen is noise.

The API had returned hundreds of 503s over the same period.

The gap wasn't a bug in the event pipeline. It was that funnel_error only fires on certain frontend paths — the ones where a component was wired to report it. Anything that failed outside those paths (API clients, the MCP server, the retry that happened before the component mounted) failed silently as far as analytics was concerned.

This is the lesson I'd generalize: product analytics measures what you instrumented, which is a subset of what your users did, which is a subset of what your system did. It is a fine tool for asking "did people click this?" It is not a health signal, and the moment you treat it as one you've built a monitor whose blind spots exactly match the code paths nobody thought about.

Lie #2: the container logs said 4%

Analytics undercounting is at least an honest kind of wrong — it never claimed to be complete. The second tool was worse, because it gave me a confident, specific, wrong answer.

I pulled the request log out of the container and counted status codes:

docker logs uct_app --tail 100000 2>&1 | grep "POST /api/v1/mailboxes" \
  | grep -oE "(503|201)" | sort | uniq -c
Enter fullscreen mode Exit fullscreen mode
13451 201
  314 503
Enter fullscreen mode Exit fullscreen mode

A 2.3% failure rate. I narrowed to the most recent requests — last 3000, last 1000, last 200 — and got 3.7%, 4.7%, 4%. Consistent, plausible, and completely wrong. The real rate at that moment was around 90%.

What happened is that --tail N did not return the actual tail. With log rotation in play, the window I was slicing wasn't the window I thought I was slicing, so the lines I was confidently calling "the last 200 requests" were nothing of the sort.

I want to be precise about the failure mode here, because "logs lied" is too glib. The log file wasn't corrupt. Every line in it was true. The command returned a real slice of real data — it just wasn't the slice I asked for, and nothing in the output said so. A tool that returns plausible wrong data with no error is far more dangerous than one that fails loudly.

Trusting the instrument over the arithmetic

Here's the part I got wrong, and it cost me most of the debugging time.

I had two sources. One was a simulation: I pulled every occupied address out of the database, generated 20,000 candidate addresses using the exact production generator, and applied the exact five-retry logic.

ocupadas en DB: 33590
espacio total : 36000
libres        : 2410
fallo simulado: 70.4% de 20000 peticiones (5 reintentos)
Enter fullscreen mode Exit fullscreen mode

The other was the log, saying 4%.

These cannot both be true, and I spent a long stretch assuming the simulation was the thing that must be wrong. I re-checked the word lists. I diffed the deployed generator against my local copy in case the container was running older code. I checked whether multiple domains were splitting the namespace. I verified the enum of every adjective and noun actually present in the database. All of it came back confirming the model: 36,000 slots, 33,590 taken, 2,410 free.

I was interrogating the calculation because the calculation was the thing I could inspect. The log was just... output. It felt like data in a way the model didn't.

That instinct was backwards. The arithmetic was four lines long and independently verifiable. The log was the output of a distributed system, a container runtime, a rotation policy, and a CLI flag whose semantics I had never actually read. The simple, checkable thing was the trustworthy one. When a model and a measurement disagree, "which of these has more moving parts I haven't verified?" is a better question than "which of these feels more like a fact?"

What finally settled it was refusing to arbitrate between the two and going to get a third source: ten real HTTP requests against production. Nine failed. The model had been right the whole time.

The root cause: nothing was ever released

Filling 93% of a 36,000-address space requires either enormous traffic or a leak. It was a leak.

Mailboxes are disposable. They expire — most within an hour. A background loop handled expiry like this:

async def _expire_mailboxes() -> int:
    """Marca is_active=False en buzones expirados."""
    now = datetime.now(timezone.utc)
    async with AsyncSessionLocal() as db:
        result = await db.execute(
            update(Mailbox)
            .where(Mailbox.expires_at <= now, Mailbox.is_active == True)
            .values(is_active=False)
        )
        await db.commit()
        return result.rowcount
Enter fullscreen mode Exit fullscreen mode

It marks. It never deletes.

Meanwhile the uniqueness check in the creation path was:

select(Mailbox.id).where(Mailbox.address == address)
Enter fullscreen mode Exit fullscreen mode

No is_active filter — and correctly so, because an address that was used an hour ago shouldn't be reissued to a different person while mail might still arrive for it. But the consequence is that an address is consumed permanently the first time it's handed out. Expiry freed the mailbox. It never freed the name.

The numbers made this vivid once I looked: of 33,590 mailboxes in the table, 5 were alive. The other 33,585 were expired, inert, and still holding their address hostage. The oldest had expired on March 2nd — six months earlier.

There was also a human cost sitting in the data. One user registered, tried to create a mailbox 38 seconds later, got a 503, tried again with a different TTL, got another 503, and never came back. That's the entire lifetime of that account: sign up, hit a wall twice, leave.

The fix

Three changes, in order of how much they matter:

Purge expired mailboxes. This is the actual fix; everything else buys time. A retention window — 72 hours, long enough that a user who wanders back can still read their mail — after which the row is deleted and the address returns to the pool.

MAILBOX_RETENTION_HOURS = int(os.getenv("MAILBOX_RETENTION_HOURS", "72"))

async def _purge_expired_mailboxes() -> int:
    cutoff = datetime.now(timezone.utc) - timedelta(hours=MAILBOX_RETENTION_HOURS)
    async with AsyncSessionLocal() as db:
        result = await db.execute(
            delete(Mailbox).where(
                Mailbox.is_active == False,
                Mailbox.expires_at < cutoff,
            )
        )
        await db.commit()
        return result.rowcount
Enter fullscreen mode Exit fullscreen mode

Widen the space. The numeric suffix went from two digits to four: 20 x 20 x 9,000 = 3,600,000 combinations, a 100x increase for a one-line change and no loss of readability (suave-delfin-1608 reads exactly as well as suave-delfin-16).

Raise the retry ceiling from 5 to 10. Cheap insurance that does nothing on its own — at 93% occupancy, ten retries still fails half the time. Widening the space without purging would have bought about ten days at our traffic. Purging without widening would have worked, but with no headroom for a traffic spike. Both together give a steady state around 10,000 occupied addresses: 0.3% of the new space.

We also ran a one-off purge of the backlog: 33,596 mailboxes deleted, cascading to 12,982 messages, all of them long expired. Occupancy went from 93.3% to 0.005%, and a burst of 30 creation requests came back 30 for 30.

Address space exhaustion is not an email problem

Any system that hands out human-readable identifiers from a finite pool has this bug latent in it. Invite codes. URL slugs. Short links. Room codes for a video call. Coupon codes. Anything where somebody chose a format because it reads well, computed the combination count once, and wrote it in a comment.

Three things make it dangerous:

  1. The failure curve is a cliff, not a slope. With k retries, the failure rate is occupancy^k, which is negligible right up until it isn't. There's no gradual signal to alert on.
  2. The pool usually leaks. Soft-deletes, tombstones, audit requirements, and "we might need it later" all mean identifiers get retired without being released. Check whether your uniqueness constraint and your lifecycle agree about when an ID is free — ours disagreed for six months.
  3. Generic error codes hide it. A 503 from an exhausted namespace looks identical to a 503 from a database hiccup. Ours said "No se pudo generar una direccion unica" in the response body and nobody was reading response bodies.

If you have such a pool, two things are worth doing this week: compute your current occupancy as a percentage, and confirm that something actually deletes retired identifiers. Both are ten-minute queries. Ours would have caught this in March.

What I'd take away

The bug was trivial and the arithmetic was in a docstring from day one. What made it an outage was that every instrument pointed away from it — analytics because it only measured instrumented paths, container logs because a flag didn't mean what I assumed, and my own judgment because I trusted output that looked like data over a calculation I could verify in four lines.

The habit I'm keeping from this one is cheap: when two sources disagree, don't arbitrate between them. Go get a third that touches the real system. Ten curl requests settled in thirty seconds what I'd spent twenty minutes reasoning about, and they'd have settled it on day one just as well.

If you want to see what the service actually does when it isn't returning 503s, it's at uncorreotemporal.com.

Top comments (3)

Collapse
 
raju_dandigam profile image
Raju Dandigam •

@francofuji The rotated-log slice is the detail I'd turn into checkable diagnostic metadata: show the actual first/last timestamps and sample count, not just “last N requests.” That makes stale-but-plausible evidence easier to challenge. Alongside the namespace fixes, are you tracking occupancy and retry exhaustion at the API boundary, independent of frontend analytics? Those two signals would expose shrinking headroom before the signup funnel starts losing users.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.