DEV Community

yureki_lab
yureki_lab

Posted on

How I Tracked Down a Race Condition in Async Python With Claude Code

TL;DR

A background worker in one of my Python services was sending duplicate webhooks about once every 3,000 jobs. It never happened locally, never in tests, and only sometimes in production. I used Claude Code to turn a "can't reproduce" ticket into a deterministic, 100%-failing test in one afternoon, and the actual fix ended up being 11 lines. Here's the exact process, plus the lessons I'd steal if I were you.


The Problem

The setup was boring, which is exactly why the bug was so annoying:

  • Python 3.12, asyncio, an async Postgres driver
  • A worker pulls "send webhook" jobs off a queue and runs them with a concurrency limit of 20
  • Each job checks whether the webhook was already delivered, and if not, sends it and records the delivery

Customers started reporting that they occasionally got the same order.paid event twice. Not often. Our logs said roughly 0.03% of jobs produced a duplicate. Downstream, a few customers' systems treated each webhook as a new payment notification, so "rare" still meant real support tickets.

The kicker:

  • ❌ Never reproduced on my laptop
  • ❌ Never showed up in the test suite (~900 tests, all green)
  • ❌ Adding debug logging made it less frequent (classic heisenbug)

I'd already burned most of a day staring at the code. Everything looked correct. That's the thing about race conditions: each line is fine, it's the gaps between them that kill you.


How I Solved It

Step 1: Give the agent the symptoms, not my theory

My first instinct was to tell Claude Code "I think it's the retry logic." I stopped myself. I'd been wrong about the retry logic for four hours already, and if I fed the agent my theory, it would happily go confirm it.

Instead, I gave it raw evidence:

Symptom: duplicate webhook deliveries, ~0.03% of jobs, production only.
Evidence: attached 40 log lines for two duplicate pairs (job IDs + timestamps).
Constraints: worker concurrency = 20. Jobs can be re-enqueued on timeout.
Task: list every code path where the same delivery could be sent twice.
Do NOT fix anything yet. Rank the paths by likelihood using the timestamps.
Enter fullscreen mode Exit fullscreen mode

That last line matters. "Do not fix anything yet" keeps the agent in investigation mode instead of jumping straight to a plausible-looking patch.

It came back with four candidate paths. Path #2 was the one I'd been ignoring: two different jobs for the same delivery (one original, one re-enqueued after a timeout) running concurrently in the same worker process. The timestamps in my logs showed both copies starting within ~40 ms of each other. I had been assuming re-enqueued jobs only ran after the original died. They didn't.

Step 2: Find the check-then-act gap

Here's a simplified version of the handler (names changed, logic intact):

async def handle_delivery(job: Job) -> None:
    delivery = await repo.get_delivery(job.delivery_id)

    if delivery.status == "sent":
        return  # already done, skip

    await http.post(delivery.url, json=delivery.payload)  # ~100-800 ms

    await repo.mark_sent(delivery.id)
Enter fullscreen mode Exit fullscreen mode

Read it top to bottom and it seems fine. But every await is a point where the event loop can switch to another task. So with two jobs for the same delivery:

sequenceDiagram
    participant A as Job A (original)
    participant B as Job B (re-enqueued)
    participant DB as Postgres
    participant C as Customer endpoint
    A->>DB: get_delivery → status = pending
    B->>DB: get_delivery → status = pending
    A->>C: POST webhook
    B->>C: POST webhook (duplicate!)
    A->>DB: mark_sent
    B->>DB: mark_sent

Classic check-then-act race. The check (status == "sent") and the act (http.post) are separated by awaits, and nothing makes them atomic.

Claude Code spotted this within minutes once it had the "two jobs, same delivery" framing. Honestly, I could have spotted it too. The value wasn't the insight; it was being forced to stop assuming.

Step 3: Make the race deterministic

A theory isn't a fix. I've shipped "fixes" for race conditions before that just shifted the timing and made the bug rarer. I wanted a test that failed every single time before the fix and passed every time after.

The trick is to control the interleaving instead of hoping for it. I asked Claude Code to write a test that pauses both tasks at the exact gap using asyncio.Event:

import asyncio
import pytest

@pytest.mark.asyncio
async def test_concurrent_jobs_send_only_once(fake_repo, fake_http):
    both_checked = asyncio.Event()
    checks = 0

    original_get = fake_repo.get_delivery

    async def gated_get(delivery_id):
        nonlocal checks
        result = await original_get(delivery_id)
        checks += 1
        if checks == 2:
            both_checked.set()
        await both_checked.wait()  # hold both tasks right after the check
        return result

    fake_repo.get_delivery = gated_get

    job = Job(delivery_id="d-1")
    await asyncio.gather(handle_delivery(job), handle_delivery(job))

    assert fake_http.post_count == 1
Enter fullscreen mode Exit fullscreen mode

Before the fix: AssertionError: assert 2 == 1, 50 out of 50 runs. 🎯

This was the moment the ticket went from "intermittent, can't reproduce" to "known bug with a red test." That shift is worth more than any amount of log staring.

A note on version drift: this used pytest-asyncio 0.24 in strict mode. Older versions handle event loops differently, so if you copy this pattern, check which loop your fixtures run on.

Step 4: Fix it at two layers

Claude Code's first proposal was a single asyncio.Lock per delivery ID. That fixes the test, but it only protects one process. We run three worker replicas, so two copies of the same job on different machines would still race.

I pushed back, and we landed on two layers:

Layer 1: claim the delivery atomically in the database.

async def handle_delivery(job: Job) -> None:
    claimed = await repo.claim_delivery(job.delivery_id, worker_id=WORKER_ID)
    if not claimed:
        return  # someone else owns it

    delivery = await repo.get_delivery(job.delivery_id)
    await http.post(
        delivery.url,
        json=delivery.payload,
        headers={"Idempotency-Key": delivery.id},
    )
    await repo.mark_sent(delivery.id)
Enter fullscreen mode Exit fullscreen mode
-- claim_delivery: only one caller can win
UPDATE deliveries
SET status = 'sending', claimed_by = $2, claimed_at = now()
WHERE id = $1
  AND (status = 'pending'
       OR (status = 'sending' AND claimed_at < now() - interval '5 minutes'))
RETURNING id;
Enter fullscreen mode Exit fullscreen mode

The UPDATE ... WHERE status = 'pending' is atomic in Postgres. Two concurrent callers can't both match the row, so exactly one gets a result back. The 5-minute clause lets a crashed worker's claim expire so deliveries don't get stuck forever.

Layer 2: idempotency key for the receiver. Even with a perfect claim, there's still the case where the POST succeeds but we crash before mark_sent. Then the claim expires and we send again. You can't fully eliminate that on the sender side (it's the old "exactly-once delivery" problem), so the Idempotency-Key header lets well-behaved receivers drop duplicates. We also documented it for customers.

After deploying: zero duplicates in 21 days, down from roughly 6 per day. The deterministic test now lives in CI as a regression guard.


Lessons Learned

1. Don't hand your AI agent your hypothesis

This was the biggest one. When I'm stuck, my theory is usually the reason I'm stuck. Coding agents are very good at building a convincing case for whatever you point them at. Give them symptoms, evidence, and constraints, and ask for a ranked list of possibilities. Then say "don't fix yet."

2. A race you can't reproduce on demand isn't understood yet

If your test passes "most of the time" before the fix, you haven't proven anything. Gate the interleaving with asyncio.Event, barriers, or injected sleeps until the failure is 100% deterministic. AI agents are great at writing this kind of scaffolding, which is tedious for humans and exactly the kind of thing you'd skip at 6 p.m.

3. Every await is a door

In async Python, it's tempting to think "single-threaded means no races." Wrong. Any await between a check and the act it guards is an opening for another task. I now ask the agent, during review, to list every await between a read and a dependent write. It's a cheap, mechanical check that catches real bugs.

4. In-process locks are a trap for distributed workers

The agent's first fix (asyncio.Lock) was correct for the code it could see and wrong for the system it couldn't. It didn't know we ran three replicas. That's not the agent's fault, it's a context gap. Lesson: tell the agent about your deployment topology when the bug smells like concurrency. One sentence ("we run N replicas behind a shared queue") changes the answer.

5. Push correctness down to the database

Application-level checks are suggestions. A conditional UPDATE ... RETURNING or a unique constraint is a guarantee. When I need "exactly one winner," I now reach for the database first and treat app-level locks as a performance optimization, not a correctness mechanism.


What's Next

A few things I'm working on now that this bug is dead:

  • A concurrency review checklist that the agent runs on any PR touching job handlers: list awaits between check and act, flag in-process locks, ask about replica count.
  • A small helper for gated interleaving tests, so writing the "pause both tasks right here" scaffolding takes 3 lines instead of 20.
  • Auditing the other 14 job handlers in the same service for the same pattern. The first pass already flagged two more check-then-act gaps (lower impact, but still real).

If the audit turns up anything interesting, I'll write a follow-up with the numbers.


Wrap-up

Race conditions feel like bad luck, but they're usually just a missing guarantee. The combination that worked for me: give the agent raw evidence instead of my theory, force a deterministic repro, then fix it where the guarantee actually lives.

If this was useful:

  • 👉 Follow me on Dev.to for more build logs on working with AI coding agents in real codebases
  • 💬 Drop a comment with the nastiest race condition you've hunted down. I'm collecting war stories
  • 🚀 If you haven't tried Claude Code for debugging yet, start with a bug you've already given up on. That's where it surprised me most

Top comments (0)