DEV Community

Cover image for Free-threaded Python in production: what actually broke
Ahmet Zeybek
Ahmet Zeybek

Posted on Originally published at zeybek.dev

Free-threaded Python in production: what actually broke

For fifteen years, "why is my Python threaded code not faster" had a one word answer. The global interpreter lock let one thread run Python bytecode at a time, so for CPU bound work the threads took turns on the cores instead of sharing them. Everyone worked around it with multiprocessing, Celery workers or a hot loop rewritten in C, or just lived with it.

PEP 703 removed the lock behind a build flag in 3.13, and last October 3.14 made the free-threaded build officially supported under PEP 779. The single-threaded overhead was around 40 percent in 3.13, which is why nobody ran it for real. It's now down to single digits on Linux and macOS. It still isn't the default build: you install python3.14t and opt in.

I did that for one service, and I want to write down what happened. The benchmark posts tell you about the speedup and skip the part where your code was never really thread safe and the lock was covering for it.

The service

It's an ingestion worker. It pulls batches of documents off a queue, parses them, normalises the text, computes some statistics and writes the results to Postgres. Parsing and normalising are pure Python and CPU bound. It ran as eight single threaded processes under a supervisor, each with its own database connection pool and its own in-memory copy of a 600 MB lookup table. Eight copies of 600 MB is most of the box.

What I hoped free threading would give me was simple: one process, eight threads, one copy of the table, one pool. Same throughput on a fifth of the memory.

The number

On the parsing benchmark, one process with eight threads on 3.14t finished the batch in 31 percent of the time a single thread took.1 It wasn't 8x, partly because the parser has a shared cache with a lock in it that the threads fight over, and partly because memory bandwidth on that box is what it is. Still, I'd never been able to get a number like that out of threads before.

Single threaded, the same code ran about 6 percent slower on 3.14t than on the standard 3.14 build. That matches the documented overhead, and it's the tax the rest of the process pays for parallelism in the hot loop. It was worth it for this service. For a service that mostly waits on I/O it isn't, and asyncio was already the right tool there.

Memory went from 8 processes at about 900 MB each to one process at 1.4 GB, and that was the real win.

Bug one: the counter that lied

The service keeps per-document-type counts for a metrics endpoint. This was the code:

counts: dict[str, int] = defaultdict(int)

def record(doc_type: str) -> None:
    counts[doc_type] += 1
Enter fullscreen mode Exit fullscreen mode

In CPython that line has been "thread safe" for a very long time, in the sense that the GIL made the read-modify-write on the dict entry atomic most of the time. Nothing ever guaranteed it. It just never bit anyone, because the interpreter switched threads at bytecode boundaries and this operation was short enough that it nearly always fit between them.

On the free-threaded build, two threads recording the same type at the same moment both read 41, both write 42, and one increment disappears. After a day the metrics endpoint was under-reporting by about 2 percent. There was no exception and no crash, only a wrong number.

You can fix it with a lock, with itertools.count, or by moving the counter into each thread and merging at the end. I went with the last one, since a lock on a hot counter would have been contended.

_local = threading.local()

def record(doc_type: str) -> None:
    if not hasattr(_local, "counts"):
        _local.counts = defaultdict(int)
        _register(_local.counts)
    _local.counts[doc_type] += 1
Enter fullscreen mode Exit fullscreen mode

▶ Two threads and one counter, with and without the GIL: an animation that plays in the original post.

The counter itself isn't the lesson. A lot of Python code is correct under the GIL by accident, and the free-threaded build switches the accident off.2 Single dict and list operations are protected by per-object locks. Compound ones, like a read followed by a write, aren't, and never were.

Bug two: the C extension that said it was fine

The normaliser uses a Unicode library with a C extension. The wheel had the cp314t tag, it installed cleanly, and it declared free-threading support through Py_mod_gil. I took that as a yes.

Under load, about one batch in ten thousand came back with a string mixed together from two inputs. The extension kept a static buffer that it reused between calls. Under the GIL two calls could never overlap. Without it they could, and they did.

The wheel tag means the extension was built for the free-threaded ABI. It doesn't mean the author went through the code looking for shared state. Those are two different claims, and this year people mix them up all the time. The good news is that upstream had a fix within a week of the issue. The compatibility tracker the community runs has been the most useful page on the internet for this migration, and before you trust an extension, check whether its entry there says "builds" or "tested".

Until the fix landed I wrapped the call in a lock. That put a GIL back around one function, and that's really what you do when you find one of these: put the lock back in the smallest scope that works, and carry on.

Bug three: the test suite that was single threaded

This one's on me. The test suite ran under pytest in a single thread and passed on 3.14t on the first try. That gave me confidence I shouldn't have had, because a test suite that never runs two things at once can't find a race.

Production traffic found the first two bugs. A stress test should have found them: one that runs the real code paths from many threads and asserts on the aggregate results. I wrote one afterwards, and there's nothing clever about it:

def test_record_is_consistent_under_threads():
    n_threads, n_each = 16, 50_000
    barrier = threading.Barrier(n_threads)

    def worker():
        barrier.wait()
        for _ in range(n_each):
            record("invoice")

    threads = [threading.Thread(target=worker) for _ in range(n_threads)]
    for t in threads: t.start()
    for t in threads: t.join()

    assert total("invoice") == n_threads * n_each
Enter fullscreen mode Exit fullscreen mode

The barrier is what makes it work. Without it the threads start staggered and the window for the race is small. With it, sixteen threads hit the same line at the same instant and the lost updates show up within seconds. Against the original code on 3.14t this test fails every time, and against the fixed code it passes. On the GIL build it also passes against the original code, and that's exactly why you write it.

▶ Why a barrier makes the race show up: an animation that plays in the original post.

If you're moving anything to the free-threaded build, write a barrier test for every piece of shared mutable state before you flip the switch. No other kind of test tells you anything here.

What did not break

Most of it. The service's web layer, a small FastAPI app for health and metrics, ran without changes. SQLAlchemy 2 and psycopg 3 both support the build, and the connection pool behaved. NumPy has been fine since 2.1. The logging module is safe. concurrent.futures did what it always did, except that now the thread pool executor is actually parallel for CPU work, which is a strange sentence to type after fifteen years.

The interpreter didn't crash once in six weeks. What people were afraid of with the lock gone, that the runtime itself would be unstable, didn't happen to me. All the instability was in my code and in one extension, which is where the free-threading team said it would be.

One more tool

The closest thing to a free-threading linter right now is python -X dev together with the free-threaded build's PYTHON_GIL=0 environment variable, plus a run of the suite under pytest-run-parallel, which runs each test body concurrently from several threads. Once I knew to run it, it found the counter bug on the first run.3 It gives you the barrier test idea for every test without having to write one.

Should you

If your service is I/O bound, no. asyncio already gives you concurrency without threads, and the free-threaded build would only cost you the single-thread tax.

If it's CPU bound and multiprocessing already works for you, probably not yet, unless the duplicated memory hurts the way it was hurting me. Processes are still the safest way to get parallelism in Python, because they share nothing and so can't race.

If you've got CPU bound work, read-mostly shared state you're duplicating per process, and a codebase small enough that you can audit every piece of shared mutable state, then yes, and 3.14t is stable enough to run. Read the free-threading HOWTO, check every C extension on the compatibility tracker, and write the barrier tests. Expect two or three places where the lock was doing your job for you.

3.15 is expected to finish the ABI work, and the talk is that the free-threaded build becomes the default some time after that. When it does, every Python codebase in the world goes through what I went through in March, and code that was correct by accident stops being correct. I'd rather find that out on one service, on purpose, with a test that fails on the right line.


Originally published at zeybek.dev.


  1. That's roughly the 3.1x the release notes quote for multi-threaded CPU work. ↩

  2. The docs have a page on this, "Python support for free threading", and the section on which operations are still atomic is worth reading twice. ↩

  3. It wouldn't have found the C extension one, which needs the real workload's overlap pattern. ↩

Top comments (0)