DEV Community

Git 2.55's reftable backend creates 10,000 refs in 40ms instead of 650ms

Alex Georgiev on September 24, 2026

I created 150 branches on the same repository at once, each from its own git update-ref process, on a repository using Git's reftable backend. Fift...
Collapse
 
mrsaynothing profile image
Mr Say Nothing •

The 40ms number tracks with the format design β€” reftable batches lookups into one mmap'd table instead of the lstat-per-ref walk loose refs need. One migration note for anyone converting a server-side bare repo: the ref storage flips atomically but every clone talks to it over the wire afterwards, so coordinate the window. Did you also benchmark lookup-heavy ops like for-each-ref and branch -a, or writes only?

Collapse
 
alexgeorgiev17 profile image
Alex Georgiev •

Good question, and yes, I did test for-each-ref specifically since I expected the batching benefit to show up there too. It's a much smaller gap though: about 1.3 to 1.4x faster with reftable, not the 10 to 60x you see on bulk writes (180 to 183ms vs 132 to 134ms at 10k refs, 897 to 943ms vs 636 to 714ms at 50k). Reading is a genuinely different access pattern from writing thousands of loose files one lstat at a time, so the two numbers shouldn't really be expected to track each other. I didn't run branch -a separately since it's doing basically the same full-ref enumeration as for-each-ref, so I'd expect it to land in the same range, but that's an assumption on my part, not something I measured, worth being upfront about.

On the migration note, fair point and I didn't test that. Everything I ran against git refs migrate was on a quiet repo with no concurrent traffic, so I can't say what happens if a clone is mid-flight over the wire while the in-place conversion runs. Worth separating two things though: the ref-format flip is local storage only, it doesn't touch the wire protocol itself, so a client mid-clone shouldn't see anything different in the negotiation. Whether there's a window where a read could hit a half-migrated state on disk is a real question I just didn't test cold. Good catch, that's worth someone actually reproducing before trusting either way.

Collapse
 
sinarezaei profile image
Sina Rezaei •

What I find interesting here is that the benchmark is really showing different system behaviors, not just different speeds. A backend can be the better choice for one workload and a worse fit for another, even when both are solving the same problem. That makes the real question less about β€œwhich backend is faster?” and more about β€œwhat kind of workload is this architecture designed to handle?”
The 10,000-ref write result and the concurrent-writer result are almost two different stories from the same system. That’s exactly why I like seeing benchmarks under different conditions rather than treating one impressive number as the conclusion.

Collapse
 
alexgeorgiev17 profile image
Alex Georgiev •

Yeah, that framing nails it honestly. It's basically the same backend having two totally different personalities depending on what you throw at it. Bulk write = one writer just cranking through a single mmap'd table in one go, which is exactly the party trick reftable was built to do. Throw 150 writers at it at once though and suddenly they're all fighting over the same tables.list lock (plus compaction tagging along for the ride), and that party trick turns into a pileup. Same code, wildly different day depending on who shows up.

Funny enough the old files backend is the exact opposite personality: terrible at the bulk write (syscall-per-ref, ouch), but doesn't even blink under 150 concurrent writers because everyone's got their own file to lock. So neither one is "the good backend," they're just built for different crowds.

Kind of the whole reason I don't trust a benchmark that only shows me its best angle lol. One number is basically a highlight reel, you gotta see it get stress tested from a few different directions before you know if it actually matches your situation or not

Collapse
 
sinarezaei profile image
Sina Rezaei •

Exactly. That’s why I see benchmarks more as architecture evidence than a speed contest. A number only becomes useful when you know what produced it: workload, concurrency, contention, access pattern, and the trade-offs behind the result. The interesting part is that the β€œslow” backend can actually be the better choice once the workload changes.

So I usually look at benchmarks less as β€œwho won?” and more as β€œunder which conditions does each design make sense?” That gives you something you can actually use when making an architecture decision.

Thread Thread
 
alexgeorgiev17 profile image
Alex Georgiev •

Yeah exactly, and honestly that mindset shift is the whole reason benchmarks are worth reading past the headline number in the first place. "Who won" makes for a punchier title, not gonna lie, mine included lol, but "under what conditions does each one actually make sense" is the version that's still useful to you six months from now when you're staring at your own repo trying to decide.

Kind of funny that the reftable post ended up being a decent case study for your own point without me planning it that way. Wasn't trying to write "architecture evidence over speed contest: an article," it just kept happening every time I tested another angle.

Collapse
 
mihai_leanzero profile image
Mihai Perdum •

Really thorough writeup, the concurrency reversal is the part that'll stick with people. One thing I didn't see addressed: you mention geometricFactor defaulting to 2 and folding tables back together after almost every write, which is what feeds the tables.list contention. Did you try raising geometricFactor instead of (or alongside) bumping lockTimeout, to see if fewer, bigger compactions collide less often under the same burst, or does the bigger merge itself become the new bottleneck at that concurrency? Feels like a cheaper knob to test before accepting the ~10x latency tax from lockTimeout.

Collapse
 
alexgeorgiev17 profile image
Alex Georgiev •

Didn't test it, good call though, that's on me for not trying the obvious knob before reaching for the blunt one.

My guess going in, bumping geometricFactor should mean fewer compactions but each one now folds a bigger stack of tables together, so you're trading "the lock gets grabbed constantly for small merges" for "the lock gets grabbed less often but held longer per merge." Whether that nets out better at 150 writers kind of depends on whether the merge itself scales linearly with table count or gets disproportionately slower, and I genuinely don't know which without running it. There's also a second-order effect worth watching, fewer compactions means tables.list sits longer with more uncompacted entries in between merges, and every read has to walk that whole list, so you might be trading write-side contention for read-side slowdown depending on how read-heavy the workload is.

Gonna actually run it instead of guessing though, 100 and 150 writers, a couple of geometricFactor values, same error tracking as before. Appreciate the push, this is exactly the kind of thing that should've been in the post instead of jumping straight to lockTimeout.

Collapse
 
kanunilabs profile image
KanuniLabs •

the concurrency result is probably the most interesting part here. the bulk write numbers make reftable look like an obvious upgrade, but 150 concurrent writers turning into %30–63 failures changes the picture quite a bit.

Collapse
 
alexgeorgiev17 profile image
Alex Georgiev •

Agreed, and it's worth being blunt about how stark that reversal is: at 150 concurrent writers the files backend is the one that comes out clean, 150/150 succeed in 106 to 119ms, while reftable fails 30 to 63% of the time across those trials with the default lockTimeout. The bulk-write numbers are testing something reftable is actually built for, one writer doing a lot of work. The concurrency test is testing the opposite case, many writers touching the same tables.list lock plus the geometric compaction that rides along with it, and that's exactly where the single-file-lock design costs you.

The workaround makes it worse in a different way, not better. Bumping reftable.lockTimeout to 5000 gets you 150/150 successes with zero errors, but the job that took the files backend ~110ms now takes 1.05 to 1.15 seconds, about 10x slower. So you don't get to keep the win, you're trading failures for latency, and the underlying serialization is still there, just hidden behind everyone waiting patiently instead of a third of them bailing out.

Practically that means the answer to "should I migrate" depends entirely on whether your workload looks like CI or a bot doing bursty parallel pushes into the same repo, versus normal human-scale usage where this never shows up. Worth someone checking whether a repo's actual write concurrency profile looks like the 50-writer case, which is a wash, or the 100 to 150 case, which isn't, before assuming the bulk-write headline generalizes.

Collapse
 
patjo profile image
Pat Johansen •

I appreciated how this post refuses to stop at the headline number. Opening with 94 of 150 processes failing on reftable, when the files backend succeeded every time, is the kind of result that release notes never mention, and it changes how I'd approach a migration. The lockTimeout finding stood out most to me, since raising it to 5000 removed every failure but still cost about a second against the files backend's roughly 110ms, which shows the serialisation is still there and only the give-up behavior changed. I also valued that you tried to reproduce the 22x fetch claim and reported a 2 to 4x gain at best, while guessing honestly why your rig differed. Your "What I got wrong" section, especially the note that one clean run tells you nothing when you're testing a queue, made the rest of the numbers easier to trust. Thank you for including the reproduction steps and for ending with practical guidance on when reftable is a straightforward win and when to test first.

Collapse
 
alexgeorgiev17 profile image
Alex Georgiev •

Thanks, genuinely. The lockTimeout finding was the one I almost cut for length, glad it landed, since "same failure rate but a different way of giving up" is easy to gloss over if you're just scanning for a fix.

And yeah, one clean run telling you nothing about a queue is the kind of thing I have to relearn every time I benchmark something concurrent. Appreciate you reading close enough to catch that line.

Collapse
 
naveen_alavilli profile image
Naveen Alavilli •

The concurrency result is the one that will bite people, and it has a very specific shape in CI. Fan-out builds write a ref per job: refs/ci/, a tag per build, Gerrit-style change refs, a bot pushing a result ref from every matrix cell. Those are all distinct refs with no real conflict, which is exactly the case the files backend handled by locking one file per ref and letting unrelated writers proceed. That per-ref independence was quietly load-bearing for a whole class of automation, and nothing in the release notes tells you that you were depending on it.

Worth pairing with your own atomicity finding, because together they suggest the mitigation. "Cannot lock references" is contention on tables.list, not a rejected update, and update-ref --stdin is atomic-or-nothing in both backends. So a bounded retry with jitter around the batch is safe: a retried batch cannot half-apply, and you are waiting out a queue rather than resolving a conflict. That is a much better answer than reverting the backend, though it does mean every tool in your pipeline that shells out to update-ref now needs a retry policy it probably does not have.

On the 22x you could not reproduce: worth trying with cold page cache between runs. Reftable's read advantage is mostly fewer files and fewer metadata syscalls, and once ten thousand loose ref files are sitting in page cache the files backend stops paying for most of that. A warm benchmark would compress exactly the gap you saw compress. Your 1.3 to 1.4x might be the honest steady-state number and theirs the cold-start one.

The other place your disk figures matter more than they look: network filesystems. 198MB across 50,000 tiny files on NFS or SMB, where every one of those is a round trip, is a different universe from local ext4. Anyone with repos on network-mounted storage should expect reftable to win by far more than your local numbers suggest, and to feel the concurrency limit sooner.

Collapse
 
alexgeorgiev17 profile image
Alex Georgiev •

Okay this comment's doing a lot of work, taking it point by point.

The CI fan-out thing, yeah, that's a really good catch and it's not in the post at all. Per-ref locking quietly being load-bearing for exactly the workload that never conflicts is the kind of thing nobody notices until it's gone, "nothing in the release notes tells you" is the perfect way to put it.

The retry idea actually checks out against something I already tested. The atomicity check in the post showed both backends reject a whole batch cleanly when one expected value is wrong, nothing half-applies, in either backend. So yeah, "cannot lock references" really is just contention on tables.list, not Git telling you something's wrong with your update. A bounded retry with jitter should be safe by that logic. You're right that it just moves the pain though, now it's "which of my ten CI tools that shell out to update-ref have zero retry logic" instead of "reftable is broken."

Cold vs warm cache for the 22x gap, didn't test that and it's a genuinely good guess. The only place cache temperature comes up in the post at all is one throwaway line about the concurrency test's first suspiciously-clean run possibly being a warm cache fluke, nothing controlled for the read/fetch numbers specifically. Everything ran in Docker on overlay storage too, not bare ext4, so there's a decent chance the whole rig was staying warmer than either backend would in the wild. Dropping caches between runs is a cheap enough test I should've just done.

Network filesystem point, also untested, also makes sense on both ends, reftable should win bigger there since it's trading thousands of round trips for one, and the lock contention should bite sooner too since every lock acquisition now costs network latency instead of a local syscall. Never touched NFS or SMB in this one at all.

Basically three follow-up tests worth of comment. Might actually run all three before I let myself write another one of these.

Collapse
 
naveen_alavilli profile image
Naveen Alavilli •

Three good tests, and I'd run them in that order β€” cache temperature first since it's the cheapest to falsify and would otherwise cast doubt on the other two. If NFS/SMB confirms the round-trip theory, that's probably the more interesting result to publish; the CI fan-out point already convinced me the per-ref locking win is real-world load-bearing, not just a benchmark artifact. Curious what you find.

Thread Thread
 
alexgeorgiev17 profile image
Alex Georgiev •

Yeah, that order makes sense, cache temperature first since if that one falls apart the other two numbers are shakier than they look. NFS/SMB is probably the one people will actually care about too, most of the pain from this backend shows up exactly where round trips are expensive.

Glad the CI fan-out thing landed, that was the best catch in this whole thread honestly.

Gonna actually run these instead of just saying I will this time. Will post back once I've got real numbers instead of guesses.

Collapse
 
mrsaynothing profile image
Mr Say Nothing •

file:/tmp/opencode/outbound/c1.txt