On paper, ext4's fast_commit feature is a sysadmin's dream: one tune2fs -O fast_commit lowers fsync latency, your data layout stays untouched, and reverting is just as easy. The tune2fs command returned no error, the exit code was zero, and dumpe2fs showed the flag on disk.
Then I looked at the kernel's counter: 0 commits. Across four thousand fsyncs, not a single fast commit had happened.
The command wasn't lying — the flag really was written to disk. But on ext4, being written to disk and being in effect are not the same thing, and nothing in between says so: no warning, no error line. This article is about exactly where that gap sits, which two files tell you the truth, and how the same counters, once you learn to read them, predict the number of full commits almost exactly.
What fast_commit does, briefly
In its default data=ordered mode, ext4 writes metadata changes to the JBD2 journal. When an fsync() arrives, the journal performs a "full commit": it copies every changed metadata block into the journal, orders it, puts up a barrier. Even if you only appended 4 KiB to a small file, the price is a whole transaction.
The idea behind fast commits is this: on most fsyncs very little has actually changed — an inode's size, one added extent. Instead of hauling entire blocks, you can write that delta into a small tag-length-value record and be done. The kernel documentation describes this area as a log shared with JBD2 and lists tags such as EXT4_FC_TAG_ADD_RANGE, EXT4_FC_TAG_CREAT and EXT4_FC_TAG_UNLINK.
The fallback condition is what matters: when the fast commit area fills up, when a fast commit isn't possible, or when the JBD2 commit timer fires, ext4 falls back to a traditional full commit — and that full commit invalidates every fast commit before it. So fast_commit doesn't eliminate full commits, it thins them out. The prediction model in the last third of this article is built directly on that sentence.
The measurement rig
There is no point sacrificing my real servers' filesystems; because fast_commit is an mkfs/tune2fs feature, working on throwaway loop files is both safe and repeatable. The rig is a privileged Debian 12 container on Docker Desktop's linuxkit VM:
kernel: 6.10.14-linuxkit
mke2fs: 1.47.0 (5-Feb-2023)
image: 1 GiB loop file, mkfs.ext4 defaults
journal: 32M (Total journal blocks: 8320)
fc area: Fast commit length: 128 blocks
The writer side is as plain as it gets — write 4 KiB, fsync, repeat:
f = os.open(path, os.O_CREAT | os.O_WRONLY | os.O_TRUNC, 0o644)
os.fsync(f)
t = time.monotonic()
for i in range(n):
os.write(f, b'x' * 4096)
os.fsync(f)
That block holds for most of this article; two sections use a different rig and say so on the spot — don't trip over the numbers when comparing them in your head.
One caveat up front: these are times measured on a loop device, not production NVMe numbers. Don't read the absolute milliseconds. Every claim in this article rests on counters, and counters are device-independent: it isn't the stopwatch that tells you which path was taken, it's the kernel's own ledger.
There are two ledgers:
/proc/fs/ext4/<dev>/fc_info # fast commit counters
/proc/fs/jbd2/<dev>-8/info # JBD2 transaction counters
First measurement: the stopwatch barely noticed
Two filesystems went up — one with -O fast_commit, the other with -O ^fast_commit — and each took 2,000 write+fsync pairs across six rounds. 12,000 fsyncs per filesystem.
round ON OFF
1 1.017 1.494
2 1.028 1.242
3 1.040 1.146
4 0.878 1.222
5 0.957 1.308
6 1.063 1.361
ON median=1.022ms min=0.878 max=1.063
OFF median=1.275ms min=1.146 max=1.494
median ratio OFF/ON = 1.25x
Not even one and a half times. Show this table to a colleague and they'll ask whether it could be noise, and they'd have a point: OFF's best round (1.146) comes close to ON's worst (1.063).
Now look at the ledger from the same run:
jbd2 transaction count: fc=94 nofc=12006
fast commit count (fc) = 11912
This table is not noise. The filesystem with fast_commit off opened 12,006 full transactions for its own 12,000 fsyncs — one per fsync. The one with it on opened 94. The other 11,912 fsyncs were served by fast commits.
The mechanism is unambiguously working; the stopwatch struggles to see it. That distinction deserves to be taken seriously, because the reverse happens too: enabling a knob, seeing the time improve, and declaring victory is the same noise read in the other direction. I once fell into the mirror image of this with wbt_lat_usec — I was convinced the knob was off, and ftrace said otherwise. What needs measuring isn't duration, it's the path taken.
Keep that number 94 in mind. We'll come back to it.
The real trap: tune2fs says yes, the mount doesn't hear it
Here's the event in the title. The rig differs from the previous section — a 512 MiB image with a 16M journal, built with fast_commit off and switched on later; its fc area is therefore not 128 but 256 blocks (why, in a moment). A filesystem with fast_commit off went up, got mounted and timed; then the flag got flipped while the filesystem was mounted, and the timing continued.
[1] fast_commit OFF, mounted, 2000 fsyncs
1.503 ms/fsync
fc_info commits=0 jbd2 tx=2001
[2] tune2fs -O fast_commit /dev/loop7 (fs MOUNTED)
exit code=0
feature on disk: fast_commit
[3] SAME mount, 2000 fsyncs (NO remount)
1.350 ms/fsync
fc_info commits=0 jbd2 tx=4002
[4] mount -o remount,rw
1.314 ms/fsync
fc_info commits=0 jbd2 tx=6003
[5] umount + mount (clean)
1.037 ms/fsync
fc_info commits=1993 jbd2 tx=8
Running these five steps a second time from scratch, the timings didn't hold — 1.073 / 1.008 / 1.006 / 0.982 ms in that order. The counter column was identical: 0, 2001, 4002, 6003, then 1993 and 8. The whole argument of this article lives in the gap between those two lines.
Three things are happening at once here, and all three are silent.
First: tune2fs wrote a feature flag into the superblock of a mounted filesystem and did it without a word. Exit code zero, no warning, no trace in dmesg. The version I tested is 1.47.0; I gave the command the loop device itself, so the tool was in a position to check /proc/mounts and see the mount.
Second: the running mount never noticed. All 4,000 fsyncs in steps 2, 3 and 4 went out as full transactions. fc_info still said 0 commits — and no ineligible counter had moved either. So when you ask "why aren't fast commits happening?", even the counter that would tell you the reason is empty.
Third, and the part that annoyed me most: mount -o remount wasn't enough. A sysadmin's reflex is remount; here remount changes nothing. The flag only took effect after a clean umount + mount: 1,993 fast commits, 8 transactions.
The reason is sensible, even unavoidable: fast commits require JBD2's own journal feature bit, and the kernel sets that up during mount. But being sensible doesn't excuse being silent. A maintenance window where someone runs tune2fs remotely and ticks "enabled" closes having enabled nothing.
In fairness, the kernel documentation already says part of this: the ext4 journal page states that the feature needs to be enabled at mkfs time. The measurement refines that slightly — tune2fs plus a clean umount/mount works too, so the feature itself can be turned on later. What cannot be set later is the size of the area; that comes in the last section. What the documentation doesn't mention at all is the silence in between.
Look at that 8 in the final step too, because it confirms this article's closing model in advance: this filesystem's fc area is 256 blocks and 2,000 fsyncs went at it. 2000 ÷ 256 = 7.8. Observed full transactions: 8.
The practical upshot is short: the only honest way to turn on fast_commit is at mkfs time or during a planned unmount. Writing a feature flag onto a mounted filesystem with tune2fs isn't merely useless, it's something not to do: the running kernel holds its own in-memory copy of the superblock and can overwrite your change on its next write. And after turning it on, don't say "it's on" before you've looked at fc_info.
Don't add the counters up: 150 operations, 200 reasons, 50 fallbacks
On to reading the counters, because there's a second trap there — and the first draft of this article fell into it.
A run deliberately triggering ineligible operations: 50 xattr writes, then 50 RENAME_EXCHANGE calls, then 50 directory renames, reading the counters after each step. What matters is that the baseline got read too, before any of it:
[B0] baseline (before the experiments) 8933 commits 70 ineligible (ALL named reasons 0)
[B2] after 50 setxattr 8983 commits 120 ineligible "Extended attributes changed": 100
[B4] after 50 RENAME_EXCHANGE 9036 commits 120 ineligible "Cross rename": 50
[B5] after 50 directory renames 9036 commits 120 ineligible "Dir renamed": 50
Three independent things live in those four lines, and all three mislead at first glance.
First: 50 operations moved the counter by 100. The loop calls setxattr once per file and goes round 50 times; the counter says 100. So the reason counters don't count syscalls — they count how many times the kernel called ext4_fc_mark_ineligible(). In mainline, fs/ext4/xattr.c calls it with EXT4_FC_REASON_XATTR from three separate places, and a single xattr write can pass through more than one. The same reason code carries a surprise too: in fast_commit.c, an inode with inline data is also marked with EXT4_FC_REASON_XATTR. So seeing "Extended attributes changed" in fc_info doesn't even mean you wrote an extended attribute.
Second: the sum of the reasons and the number above measure different things. 150 user operations produced 200 reason increments, but the ineligible line rose by only 50 (70 to 120), and all 50 came from the xattr step. The RENAME_EXCHANGE and directory-rename steps each raised their reason by 50 and the fallback count not at all.
In the source these really are separate: calling ext4_fc_mark_ineligible() increments fc_ineligible_reason_count[reason], while on the commit path a status of EXT4_FC_STATUS_INELIGIBLE runs fc_ineligible_commits++. One means "something made the filesystem ineligible", the other means "an fsync fell back to a full commit because of it".
So why did B4 and B5 record no fallbacks? Because both loops take their durability from a directory fsync — and as the next section shows, directory fsyncs never enter the fast commit path at all. You don't fall off a path you were never on.
Third: 70 of them have no reason whatsoever. Those 70 fallbacks sat in the baseline while every named reason was zero; what produced them was the previous section's 12,000 fsyncs. The source has an explanation, and it's an uncomfortable one: when the status is FAILED, the kernel increments both fc_failed_commits++ and fc_ineligible_commits++. So the "ineligible" line in fc_info pours genuine ineligibility and outright failure into one bucket — and fc_failed_commits isn't printed in this output on any kernel version, mainline included.
The reading rule, then, isn't "add up and compare": read both sides as deltas. The reasons tell you how many times the kernel said "this work doesn't suit a fast commit"; the number above tells you how many fsyncs paid for it. There's no reason to expect them to be equal, and their not matching isn't a bug.
Also know that the list of counters isn't fixed. My 6.10 kernel prints ten reasons; fs/ext4/fast_commit.h in mainline carries thirteen definitions today — "Inode format migration", "fs-verity enable" and "Move extents" were added later. Read your own fc_info output, not a blog post's list.
"Falloc range op" is badly named: two of six modes
One counter name misled me in particular. Reading the name Falloc range op, I assumed anything using fallocate() breaks fast commits. The counters say otherwise. Each mode ran 20 times with an fsync after each:
mode commits+ ineligible+ FallocRangeOp+
punch-hole 20 0 0
zero-range 19 1 0
zero-range|keep-size 20 0 0
collapse-range 0 20 20
insert-range 0 20 20
plain fallocate (mode 0) 20 0 0
Note the zero in the commits+ column on the collapse-range and insert-range rows: these two modes don't just bump a counter, they push every one of those fsyncs to a full commit. The other four modes stay on the fast commit path.
The ineligible+ 1 on the zero-range row bothered me, because none of the named reasons had moved. Raising the repetitions to 200 and running it on its own made it clear: 199 fast commits, 1 fallback, no named reason. So that single one isn't zero-range's own fault — it's a live example of the "operation counter and commit counter are not the same thing" situation from the previous section.
The source confirms the measurement. EXT4_FC_REASON_FALLOC_RANGE is marked from only two places in mainline: ext4_collapse_range() and ext4_insert_range(). ext4_punch_hole() does the opposite — it calls ext4_fc_track_inode() and ext4_fc_track_range(), including the change in the fast commit.
The practical benefit: there's no reason to give up on fast_commit because a database or a virtual disk image works with holes (punch-hole). The reason to give up would be running something that uses collapse-range/insert-range — which is far rarer.
Directory fsyncs get no share at all
This is where the debt left in the previous section gets paid. The next loop deletes files and takes its durability from a directory fsync:
100x (unlink + dir fsync) -> fc=+0 ineligible=+0 jbd2 tx=+100
100x (unlink + PARENT dir fsync) -> fc=+0 ineligible=+0 jbd2 tx=+100
100x (unlink + OTHER file fsync) -> fc=+99 ineligible=+1 jbd2 tx=+1
The third row is the control group and it settles the question: the culprit isn't unlink. The same deletions, when durability is taken via an ordinary file's fsync, are served by 99 fast commits. The moment you fsync a directory, a full transaction opens every time — and it isn't counted as "ineligible" either, because technically nothing is ineligible; the work simply goes another way.
I didn't follow the code path to the end, so I won't make a mechanism claim. But the measurement and the control are clear enough: a maildir-style queue, a library that writes with rename plus a directory fsync, or code that syncs the directory after every operation gets no share of fast_commit. And no counter will tell you — only the JBD2 transaction count failing to drop gives it away.
Prediction: the fc area sets the number of full commits
Now back to that number 94. In the first measurement there were 11,912 fast commits and 94 full transactions for 12,000 fsyncs. The ratio: 126.7. The line sitting in that same filesystem's dumpe2fs output: Fast commit length: 128.
If that isn't a coincidence, enlarging the fc area should pull the number of full commits down proportionally. Five separate filesystems went up, each with a different area via -J fast_commit_size=, and each took 6,000 fsyncs:
fc_size_KB FCLEN_blk journal fsync fastcommit fulltx observed_ratio predicted
128 32 32M 6000 5813 188 30.9 32
256 64 32M 6000 5907 94 62.8 64
512 128 32M 6000 5954 47 126.7 128
1024 256 33M 6000 5977 24 249.0 256
2048 512 34M 6000 5989 12 499.1 512
The predicted column isn't derived from the measurement; it's the Fast commit length value straight off the disk. At all five points the observed ratio lands within 4% of the prediction; the worst point is 30.9 against 32, a 3.4% deviation. The rule is simple:
full commits ≈ fsyncs ÷ (fc area blocks)
Why it works is clear too: in this workload every fast commit takes exactly one block. In every run where commits and numblks were both recorded, the two were identical. If the area is 128 blocks, it fills after 128 fast commits, and when it fills ext4 falls back to a full commit and refreshes the area.
I also tested where this model can break, because "one block per commit" is an assumption and assumptions don't get left unmeasured. I repeated it writing and fsyncing 1, 4 and 16 files per round:
1 file/round: fsync= 1500 fastcommit= 1488 blocks/commit=1.00 fulltx= 12 fc/tx=124.0
4 file/round: fsync= 6000 fastcommit= 5953 blocks/commit=1.00 fulltx= 47 fc/tx=126.7
16 file/round: fsync=24000 fastcommit=23813 blocks/commit=1.00 fulltx=187 fc/tx=127.3
Blocks per commit didn't budge. So the number of files doesn't break the model — each fsync writes one inode's delta and it fits in one block. The model breaks when a single fast commit needs more than one block; the ratio then drops by whatever numblks/commits is. Rather than memorising the formula, I'd suggest watching those two lines.
mkfs and tune2fs don't do the same thing
A small difference with a bill attached. Turning fast_commit on at mkfs time versus later with tune2fs pays for the journal budget out of different pockets:
[before tune2fs] Total journal size: 32M blocks: 8192 Fast commit length: 0
[mkfs -O fast_commit] Total journal size: 32M blocks: 8320 Fast commit length: 128
[tune2fs -O fast_commit] Total journal size: 32M blocks: 8192 Fast commit length: 256
On first reading this went into the draft as "the two tools have different defaults". Wrong: 256 is nobody's default.
The area's size lives in the s_num_fc_blks field of the JBD2 journal superblock. When mkfs builds the filesystem with fast_commit, it writes 128 into that field and grows the journal by exactly that much — the 32M you asked for stays in place as the full-commit journal and the fc area is added on top (8192 → 8320 blocks). tune2fs -O fast_commit writes nothing to the journal superblock at all; it sets the feature flag in the filesystem superblock and walks away. s_num_fc_blks stays zero. And JBD2 has a fallback for exactly that case:
int num_fc_blocks = be32_to_cpu(jsb->s_num_fc_blks);
return num_fc_blocks ? num_fc_blocks : JBD2_DEFAULT_FAST_COMMIT_BLOCKS;
JBD2_DEFAULT_FAST_COMMIT_BLOCKS is 256. So the 256 you see doesn't mean "tune2fs chose this", it means "nobody wrote anything, so the kernel constant applies". The number dumpe2fs prints comes from the same place.
That has two practical consequences. First: when you enable it later, the fc area comes out of the existing journal, so there's less room left for full commits. On the 512 MiB filesystem where the trap from the title was measured, the journal is 16M — 4,096 blocks — and the fc area is 256 of them, about 6% of the journal. On a write-heavy system with a tight journal, that's something to do knowingly. Second, and more insidious: every filesystem enabled after the fact gets the same 256 blocks regardless of its journal size, so the size stops being a decision related to your workload at all.
As for growing the area later, the story from the start of this article repeats itself a third time. Running tune2fs -J fast_commit_size=2048 against an unmounted filesystem gives exit code zero, not one line of output, and a clean e2fsck. Then check it — Fast commit length is still 128. Pushing 6,000 fsyncs at that same filesystem, the ratio hadn't budged: 5,954 fast commits, 47 full transactions, 126.7 again. The request was silently ignored.
The reason is easy to find in the e2fsprogs source: in tune2fs.c, the fast_commit_size value only reaches the figure_journal_size() call on the add_journal() path, and that call only happens when a journal_size was given — while add_journal() turns away a filesystem that already has one with "The filesystem already has a journal." There is no code path that resizes an existing journal's fc area. Without tearing the journal down and rebuilding it the size doesn't change; I didn't try that, so I'm not telling you to — only that -J fast_commit_size= on its own does nothing.
And do learn where dumpe2fs shows what, because this ate half an hour of mine: fast_commit appears on the Filesystem features line and not on the Journal features line. That isn't a display quirk either: the fast commit bit on the JBD2 side isn't a permanent flag, it gets set while fast commit blocks are actually present in the journal — which is how the ext4 documentation describes it. The area size only shows up in the full output, without -h, on the Fast commit length line. It's easy to look with -h and conclude it isn't there.
Before production: reverting, recovery, older kernels
Everything so far has been the question "is it working?". In production there's also "what happens when it goes wrong", and the rest of this article wasn't answering it.
Reverting is symmetrically silent. Running tune2fs -O ^fast_commit against a mounted filesystem cleared the flag from disk — but the running mount kept producing fast commits, 985 of the next 1,000 fsyncs. As with enabling, the decision is made at mount time. The introduction called reverting just as easy; it is, with the same condition attached: it wants an unmount.
Compatibility is where the real risk sits. On the filesystem side fast_commit is a COMPAT feature — an older kernel that doesn't know it will ignore it and mount read-write. But while fast commit blocks are actually present in the journal, the bit on the JBD2 side is in the INCOMPAT class. So carrying a cleanly unmounted filesystem to an older kernel is fine; carrying one with a dirty journal means a kernel that doesn't know fast commits will refuse to recover it. In an environment where a downgrade is on the table, decide accordingly.
On the e2fsck side there were no surprises: e2fsck -fn came back clean both after enabling and after clearing. That does not mean a dirty-journal recovery scenario was tested — it wasn't.
A root filesystem is its own job. For /, "a planned unmount" means a rescue environment or initramfs; it isn't something a remote tune2fs settles.
And set the expectation properly. The 1.25x here is a loop-device measurement; it says nothing about what your disk will do. What it does say is where the win comes from: a drop in full transactions per fsync. Don't put a number in a capacity plan before measuring that on your own server.
Checklist
If you sit down to try this on your own system, the order is:
-
Verify the flag; don't trust the command's exit code.
dumpe2fs -h <dev> | grep fast_committells you the on-disk state. -
Is it working?
/proc/fs/ext4/<dev>/fc_info→ iscommitsrising? If not, the mount isn't using it even though the flag is on disk. -
If the flag was just written, don't trust remount. The only valid path is
umount+mount. -
Read the two counters as deltas. The
ineligibleline above counts fsyncs that fell back; the named reasons below countext4_fc_mark_ineligible()calls — a single syscall can bump several. Compare before and after, not the totals. -
Know the counters' lifetime.
mount -o remountdoes not reset them,umount+mountdoes. If you're comparing before and after a maintenance window, take both readings within the same mount lifetime. -
Read the win from JBD2. A drop in the transaction count in
/proc/fs/jbd2/<dev>-8/infois the real evidence; ms/fsync isn't. - Does your workload get a share? A directory-fsync-heavy workload (maildir, queue directories) got none in my measurements. File-fsync-heavy ones did.
-
If the area is tight, grow it — at mkfs time. Full commits are fsyncs ÷ fc blocks. Growing it with
mkfs -J fast_commit_size=works directly; handing the same option totune2fson a filesystem that already has a journal is silently ignored. If you enabled it after the fact, your area is sized by JBD2's 256-block constant, not by your workload. - Plan the revert around an unmount too. Clearing the flag while mounted works on disk, but the running mount keeps producing fast commits.
And in capitals: fast_commit does not work at all together with data=journal. Mounted with data=journal, fc_info showed zero commits and zero ineligible after 1,000 fsyncs — including the named "Data journalling" reason counter. The counter side is entirely silent.
This is the article's one exception: here the kernel does speak. super.c carries a warning, issued when mounting with data=journal, that delayed allocation, dioread_nolock, O_DIRECT and fast commit support are all disabled. But it goes out via printk_once — once since the machine booted. On a server that has been up for months that line may be long gone from the dmesg ring, and it didn't appear on the second run in my lab either. Look at the counter anyway.
Conclusion
In this rig, fast_commit's mechanism unambiguously works: 11,912 of 12,000 fsyncs were served without opening a full transaction. Yet the stopwatch showed that same win as 1.25x — a difference that would vanish into the noise on a server graph. The two don't contradict each other; they measure different things. Whether the mechanism runs is a choice of path, and you read it off a counter; how big the win is depends on your device, your workload and the size of your fc area.
That distinction is what pushed me to write this. tune2fs said yes, the flag was on disk, the timing ticked up slightly — and all three of those were equally true in a state where nothing was working. Only commits=0 told the difference.
The kernel doesn't expose these counters for nothing. Measuring that an optimisation works, rather than that you enabled it, is usually one extra cat of effort. The real cost is not running that cat: building a capacity plan around a feature that has been off for months, then spending a night wondering why fsync latency never dropped.
Top comments (0)