DEV Community

Cover image for I Enabled casefold, and the Directory Stayed Case-Sensitive in Turkish
Mustafa ERBAY
Mustafa ERBAY

Posted on Originally published at mustafaerbay.com.tr

I Enabled casefold, and the Directory Stayed Case-Sensitive in Turkish

Anyone who has moved a file server from Windows to Linux hits the same wall: Rapor.docx and rapor.docx are one file on Windows and two on ext4. The casefold feature that arrived in ext4 with Linux 5.2 exists precisely to fix that — per-directory case insensitivity, implemented inside the filesystem itself.

This week I sat down and turned it on. It worked: I found Rapor.txt by asking for RAPOR.TXT, and by asking for RaPoR.TxT. Then I tried Turkish, and saw this — in the same directory, at the same time, IŞIK.TXT and işik.txt are one file, while IŞIK.TXT and ışık.txt are two.

So my directory became case-insensitive in English and case-sensitive in Turkish. Not half-broken; perfectly consistent, just not with the rule I expected. The reason is a single line of kernel source, and I found it.

The second finding bothers me more: cp, rsync and tar — all three — copy two files into that directory, return zero, and leave one file behind. No error, no warning, clean exit code.

A note on the terminal blocks below: they are the output of the runs as they happened, on a Turkish-locale shell, so some labels are Turkish (BULDU = found, YOK = missing, ESLESTI = matched, kalan_ad = surviving name, icerik = content). I have left machine output exactly as it came out rather than retyping it in English.

The lab

Every measurement below comes from this environment, on fresh ext4 images on loop devices:

$ . /etc/os-release; echo "$PRETTY_NAME"
Ubuntu 26.04 LTS
$ uname -r
7.0.0-34-generic
$ mke2fs -V
mke2fs 1.47.2 (1-Jan-2025)
Enter fullscreen mode Exit fullscreen mode

Setup is two commands:

sudo mkfs.ext4 -q -O casefold -E encoding=utf8 /var/tmp/cf.img
sudo mount -o loop /var/tmp/cf.img /mnt/cf
Enter fullscreen mode Exit fullscreen mode

casefold is a filesystem feature, but on its own it makes nothing insensitive. All it does is record in the superblock how text is encoded on this filesystem:

$ sudo tune2fs -l /var/tmp/cf.img | grep -i "character encoding"
Character encoding:       utf8-12.1
Enter fullscreen mode Exit fullscreen mode

Insensitivity is per directory, switched on with the +F inode flag. The kernel documentation says it in one sentence: "It is enabled by flipping the +F inode attribute of an empty directory." The empty-directory requirement is a real requirement, not advice:

$ mkdir /mnt/cf/dolu && touch /mnt/cf/dolu/x
$ chattr +F /mnt/cf/dolu
chattr: Directory not empty while setting flags on /mnt/cf/dolu
Enter fullscreen mode Exit fullscreen mode

On an empty directory it goes through, and once set, every directory you create underneath inherits the flag:

$ mkdir /mnt/cf/ci && chattr +F /mnt/cf/ci
$ mkdir /mnt/cf/ci/alt && lsattr -d /mnt/cf/ci/alt
--------------e--F---- /mnt/cf/ci/alt
Enter fullscreen mode Exit fullscreen mode

In practice those two rules combine into this: you have to catch the top of the tree you want insensitive while it is still empty. You cannot convert a populated share after the fact — you have to create a new tree and move things into it. The move itself is its own problem, and I will get to it.

The mount point itself can never be insensitive, because the root of a fresh ext4 is never empty:

$ ls -a /mnt/cf
.  ..  ci  dolu  lost+found
$ sudo chattr +F /mnt/cf
chattr: Directory not empty while setting flags on /mnt/cf
Enter fullscreen mode Exit fullscreen mode

lost+found closes that door from the start. Your insensitive tree will always be a subdirectory of the mount point.

First, the portability trap: mke2fs does not consult your kernel

I walked into this one myself. I created the casefold image inside Docker Desktop's VM, mkfs ran happily, and then it would not mount:

# uname -r
6.10.14-linuxkit
# mkfs.ext4 -q -O casefold -E encoding=utf8 /tmp/t.img && echo BASARILI
BASARILI
# mount -o loop /tmp/t.img /mnt
mount: /mnt: wrong fs type, bad option, bad superblock on /dev/loop0,
       missing codepage or helper program, or other error.
# dmesg | tail -1
[309036.421209] EXT4-fs (loop0): Filesystem with casefold feature cannot be mounted without CONFIG_UNICODE
Enter fullscreen mode Exit fullscreen mode

The cause is plain:

# zcat /proc/config.gz | grep -i CONFIG_UNICODE
# CONFIG_UNICODE is not set
Enter fullscreen mode Exit fullscreen mode

mke2fs does not check whether your kernel supports the feature; it writes the flag into the superblock and calls it a day. So when you move a casefold disk to a machine whose kernel was built without Unicode support, the disk is not corrupted — it simply never mounts. Rescue environments, embedded devices, minimal VM images: all candidates.

The picture across my own fleet looks like this. Docker Desktop's linuxkit kernel does not support it; VPS3 does, but nothing uses it:

# VPS3
$ uname -r; grep -i "^CONFIG_UNICODE" /boot/config-$(uname -r)
6.8.0-142-generic
CONFIG_UNICODE=y
$ for d in sda1 sda16; do printf "%s: " $d; \
    sudo dumpe2fs -h /dev/$d 2>/dev/null | grep -o casefold || echo "casefold YOK"; done
sda1: casefold YOK
sda16: casefold YOK
Enter fullscreen mode Exit fullscreen mode

Make dumpe2fs -h | grep casefold a habit before you take a backup image. Ever since the evening df reported zero while root kept writing, I stopped doing filesystem work without looking at the feature flags.

Baseline behaviour: names are preserved, lookups flex

Here is the part that works the way I expected — and it works well:

$ touch /mnt/cf/ci/Rapor.txt
$ for n in Rapor.txt rapor.txt RAPOR.TXT RaPoR.TxT; do
    [ -e "/mnt/cf/ci/$n" ] && echo "BULDU  $n" || echo "YOK    $n"; done
BULDU  Rapor.txt
BULDU  rapor.txt
BULDU  RAPOR.TXT
BULDU  RaPoR.TxT
$ ls /mnt/cf/ci
Rapor.txt
alt
Enter fullscreen mode Exit fullscreen mode

All four resolve to the same file, yet ls still says Rapor.txt (alt is the inherited subdirectory from the previous step). In the kernel documentation's words, the behaviour is "name-preserving on the disk" — the name written to disk is a byte-per-byte match of what the user supplied. Only the comparison is flexible, not the storage. Windows and NTFS do the same thing: keep the name as given, flex the comparison.

This is where I saw the first silence:

$ touch /mnt/cf/ci/RAPOR.TXT; echo "exit=$?"
exit=0
$ ls /mnt/cf/ci
Rapor.txt
alt
Enter fullscreen mode Exit fullscreen mode

touch did not create a new file; it opened the existing one and updated its mtime. Exit code zero. If you want to state that your intent was "a new file", you need O_EXCL — and then you learn the truth:

>>> os.open('/mnt/cf/ci/RAPOR.TXT', os.O_CREAT|os.O_EXCL|os.O_WRONLY, 0o644)
FileExistsError: [Errno 17] File exists: '/mnt/cf/ci/RAPOR.TXT'
Enter fullscreen mode Exit fullscreen mode

Keep that in mind: only the caller who asks gets to see EEXIST.

Turkish: one directory, two different rules

Now the part that surprised me. I created a file called IŞIK.TXT and looked for its variants:

Name on disk Lookup Result
IŞIK.TXT IŞIK.txt match
IŞIK.TXT işik.txt match
IŞIK.TXT ışık.txt no match
IŞIK.TXT Işık.txt no match
İŞLEM.TXT İŞLEM.txt match
İŞLEM.TXT işlem.txt no match
İŞLEM.TXT i̇şlem.txt (i + U+0307) match

Read that again. As far as this directory is concerned, the lowercase of IŞIK.TXT is işik — so the spelling that is wrong in Turkish matches, and the correct one does not. The consequence shows up immediately: when I tried to create ışık.txt, I got a second file.

>>> os.open(b+'ışık.txt', os.O_CREAT|os.O_EXCL|os.O_WRONLY, 0o644)   # created
>>> os.open(b+'işik.txt', os.O_CREAT|os.O_EXCL|os.O_WRONLY, 0o644)   # EEXIST
>>> sorted(os.listdir(b))
['IŞIK.TXT', 'İŞLEM.TXT', 'ışık.txt']
Enter fullscreen mode Exit fullscreen mode

The directory now holds, to a Turkish reader, the same word twice in two cases — and the filesystem counts them as two different things.

Why: the kernel drops the T rows

This is not a bug; it is a deliberate choice. The kernel's Unicode tables are generated from the Unicode Character Database, and fs/unicode/README.utf8data lists exactly which files feed them — CaseFolding.txt among them. Every row in that file carries a status code: C (common), F (full), S (simple), T (Turkish/Azeri special case).

Inside fs/unicode/mkutf8data.c, which generates the tables, a single line tells the whole story:

/* Use the C+F casefold. */
if (status != 'C' && status != 'F')
        continue;
Enter fullscreen mode Exit fullscreen mode

The T rows are discarded. Here are the CaseFolding.txt entries for the letters involved — the annotations on the right are mine:

0049; C; 0069;       # LATIN CAPITAL LETTER I                 -> kept
0049; T; 0131;       # LATIN CAPITAL LETTER I                 -> DROPPED
0130; F; 0069 0307;  # LATIN CAPITAL LETTER I WITH DOT ABOVE  -> kept
0130; T; 0069;       # LATIN CAPITAL LETTER I WITH DOT ABOVE  -> DROPPED
015E; C; 015F;       # LATIN CAPITAL LETTER S WITH CEDILLA    -> kept
Enter fullscreen mode Exit fullscreen mode

The detail that closes the table is the one not in that list: CaseFolding.txt has no row at all for ı (U+0131) or ş (U+015F), so both fold to themselves. UnicodeData.txt does record I as the uppercase of ı, but folding and uppercasing are different operations, and the filesystem uses folding.

From there: IŞIK folds to işik (I → i, Ş → ş) while ışık folds to ışık. Different. İŞLEM folds to i̇şlem, that is i plus a combining dot, whereas işlem is a plain i. Different again. All seven rows of the table come out of these entries.

The İ row has a subtlety. The canonical decomposition of U+0130 is already 0049 0307, that is capital I plus the dot — and the C row folds that I to i. So even without the 0130; F; row the result would land in the same place; two paths meet at one point.

The reasoning behind the choice is defensible: a filesystem has no locale. When you mount the same disk from a Turkish machine and an English one, the answer to "which file is which" cannot change. Had they kept the T rows, the file that IŞIK.TXT resolves to would depend on the machine's LANG. Between correctness and stability, they picked stability.

This is not an ext4 quirk

I ran the identical test on my own Mac (macOS 26.6.2, APFS, the default insensitive volume):

dosya 'IŞIK.TXT' varken 'ışık.txt' -> -
dosya 'IŞIK.TXT' varken 'işik.txt' -> ESLESTI
dosya 'İŞLEM.TXT' varken 'işlem.txt' -> -
dosya 'İŞLEM.TXT' varken 'i̇şlem.txt' -> ESLESTI
NFC 'rapor-é.txt' varken NFD arama -> ESLESTI
Enter fullscreen mode Exit fullscreen mode

Line for line the same. So there is no story here about Linux not knowing Turkish; this is the shared stance of case-insensitive filesystems. It has worked this way on your Mac for years and you probably never noticed — because nobody wakes up wanting to put both IŞIK and ışık in the same folder. On a migrated file server with 200,000 files, it happens for you.

casefold does normalization too, not just letters

The sentence that caught my eye in the documentation: "The comparison algorithm is implemented by normalizing the strings to the Canonical decomposition form, as defined by Unicode, followed by a byte per byte comparison." So NFD normalization is part of the comparison. I tested it — wrote é as a single code point (NFC, U+00E9) and looked it up as two (NFD, e + U+0301):

olusturuldu NFC 'rapor-é.txt' (12 bayt)
  NFD ile arama 'rapor-é.txt'  -> BULDU
  NFD buyuk harf 'RAPOR-É.TXT' -> BULDU
  NFC buyuk harf 'RAPOR-É.TXT' -> BULDU
Enter fullscreen mode Exit fullscreen mode

The control group is a directory on the same filesystem without +F — there, all four are separate files:

>>> # /mnt/cf/plain (plain directory)
>>> os.path.exists(b+'RAPOR.TXT')          # with 'Rapor.txt' present
False
>>> os.path.exists(b+'rapor-é.txt')  # with the NFC version present
False
>>> sorted(os.listdir(b))
['RAPOR.TXT', 'Rapor.txt', 'rapor-é.txt', 'rapor-é.txt']
Enter fullscreen mode Exit fullscreen mode

Look at that last line: the directory holds two names that look identical. One is NFC, the other NFD. ls prints both as rapor-é.txt. This bites you with archives coming from macOS, quite independently of case, and it drives people up the wall — this may well be the more valuable problem casefold solves.

The full lookup path works like this. The kernel runs the folding and the decomposition from one combined table; the split below is for readability:

Diagram

Three tools, zero exit codes, one file

Everything so far was "interesting". This section is "be careful in production".

On a plain ext4 I prepared two files whose names differ only in case, with different contents:

$ cd /var/tmp/kaynak && ls -1 && md5sum Rapor.txt RAPOR.TXT
RAPOR.TXT
Rapor.txt
a336ddc0133b9f2219b73d9fbdc43239  Rapor.txt
2b40728022b6f578b4f6a93288dce045  RAPOR.TXT
Enter fullscreen mode Exit fullscreen mode

Then I copied them into three separate, freshly created casefold targets:

[cp]    exit=0  kalan_ad=Rapor.txt  icerik=YENI SURUM
[rsync] exit=0  kalan_ad=RAPOR.TXT  icerik=ESKI SURUM
[tar]   exit=0  kalan_ad=RAPOR.TXT  icerik=YENI SURUM
Enter fullscreen mode Exit fullscreen mode

All three returned zero. All three left one file in the target. rsync went further and printed >f+++++++++ on two separate lines — it claimed it had transferred two new files.

Which file survives depends on the tool, and that is a trap of its own. cp and rsync keep the name they created first and overwrite the content from the second file; tar unlinks and recreates on a collision, so it takes both the name and the content from the last entry. Run the same source into the same kind of target with three tools and you get three different results — and with cp and rsync you are left with a hybrid whose name came from one file and whose content came from the other.

The one thing that does not change: one of your two files is gone, and no exit code said so.

Before migrating: scan, but scan correctly

My first instinct was a tolower-based awk one-liner. I tested it against deliberate collisions and it came up short. There are six names in the source directory; three of them collide:

$ ls -1 /var/tmp/tara
IŞIK.TXT
RAPOR.TXT
Rapor.txt
işik.txt
rapor-é.txt
rapor-é.txt
Enter fullscreen mode Exit fullscreen mode

The awk/tolower script saw only two of them — it missed the NFC/NFD pair. What the kernel actually does came out of copying the same names into a casefold directory: six names collapsed to three in the target. So the pair it missed is real data loss.

The right key is the one the kernel uses: casefold first, then NFD.

python3 - <<'PY'
import os, unicodedata, collections, sys
kok = sys.argv[1] if len(sys.argv) > 1 else '.'
g = collections.defaultdict(list)
for d, _, dosyalar in os.walk(kok):
    for f in dosyalar:
        g[(d, unicodedata.normalize('NFD', f.casefold()))].append(f)
for (d, _), v in sorted(g.items()):
    if len(v) > 1:
        print('CAKISMA:', d, sorted(v))
PY
Enter fullscreen mode Exit fullscreen mode

In the same directory it finds all three:

CAKISMA: ['IŞIK.TXT', 'işik.txt']
CAKISMA: ['rapor-é.txt', 'rapor-é.txt']
CAKISMA: ['RAPOR.TXT', 'Rapor.txt']
Enter fullscreen mode Exit fullscreen mode

Do not use this script without planting a deliberate collision first, or you will not be able to tell "found nothing" apart from "does not work". What it costs to over-trust a single measurement is laid out in that "very small time slice" piece.

One more caveat: Python's casefold uses the Unicode tables of your installed interpreter, not 12.1, and it may not line up with the filesystem on ı/İ. The script gives you a candidate list, not a verdict — make the call by looking at the names.

strict: keeping invalid UTF-8 at the door

casefold has another quiet side. In the default configuration a filename that is not valid UTF-8 is accepted — but insensitivity does not work for it. The documentation says so outright: when invalid strings are encountered, "it falls back to considering the entire string as an opaque byte sequence, which still allows the user to operate on that file, but the case-insensitive lookups won't work."

Measured, exactly so:

>>> os.open(b'/mnt/cf/op/bozuk-\xff\xfe.txt', os.O_CREAT|os.O_WRONLY, 0o644)  # created
>>> os.path.exists(b'/mnt/cf/op/bozuk-\xff\xfe.txt')
True
>>> os.path.exists(b'/mnt/cf/op/BOZUK-\xff\xfe.TXT')
False
Enter fullscreen mode Exit fullscreen mode

The result is a directory in which some names are insensitive and some are not. You cannot tell which is which without inspecting the bytes of the name. Closing that door is only possible at mkfs time, with the strict flag — and the mke2fs man page notes the default is off:

sudo mkfs.ext4 -q -O casefold -E encoding=utf8,encoding_flags=strict /var/tmp/st.img
Enter fullscreen mode Exit fullscreen mode

The outcome splits in two, and the split matters:

strict +F dizin             : EINVAL (Invalid argument)
strict ama +F OLMAYAN dizin : olusturuldu (EVET)
Enter fullscreen mode Exit fullscreen mode

strict does not cover the whole filesystem; it only applies inside insensitive directories.

And here a gap turned up that I had not expected: you cannot read back whether strict is on. I compared a strict and a non-strict image with three different tools, and all three said the same thing:

$ sudo tune2fs -l  /var/tmp/s2.img | grep -i "encoding"
Character encoding:       utf8-12.1
$ sudo dumpe2fs -h /var/tmp/s2.img | grep -i "encoding"
Character encoding:       utf8-12.1
$ sudo debugfs -R "features" /var/tmp/s2.img
Filesystem features: has_journal ext_attr ... casefold sparse_super ...
Enter fullscreen mode Exit fullscreen mode

No Encoding flags line is ever printed. The only way to answer "is strict on for this disk" is to try writing an invalid UTF-8 name into a +F directory and see whether you get EINVAL. If you are building a new share, turn encoding_flags=strict on from the start and write it down somewhere — you can neither add it later nor ask for it later.

Cost: the first measurement was wrong

I was curious about lookup cost. I created two directories of 20,000 files — one +F, one plain — and timed 20,000 random stat calls. The first result was this:

casefold (+F)  tam yazim     6.93 us/arama
duz dizin      tam yazim     1.65 us/arama
Enter fullscreen mode Exit fullscreen mode

A factor of 4.2. I was about to write "folding is expensive", except in the same run a lookup against the casefold directory with a different spelling came out at 1.86 µs. That would mean the case that requires folding is faster than the case that does not. Impossible; the measurement order had leaked in — whichever directory was measured first was paying for a cold dentry cache.

I added warm-up passes, interleaved the scenarios, and took the median of three rounds:

casefold (+F)  tam yazim     medyan   1.25 us/arama  (1.36, 1.16, 1.25)  isabet=20000
casefold (+F)  farkli yazim  medyan   1.63 us/arama  (1.63, 1.54, 1.69)  isabet=20000
duz dizin      tam yazim     medyan   1.19 us/arama  (1.08, 1.19, 1.21)  isabet=20000
duz dizin      farkli yazim  medyan   0.87 us/arama  (0.87, 1.04, 0.86)  isabet=0
Enter fullscreen mode Exit fullscreen mode

That is the real table, and even the comparison in its first two rows is weaker than I thought. For exact-spelling lookups, the casefold directory's best round (1.16) beats the plain directory's worst (1.21); with three rounds I cannot claim to have measured that difference. The honest statement: there is no measurable difference for exact spellings.

When the name genuinely has to be folded, the gap does clear the rounds: 1.63 µs against 1.19, roughly 37%. The last row is not comparable, and let me underline that — all 20,000 lookups there are misses, and a negative dentry is cheaper than a hit.

Directory metadata grew negligibly too at 20,000 files: 708,608 bytes against 696,320, that is 1.8%. Do not go looking for a performance argument to reject this feature; the arguments live elsewhere.

Nobody would have noticed if I had published the wrong measurement, so let me note it: a micro-benchmark taken in a single round and a single order tells you about your ordering, not about the thing you measured.

Your backup does not carry +F

This was nowhere in my plan, and I think it is the most expensive finding in the article. +F is an inode flag, and archive formats have no place for it:

$ lsattr -d /mnt/y/canli
--------------e--F---- /mnt/y/canli

$ tar -C /mnt/y -cf /var/tmp/yedek.tar canli
$ tar -C /mnt/y/geri1 -xf /var/tmp/yedek.tar
$ lsattr -d /mnt/y/geri1/canli
--------------e-------

$ rsync -aAX /mnt/y/canli /mnt/y/geri2/
$ lsattr -d /mnt/y/geri2/canli
--------------e-------
Enter fullscreen mode Exit fullscreen mode

Not even rsync's -aAX carries it — those flags are for ACLs and extended attributes, not inode flags. So when you back up your insensitive share and restore it, what you get back is a sensitive directory, and no command reports an error.

What follows is worse: in the restored tree, Rapor.txt and RAPOR.TXT can sit side by side again. If the next restore target is +F, the "three tools ate a file" trap from earlier in this article fires during recovery — that is, while you are coming back from a disaster.

The remedy is simple but has to be remembered: your restore procedure must create the empty +F directories first and unpack into them. If you have not measured this in a recovery drill, you do not know it.

Security: anyone who can choose a name can choose a collision

Insensitivity changes the answer to "which file is which". Anywhere an attacker controls a filename, that becomes a question about authority. Measured:

$ chattr +F /mnt/y/guv
$ printf 'ssh-ed25519 AAAA...SAHIBI\n'   > /mnt/y/guv/authorized_keys
$ printf 'ssh-ed25519 AAAA...BASKASI\n'  > /mnt/y/guv/AUTHORIZED_KEYS
$ ls -1 /mnt/y/guv
authorized_keys
$ cat /mnt/y/guv/authorized_keys
ssh-ed25519 AAAA...BASKASI
Enter fullscreen mode Exit fullscreen mode

Someone with permission to write AUTHORIZED_KEYS changed the contents of authorized_keys. The filename stayed the same, sshd reads the same path, and the content belongs to someone else. The same logic applies to every fixed-name file: .htaccess, .env, index.php.

The rule that falls out: enable casefold on shares that hold user data, not on trees that hold configuration and credentials. If they live in the same tree, give +F to the data subtree rather than to a common ancestor.

Backing out: two separate doors

If your plan is "I will turn it off if I don't like it", both doors are tighter than you think.

The directory flag only comes off while the directory is empty — the same condition as putting it on:

$ mkdir /mnt/cf/bos && chattr +F /mnt/cf/bos && chattr -F /mnt/cf/bos
$ lsattr -d /mnt/cf/bos
--------------e------- /mnt/cf/bos          # cleared

$ mkdir /mnt/cf/dolu2 && chattr +F /mnt/cf/dolu2 && touch /mnt/cf/dolu2/a
$ chattr -F /mnt/cf/dolu2
chattr: Directory not empty while setting flags on /mnt/cf/dolu2
Enter fullscreen mode Exit fullscreen mode

The filesystem feature itself is more stubborn. On a fresh filesystem where no +F directory was ever created, it clears without complaint. But once you have created one — even if you deleted it — tune2fs refuses:

$ mkdir /mnt/m/ci && chattr +F /mnt/m/ci && touch /mnt/m/ci/a
$ rm -rf /mnt/m/ci && sudo umount /mnt/m
$ sudo tune2fs -O ^casefold /var/tmp/m.img
The casefold feature can't be cleared when there are inodes with +F flag.
Enter fullscreen mode Exit fullscreen mode

No directory, no file, and the flag still counts as present. The fix is e2fsck:

$ sudo e2fsck -fy /var/tmp/m.img >/dev/null
$ sudo tune2fs -O ^casefold /var/tmp/m.img
tune2fs 1.47.2 (1-Jan-2025)
$ sudo tune2fs -l /var/tmp/m.img | grep "^Filesystem features"
Filesystem features:      has_journal ext_attr resize_inode dir_index orphan_file
filetype extent 64bit flex_bg metadata_csum_seed sparse_super large_file huge_file
dir_nlink extra_isize metadata_csum
Enter fullscreen mode Exit fullscreen mode

casefold is off the list. The interesting part is that e2fsck does not clear the flag. I read the deleted directory's inode before and after:

-- after deletion --   Inode: 13   Type: directory   Flags: 0x40080000
-- after e2fsck --     Inode: 13   Type: directory   Flags: 0x40080000
Enter fullscreen mode Exit fullscreen mode

0x40000000 is exactly FS_CASEFOLD_FL. The flag byte sits in the inode table and stays there after e2fsck too; what changes is that the inode no longer counts as in use, so tune2fs's scan skips it from then on. I did not read what e2fsck thought it repaired; the behaviour is what the lab recorded.

The error message is correct, just incomplete — it does not say "run e2fsck -f". I hit that same wall on four separate attempts; I am writing it down so you do not hit it once.

A decision frame

If I am building a new insensitive share:

  • At mkfs time, -O casefold -E encoding=utf8,encoding_flags=strict. You can neither add strict later nor read it back; write the decision down.
  • chattr +F the root of the tree while it is empty; subdirectories inherit it.
  • The mount point is never empty because of lost+found, so you will always work with a subdirectory.
  • Verify CONFIG_UNICODE=y on every kernel that will mount this disk. Count the rescue environment.
  • Keep configuration and credential files outside the insensitive tree.

If I am migrating an existing tree:

  • Scan with a casefold + NFD key; plain tolower misses normalization collisions.
  • Prove the scanning script with a deliberate collision.
  • Do not trust the exit codes of cp/rsync/tar. Compare the file count between source and target afterwards; that is the only check that works.
  • Know that I/ı and İ/i pairs in Turkish names do not count as collisions as far as the filesystem is concerned.

For backup and recovery:

  • The restore procedure has to create the +F directories empty first. The archive does not carry the flag.
  • In a drill, check lsattr on the restored tree, and the file count too.
  • Backing out with ^casefold needs umount plus e2fsck -f; in production that means planned downtime.

Places I would not touch at all:

  • Applications that compare filenames themselves. readdir returns the name on disk; insensitivity stays in the kernel.
  • Backup targets under continuous sync. Every sync from a sensitive source into an insensitive target carries a silent-merge risk.

Conclusion

casefold is a good feature and it works. My mistake was reading the phrase "case-insensitive" with the meaning it has in my own language. The filesystem never promised me insensitivity; it promised an equivalence defined by Unicode, independent of locale. The two coincide in English and diverge in Turkish, and the filesystem does not regard the second case as an exception — that is simply the rule.

The lesson I take from this is bigger than filesystems: when a system says two things are "the same", do not build on top of it before asking whose "same" it means. tolower has a locale; casefold does not. cp's exit code says "I transferred", not "I transferred all of them". rsync -aAX reads like "preserve everything" but does not cover inode flags. The gaps between those are silent — right up until one file is missing from a 200,000-file migration.

Two current notes. casefold is spreading: tmpfs now supports casefold=utf8-12.1.0 and strict_encoding as mount options, with the same +F and inheritance rules. And the "your application does not know" gap above is closing — FS_XFLAG_CASEFOLD and FS_XFLAG_CASENONPRESERVING have landed in the kernel's userspace interface, so an application can now ask FS_IOC_FSGETXATTR whether a directory is insensitive. The comment above the definition carries its own warning: do not assume the bit belongs to directories only.

One more: the kernel's Unicode tables are still generated from 12.1.0, the release from 2019. For something like filename equivalence, a table frozen for seven years is, I suspect, a choice rather than an oversight.

Official Sources

Top comments (0)