The .gitignore in this blog's repository looks like years of accumulated habits: build output, a local database, a directory left over from a tool nobody uses any more, temporary files from the shorts pipeline. To see what is actually being left out, the most obvious command went in:
git status --porcelain --ignored
The answer is three lines:
!! data/
!! dist/
!! node_modules/
Three lines. Not bad at all — except it took noticeably longer than plain git status. Curiosity won, out came the stopwatch, and then git's own internal counters. To print those same three lines, git had walked 36,811 entries and opened 3,197 directories. The run without the flag visited 81. (What I call an entry is git's paths-visited counter.)
More than four hundred and fifty times the work, hiding behind three lines of output. One caveat straight away: that 81 was taken with the untracked file cache on. With the cache off, the flagless run walks 6,250 paths and the ratio drops to 5.9 times — which is what a fresh repository shows you, because core.untrackedCache defaults to keep and the cache is not set up on its own. Both ends tell the same story; only the magnitude changes.
This piece follows that gap: where it comes from, which mode removes it, and why "just always use the cheap mode" is advice I cannot give you. It ends at a two-line decision inside dir.c.
A similar scene showed up in a different tool once: fstrim said 2 GiB again, and nothing was freed looked at a command whose reported number said nothing about the work performed. This one is the mirror image: tiny output, enormous work. Two ends of the same lesson.
The lab: 33 thousand files, one repository
None of the measurements touched the main working tree. A separate worktree came off origin/main, and then the ignored directories from the main repository went into it through APFS's clonefile support — cp -c -R, four seconds, almost no disk space:
git worktree add --detach /tmp/lab/wt origin/main
cp -c -R ~/IdeaProjects/mustafaerbay/node_modules /tmp/lab/wt/node_modules
cp -c -R ~/IdeaProjects/mustafaerbay/dist /tmp/lab/wt/dist
cp -c -R ~/IdeaProjects/mustafaerbay/data /tmp/lab/wt/data
The resulting landscape: 33,572 files in the working tree, 6,165 of them tracked. The overwhelming majority of the rest is node_modules — 27,135 files on its own. Four fifths of the tree is the part git has been told to skip. Most repositories look like this; in yours the ratio may be sharper still.
Environment: git 2.48.1, macOS 26.6.2, SSD on APFS. Timing came from a small Python wrapper — two warm-up runs, then the median of fifteen:
import subprocess, time, statistics
def measure(cmd, n=15, cwd=None):
for _ in range(2):
subprocess.run(cmd, cwd=cwd, stdout=subprocess.DEVNULL)
s = []
for _ in range(n):
t = time.perf_counter()
subprocess.run(cmd, cwd=cwd, stdout=subprocess.DEVNULL)
s.append((time.perf_counter() - t) * 1000)
return statistics.median(s)
Median, min and max were all recorded; the table below shows the median. The page cache is warm on every run — these are not cold-disk numbers, but the ratios are the portable part.
For git's own counters, GIT_TRACE2_PERF is enough:
GIT_TRACE2_PERF=1 git status --porcelain --ignored 2>&1 >/dev/null | grep visited
That tells you how long the read_directory region took and how many directories and paths it walked. Note what this counts: not syscalls, but git's own traversal counter.
The measurement: same three lines, four times the clock
Four flag variants, two states each — with the untracked file cache (core.untrackedCache) off and on. Why the cache enters the story at all comes a few sections down.
| command | cache off | cache on | dirs/paths (off) | dirs/paths (on) |
|---|---|---|---|---|
git status --porcelain |
20.7 ms | 15.3 ms | 82 / 6,250 | 82 / 81 |
--ignored=no |
20.0 ms | 15.3 ms | 82 / 6,250 | 82 / 81 |
--ignored=matching |
18.8 ms | 20.2 ms | 82 / 6,250 | 82 / 6,250 |
--ignored |
79.0 ms | 78.2 ms | 3,197 / 36,811 | 3,197 / 36,811 |
Three things stand out.
First: bare --ignored is roughly four times everything else. With the cache on, 5.1 times the plain run.
Second: --ignored=matching and bare --ignored print the same three lines. In this tree the outputs are identical character for character. One walks 36,811 paths, the other 6,250.
Third, and the sneakiest: with the cache on, --ignored=matching is slower than plain git status (20.2 against 15.3 ms). The mode I took to be the cheap one wipes out everything the cache was giving me.
The output collapses, the work does not
Let me make that "same output" claim concrete. The two modes side by side:
$ git status --porcelain --ignored
!! data/
!! dist/
!! node_modules/
$ git status --porcelain --ignored=matching
!! data/
!! dist/
!! node_modules/
Three lines, three lines. Not one of the 27,135 files inside node_modules is listed; both collapse the directory into a single line.
But the bare mode cannot write that single line without looking at all 27 thousand of those files. It walks them one by one, adds them to its list of ignored entries, concludes at the end that "everything in this directory turns out to be ignored", and then throws the entire collected list away and prints just the directory name. You cannot learn that it does this much work from the documentation; it shows up in the counters and in the source.
The official docs only tell you what you will see. git-status(1) says the mode parameter for --ignored is optional and defaults to traditional. In other words, the expensive one is also the default one.
Why it descends: a two-line decision in dir.c
Counters tell you something happens; they do not tell you why. The why sits in two assignments inside treat_directory in dir.c (v2.48.1 source):
check_only = ((dir->flags & DIR_HIDE_EMPTY_DIRECTORIES) &&
!(dir->flags & DIR_SHOW_IGNORED_TOO));
...
stop_early = check_only && excluded;
Because git status does not report empty directories, DIR_HIDE_EMPTY_DIRECTORIES is on in the default untracked-files mode — wt-status.c sets it under the condition show_untracked_files != SHOW_ALL_UNTRACKED_FILES, so with -uall it does not. In the default mode, then, a run without the flag has check_only set to one, and when an ignored directory comes up, stop_early is one too. The comment above it spells the shortcut out: if the directory matches an exclude pattern, git can stop at the first file it finds underneath, because the only remaining question is whether the directory is empty.
The moment you pass --ignored, the DIR_SHOW_IGNORED_TOO flag lights up and check_only drops to zero. The shortcut closes; git walks the whole tree. What happens after the walk is in the same function:
if (want_ignored_subpaths) {
...
state = path_none;
} else {
for (int i = old_ignored_nr; i < dir->ignored_nr; i++)
FREE_AND_NULL(dir->ignored[i]);
dir->ignored_nr = old_ignored_nr;
}
That FREE_AND_NULL loop releases every ignored entry collected under the directory and winds the counter back to where it was. The information produced gets destroyed; !! node_modules/ is what remains. Most of the 79 milliseconds goes into building a list that will be discarded.
The matching mode leaves the same function earlier, without entering the directory at all:
if ((dir->flags & DIR_SHOW_IGNORED_TOO) &&
(dir->flags & DIR_SHOW_IGNORED_TOO_MODE_MATCHING))
return path_excluded;
node_modules/ matches a pattern; a pattern-matched directory is reported as ignored and closed. That is where 6,250 paths instead of 36,811 comes from.
Here is the interesting part: git had cut this cost once before. The 2.15.0 release notes record that an ignored directory containing no tracked paths used to have all of its ignored paths enumerated unnecessarily, and that the code path was relieved of that overhead. My first draft reconciled this with "that shortcut only ever applied when you don't pass --ignored". That was wrong; putting the tags side by side showed the opposite.
The 2.15.0 commit itself (5aaa7fd3, "Improve performance of git status --ignored") targets exactly this command and gives a comparison on a repository with 196,000 files in 400 ignored directories: git status --ignored at 3.9 seconds before, 1.4 seconds after. In that release treat_directory passed check_only into the recursion as a constant 1; today's check_only = (… && !(dir->flags & DIR_SHOW_IGNORED_TOO)) condition and the stop_early variable do not exist in that code at all. Downloading five tags and checking: stop_early first appears in v2.27.0.
The commit that introduced it is 8d92fb29 ("dir: replace exponential algorithm with a linear one", April 2020), and its goal was not to create a regression but to remove a worse problem — recursion on the --ignored path that doubled at every directory level. The same commit describes today's cost in its own words:
…this forces us to walk all untracked files underneath the directory as well…
and adds that what it finds is stripped from the output. The 36,811 in my table is that sentence. So the strongest source for this piece's thesis is not the release note, but the admission inside the commit that traded the note's gain away. The 2.16.0 notes, three months after 2.15.0, announce that the output stopped being tightly tied to the --untracked-files mode and became more flexibly controllable — that is where matching comes from, and today it is your only way out.
The two modes are not the same thing
Everything so far reads like an invitation to write =matching everywhere. But the outputs matching in this tree is a coincidence; the two modes are saying different things. Three lines are enough to see the difference:
mkdir ignored-logs
echo x > ignored-logs/run.log # *.log is in .gitignore, the directory is not
Now the directory itself matches no pattern, while everything inside it is ignored. The two modes answer this case differently:
$ git status --porcelain --ignored=traditional
!! ignored-logs/
$ git status --porcelain --ignored=matching
!! ignored-logs/run.log
traditional collapses the directory and gives you its name; matching does not report the directory at all and counts off its contents instead. The documented definition says exactly this: in matching mode, if a directory matches an exclude pattern the directory is shown and the paths inside it are not; if it does not match a pattern but all of its contents are ignored, the directory is not shown and all of the contents are.
The comment on the source side names the same thing an "implicitly ignored" directory (dir.h): directories that do not match a pattern but whose contents are entirely ignored are not reported, and the contents are reported instead.
The practical consequence: if a script consumes this output, switching modes changes the meaning of what it consumes. If your .gitignore leans on extension patterns (*.log) rather than directory patterns (node_modules/), matching can dump hundreds of lines on you — expensive and noisy at once. In this tree all three ignored entries come from directory patterns, which is why the outputs coincided. There is no guarantee they coincide in your repository; look first, switch second.
And then there is this combination, sitting in the docs as a short subordinate clause: given together with --untracked-files=all, traditional displays individual files inside ignored directories. Timed:
$ git status --porcelain -uall --ignored | wc -l
27449
108 milliseconds, 27,449 lines, and the same 36,811 paths walked. Here at least the work reaches the output — the collected list is not thrown away. If you genuinely want to see ignored files one by one, this is the right command; if you typed it by accident, twenty-seven thousand lines land in your terminal.
The cache leaves the table the moment you say --ignored
The strangest cell in that table was matching coming out slower than the bare run with the cache on. The reason is that the untracked file cache is never used on this path.
core.untrackedCache is an index extension that trusts directory modification times to skip the untracked file scan. The docs say the default is keep, that true adds it and false removes it, and that you should check that mtime works properly on your system before turning it on. The check has its own command:
$ git update-index --test-untracked-cache
Testing mtime in '/tmp/lab/wt' ...... OK
Once it is on, the effect is plain in the counters: plain git status drops from 6,250 walked paths to 81, and from 20.7 ms to 15.3 ms. A real gain.
Then you add --ignored and the gain evaporates. The reason is eight lines in wt-status.c (v2.48.1 source):
if (s->show_ignored_mode) {
dir.flags |= DIR_SHOW_IGNORED_TOO;
if (s->show_ignored_mode == SHOW_MATCHING_IGNORED)
dir.flags |= DIR_SHOW_IGNORED_TOO_MODE_MATCHING;
} else {
dir.untracked = istate->untracked;
}
The single line that wires the cache up — dir.untracked = istate->untracked; — lives in the else branch. Ask for any ignored mode, matching included, and the cache never comes into play. That is why matching walks 6,250 paths where the flagless run was down to 81.
The rule that falls out is short: a script that adds --ignored hands back everything the cache was giving it, in the same call.
"What about fsmonitor?" — it went on, and nothing changed
The question on everybody's mind by now: does git's other accelerator for large repositories, the filesystem monitor, rescue this walk? With core.fsmonitor on, git runs a background process that listens to filesystem events and learns which paths changed from there.
An A/B run in the same tree state, changing only that one setting. In both rows the untracked file cache is on — since this piece argues that the monitor can only do its work through that cache, the variable has to be held fixed:
git status |
git status --ignored |
|
|---|---|---|
| fsmonitor off | 16.4 ms | 81.8 ms |
| fsmonitor on | 12.5 ms | 83.1 ms |
(This A/B ran in its own session with its own baseline; the gap between 16.4 ms and the first table's 15.3 ms is a session difference, and what is being measured is the comparison within each row.)
On the plain run the gain is real: 16.4 down to 12.5 milliseconds. On the --ignored side nothing moves; the 1.3 millisecond difference sits inside the measurement noise, and the walked-path count does not twitch.
The reason continues the previous section's eight lines. Inside dir.c, the monitor's only contribution to the directory walk goes through the untracked file cache:
if (!untracked)
return 0;
/*
* With fsmonitor, we can trust the untracked cache's valid field.
*/
refresh_fsmonitor(istate);
if (!(dir->untracked->use_fsmonitor && untracked->valid)) {
The first line decides it: no cache, immediate return. Pass --ignored and dir.untracked stays empty, so the monitor is never even consulted on this path. Both of the mechanisms git offers for "make status fast" hang off the same else branch; close that branch and the two of them leave the table together.
A note in the interest of honesty: this repository tracks 6,165 files. The monitor's gain on the plain run grows with the size of the tracked set, and the roughly 4 milliseconds here is the small-scale version of it. The --ignored result, though, is scale-independent, because it is a "never asked" result.
How the cost grows
The ratio is interesting, but it is worth knowing whether it travels. Two and four copies of node_modules later, through clonefile, the same measurement:
| ignored tree | paths walked |
--ignored time |
per path |
|---|---|---|---|
| 1× | 36,811 | 85.7 ms | 2.33 µs |
| 2× | 67,070 | 147.9 ms | 2.21 µs |
| 4× | 127,588 | 278.1 ms | 2.18 µs |
These three rows were timed in a session of their own, back to back, in a freshly built clean worktree — and in that session the 1× figure came out at 85.7 ms rather than the 79.0 ms of the earlier table. The reason is other work running on the machine at the time. Rather than hide the discrepancy I am writing it down, because it is exactly why the portable quantity is not the absolute time: it is the per-path cost and the linearity of the scaling.
Path count rises by a factor of 1.82 while time rises by 1.73; then by 1.90 while time rises by 1.88. Roughly 2.2 microseconds per path. That number belongs to this machine, this filesystem and a warm cache — but the shape of the scaling will be the same for you. The rough estimate for your own tree goes like this: read the paths-visited counter from an --ignored run, measure your own per-path cost once, multiply. A fifty-thousand-file node_modules tree and a spinning disk move this arithmetic from hundreds of milliseconds into seconds.
What matters is not the size of a single measurement but its product with frequency. Seventy-nine milliseconds is nothing for a command you type once a day. Inside a watcher that fires on every file save, or called before every step in CI, the same seventy-nine milliseconds is a different animal.
Neighbouring commands: a similar bill, a different mechanism
The dry run of git clean is expensive too. My first instinct was to write "so it takes the same walk"; the counters said otherwise:
| command | dirs / paths walked | time |
|---|---|---|
git clean -nd |
82 / 6,250 | 15.8 ms |
git clean -ndX |
82 / 6,250 | 190.5 ms |
git clean -ndx |
3,197 / 36,811 | 275.1 ms |
Look at the first two rows: -ndX walks exactly as many paths as -nd. So -X does not take the ignored-tree walk that git status --ignored takes; the read_directory region finishes in six milliseconds. Even so, the process runs twelve times as long. The cost is not in the directory walk.
builtin/clean.c says where it is. remove_dirs is called for every directory to be removed, and that function descends the directory recursively with opendir/readdir/lstat — even on a dry run, because dry_run only clips the final step:
res = dry_run ? 0 : rmdir(path->buf);
So git clean -ndX walks into node_modules knowing full well it will not delete it; it has no other way to learn whether the directory can be emptied and whether a separate repository sits inside. -x pays both bills at once: with the ignore rules switched off nothing counts as "ignored", read_directory walks the whole tree (3,197 directories / 36,811 paths), and then remove_dirs descends once more.
This section says the opposite of the rest of the piece, and that is exactly why it is here: do not assume two commands do the same work because they finish in similar time. The clock labels both of them "expensive"; the counter tells you at a glance that -ndX never walked the ignored tree.
A separate repository inside an ignored directory
One more trap, and it concerns everyone who keeps a vendor directory: what happens when a separate git repository lives inside an ignored directory? The experiment — vendor/ in .gitignore, with a repository and its own history inside:
mkdir -p vendor && git init -q vendor/library
echo data > vendor/library/a.txt
git -C vendor/library add -A && git -C vendor/library commit -qm first
The modes stop in different places:
$ git status --porcelain --ignored | grep '^!! vendor'
!! vendor/
$ git status --porcelain -uall --ignored | grep '^!! vendor'
!! vendor/library/
The -uall mode was listing everything one by one; here it does not. It stops at the nested repository's boundary and reports the directory rather than its files. That behaviour was not free: the 2.16.0 release notes record that git status --ignored -u used not to stop at the working tree of a separate project embedded in an ignored directory and listed that project's files, and that it was corrected. What 2.48.1 shows is today's shape of that correction.
On the cleanup side the same boundary protects you:
$ git clean -ndX vendor
Would skip repository vendor/library
What prints this line is the previous section's remove_dirs: while descending it checks for a separate repository with is_nonbare_repository_dir. git clean will not remove a directory containing a separate repository on its own — it tells you it is skipping it. If you have written a script that cleans ignored directories, this matters in both directions: your vendor repository does not vanish by accident, but your assumption that it "was cleaned" is also wrong. Writing -ff lifts the boundary, which is exactly why you should look at what is there with -n before you write -ff.
The better question: which rule caught it
Most of the time my real reason for typing --ignored is not to see the list but to answer one question: why is this file not being committed? That question needs no tree walk; git has a direct answer:
$ git check-ignore -v node_modules/.package-lock.json
.gitignore:20:node_modules/ node_modules/.package-lock.json
10.3 milliseconds. And it tells you the thing --ignored never says: which file, on which line, holds the rule. If you are hunting through five different .gitignore files and an info/exclude to find out which one wins, this is the command you want.
A small warning: those 10.3 milliseconds are per process. If a script is going to ask about five hundred paths, do not call it five hundred times; there is --stdin. Timed: 500 paths in a single call took 35.0 milliseconds, 0.07 milliseconds per path.
If you need the list itself, git status is not the only way:
$ git ls-files -o -i --exclude-standard --directory
data/
dist/
node_modules/
82 directories / 6,250 paths, 15.5 milliseconds — a fifth of --ignored, with identical output in this tree. The reason is the bit this piece has been talking about from the start: ls-files does not set DIR_HIDE_EMPTY_DIRECTORIES, so treat_directory returns from a pattern-matched directory without going inside.
But this is a third semantics. Repeat the implicitly-ignored directory experiment with it:
$ git ls-files -o -i --exclude-standard --directory | grep ignored-logs
ignored-logs/
ignored-logs/run.log
traditional prints only the directory, matching only the file, ls-files both. Three commands, one tree, three different answers. There is no single right "list the ignored files"; you have to pick which definition you want.
The bad news: you cannot change the default mode through configuration. Among the keys beginning with status. in git help --config there is none that selects the ignored mode — there is showUntrackedFiles, there is no showIgnoredFiles. The expensive mode is both the default and unconfigurable. Your only lever is an alias:
git config --global alias.ign 'status --porcelain --ignored=matching'
Checklist
The command follows from the question you are asking:
-
"Why is this file ignored?" →
git check-ignore -v <path>. Cheapest, and the only one that names the rule. -
"Which pattern catches what? I need a list." →
--ignored=matching, knowing that the meaning of the output differs: implicitly ignored directories come out as their contents, not as a directory. -
If a directory-level list is enough →
git ls-files -o -i --exclude-standard --directory: the same three lines at a fifth of the cost (remember the third semantics). -
If you are asking about many paths →
git check-ignore -v --stdin; pay the per-process cost once. -
"I want the full inventory of ignored files." →
-uall --ignored. Expensive, but you get what you pay for. -
"I am about to clean up." → accept the price of
git clean -ndX; this walk has no shortcut. -
If a separate repository sits in an ignored directory,
git cleanskips it; verify your "it was cleaned" assumption with-n. -
In scripts and automation → do not put
--ignoredinside a loop; call it once and keep the output. And count on that call cancelling thecore.untrackedCachegain. - You cannot change the default; write an alias.
What three lines cost
What this measurement taught me is not specific to git. The size of a tool's output says nothing at all about the size of the work behind it. Three lines can be the residue of a thirty-seven-thousand-entry walk; the line !! node_modules/ can be the summary of 27 thousand entries that were collected and then released.
This is why I like turning on a tool's own counters. GIT_TRACE2_PERF is a one-line flag and it converts "that felt four times slower" into "3,197 directories opened, 36,811 paths walked". A feeling is arguable; a counter is not. What closes the remaining gap is that the sentence can be traced in the source down to a single assignment: the instant check_only becomes zero, the shortcut closes and the rest is arithmetic.
So for the next "why is this command slow": turn on the tool's counter, count the work, then read the source. Measuring before guessing cost, in this case, exactly as much as running four flags fifteen times each.
Official Sources
The code quotes and line references in the body use links pinned to the v2.48.1 tag; the list below points at the current main branch of the same files.
- git-status(1) —
--ignoredmodes and--untracked-files - git-clean(1) — behaviour of
-xand-X - git-config(1) —
core.untrackedCache - git/git · dir.c —
treat_directory,check_onlyandstop_early - git/git · dir.h — definition of
DIR_SHOW_IGNORED_TOO_MODE_MATCHING - git/git · wt-status.c — the branch where the untracked cache is wired
Top comments (0)