DEV Community

Remdore
Remdore

Posted on AI-assisted

tar checksums its headers and never your files

tar is older than most of the people who type it, and the format underneath has barely changed since the tape drives it was named after. I had used it for years without ever looking inside, so I wrote one by hand from the POSIX specification using nothing but Python's struct module, and checked every claim below against GNU tar 1.35 and Python's own tarfile.

The format turns out to be simple enough to write in forty lines and odd enough that most of its consequences are not what I would have guessed. The oddest one is that tar carefully checksums every header and does nothing at all to protect your data.

Forty lines and a tape drive

A tar archive is a sequence of 512-byte blocks. Each file gets one block of header, followed by its contents padded out to the next multiple of 512, and the archive ends with two blocks of zeros. That is the whole structure. There is no magic number at the start of the file, no table of contents, and nothing at the end except those zeros.

The header has fixed offsets, and almost every number in it is written as octal digits in ASCII text rather than as a binary integer:

The 512-byte tar header, and the cost of finding one file

put(124, 12, b"%011o\0" % size)      # file size: eleven octal digits, as text
put(136, 12, b"%011o\0" % mtime)     # modification time: same again
put(148, 8,  b" " * 8)               # checksum field holds spaces while summing
chksum = sum(h) & 0o777777           # add up all 512 bytes
put(148, 8,  b"%06o\0 " % chksum)    # then write the total back, in octal
Enter fullscreen mode Exit fullscreen mode

Two files built that way, plus the two zero blocks, came to 3,072 bytes: six blocks of 512. GNU tar listed it and extracted both files correctly, from an archive that no copy of tar had ever touched:

-rw-r--r-- 0/0              11 1970-01-01 02:00 hello.txt
-rw-r--r-- 0/0              12 1970-01-01 02:00 world.txt
Enter fullscreen mode Exit fullscreen mode

The checksum guards the header and nothing else

That checksum is a plain sum of the 512 header bytes, not a CRC and certainly not a hash. More importantly, it only covers the header.

I changed one byte inside the contents of hello.txt, turning first file into Xirst file, and extracted the archive:

exit code: 0   extracted hello.txt -> 'Xirst file'
Enter fullscreen mode Exit fullscreen mode

No warning, no error, exit status zero, and a corrupted file on disk. Then I changed a single byte of the header instead, the first letter of the file name:

tar: This does not look like a tar archive
tar: Skipping to next header
tar: Exiting with failure status due to previous errors
Enter fullscreen mode Exit fullscreen mode

So tar refuses an archive whose file name has been damaged, and cheerfully writes out an archive whose file contents have been damaged. Once you know what the checksum is for, the asymmetry makes sense. It exists so a tape drive could tell a header block from a data block and resynchronise after a bad read, not so you could trust what came out.

The practical consequence is that a bare .tar gives you no integrity guarantee whatsoever. If you compress it, gzip and xz both carry their own checks over the whole stream, so a .tar.gz is protected by the compression rather than by tar, and that is a dependency worth knowing about if you ever store uncompressed archives and assume otherwise.

Eleven octal digits

Writing numbers as text has a hard ceiling built into it. The size field holds eleven octal digits, and the largest eleven-digit octal number is 8,589,934,591, which is one byte short of 8 GiB.

A sparse 9 GiB file costs nothing on disk, so I asked GNU tar to archive one in each of its three formats.

The strict POSIX ustar format refuses outright, and the error message states the exact limit:

tar: value 9663676416 out of off_t range 0..8589934591
Enter fullscreen mode Exit fullscreen mode

The GNU format sets the top bit of the size field to signal that the remaining bytes are a binary integer instead of octal text:

size field = b'\x80\x00\x00\x00\x00\x00\x00\x02@\x00\x00\x00'
decoded as big-endian binary = 9663676416
Enter fullscreen mode Exit fullscreen mode

The POSIX pax format does something more interesting. It writes an extra header in front of the file containing plain key=value records, puts the real size there as decimal text with no length limit, and leaves zero in the ordinary size field:

pax records: 19 size=9663676416 | 30 mtime=1791179820.339886948 | ...
Enter fullscreen mode Exit fullscreen mode

Three formats and three different answers to the same 1980s decision, all still in use. When an old tool chokes on a large archive, this is usually why.

There is no index

Because there is no table of contents, the only way to find a file in a tar archive is to start at the beginning and walk forward one header at a time. Each header's size field tells you how far to jump to reach the next one, and that jump is the only navigation the format has.

To see what that costs, I built an archive of 2,000 files of 50 KB each, about 100 MB, and the same 2,000 files as an ordinary uncompressed ZIP. Then I measured how much of each file a reader had to touch to extract only the last member:

extracting the last of 2,000 files read seeks time
tar, uncompressed 3.0% 2,001 34 ms
zip 0.2% 7 3 ms

Python's tarfile did not read all 100 MB, because it used each size field to seek straight over the file data. It did have to visit every one of the 2,000 headers, with a seek per header, and it was about ten times slower than the ZIP reader, which keeps a central directory at the end of the file and jumped straight to the entry it wanted.

Compression takes away even that shortcut, because a gzip stream cannot be seeked. To reach the last file in a .tar.gz you have to decompress everything in front of it:

.tar.gz, streaming read read time
first file 0.1% 0.3 ms
last file 100% 180 ms

Which file you ask for decides the cost. Asking for the first is nearly free, and asking for the last costs the whole archive.

Appending, and why GNU tar reads to the end

The lack of an index has an upside that I had never thought about. Appending a file to a tar archive does not require rewriting anything.

tar -r added a file to the 103 MB archive in no measurable time, and the first 100 MB were byte-for-byte identical afterwards. The archive did not even grow. Tar pads archives out to whole 10,240-byte records, and this one ended in twenty empty blocks of padding, so the new member was simply written over the zeros where the end marker used to be.

The same mechanism lets you append an updated version of a file that is already in the archive:

echo "version 1" > config.txt && tar -cf app.tar config.txt
echo "version 2" > config.txt && tar -rf app.tar config.txt
Enter fullscreen mode Exit fullscreen mode

The archive now holds two entries called config.txt, and extracting it gives you version 2, because later entries overwrite earlier ones as tar writes them to disk. Adding --occurrence=1 gives you version 1 instead.

That explains something that had always puzzled me slightly about GNU tar. Asked for a single file from the .tar.gz above, it took 0.22 seconds whether I asked for the first file or the last. It cannot stop when it finds a match, because a later entry with the same name would be a newer version. With --occurrence=1 the first file came back in 0.00 seconds.

Same files, different archive

Tarring the same three files twice, two seconds apart, produced two archives with different hashes:

one.tar two.tar differ: byte 659, line 1
Enter fullscreen mode Exit fullscreen mode

Byte 659 falls inside the modification time field of the second entry's header. The order of the entries was not the order I had created the files in, and not alphabetical either. It was b, a, c, which is whatever order the filesystem happened to return when tar listed the directory. Owner names, group names and access times can all leak in the same way.

For a build system, or anything that compares archives by hash, that is a real problem, and it has a known fix:

tar --sort=name --mtime='2024-01-01 00:00Z' --owner=0 --group=0 --numeric-owner \
    --pax-option=exthdr.name=%d/PaxHeaders/%f,delete=atime,delete=ctime \
    --format=posix -cf repro.tar -C src .
Enter fullscreen mode Exit fullscreen mode

Run twice with the files touched in between, that produced identical archives byte for byte.

What I got wrong

Three things, and the third would have put a false number into the figure.

My first test of the 8 GiB limit told me that ustar refused the file, which was what I expected, so I very nearly wrote it down. The error was actually GNU features wanted on incompatible archive format, because I had passed --sparse to save disk space and sparse files are themselves a GNU extension. Tar had rejected the flag before it ever looked at the file's size. Without --sparse, and with the output piped into head so tar would stop after writing the header rather than writing 9 GB, I got the real limit and the real error text.

The second was publishing scripts that did not run. The read-cost measurements depended on test archives that I had generated with a throwaway snippet and never saved, so a fresh clone of the repository could reproduce every claim except the ones in the table. I only found out because I cloned it and ran it, and I have now made the same mistake in two different repositories this month.

The third is the one I most want to flag. My first measurement of Python's convenient getmember() API, on a .tar.gz, said that extracting the last file read 200% of the archive, meaning it decompressed the whole thing twice. That went into the first version of the figure. When I reran it from the clean clone, it read 100%.

The difference turned out to be which program had compressed the archive. In both cases the uncompressed bytes were identical, getmember() scanned the entire archive before extracting anything, and it then needed to seek about 50 KB backwards to reach the last file. With an archive written by the gzip command-line tool, that backward seek restarted decompression from the beginning. With one written by Python's own gzip module, it did not. I have not established why, so the figure shows both, and the useful conclusion does not depend on the answer: if you need one file from a compressed tar, use streaming mode and stop as soon as you find it.

What to take from it

Tar was designed for tape, where the only operation available is reading forward, and almost all of its behaviour follows from that. The headers are checksummed so a reader can find its place again after a bad block. Numbers are text because that was portable between machines that disagreed about binary integers. There is no index because a tape cannot jump to the end to read one, and appending is cheap for the same reason.

Most of the practical advice falls out of it. Do not rely on an uncompressed tar to tell you that your data is intact, because it will not. Expect old tools to fail on files over 8 GiB unless the archive uses the GNU or pax extensions. Use a ZIP, or a tar with an external index, if you need to pull individual files out of a large archive often. And if anything compares your archives by hash, pass the reproducibility flags, because two archives of the same files will otherwise differ even when nothing in them has changed.

The scripts, the hand-built archive and the figure are in a small repository, and reproduce.sh reruns every claim above in order.

Top comments (0)