Hey all ๐
Part 1 worked out how to move twenty-five container images across an air gap as one file: pack them into a single OCI layout so shared layers are stored once, and pick between two modes - pristine, where layers travel exactly as the registry stores them and every digest survives, and repack, where layers travel uncompressed and the whole bundle is compressed as one zstd stream.
Nice theory. This half is the part where the theory meets a laptop: how I measured it, the several ways my measurements were wrong before they were right, what four reviewers found when they went looking for holes, and the numbers that came out - including the one that overruled a decision I had just finished defending.
๐ฌ The process: build it small, measure it honestly
I wrote the prototype in bash, on purpose. The final implementation is planned to be in Go (a static binary you can drop on the far side), but the question M0 had to answer was "how big is the file, and what does it cost", not "is my Go code nice". Bash plus skopeo, zstd, jq and GNU tar reaches that answer in a day or two, and every step stays visible.
One rule was meant to keep it honest: anything that decides a digest lives in Go, in the pinned cz1 helper, with bash owning only the outer compression (digest-irrelevant) and the glue. I half-kept it. Layer bytes are Go; the manifest rewrite is jq - and the manifest digest is the image digest, so jq is deciding identity too. That shortcut is exactly where a reviewer later found a hole, and it is why the manifest rewrite rules have to be frozen alongside the encoder. (I also learned the hard way that you cannot reproduce klauspost's output with the zstd CLI. Different implementations, different bytes, same settings.)
Measuring, and the several ways I got it wrong first
The first version of my measurement rig was wrong in ways that flattered the results:
-
CPU was the change in the shell's
cutime, which swept up the tar, tee and hashing processes alongside the compressor. -
Peak memory was sampled from
/proc/<pid>/statusevery 100 ms. On short runs that read 24-38% low, and sometimes sampled beforeexec. - The report said zstd used "12 threads".
zstd -T0uses one worker per physical core: 6 on this machine.
The fix was a 58-line Go helper that runs one command and reports what the kernel says via wait4: exact user+sys CPU and exact ru_maxrss. No sampling, no guessing.
Then, because a laptop is not a lab, every run also records how much CPU the rest of the machine used during it (system-wide busy time minus our own) and how many pages swapped. Rows where outside load exceeded 10% get flagged in the report, and cached measurements taken under load are automatically re-timed. An idle desktop here sits at 2-5%, so the threshold had to be calibrated rather than guessed - at my first attempt (3%) every single row got flagged, which is the same as flagging none.
There is also a guard that estimates zstd's peak RAM from level, window and worker count, and skips a variant the machine cannot hold. It puts a 2 GiB window at level 19 at about 4.8 GiB: roughly 3.8 GiB of estimated zstd footprint - it actually peaked at 3.90 - plus a deliberate 1 GiB margin for everything else on the machine. It still swapped 68 MiB on that run, which is why the timing carries a flag.
Four reviewers and a skeptic
Before trusting any of it, I had the prototype reviewed along four independent lenses (AI Agents ๐ซฃ) - OCI correctness, bash robustness, measurement validity, architecture fidelity - with a separate pass trying to refute each finding. 29 findings, 22 confirmed. Then the same treatment for the report itself: fact-checking every number against the raw data, technical claims against the sources, and a read-through as the person receiving it. That pass produced about 40 corrections, including a percentage stated backwards and a claim that three of four bundles were round-tripped when I had written "all".
The bash bugs are worth listing, because they are the kind that pass tests and lie to you:
-
$(...)strips trailing newlines. I hashed a manifest through a command substitution; two images in my test set have manifests ending in\n; those two "landed with the wrong digest". The registry was right, my check was wrong. -
[[ $a == $(cmd) ]]treats the right side as a glob pattern. JSON is full of[and], so a comparison that looked fine silently never matched, and the code path it guarded was never exercised. -
set -e+pipefail+[[ cond ]] && cmdas the last line of a function exits the whole script silently when the condition is false, because the function returns 1. Hunting that one is how I learned to add anERRtrap that prints the failing command and line. -
A backgrounded command gets
/dev/nullas stdin unless it has its own redirect, so my measured compressor was cheerfully compressing nothing. -
FIFO deadlock: a reader blocked on
open()waits forever if the writer never starts. Always check the producer's status before waiting on consumers.
๐ Results
The size and time tables below are from a clean-slate run: every generated artifact deleted, all 25 images re-pulled, measured on a 6-core/12-thread laptop. The window/level sweep and the layer-order result after them come from the earlier full measurement matrix - same machine, same images, a week earlier - because the clean-slate re-run only rebuilt the settings I would recommend. The dataset is 190 layer references collapsing to 123 unique layers, 7.00 GiB uncompressed.
Size
| Packaging | gzip -9 โ | zstd -6 | zstd -9 |
|---|---|---|---|
One podman save archive per image (today) |
3.43 GiB | 3.44 GiB | 3.43 GiB |
| Registry blobs per image, no dedup | 3.30 GiB | 3.30 GiB | 3.30 GiB |
| Deduplicated, registry gzip layers (pristine, before ingest) | 2.30 GiB | 2.30 GiB | 2.29 GiB |
Deduplicated, cz1 zstd layers (pristine, after ingest) โก |
1.82 GiB | ||
| Deduplicated, uncompressed layers, one zstd stream (repack) | 2.26 GiB | 1.93 GiB | 1.84 GiB |
| โฆ same, with a 1 GiB window | 1.48 GiB |
โ Every gzip -9 figure here was produced with pigz -9 - same format, deflate spread across all twelve threads, so gzip gets the same cores zstd does. Plain gzip -9 produces the same bytes and spends the whole CPU figure on the clock.
โก From the spike matrix rather than the clean-slate re-run: this is the mode where each layer is compressed individually with cz1, so there is no single-stream column for it.
(The zstd -9 row of the first line lands on exactly the size of my original images.tar.zst - 3,681,162,191 bytes - so the baseline is anchored to a re-measurement, not remembered from a week ago.)
The savings stack in four steps: deduplication ~30%, plus 3% for not letting podman re-compress (3.43 โ 2.30 GiB), then re-encoding each layer with cz1 instead of gzip to reach 47% (2.30 โ 1.82 GiB), then compressing the whole bundle as one stream with a big window to 57% (1.82 โ 1.48 GiB).
Time
| Packaging | Codec | Compress (wall / CPU) | Decompress |
|---|---|---|---|
| One archive per image (today) | gzip -9 (via pigz) |
11.7 s / 112.9 s | 5.8 s |
| One archive per image (today) | zstd -9 | 6.2 s / 42.1 s | 1.3 s |
| Deduplicated, registry gzip layers (pristine) | zstd -9 | 4.0 s / 25.6 s | 0.8 s |
| Deduplicated, uncompressed layers (repack) | gzip -9 (via pigz) |
84.0 s / 898.5 s | 15.4 s |
| Deduplicated, uncompressed layers (repack) | zstd -6 | 20.3 s / 124.7 s | 5.2 s |
| Deduplicated, uncompressed layers (repack) | zstd -9 | 34.4 s / 203.1 s | 5.1 s |
| Deduplicated, uncompressed layers (repack) | zstd -9, 1 GiB window | 32.9 s / 188.7 s | 4.8 s |
gzip loses on every axis that matters here: on uncompressed layers it burns 899 CPU-seconds to produce a larger file (2.26 GiB) than zstd -9 does in 203 (1.84 GiB), and it decodes 3ร slower - and that is with parallel deflate; single-threaded gzip -9 would take fifteen minutes of wall clock for the privilege.
There is no budget at which gzip is the right answer for the bundle itself. In registries it survives for compatibility, and that part is not nothing: both modes land tar+zstd layers, so every registry on the path has to accept OCI 1.1 media types and every puller has to be containerd โฅ 1.5 or Docker โฅ 23. That is a precondition to check before adopting any of this, not a footnote.
The finding that overruled my design
I had specified a 128 MiB long-distance window as the repack default. The data says that is the wrong knob setting, and it is not close:
| Setting | Bundle | Compress |
|---|---|---|
| level 6, 128 MiB window | 1.78 GiB | 24.7 s |
| level 9, 128 MiB window | 1.70 GiB | 36.8 s |
| level 9, 1 GiB window | 1.48 GiB | 35.3 s |
| level 9, 2 GiB window | 1.43 GiB | 141 s ยง |
| level 19, 128 MiB window | 1.46 GiB | 625 s ยง |
| level 19, 2 GiB window | 1.21 GiB | 869 s ยง |
Those three runs hit the 15 GiB laptop's memory ceiling rather than measuring zstd cleanly: they swapped 18, 538 and 68 MiB, and 19%, 5% and 10% of the machine's CPU went somewhere other than the compressor. The sizes are deterministic and unaffected; the wall times are inflated by an unknown amount. A repeat of the 2 GiB/level 9 run in digest order took 132 s, so the roughly 4ร gap to the 1 GiB window is real, not an artifact.
Moving through levels 6โ9 buys 4%. Moving the window from 128 MiB to 1 GiB buys 13%, at the same speed. Level 19 gets you to roughly the same place as the bigger window, 17ร slower.
The window is not free at the far end, though, and this is worth budgeting before adopting it: peak decode memory goes from 133 MiB to 1.01 GiB, the zstd CLI refuses a frame needing more than a 128 MiB window unless you pass --long=31, and klauspost's Go decoder rejects anything above 512 MiB until you call WithDecoderMaxWindow - which has to happen before decoding starts, so the importer cannot read the setting out of the bundle it has not opened yet.
The mechanism is visible in the data. Remember those six near-identical revisions of the Debian base - same 74.3 MiB of files, different digests, so deduplication cannot touch them. In first-use order they sit at offsets 142, 216, 403, 549, 5367 and 5563 MiB in a 7 GiB stream. A 128 MiB window reaching back from 403 MiB sees nothing but itself; a 1 GiB window covers the whole early cluster at once, and separately pairs the two stragglers at 5367 and 5563. That is where the 13% comes from: not from any single image compressing better, but from the compressor finally being able to look far enough back to notice it has seen this Debian base several times already.
(On a 5-image test set the whole tar is 181 MiB, everything related is already within reach, and larger windows add exactly nothing - which is how a small test set talks you out of the right answer.)
One more small thing with a mechanism behind it: writing layers in first-use order beat digest order by 2.8% at a 2 GiB window, and made no difference at 128 MiB. Related images' layers end up near each other only if you don't scatter them by hash.
๐งช The clean-slate test
Numbers in a table are a claim. This was the test:
- Delete every generated artifact and re-pull all 25 images.
- Export the 1.48 GiB bundle, verify it offline.
-
podman rmi --all- 28 images gone,images left: 0. That includes the skopeo andregistry:3images the import tooling itself runs, which the import then re-pulls. So this is not a test of the air gap; it is a test that nothing about the twenty-five payload images was left cached anywhere on the machine. - Import from that one file into a fresh registry.
- Pull all 25 back and compare every image ID against the original config digest.
- Run containers.
image-ID mismatches: 0
alpine:3.24 Alpine Linux v3.24 x86_64 ok
debian:bookworm-slim 12.15
python:3.13-slim-bookworm 3.13.15 2689367b205c
node:22-bookworm-slim v22.23.3 x64
redis:7-alpine Redis server v=7.4.11 ...
postgres:16-bookworm postgres (PostgreSQL) 16.15
busybox:1.36 busybox ok 406
25 of 25 image IDs matched, podman verified every layer digest on pull, and the containers ran. 3.43 GiB of per-image tarballs became a 1.48 GiB file that fully reconstructs the fleet.
One honest gap, since it is the premise of the whole post: the destination here was a local registry:3, not the Artifactory this is ultimately for. Whether Artifactory accepts OCI 1.1 zstd manifests, exposes the referrers API and allows cross-repo mounts is the first thing on the list, and until that run happens these numbers describe a design that has been proven against a reference registry, not against production.
๐งญ What I would tell the next person
- Find the decision that forks the design, and make the fork explicit. Here it was digest fidelity. Everything else - bundle format, compression, import logic - followed from it. If I had started with "which compression level?", I would have optimised the wrong axis for a week.
- Make the expensive property a property of the format, not of the code. Dedup came from content addressing. Streaming import came from entry ordering. Reproducibility came from tar flags. No algorithms required.
- If a hash of your output is an identity, freeze everything that can change a byte - library version, API call, serialisation, key order - and add a test that fails when it moves.
-
Measure before you tune. My design said a 128 MiB window; the data said 1 GiB, for 13% at the same speed. My instinct said single-stream compression was the big win; the data said deduplication was (3.43 โ 2.30 GiB), that re-encoding each layer once with
cz1was the next step (2.30 โ 1.82 GiB), and that compressing the bundle as one stream added 18.5% on top (1.82 โ 1.48 GiB) only with a window large enough to see across it. - Then measure your measurements. Sampled memory read a third low; CPU accounting swept in helper processes; a "quiet" laptop is 2-5% busy. If a number is going to justify a decision, know how it was produced.
- Adversarial review earns its keep. Reviewing the code, then the numbers, then the write-up, each with someone trying to refute the findings, caught a reversed percentage, an encoder that was not as frozen as I claimed, and a signature claim I could not support.
The code, the architecture document, the delta design and the full measurement report - including the bits that did not flatter me - are on GitHub: snskArora/oci_export. The bash prototype is complete end to end (fetch โ transform โ bundle โ verify โ import) and the Go implementation is next.
Until then: one file, twenty-five images, and a very quiet air gap.
Top comments (0)