DEV Community

sanskar arora
sanskar arora

Posted on

๐Ÿ“ฆ One File, Twenty-Five Images, Part 1/2 - The Design

Hey all ๐Ÿ‘‹

Here is a problem that sounds boring until you try it: move twenty-five container images from one world to another, where the two worlds share no network. Source registry on one side, Artifactory on the other, and between them a link, a laptop, or in the worst case a person carrying a disk.

The obvious script exists, everybody writes it, and it was the first thing I wrote too:

for image in "${images[@]}"; do
  podman pull "$image"
  podman save --format oci-archive --output "images/${safe_name}.tar" "$image"
done
tar --use-compress-program="zstd -T0 -9" -cf images.tar.zst images
Enter fullscreen mode Exit fullscreen mode

Twenty-five images in, 3.43 GiB out. It works. Ship it, go home.

Except it bothered me, because I knew what was inside that file. Nine of those images start from the exact same debian:bookworm-slim base layer - same digest, 74 MiB each - and every one carries its own private copy. Four Alpine images share a second base, the three Jupyter notebooks a third. And five more each start from a slightly different revision of that same Debian base: five near-identical 74 MiB layers with five different digests, which content addressing cannot collapse at all. The tarball is a museum of duplication, and then we compress the duplicates separately, so the compressor cannot even notice they are duplicates.

So I did what the situation deserved: designed it properly, built a prototype, measured everything, and let the numbers overrule my instincts twice. Two answers came out, not one. 1.82 GiB (47% smaller) keeps every image digest unconditionally and imports with a plain push - that is the default I would ship. 1.48 GiB (57% smaller) goes further by compressing the whole bundle as one stream, at the cost of re-encoding every layer on import, about 690 CPU-seconds for this set. Either way it is one file, and that file reconstructs all twenty-five working images after you delete every image on the machine.

This post is the whole path - every design decision, why it went that way, and what the measurements said when I finally stopped guessing. Code is at the end.


๐Ÿงจ Why the naive approach is expensive

Three separate costs stack up in that one-liner, and it is worth separating them because they have different fixes.

1. No deduplication. A container image is a manifest pointing at a config blob and a list of layer blobs, each addressed by the SHA-256 of its bytes. podman save writes one self-contained archive per image, so a layer shared by nine images is written nine times. Content addressing already gives you deduplication for free - if you put all the images in one content-addressed store instead of twenty-five separate ones.

2. Layers are gzip, and gzip is old. Registries store layers gzipped, mostly for compatibility reasons dating to Docker's first manifest format. zstd at a comparable CPU budget is roughly a fifth smaller.

3. Compressing already-compressed bytes. Wrapping gzipped layers in zstd -9 buys almost nothing - it is re-compressing entropy. In my measurements the outer level changed a pristine bundle's size by less than 0.2% between zstd -3 and zstd -9. All that CPU for nothing.

There is also a fourth cost that only showed up when I measured: podman save re-compresses layers as it writes them. Podman stores layers unpacked and re-gzips on the way out, and its output is about 3% larger than the bytes the registry served. You pay a few percent for a round trip through local storage.


๐Ÿงญ HLD, decision 1: what is the artifact?

First real decision. The bundle format could be anything - a custom container format with an index, a squashfs, a big tarball with a manifest at the front.

I picked: a valid OCI image layout inside a tar, compressed with zstd.

bundle.tar.zst
โ””โ”€โ”€ (zstd)
    โ””โ”€โ”€ (tar)
        โ”œโ”€โ”€ oci-layout
        โ”œโ”€โ”€ bundle-meta.json      <- my metadata: mode, images, inventory
        โ”œโ”€โ”€ index.json            <- the OCI index: one entry per image
        โ””โ”€โ”€ blobs/sha256/
            โ”œโ”€โ”€ <manifest digests>
            โ”œโ”€โ”€ <config digests>
            โ””โ”€โ”€ <layer digests>
Enter fullscreen mode Exit fullscreen mode

Why this and not something bespoke:

  • Deduplication falls out of the format. blobs/sha256/<digest> is content-addressed by definition. Two images sharing a layer write the same path. I did not implement dedup; I picked a layout where dedup is the only possible behaviour.
  • It is not a dead end. If my tool is unavailable on the far side - and on an air-gapped host, that is a real scenario - skopeo copy oci:unpacked:<ref> docker://registry/... pushes images straight out of an unpacked pristine bundle. That is the case I actually tried, and it wants skopeo โ‰ฅ 1.22, because older versions refuse the Docker schema-2 manifests that three of these twenty-five images still carry. A repack bundle is spec-valid too, so the same command should work, but it would push the uncompressed layers as-is. The format degrades to "standard tools work", with an asterisk.
  • The ecosystem already agreed on it. Manifests, configs, indexes, media types, annotations: all specified, all understood by every registry client. Inventing a format means inventing all of that badly.

The interesting design work is not the container, it is what goes inside it.


โš–๏ธ HLD, decision 2: the fork that surprised me

Here is the thing I did not see coming when I sketched the first version. My plan was: store layers uncompressed inside the bundle, compress the whole bundle with zstd, recompress layers at import. A bigger compression window - zstd's window is how far back in the stream it may look for a repeat, 4 MiB at level 9 by default and up to 2 GiB with long-distance matching enabled - would finally let the compressor see that those nine Debian bases are the same bytes. Cross-image redundancy, everybody wins.

It does win. It also changes every image digest, and that took me a while to accept.

A manifest references each layer by the digest of its compressed bytes:

{ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
  "digest": "sha256:e2de96513ba9...", "size": 3849738 }
Enter fullscreen mode Exit fullscreen mode

Decompress a layer for transport and you cannot reproduce those exact bytes later. gzip output depends on the implementation, the level, the window, the version. So the importer must recompress, which yields different layer digests, which changes every manifest, and the manifest digest is the image digest. Everything pinned as image@sha256:... breaks. Every cosign signature, which signs the manifest digest, stops verifying.

What survives is more subtle and worth knowing: the image config lists rootfs.diff_ids, the digests of the uncompressed layers. Those are preserved exactly, and runtimes verify them on pull. So integrity is fully intact and tags work perfectly; it is identity that changes.

I considered reconstructing the original compressed bytes (the trick pristine-tar uses for Debian source packages) and rejected it: it means capturing enough of the gzip encoder's state to replay it bit-exactly across versions, which is fragile machinery in exchange for an edge case.

One term first, because everything after this leans on it. Ingest is a one-time step that runs before an image is signed or pinned: pull it, re-encode each layer with the canonical profile described in the next section, push it back. From then on the registry itself holds canonical zstd layers, and that is the version tags, pins and signatures refer to.

The resolution I landed on is the part of the design I would defend hardest: pair the transport mode with the trust anchor.

pristine repack
What is signed each image, individually the bundle, once
Layers travel exactly as stored uncompressed, deduplicated by diff_id
Compression happens once per layer, at ingest over the whole bundle, at export
Digests preserved byte for byte rewritten (but see the twist below)

If images are individually signed, their bytes may not be touched, so the bundle carries them verbatim and the outer zstd does almost nothing (the layers are already compressed). If only the bundle is signed, layers can travel raw and the outer compressor can do real work across image boundaries. One flag, two coherent worlds, and no mode where a signature quietly stops verifying.

This is the kind of decision I now look for in designs generally: not "which is better", but "what do these two options actually depend on, and can I make the dependency explicit in the interface?"


๐ŸงŠ HLD, decision 3: a frozen encoder, and why determinism is the whole game

If images are going to be re-encoded (once at ingest for pristine, at every import for repack), the encoder's output must be reproducible byte for byte, forever. If it is not:

  • the same image ingested twice gets two different digests;
  • the importer cannot tell "already present" from "different";
  • deduplication against the destination stops working;
  • and repack's digests drift every time a library is upgraded.

So cz1 (canonical zstd, profile 1) is a pinned set of settings, not "zstd -19":

enc, err := zstd.NewWriter(out,
    zstd.WithEncoderLevel(zstd.SpeedBestCompression),
    zstd.WithWindowSize(8<<20),
    zstd.WithEncoderConcurrency(1),   // concurrency changes block boundaries
    zstd.WithEncoderCRC(true),
    zstd.WithAllLitEntropyCompression(true),
    zstd.WithZeroFrames(true),
    zstd.WithLowerEncoderMem(false),
)
Enter fullscreen mode Exit fullscreen mode

The library version is pinned in go.mod. Upgrading it is not a dependency bump, it is a profile change (cz2) and a deliberate, scheduled re-canonicalisation.

Then a hole is uncovered, and it is a good one. With identical options, klauspost/compress writes different bytes depending on how you drive it:

  • NewWriter + streaming, for input larger than one block โ†’ frame with no Frame_Content_Size;
  • ResetContentSize โ†’ content size in the header;
  • EncodeAll โ†’ content size, plus the single-segment flag when the input fits the window;
  • and NewWriter + streaming for input smaller than one block โ†’ the library takes the EncodeAll path internally and emits that header shape anyway.

Three different digests for the same layer and the same settings. And a Go implementation written from my architecture document would plausibly reach for EncodeAll, because tar headers hand you the size up front. My own doc would have produced a tool that disagreed with my prototype.

Fix, in two parts. The frame shape became part of the profile definition: one frame, driven through ReadFrom, with whatever header the library writes for that input size. Naming an API is not enough, because the library crosses between them on its own - so the profile also got golden vectors, checked on every run and in CI:

# sha256 of 'seq 1 N | cz1' - frozen outputs of the cz1 profile
8adb12de4536b3c50b4fed4c43fa4411088d61374243b51ae0b8237807d33d92 10000
719f1575295ca6deeb551131294ebc3c20d7c186ea76ad94e9e82726ffdcc878 500000
4ed7f465a721e86af9d211023cc8e473cf1f8787550a5f58d9db5a31504eaaae 3000000
Enter fullscreen mode Exit fullscreen mode

Three sizes on purpose. N is the argument to seq, so the inputs are 48 KB, 3.4 MB and 22.9 MB: one that fits in a single 128 KiB block, one below the 8 MiB window, one spanning several windows. They exercise different frame-header paths, and you can see it in the first byte after the magic: the smallest vector gets 0x64 (content size present, single segment), the larger two 0x04 (neither). That is the library switching paths on input size - precisely the behaviour a hand-written implementation would get wrong.

If a library upgrade changes a single byte of output, CI fails with "cz1 drift" instead of silently giving every future image a new identity.

LLD lesson worth generalising: when a hash of your output is part of your system's identity, your dependency's default behaviour is not a detail you can leave to the library. Pin it, and prove the pin with a test that would fail if it moved.


๐Ÿงฑ LLD: the bundle, in detail

Entry order is a design decision

The tar is written in a fixed order:

oci-layout โ†’ bundle-meta.json โ†’ index.json โ†’ manifests โ†’ configs โ†’ layers
Enter fullscreen mode Exit fullscreen mode

Because a tar is a stream. An importer reading it sequentially meets the metadata first: it can plan, check the format version, verify sizes, and decide what to skip before gigabytes of layers arrive. Put the index last and streaming import becomes impossible - you would have to buffer everything to disk first. Ordering is free at write time and decides what the reader is capable of.

The tar must be deterministic

Same images in, byte-identical file out. That makes bundles cacheable, resumable, comparable, and auditable:

tar --create --file=- --format=posix \
    --pax-option='exthdr.name=%d/PaxHeaders/%f,delete=atime,delete=ctime' \
    --mtime='@0' --owner=0 --group=0 --numeric-owner --mode='a=r,u+w,a+X' \
    --no-recursion --verbatim-files-from -C "$layout" --files-from="$list"
Enter fullscreen mode Exit fullscreen mode

Every flag there is load-bearing, and one guards against a trap: the pax extended-header name. POSIX specifies %d/PaxHeaders.%p/%f, and %p is the process ID - so on GNU tar 1.32 and older, or on any version with POSIXLY_CORRECT set, a "deterministic" archive quietly embeds the PID of the process that made it. Two runs, two different files, identical content. GNU tar 1.35 (what I measured with) already defaults to the PID-free form, which means this is exactly the kind of bug that works on your laptop and breaks on a build box.

(Bonus finding: zstd's compressed output is byte-identical at 1, 2, 4 and 12 threads. So the compressed bundle is reproducible too, not just the tar.)

bundle-meta.json carries the inventory

{ "formatVersion": 1, "mode": "repack", "profile": "cz1",
  "platform": "linux/amd64",
  "outer": { "codec": "zstd", "level": 9, "windowLog": 30 },
  "images": [ { "ref": "...", "digest": "sha256:...", "sourceDigest": "sha256:..." } ],
  "inventory": [ { "digest": "sha256:...", "kind": "layer", "size": 41234567, "included": true } ],
  "exceptions": [ { "digest": "...", "mediaType": "...", "reason": "unrecognized layer media type" } ] }
Enter fullscreen mode Exit fullscreen mode

Two pieces of foresight in there:

  • inventory with included: true - today everything is included. But it is the hook that makes delta bundles possible later: ship only the blobs the destination lacks, and mark the rest included: false for the importer to find locally. Designing the field in now costs nothing; adding it later is a format break. (The full delta design is a separate document in the repo, deliberately not built yet.)
  • exceptions - the safety rule. Only recognised image layers are ever transformed. Helm charts, ORAS artifacts, foreign/nondistributable layers, anything with a media type I don't know: carried verbatim and listed here. A converter that touches things it doesn't understand is a data-corruption bug waiting for a slow afternoon.

Verify while you are already touching the bytes

Every layer is decompressed exactly once during a repack export, and its diff_id is checked in the same pass, not in a second read:

mkfifo "$tmp.fifo"
sha256sum <"$tmp.fifo" >"$tmp.sha" &
shapid=$!
gzip -dc <"$src" | tee "$tmp.fifo" | "${encode[@]}" >"$tmp"
wait "$shapid"
[[ "sha256:$(cut -d' ' -f1 "$tmp.sha")" == "$diff_id" ]] || die "layer does not match its diff_id"
Enter fullscreen mode Exit fullscreen mode

Verification that costs a second pass over 7 GB gets switched off "temporarily" by someone in a hurry. Verification that rides along a pass you are already making is free, so it stays on. Design your integrity checks to be cheap enough that nobody has an argument for disabling them.


๐Ÿ” LLD: the import side, and a twist I did not plan

Import is the mirror, with one ordering rule that comes from the registry protocol: blobs before manifests. A registry rejects a manifest whose blobs are not uploaded yet, which conveniently means a half-finished import is invisible - unreferenced blobs, no tags, and re-running is idempotent.

For repack, the importer re-encodes every layer with cz1 and rewrites the manifests. And here is where the measurement produced something better than the design:

Because cz1 is deterministic, a repack import lands on exactly the digests that the ingest step would have produced. Both modes converge on one canonical digest per image. I tested this at 25 images: 25/25 identical.

Which leads to the twist. If the images were already ingested - canonical cz1 layers - then the importer's re-encode reproduces the original layer blobs bit for bit. So the importer can put the original manifest back, byte for byte, and the image keeps its digest through repack transport. The prototype does this: it carries the pre-repack manifest in the bundle, and restores it when the re-encode reproduces every blob it references. 25/25 restored.

The consequence is a change to the architecture I wrote: repack no longer has to break signatures. For already-ingested images it preserves digests, so a per-image signature made at ingest would still verify - if the bundle also carried the signature artifact, which my repack design deliberately doesn't, and which the prototype never tested, and as long as the importer runs the same cz1 profile the ingest did. My trust table was more pessimistic than reality; it is still the honest version that says "digests preserved", not "signatures shipped".

But a second reviewer caught me overstating it, and the correction is the honest version: digest preservation here depends on the manifest rewrite rules as much as on the encoder. Annotation set, key order, JSON serialisation - jq preserves insertion order; Go's encoding/json sorts map keys and HTML-escapes by default. Two implementations following my architecture doc would produce different manifest bytes and therefore different digests. So the rewrite rules have to be frozen alongside the encoder, with golden manifests in the test suite. The digest is a hash of the whole manifest; every byte of it is part of your public interface.


๐Ÿ”œ Next

That is the design: two transport modes tied to their trust anchors, a frozen encoder, a bundle whose format does the deduplication for you, and an importer that turned out to preserve more than I designed it to.

Designs are cheap. Part 2 is where I stop reasoning and start measuring - the bash prototype, the several ways my own measurements lied to me, four reviewers trying to break it, and the numbers. Including the one that overruled a decision I had just finished making.

Top comments (0)