DEV Community

Cover image for Stripping EXIF at the byte level: no re-encoding, no quality loss
Danilo Fortunato
Danilo Fortunato

Posted on Originally published at danyblitz-bit.github.io

Stripping EXIF at the byte level: no re-encoding, no quality loss

The tool everyone's using is quietly damaging your photos

If you strip metadata from an image, the usual approach is to decode it, delete the fields, and encode it again. That is the natural way to do it, because every image library exposes exactly that: a decode step, a field dictionary, an encode step.

The problem is that JPEG is a lossy format, and the encoder you use is not the encoder that made the file. So a round trip through Pillow changes the bytes of the image data itself, not just the metadata:

from PIL import Image

im = Image.open("photo.jpg")
clean = Image.new(im.mode, im.size)     # a fresh, empty image
clean.putdata(list(im.getdata()))        # copy every pixel
clean.save("clean.jpg", quality=95)      # re-encode at a quality you invented
Enter fullscreen mode Exit fullscreen mode

You now have a file that is not your original. It is a lossy re-compression of your original, and the damage is not recoverable. The quality=95 is not a neutral default — it is a lossy re-compression of your original. If your original was quality=88, you just paid twice.

Stripping metadata should not change the picture. It is a container edit, not an image edit.

JPEG is a sequence of segments, and only some of them are the image

A JPEG file is a stream of markers. Each is FF xx, followed by a 2-byte big-endian length, followed by payload. The pixel data lives in exactly one of them: SOS (start of scan). Everything before SOS is metadata containers, and everything from SOS to FF D9 (EOI) is the compressed scan data.

That means the whole operation is: copy the bytes, skip the segments you do not want.

import struct

# APPn = the metadata containers; APP0 is JFIF, APP1 is where Exif/XMP live
DROP = {0xE1} | set(range(0xE2, 0xF0))


def strip_jpeg_metadata(data: bytes) -> bytes:
    out = bytearray(data[:2])          # SOI
    i = 2
    while i < len(data) - 1:
        if data[i] != 0xFF:
            # we are inside the entropy-coded scan: copy the rest verbatim
            out += data[i:]
            break
        marker = data[i + 1]
        if marker == 0xDA:             # SOS: the image itself starts here
            out += data[i:]
            break
        if marker == 0xD9:             # EOI
            out += data[i:]
            break
        seg_len = struct.unpack(">H", data[i + 2 : i + 4])[0]
        end = i + 2 + seg_len
        if marker not in DROP:
            out += data[i:end]
        i = end
    return bytes(out)
Enter fullscreen mode Exit fullscreen mode

Two properties fall out of this for free:

  1. The scan data is never touched, so there is no quality loss. Not "low loss" — none. The compressed bytes are the same bytes.
  2. It cannot leave a partial segment behind, because lengths are read from the file rather than guessed.

PNG needs a different, simpler rule

PNG is chunks, not segments: an 8-byte length, a 4-byte type, the payload, a 4-byte CRC. Metadata lives in the ancillary chunks tEXt, zTXt, iTXt and eXIf. The critical chunks (IHDR, PLTE, IDAT, IEND) and the ancillary ones that affect rendering (tRNS, gAMA, cHRM) must stay.

import struct

PNG_DROP = {b"tEXt", b"zTXt", b"iTXt", b"eXIf"}


def iter_png_chunks(data: bytes):
    pos = 8                                   # skip the PNG signature
    while pos < len(data):
        (length,) = struct.unpack(">I", data[pos : pos + 4])
        ctype = data[pos + 4 : pos + 8]
        total = 12 + length                   # len + type + data + crc
        yield ctype, data[pos : pos + total]
        pos += total


def strip_png_metadata(data: bytes) -> bytes:
    out = bytearray(data[:8])
    for ctype, chunk in iter_png_chunks(data):
        if ctype not in PNG_DROP:
            out += chunk
    return bytes(out)
Enter fullscreen mode Exit fullscreen mode

Because every chunk carries its own CRC, dropping whole chunks keeps the file valid — there is no checksum to recompute.

The part that matters: assert it, do not hope

A metadata stripper that silently fails is worse than none, because you believe you are covered. The cheapest possible insurance is a self-check that builds a file known to contain the data, strips it, and asserts nothing survives:

import io, sys
from PIL import Image, TiffImagePlugin

def make_jpeg_with_exif() -> bytes:
    # a real 64x64 JPEG carrying a genuine GPS block, written by PIL so the
    # fixture is a file any image library can actually open
    im = Image.new("RGB", (64, 64), (40, 90, 160))
    exif = im.getexif()
    gps = exif.get_ifd(0x8825)
    gps[1] = "N"
    gps[2] = (TiffImagePlugin.IFDRational(48, 1), TiffImagePlugin.IFDRational(51, 1), TiffImagePlugin.IFDRational(2952, 100))
    gps[3] = "E"
    gps[4] = (TiffImagePlugin.IFDRational(2, 1), TiffImagePlugin.IFDRational(17, 1), TiffImagePlugin.IFDRational(4024, 100))
    buf = io.BytesIO()
    im.save(buf, format="JPEG", quality=90, exif=exif)
    return buf.getvalue()

def jpeg_scan_data(buf: bytes) -> bytes:
    """The compressed image itself: from the start-of-scan marker to the end."""
    i = buf.find(b"\xff\xda")
    assert i != -1, "no SOS marker, the fixture is not a JPEG we understand"
    return buf[i:]

def main() -> int:
    raw = make_jpeg_with_exif()
    # the fixture must really contain the data, or the test proves nothing
    assert Image.open(io.BytesIO(raw)).getexif().get_ifd(0x8825), "fixture is wrong"
    clean = strip_jpeg_metadata(raw)
    # and the stripper must really remove it
    assert not Image.open(io.BytesIO(clean)).getexif().get_ifd(0x8825), "GPS survived"
    assert len(clean) < len(raw)
    # the image data must be byte-identical, not merely the same size
    assert jpeg_scan_data(clean) == jpeg_scan_data(raw), "pixels were re-encoded"
    print("self-check OK: EXIF/GPS stripped, pixels preserved")
    return 0

if __name__ == "__main__":
    sys.exit(main())
Enter fullscreen mode Exit fullscreen mode

Both halves have to be asserted. "The metadata is gone" is the easy claim; "the image data is byte-identical" is the claim that actually matters, and it is the one that is normally never checked.

Where this lives

I packaged this as Photo MetaClean — standard library only, no dependencies, offline by design, MIT licensed:

The same byte-level treatment for documents is in PDF MetaClean, and the landing page is here.

What to check on your own photos

Run any of these on a file before you upload it anywhere:

exiftool photo.jpg
strings photo.jpg | head -40
Enter fullscreen mode Exit fullscreen mode

If you see a GPS block, a camera serial number, your name, or an original filename, it is in the file. Social platforms strip some of it; not all of it, not consistently, and not for every upload path. The part you control is the file you hand over.

One caveat worth being honest about

Byte-level copying means the tool does not rewrite anything it does not understand. That is a feature for your image data and a limitation for thumbnails, colour profiles or ICC data you may want to keep. Decide per file: a public upload and a print-ready master have different requirements, and the safe default is to strip for the upload and keep the master untouched.

Top comments (1)

Collapse
 
florian131313 profile image
G •

Dropping APP1 also drops the Orientation tag. Phone JPEGs are often stored unrotated, and that tag is what turns them. The scan bytes stay the same, and the preview can still come out sideways. COM (marker 0xFE) sits outside the DROP set, so a comment segment survives the pass.