The tool everyone's using is quietly damaging your photos
If you strip metadata from an image, the usual approach is to decode it, delete the fields, and encode it again. That is the natural way to do it, because every image library exposes exactly that: a decode step, a field dictionary, an encode step.
The problem is that JPEG is a lossy format, and the encoder you use is not the encoder that made the file. So a round trip through Pillow changes the bytes of the image data itself, not just the metadata:
from PIL import Image
im = Image.open("photo.jpg")
clean = Image.new(im.mode, im.size) # a fresh, empty image
clean.putdata(list(im.getdata())) # copy every pixel
clean.save("clean.jpg", quality=95) # re-encode at a quality you invented
You now have a file that is not your original. It is a lossy re-compression of your original, and the damage is not recoverable. The quality=95 is not a neutral default — it is a lossy re-compression of your original. If your original was quality=88, you just paid twice.
Stripping metadata should not change the picture. It is a container edit, not an image edit.
JPEG is a sequence of segments, and only some of them are the image
A JPEG file is a stream of markers. Each is FF xx, followed by a 2-byte big-endian length, followed by payload. The pixel data lives in exactly one of them: SOS (start of scan). Everything before SOS is metadata containers, and everything from SOS to FF D9 (EOI) is the compressed scan data.
That means the whole operation is: copy the bytes, skip the segments you do not want.
import struct
# APPn = the metadata containers; APP0 is JFIF, APP1 is where Exif/XMP live
DROP = {0xE1} | set(range(0xE2, 0xF0))
def strip_jpeg_metadata(data: bytes) -> bytes:
out = bytearray(data[:2]) # SOI
i = 2
while i < len(data) - 1:
if data[i] != 0xFF:
# we are inside the entropy-coded scan: copy the rest verbatim
out += data[i:]
break
marker = data[i + 1]
if marker == 0xDA: # SOS: the image itself starts here
out += data[i:]
break
if marker == 0xD9: # EOI
out += data[i:]
break
seg_len = struct.unpack(">H", data[i + 2 : i + 4])[0]
end = i + 2 + seg_len
if marker not in DROP:
out += data[i:end]
i = end
return bytes(out)
Two properties fall out of this for free:
- The scan data is never touched, so there is no quality loss. Not "low loss" — none. The compressed bytes are the same bytes.
- It cannot leave a partial segment behind, because lengths are read from the file rather than guessed.
PNG needs a different, simpler rule
PNG is chunks, not segments: an 8-byte length, a 4-byte type, the payload, a 4-byte CRC. Metadata lives in the ancillary chunks tEXt, zTXt, iTXt and eXIf. The critical chunks (IHDR, PLTE, IDAT, IEND) and the ancillary ones that affect rendering (tRNS, gAMA, cHRM) must stay.
import struct
PNG_DROP = {b"tEXt", b"zTXt", b"iTXt", b"eXIf"}
def iter_png_chunks(data: bytes):
pos = 8 # skip the PNG signature
while pos < len(data):
(length,) = struct.unpack(">I", data[pos : pos + 4])
ctype = data[pos + 4 : pos + 8]
total = 12 + length # len + type + data + crc
yield ctype, data[pos : pos + total]
pos += total
def strip_png_metadata(data: bytes) -> bytes:
out = bytearray(data[:8])
for ctype, chunk in iter_png_chunks(data):
if ctype not in PNG_DROP:
out += chunk
return bytes(out)
Because every chunk carries its own CRC, dropping whole chunks keeps the file valid — there is no checksum to recompute.
The part that matters: assert it, do not hope
A metadata stripper that silently fails is worse than none, because you believe you are covered. The cheapest possible insurance is a self-check that builds a file known to contain the data, strips it, and asserts nothing survives:
import io, sys
from PIL import Image, TiffImagePlugin
def make_jpeg_with_exif() -> bytes:
# a real 64x64 JPEG carrying a genuine GPS block, written by PIL so the
# fixture is a file any image library can actually open
im = Image.new("RGB", (64, 64), (40, 90, 160))
exif = im.getexif()
gps = exif.get_ifd(0x8825)
gps[1] = "N"
gps[2] = (TiffImagePlugin.IFDRational(48, 1), TiffImagePlugin.IFDRational(51, 1), TiffImagePlugin.IFDRational(2952, 100))
gps[3] = "E"
gps[4] = (TiffImagePlugin.IFDRational(2, 1), TiffImagePlugin.IFDRational(17, 1), TiffImagePlugin.IFDRational(4024, 100))
buf = io.BytesIO()
im.save(buf, format="JPEG", quality=90, exif=exif)
return buf.getvalue()
def jpeg_scan_data(buf: bytes) -> bytes:
"""The compressed image itself: from the start-of-scan marker to the end."""
i = buf.find(b"\xff\xda")
assert i != -1, "no SOS marker, the fixture is not a JPEG we understand"
return buf[i:]
def main() -> int:
raw = make_jpeg_with_exif()
# the fixture must really contain the data, or the test proves nothing
assert Image.open(io.BytesIO(raw)).getexif().get_ifd(0x8825), "fixture is wrong"
clean = strip_jpeg_metadata(raw)
# and the stripper must really remove it
assert not Image.open(io.BytesIO(clean)).getexif().get_ifd(0x8825), "GPS survived"
assert len(clean) < len(raw)
# the image data must be byte-identical, not merely the same size
assert jpeg_scan_data(clean) == jpeg_scan_data(raw), "pixels were re-encoded"
print("self-check OK: EXIF/GPS stripped, pixels preserved")
return 0
if __name__ == "__main__":
sys.exit(main())
Both halves have to be asserted. "The metadata is gone" is the easy claim; "the image data is byte-identical" is the claim that actually matters, and it is the one that is normally never checked.
Where this lives
I packaged this as Photo MetaClean — standard library only, no dependencies, offline by design, MIT licensed:
- Windows binary: danyblitz.gumroad.com/l/photo-metaclean (8 MB, single file, no installer)
- Source: github.com/danyblitz-bit/photo-metaclean
The same byte-level treatment for documents is in PDF MetaClean, and the landing page is here.
What to check on your own photos
Run any of these on a file before you upload it anywhere:
exiftool photo.jpg
strings photo.jpg | head -40
If you see a GPS block, a camera serial number, your name, or an original filename, it is in the file. Social platforms strip some of it; not all of it, not consistently, and not for every upload path. The part you control is the file you hand over.
One caveat worth being honest about
Byte-level copying means the tool does not rewrite anything it does not understand. That is a feature for your image data and a limitation for thumbnails, colour profiles or ICC data you may want to keep. Decide per file: a public upload and a print-ready master have different requirements, and the safe default is to strip for the upload and keep the master untouched.
Top comments (1)
Dropping APP1 also drops the Orientation tag. Phone JPEGs are often stored unrotated, and that tag is what turns them. The scan bytes stay the same, and the preview can still come out sideways. COM (marker 0xFE) sits outside the DROP set, so a comment segment survives the pass.