Short answer: re-encode the upload before it becomes public, then parse the resulting bytes and fail the publish job if GPS EXIF is still present. In a game community, that extra check matters more than where OCR runs. A screenshot can contain a player's home location, a phone model, and the original capture time even when the visible pixels look harmless.
The useful boundary is simple: the private upload is evidence; the public derivative is a new asset. Do not mutate the evidence and hope a later CDN transform removes metadata. Keep the original access-controlled, create a derivative for OCR and display, and verify that derivative before writing its public URL. The 8 MB request limit is a policy decision, too: set it from your largest supported screenshot, then test the rejection path with an oversized body so Express cannot spend worker time decoding it.
Ship the derivative.
How should Node.js Express strip EXIF location data before publishing an image?
I use an allow-list pipeline. Express limits bytes and MIME types, a worker decodes the image, and the encoder writes a fresh JPEG or WebP without copying metadata. The worker then reads the output file as if it were an untrusted download. That last step catches configuration drift.
Here is the smallest version of that boundary. sharp is only the image codec; the policy lives in the surrounding code.
import express from "express";
import multer from "multer";
import sharp from "sharp";
import { parse } from "exifr";
const app = express();
const upload = multer({ limits: { fileSize: 8 * 1024 * 1024 } });
app.post("/publish", upload.single("image"), async (req, res) => {
if (!req.file) return res.status(400).json({ error: "image required" });
try {
const publicBytes = await sharp(req.file.buffer)
.rotate()
.jpeg({ quality: 88, mozjpeg: true })
.toBuffer();
const exif = await parse(publicBytes, { gps: true, tiff: true, exif: true });
const forbidden = ["latitude", "longitude", "GPSLatitude", "GPSLongitude"];
const leaked = forbidden.some((key) => exif && key in exif);
if (leaked) return res.status(422).json({ error: "metadata policy failed" });
// Store publicBytes, enqueue OCR against these bytes, then publish its key.
return res.status(201).json({ bytes: publicBytes.length, verified: true });
} catch {
return res.status(415).json({ error: "unsupported image" });
}
});
The rotate() call is deliberate. Orientation is often stored as metadata; applying it to pixels before encoding prevents a privacy fix from turning a portrait screenshot sideways. The output format is also explicit. Letting an input extension choose the encoder is a small way to create surprising files.
What does verification after re-encoding actually prove?
It proves a property of the bytes you are about to publish, not a property of the upload. That distinction is the whole point. A unit test that inspects the incoming multipart buffer can pass while a later transform quietly copies tags back in.
I test three classes of fixtures: a phone photo with GPS, a game screenshot with an orientation tag, and an image with no EXIF at all. For each fixture, the test re-encodes, parses the output, and asserts that location fields are absent. I also assert that width, height, and pixel decode still work. Privacy that destroys the image is not a successful publish.
import assert from "node:assert/strict";
import sharp from "sharp";
import { parse } from "exifr";
export async function assertPublicImage(bytes: Buffer): Promise<void> {
const metadata = await parse(bytes, { gps: true, tiff: true, exif: true });
assert.equal(metadata?.latitude, undefined);
assert.equal(metadata?.longitude, undefined);
const info = await sharp(bytes).metadata();
assert.ok(info.width && info.height);
}
A passing parser check does not prove that a CDN, chat preview, or OCR vendor will preserve the same bytes. Hash the verified derivative, log the hash with the asset ID, and make downstream jobs consume that ID. If a later stage creates another derivative, it gets the same test.
Upload-time or on-demand OCR for game screenshots?
For moderation, search, and accessibility labels, I run OCR after the verified derivative is created. That keeps the public path deterministic: one privacy gate, one image identity, then asynchronous recognition. On-demand OCR is a better fit for an occasional “read this screenshot” button where most uploads are never analyzed.
| Choice | Strength | Cost or risk |
|---|---|---|
| At upload | Text is ready for indexing and moderation; the privacy check is naturally in front | Every upload pays CPU and queue time, including throwaway images |
| On demand | Lower background work and simpler retention for unused images | A request can wait on OCR, and an unreviewed image may be visible first |
The catch is that on-demand processing is not suitable when publication itself depends on recognized text or policy checks. Stick with upload-time OCR for those gates. Use on-demand when the feature is exploratory and latency is visible in the UI.
What I would change at scale
I would move the transform out of the Express process, add a queue with an idempotency key, and record inputHash, publicHash, codec, and policy version. A retry must not create a second public asset. Metrics should separate decode rejects, metadata-policy rejects, OCR latency, and publish latency; one blended “image failure” counter is nearly useless during an incident. A useful load test replays a mixed queue rather than a neat batch: tiny UI captures, 8 MB boundary cases, malformed headers, and camera photos with several metadata blocks, while a second worker retries the same idempotency key. That is where accidental duplicate publication and unbounded decoder memory usually surface.
There is a hard capability boundary here: a raster re-encode cannot make a malicious file safe by itself. Keep size limits, content-type sniffing, decoder isolation, and access control around it. Your mileage may vary with animated formats and color profiles, so add those fixtures before enabling them. I am not sure a single default quality value is right for every game UI; benchmark text-heavy screenshots separately from camera photos.
The practical rule is boring and reliable: private original, public re-encode, parse-and-assert, then OCR. That ordering makes the privacy claim testable instead of aspirational.
Top comments (1)
The evidence vs derivative boundary is the right framing. Treating the private upload as evidence and the public file as a new asset makes the verification step obvious. The idempotency key point is worth emphasizing too. A retry that creates a second public asset is a privacy bug, not just a duplicate.