Short answer: In a Node.js Express portfolio, save each original image behind a private storage boundary, run OCR against those untouched bytes, and publish a separately encoded, watermarked derivative. Pre-generate that derivative on upload unless every viewer needs a different mark.
| Design | Image quality | Bandwidth and compute | Better fit |
|---|---|---|---|
| Pre-generated derivative | Stable output that can be reviewed once | One processing pass; every request serves the same file | Public creator portfolios and predictable marks |
| On-demand derivative | Can adapt size or mark per request | Repeats work unless a cache absorbs it | Per-customer marks or many negotiated sizes |
Pick pre-generation for the ordinary portfolio path. It gives the public URL one boring job: return an already-approved asset. The private original remains the source for OCR, reprocessing, and future formats, while the browser never receives it by accident.
How should a Node.js Express app watermark public derivatives while originals stay private?
Treat the original and the public derivative as different records, not two names for the same blob. The upload handler accepts bytes, asks an image decoder to validate them, stores the exact input under an unguessable private key, passes that private input to OCR, then creates a fresh public image with a visible mark. Only the derivative directory is mounted with express.static. The original directory has no public route.
That boundary matters more than URL obscurity. A random filename is useful for avoiding collisions, but it does not make a file private if a web server exposes the containing directory. Access control comes from where the bytes are stored and which route, if any, can read them. For remote object storage, apply the same interface: private reads require application authorization; public derivative reads do not.
Order is important too. OCR should inspect the original before resizing, recompression, or watermark pixels alter the input. This is a B2B creator portfolio, so a photo may carry a product code, label, or caption that downstream search needs. The public copy has a different job. It should look good enough to judge the work without spending original-file bandwidth on every visit.
One warning: don't trust the upload's Content-Type field as proof of format. Decode the bytes. OWASP's file-upload guidance recommends defense in depth, including allowlisted extensions, generated filenames, size limits, storage outside the webroot, and file-content validation. The example below makes the decoder the content check and never uses the client filename as a path.
Quality is the gate; bandwidth is the budget
There is no universal JPEG quality value that preserves every portfolio image. I'm not sure such a value could exist: fine typography, flat illustration, and noisy photography fail differently. A fixed number is a starting policy, not evidence. Use a small representative corpus and keep the choice only after looking at the output.
I would benchmark each candidate width, format, and encoder setting against four things: derivative bytes, visible watermark legibility, OCR output from the untouched original, and a manual image review. The first is objective. The second and fourth need a defined review rubric. OCR accuracy belongs in the report because the pipeline must prove that text extraction used the original, even though changing the public encoder should not change that result. Start with an intentionally tiny matrix: long-edge widths of 1200 and 1800 pixels crossed with two encoder settings. Build the corpus around failure-prone material, including a wide photograph, a tall phone image, a transparent illustration, fine diagonal lines, low-contrast text, and a high-detail scene. Run every source through every candidate, save the byte count beside the output, and review the images at the display sizes the portfolio actually uses. Record results per source instead of averaging everything into one comforting number. A derivative that looks acceptable on nine photographs but destroys small lettering on the tenth is still a bad default. Format choice changes the trade-off as well. MDN's image-format guide describes JPEG as a common choice for still images and photographs, PNG as lossless with broad transparency support, WebP as supporting lossy or lossless compression and animation, and AVIF as offering both lossy and lossless compression. Browser support and input characteristics still decide what is sensible for a particular audience. Don't convert transparent artwork to JPEG unless the flattening behavior is deliberate. Finally, repeat the selected candidate through the actual HTTP path and inspect response headers and the decoded browser result; a command-line encoder benchmark alone misses delivery configuration. This is enough evidence for a first policy without pretending the policy is permanent.
No averages.
The bandwidth calculation should include cache misses and repeat views, not just the upload response. A pre-generated file can be served as an ordinary static asset and cached under an immutable key. If any output parameter changes, create a new key. Overwriting bytes behind a long-lived URL makes cache behavior hard to reason about and makes visual review less trustworthy.
Implement the private-original boundary
This example is deliberately local and explicit. It uses two filesystem roots so the security boundary is visible: data/originals is application-only, while data/public is the sole static mount. Replace the LocalStore methods with equivalent private and public object-storage operations in production; the route contract does not need to change.
The chosen limits are policy values, not universal recommendations. Here the upload cap is 15 MiB, the derivative's maximum width is 1800 pixels, and JPEG quality is 82. The handler returns 415 for bytes the decoder rejects and 413 when the upload middleware rejects an oversized body. Concrete errors beat a generic 400 when another developer has to wire the client.
import express, { NextFunction, Request, Response } from "express";
import multer from "multer";
import sharp from "sharp";
import { randomUUID } from "node:crypto";
import { mkdir, writeFile } from "node:fs/promises";
import { join } from "node:path";
type OcrResult = { text: string; confidence?: number };
type Ocr = (original: Buffer) => Promise<OcrResult>;
class LocalStore {
constructor(
private readonly privateRoot: string,
private readonly publicRoot: string,
) {}
async init(): Promise<void> {
await Promise.all([
mkdir(this.privateRoot, { recursive: true }),
mkdir(this.publicRoot, { recursive: true }),
]);
}
async putOriginal(key: string, bytes: Buffer): Promise<void> {
await writeFile(join(this.privateRoot, key), bytes, { flag: "wx" });
}
async putDerivative(key: string, bytes: Buffer): Promise<void> {
await writeFile(join(this.publicRoot, key), bytes, { flag: "wx" });
}
}
const publicRoot = join(process.cwd(), "data", "public");
const store = new LocalStore(
join(process.cwd(), "data", "originals"),
publicRoot,
);
const upload = multer({
storage: multer.memoryStorage(),
limits: { fileSize: 15 * 1024 * 1024, files: 1 },
});
function watermarkSvg(width: number, height: number): Buffer {
const fontSize = Math.max(24, Math.round(width * 0.035));
const x = width - Math.round(width * 0.035);
const y = height - Math.round(height * 0.04);
return Buffer.from(`
<svg width="${width}" height="${height}" xmlns="http://www.w3.org/2000/svg">
<text x="${x}" y="${y}" text-anchor="end"
font-family="sans-serif" font-size="${fontSize}" font-weight="700"
fill="white" stroke="black" stroke-width="2" opacity="0.78">
PORTFOLIO PREVIEW
</text>
</svg>
`);
}
function buildApp(runOcr: Ocr): express.Express {
const app = express();
app.use(
"/portfolio-images",
express.static(publicRoot, {
immutable: true,
maxAge: "1y",
fallthrough: false,
}),
);
app.post(
"/portfolio-assets",
upload.single("image"),
async (req: Request, res: Response, next: NextFunction) => {
if (!req.file) {
res.status(400).json({ code: "IMAGE_REQUIRED" });
return;
}
const id = randomUUID();
try {
const decoded = sharp(req.file.buffer, { failOn: "error" }).autoOrient();
const metadata = await decoded.metadata();
if (!metadata.width || !metadata.height) {
res.status(415).json({ code: "UNSUPPORTED_IMAGE" });
return;
}
const originalKey = `${id}.source`;
await store.putOriginal(originalKey, req.file.buffer);
const ocr = await runOcr(req.file.buffer);
const resized = await decoded
.resize({
width: 1800,
height: 1800,
fit: "inside",
withoutEnlargement: true,
})
.toBuffer({ resolveWithObject: true });
const derivative = await sharp(resized.data)
.composite([
{
input: watermarkSvg(resized.info.width, resized.info.height),
gravity: "center",
},
])
.jpeg({ quality: 82 })
.toBuffer();
const derivativeKey = `${id}.jpg`;
await store.putDerivative(derivativeKey, derivative);
res.status(201).json({
id,
publicUrl: `/portfolio-images/${derivativeKey}`,
extractedText: ocr.text,
});
} catch (error) {
if (error instanceof Error && /unsupported image format/i.test(error.message)) {
res.status(415).json({ code: "UNSUPPORTED_IMAGE" });
return;
}
next(error);
}
},
);
app.use((error: unknown, _req: Request, res: Response, next: NextFunction) => {
if (error instanceof multer.MulterError && error.code === "LIMIT_FILE_SIZE") {
res.status(413).json({ code: "IMAGE_TOO_LARGE" });
return;
}
next(error);
});
return app;
}
async function main(runOcr: Ocr): Promise<void> {
await store.init();
buildApp(runOcr).listen(3000);
}
export { buildApp, main };
There is a subtle ordering choice in that handler. It persists the original before running OCR or deriving the public image, which prevents a later processing retry from depending on the client's ability to upload again. A real deployment should also persist job state and asset metadata in a database, then move OCR and derivative generation to a queue if request latency is unacceptable. Keep the state transitions dull: received, original stored, text extracted, derivative ready, published.
The sample holds at most one 15 MiB upload in memory for each active request. That is easy to read but can become a capacity problem under concurrency because decoded images also consume memory. For a busy service, stream uploads to private temporary storage, cap concurrent image work, and run the decoder in a worker process or job runner. Do not mount that temporary location as static content.
Also decide what deletion means. If a creator removes an asset, remove public delivery first, then schedule deletion of the derivative, original, OCR text, and operational metadata according to the product's retention rules. A dangling public derivative is the failure users can see; a dangling private original is the governance failure they cannot.
When should on-demand derivatives win?
The catch is personalization. Pre-generation is not suitable when the mark contains a viewer identity, an expiring order number, or another per-request value. In that case, use an authenticated transformation service, cache only where the cache key includes every output-changing parameter, and keep its source read private. The public response can still be cached privately or for a short period if the authorization model permits it.
Stick with on-demand generation when the number of legitimate size and format combinations would create wasteful storage or when negotiation is part of the product requirement. Even then, avoid turning arbitrary query parameters into decoder instructions. Map a small set of named presets to bounded dimensions and encoder options. Config bloat is an attack surface with a maintenance bill.
Pre-generation also loses when a watermark must be revoked instantly across a very large catalog and storage rewrite time is unacceptable. The trade is real: request-time work buys flexibility, while fixed artifacts buy predictable serving cost and simpler review. Measure both with the same corpus and traffic assumptions. Your mileage may vary — especially when cache hit rate changes — so a benchmark copied from another portfolio is weak evidence.
For the common case, keep the original private, derive once, and publish only the marked output. It is easy to inspect, easy to cache, and hard to misconfigure accidentally because the private bytes never enter the static route.
References
- https://developer.mozilla.org/en-US/docs/Web/Media/Formats/Image_types
- https://expressjs.com/en/starter/static-files.html
- https://expressjs.com/en/guide/error-handling.html
- https://sharp.pixelplumbing.com/api-composite
- https://sharp.pixelplumbing.com/api-output/#jpeg
- https://github.com/expressjs/multer
- https://cheatsheetseries.owasp.org/cheatsheets/File_Upload_Cheat_Sheet.html
- https://nodejs.org/api/crypto.html#cryptorandomuuidoptions
Top comments (0)