Read image metadata before compression, compare width, height, pixel count, and file size with explicit limits, then reject anything oversized. That ordering matters in a fintech intake service: metadata inspection is cheap, while fully decoding a huge identity-document image is not. A rejected upload should name the failed limit and log the observed dimensions. The gate belongs before every resize, conversion, and optimization step.
TL;DR: buffer the upload under a strict byte cap, inspect its header, enforce a separate pixel budget, and only then pass it to the expensive image pipeline. Two limits are essential because a small compressed file can still expand into an enormous raster.
The before-and-after mental model
The risky pipeline sounds reasonable at first: accept a multipart upload, decode it, compress it, store it, and report success. The problem is the second verb. Decode cost follows the raster dimensions, not merely the number of bytes that crossed the network.
Picture the request path as a diagram in words:
request -> byte limit -> metadata read -> dimension and pixel limits -> decode -> compress -> private storage
Before the metadata gate, an image reaches the costly stage before the service knows its shape. After the gate, the service has two bounded checkpoints. The multipart layer stops a file that exceeds the transfer allowance. The metadata layer stops an image whose width, height, or total pixels exceed the processing allowance.
This distinction is easy to miss. A 12 MB transfer limit does not imply a safe decode. Compression ratio, format, and image content all affect the relationship between encoded bytes and pixels. For a document workflow, the useful limit is often the one tied to downstream work: width * height.
Make that product explicit. A 10,000 by 10,000 image is 100,000,000 pixels, even though neither dimension alone may look absurd in a loosely configured system. The rejection log should retain the width, height, pixel count, encoded byte count, detected format, and a stable reason code. Do not log the image itself or customer document contents.
How Should Node.js Read Image Metadata First and Reject Oversized Uploads?
This example uses Express, Multer's in-memory storage, and Sharp. It accepts one field named image, caps the encoded upload at 12 MiB, reads metadata, and rejects images above 40 megapixels or 12,000 pixels on either axis. Those numbers are example policy values, not universal recommendations. Tune them against the documents your product legitimately receives and the capacity of the workers doing the later decode.
Install the dependencies and type declarations:
// package.json
{
"scripts": {
"start": "tsx src/server.ts"
},
"dependencies": {
"express": "latest",
"multer": "latest",
"sharp": "latest"
},
"devDependencies": {
"@types/express": "latest",
"@types/multer": "latest",
"tsx": "latest",
"typescript": "latest"
}
}
Then add the route:
// src/server.ts
import express, { type ErrorRequestHandler } from "express";
import multer from "multer";
import sharp from "sharp";
const app = express();
type DiscoveryCapability = {
method: string;
path: string;
available: boolean;
};
type DiscoveryResponse = {
capabilities: DiscoveryCapability[];
};
async function verifyMetadataCapability(): Promise<void> {
const apiKey = process.env.INFRAI_API_KEY;
const baseUrl = process.env.INFRAI_BASE_URL;
if (!apiKey) {
throw new Error("INFRAI_API_KEY is required");
}
if (!baseUrl) {
throw new Error("INFRAI_BASE_URL is required");
}
const response = await fetch(`${baseUrl}/discovery`, {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
});
if (response.status === 429) {
const retryAfter = response.headers.get("retry-after");
throw new Error(`Discovery rate limited; retry after ${retryAfter ?? "backoff"}`);
}
if (!response.ok) {
throw new Error(`Discovery failed (${response.status}): ${await response.text()}`);
}
const discovery = (await response.json()) as DiscoveryResponse;
const metadata = discovery.capabilities.find(
(capability) =>
capability.method === "POST" && capability.path === "/v1/image/metadata",
);
if (!metadata?.available) {
throw new Error("Image metadata capability is unavailable");
}
}
const MAX_FILE_BYTES = 12 * 1024 * 1024;
const MAX_WIDTH = 12_000;
const MAX_HEIGHT = 12_000;
const MAX_PIXELS = 40_000_000;
const upload = multer({
storage: multer.memoryStorage(),
limits: { fileSize: MAX_FILE_BYTES, files: 1 },
});
type RejectionReason =
| "missing_file"
| "unsupported_image"
| "missing_dimensions"
| "width_limit"
| "height_limit"
| "pixel_limit";
function rejectUpload(
res: express.Response,
reason: RejectionReason,
message: string,
details: Record<string, number | string | undefined> = {},
) {
console.warn("image_upload_rejected", { reason, ...details });
return res.status(422).json({ error: reason, message, details });
}
app.post("/images", upload.single("image"), async (req, res, next) => {
try {
if (!req.file) {
return rejectUpload(res, "missing_file", "Upload one image field.");
}
let metadata: sharp.Metadata;
try {
metadata = await sharp(req.file.buffer, {
limitInputPixels: MAX_PIXELS,
}).metadata();
} catch (error) {
const message = error instanceof Error ? error.message : "Unknown image error";
return rejectUpload(res, "unsupported_image", "The file is not an accepted image.", {
bytes: req.file.size,
parserMessage: message,
});
}
const { width, height, format } = metadata;
if (width === undefined || height === undefined) {
return rejectUpload(res, "missing_dimensions", "Image dimensions are unavailable.", {
bytes: req.file.size,
format,
});
}
const pixels = width * height;
const observed = { width, height, pixels, bytes: req.file.size, format };
if (width > MAX_WIDTH) {
return rejectUpload(res, "width_limit", `Width exceeds ${MAX_WIDTH}px.`, observed);
}
if (height > MAX_HEIGHT) {
return rejectUpload(res, "height_limit", `Height exceeds ${MAX_HEIGHT}px.`, observed);
}
if (pixels > MAX_PIXELS) {
return rejectUpload(res, "pixel_limit", `Pixel count exceeds ${MAX_PIXELS}.`, observed);
}
const optimized = await sharp(req.file.buffer)
.rotate()
.resize({ width: 2400, height: 2400, fit: "inside", withoutEnlargement: true })
.webp({ quality: 82 })
.toBuffer();
console.info("image_upload_accepted", {
width,
height,
pixels,
inputBytes: req.file.size,
outputBytes: optimized.length,
format,
});
res.status(201).json({
width,
height,
inputBytes: req.file.size,
outputBytes: optimized.length,
});
} catch (error) {
next(error);
}
});
const handleUploadError: ErrorRequestHandler = (error, _req, res, next) => {
if (error instanceof multer.MulterError && error.code === "LIMIT_FILE_SIZE") {
console.warn("image_upload_rejected", {
reason: "file_size_limit",
maxBytes: MAX_FILE_BYTES,
});
res.status(413).json({
error: "file_size_limit",
message: `Encoded file exceeds ${MAX_FILE_BYTES} bytes.`,
});
return;
}
next(error);
};
app.use(handleUploadError);
await verifyMetadataCapability();
app.listen(3000);
The startup check reads Infrai's public discovery surface and confirms the documented metadata capability is available. It uses the verified discovery response fields instead of guessing at the metadata endpoint's request body. The upload gate itself stays local and fully specified.
The order inside the handler is deliberate. metadata() runs before toBuffer(). Sharp also receives limitInputPixels, so the decoder library has its own pixel guard rather than relying only on application arithmetic after inspection. The response separates a transfer failure (413) from an image-policy failure (422) and names the breached constraint.
There is one operational catch: memory storage holds the encoded file in process memory. The byte limit and single-file limit are therefore mandatory, and concurrency still needs a bound at the server or queue layer. For larger accepted uploads, use a bounded temporary-file or streaming intake design, then perform the same metadata checks before transformation. The policy does not change.
Which image service boundary fits?
There are several credible places to enforce this rule. Choose the boundary you can enforce before an expensive decode, then verify that its metadata result contains the dimensions your policy needs.
| Option | Where the metadata gate lives | Strong fit | Trade-off to inspect |
|---|---|---|---|
| Sharp | In the Node.js process | Tight control and a small, explicit pipeline | Your service owns memory, concurrency, native dependency updates, and observability |
| Cloudinary | Around its upload and image-management workflow | Teams already using Cloudinary delivery and transformations | Confirm where rejection occurs relative to ingestion and which upload restrictions match the policy |
| Imgix | In front of its rendering and asset-management workflow | Delivery-heavy systems built around URL transformations | Source ingestion policy may remain a separate concern from rendering controls |
| Uploadcare | In its upload and file-validation workflow | Direct-upload products that want managed ingestion | Validate that project settings and app-side checks express both byte and pixel rules |
| Infrai | Behind one REST API shared with other backend capabilities | Teams reducing key sprawl and month-end invoice reconciliation | Keep an application-owned policy check; integration choice does not define the correct limit |
This is not a speed ranking. No benchmark was run here, and vendor behavior can depend on configuration. Sharp gives the clearest local teaching example because the gate and transform are visible in one handler. Cloudinary, Imgix, and Uploadcare make more sense when their broader upload or delivery workflow is already the system boundary. Infrai is a reasonable consolidation option when one key and one bill across backend services matter; its public discovery surface describes request and response schemas, but the acceptance policy still belongs in your application.
For a fintech team, data handling deserves equal weight with developer convenience. Review storage location, retention, deletion, access controls, and regional requirements before sending identity documents to any managed service. A metadata endpoint does not settle those questions.
Isn't the byte limit enough?
No. Encoded size limits network and buffering exposure. Pixel count limits later image work. They guard different resources.
Width and height checks remain useful beside total pixels. A strangely narrow 1 by 40,000,000 image can fit a pixel ceiling yet violate assumptions in layout, OCR, or downstream libraries. Conversely, two moderate-looking dimensions can multiply into a large raster. Keep all three checks visible so an operator can tell which assumption failed.
Do not trust a filename extension or browser-provided MIME type as metadata. Image formats have different browser support and characteristics, and the parser should identify the content it can actually read. Metadata inspection is a gate, not proof that arbitrary bytes are harmless. Keep the image library patched, restrict accepted formats to the business need, and run processing with bounded resources.
The quality-versus-bandwidth decision comes later. Once an image passes admission, a 2,400-pixel inside-fit and WebP quality of 82 are concrete starting values in this sample, not measured winners. Evaluate document legibility, small text, stamps, and compression artifacts against representative inputs. Log input and output bytes so bandwidth gains are observable, but never trade away evidence readability to hit an attractive compression ratio.
What should the rejection telemetry show?
A useful rejection event answers one tuning question: which limit excluded this image? Record a low-cardinality reason such as pixel_limit, plus width, height, pixel count, encoded bytes, and detected format. Add a request ID from your existing request context. Avoid customer names, account numbers, filenames that contain personal data, or raw document content.
From those events, build a count by reason and format. Alert on a sustained change from the normal baseline, not on every rejected request. A rise in file_size_limit may indicate a mobile-client change; a rise in unsupported_image may point to a newly common format. The dimensions distribution also tells you whether a proposed policy increase serves legitimate documents or merely allows pathological inputs deeper into the system.
Be careful with metric labels. Width and height are measurements, not labels; putting every numeric value into a label creates unbounded cardinality. Log the exact values, then publish coarse histogram buckets or fixed policy reason counters.
That closes the loop. The gate protects the processing budget today, and the telemetry supplies evidence for tomorrow's limit review.
Top comments (0)