The first version of every upload endpoint I've seen looks the same. It accepts a file, converts it, and returns the result in one request. It works great on your laptop with a 200 KB test file.
Then a user uploads a 90 MB scanned document on hotel Wi-Fi, the request hangs for 40 seconds, the load balancer kills it, and your server keeps converting a file nobody is waiting for. Multiply that by a few hundred users and you have an incident.
Here's how I'd design it instead.
Stop treating it as one request
Upload, conversion, and download have completely different failure profiles. Uploads fail because of networks. Conversions fail because of bad input or heavy CPU. Downloads fail because of expired links. Cramming them into a single synchronous call means one slow step ruins the other two.
Split them into a job:
POST /v1/uploads -> returns upload URL + upload_id
POST /v1/jobs -> { upload_id, target: "docx" } returns 202 + job_id
GET /v1/jobs/{job_id} -> status, progress, result URL
DELETE /v1/jobs/{job_id} -> cancel and clean up now
The key move is returning 202 Accepted from job creation. You're telling the client "I got it, it's not done, check back." That one status code removes most timeout headaches, because no HTTP connection has to stay open while the work happens.
Keep the bytes off your API server
If files stream through your app server, your app server becomes a file proxy. That's a bad job for something that should be handling JSON.
Use presigned URLs instead. The client asks your API for an upload slot, and your API returns a short-lived URL pointing straight at object storage (S3, GCS, R2, whatever you use). The client uploads there directly.
{
"upload_id": "upl_8f2c1a",
"url": "https://storage.example.com/...signature...",
"expires_in": 600,
"max_bytes": 52428800
}
Why this works: your API only handles tiny metadata requests, storage handles the heavy transfer, and you can scale each independently. Set the expiry short (5 to 15 minutes is typical) and enforce the size limit in the signed policy, not just in your docs.
Validate the file, not the filename
An extension is a suggestion. A user can rename payload.exe to report.pdf in two seconds.
After the upload lands, check the actual content before you pass it to a converter. For PDFs, that means looking at the magic bytes (files start with %PDF-). For other formats, use a library that inspects the signature rather than trusting the Content-Type header the client sent.
def looks_like_pdf(path):
with open(path, "rb") as f:
return f.read(5) == b"%PDF-"
Then return the right error. Two that get mixed up constantly:
- 413 Payload Too Large when the file exceeds your limit
- 415 Unsupported Media Type when the format isn't one you accept
And add a stable, machine-readable error code in the body (file_too_large, unsupported_format, encrypted_file). Clients shouldn't have to string-match your error messages.
Design the job states before you write the worker
A job status endpoint is a tiny state machine. Decide the states up front and don't invent new ones later.
queued -> processing -> succeeded
-> failed
-> canceled
(any) -> expired
Two rules that save you pain. First, make states one-directional. A job never goes from failed back to processing (a retry creates a new attempt, or a new job). Second, include a reason on failure, not just the word "failed."
{
"job_id": "job_41d9",
"status": "failed",
"error": {
"code": "password_protected",
"message": "The PDF is encrypted. Provide a password or upload an unlocked copy."
}
}
That single field turns a support ticket into a self-service fix.
Polling, webhooks, or both
Polling is simple and works everywhere. Webhooks are efficient but need retries and signature verification on your side. I'd ship polling first.
If you do poll, be kind to yourself. Return a Retry-After header or a suggested interval in the response, and tell clients to back off (1s, 2s, 4s, capped at something like 10s). Otherwise you'll discover what 5,000 clients polling every 200 ms feels like.
If you add webhooks later, sign the payload with an HMAC and include a timestamp so receivers can reject replays.
Make job creation idempotent
Mobile networks drop responses, not just requests. A client sends POST /jobs, the server creates the job, the response never arrives, and the client retries. Now you've converted the same file twice.
Accept an Idempotency-Key header. If you see the same key again within a window (24 hours is common), return the original job instead of creating a new one. It's a small feature that prevents duplicate work and duplicate bills.
Cleanup is a feature, not a chore
This is the part most teams skip until the storage bill or a privacy question forces the issue.
Uploaded files are often sensitive. Contracts, invoices, IDs, medical forms. Keeping them forever is a liability, and keeping them "just in case" isn't a plan. Build expiry into the design from day one:
- Lifecycle rules on the bucket. Set uploads to auto-delete after a fixed window (say 24 hours). This is your safety net, and it works even if your own cleanup code is broken.
- Explicit delete after success. Once the client has downloaded the result, or the result link expires, remove both input and output.
- A sweeper for orphans. Uploads that never got a job attached (user closed the tab) still exist. A scheduled task that removes unattached uploads older than an hour catches them.
-
Temp files on workers. If your converter writes to local disk, clean up in a
finallyblock, not just on the happy path. A crashed conversion that leaves 300 MB in/tmpwill fill a container faster than you'd think.
Tell users the retention window in the API docs and in the response ("expires_at": "2026-09-29T10:00:00Z"). A deletion policy nobody can see doesn't build much trust.
Protect the converter itself
Converters chew through untrusted input, so treat them like it.
Run conversions in a sandboxed worker with a hard timeout (60 to 120 seconds is a reasonable ceiling for documents). Cap memory. Cap page counts if you can. A malformed PDF with thousands of nested objects can eat CPU in ways a normal file never will.
Also set a per-user or per-IP concurrency limit. One person queuing 400 conversions shouldn't push everyone else's job to the back.
Test with ugly files
Your test suite probably contains three clean PDFs. Real users bring:
- a scanned document with no text layer
- a file with a password
- a 0-byte upload
- a PDF renamed to
.docx - a 300-page manual with embedded fonts
Keep a small folder of these and run it in CI. Every bad file you collect from production goes into that folder. It's the cheapest regression suite you'll ever build.
When I want a quick baseline for what a "correct" conversion should look like, I run the same file through a browser tool first and compare against my pipeline's output. PDF Conveter is handy for that, since it's free and needs no setup. If your API's output for a table-heavy PDF looks worse than a reference conversion, you've found something worth debugging.
A quick checklist before you ship
- Uploads go directly to storage, with size limits in the signature
- Job creation returns 202 and an ID
- File content is validated, not just the extension
- Errors have stable codes and human-readable messages
- Job creation accepts an idempotency key
- Every stored file has an expiry, and you can point to where it's enforced
- The converter has timeouts, memory caps, and per-user limits
None of this is exotic. It's mostly about admitting up front that files are messy, networks are flaky, and users will do things you didn't plan for. Design for that, and the "graceful" part mostly takes care of itself.
If you're building something similar, I'd like to hear what your ugliest test file was. Mine was a 40 MB scan of a fax of a printout, and it taught me more than any spec did
Top comments (0)