I run a free PDF toolkit. Last month I checked something I had never actually measured: whether the compress tool compressed anything.
It did not. A 2,879,606-byte PDF came back at 2,938,188 bytes. Two percent larger. The meta description on that page promised image-heavy PDFs shrink 30-70%, and it had been promising that to every visitor for months. Examples written in Go
The cause
The handler sent the file to Gotenberg asking for PDF/A-2b:
writer.WriteField("pdfa", "PDF/A-2b")
PDF/A is an archival format. It embeds every font and ICC colour profile and leaves images at full resolution, so the file still renders the same way in forty years. I read "PDF/A" and pattern-matched: A for Archive, archives are small.
The fix
Ghostscript, which actually resamples images:
cmd := exec.Command(
"gs",
"-dBATCH", "-dNOPAUSE", "-dQUIET",
"-sDEVICE=pdfwrite",
"-dCompatibilityLevel=1.7",
"-dPDFSETTINGS="+preset,
"-dDetectDuplicateImages=true",
"-dCompressFonts=true", "-dSubsetFonts=true",
"-dDownsampleColorImages=true",
"-dColorImageResolution="+dpi,
"-sOutputFile="+outPath, inPath,)
Same file, measured against production: 2,879,606 bytes in, 417,480 out. Minus 85.5%, and 3.2 seconds instead of 16.2. Text stays selectable because fonts are subset rather than rasterized.
if len(compressed) >= len(fileBytes) {
compressed = fileBytes // already optimized; hand back the original
}
Without it you ship the PDF/A bug again in a different costume. Files that are already optimized get bigger when you rebuild them, and a compression tool that sometimes inflates your file is worse than no tool.
Then I tried to sell it
The web app makes no money. Anonymous people convert one PDF and leave, which is the correct behaviour for a free converter and a terrible business. But the same endpoints behind an API serve a different customer - one who pays, because their invoicing system needs to merge PDFs every night.
So: list it on RapidAPI. Before writing any code, measure whether the hardware survives paying traffic.
Capacity first
The whole thing runs on one t3.micro. That is a burstable instance:
you get a baseline CPU allocation and a credit balance you spend when you exceed it. Run out of credits and you either throttle hard or start paying surplus charges.
I ramped concurrent conversions and watched CPUCreditBalance
- 144 conversions per hour sustained without touching the credit balance
- 5 concurrent requests comfortable, degrading past that
- credit balance stayed full throughout
That is small. It is also more than enough - 144 per hour is 3,456 per day, and the paid tiers I planned top out at 25,000 per month. The go or no-go answered itself, and it cost an afternoon instead of a migration.
The bug that would have broken every paying customer
RapidAPI proxies all traffic through its own IP addresses.
My rate limiter was keyed on client IP. Ten requests per minute, per IP, which is sensible for a public web app. Behind the gateway, every paying subscriber in the world shares the same handful of proxy IPs. Customer A's requests would have exhausted Customer B's limit. The first two subscribers would have throttled each other into 429s and churned, and I would have had no idea why.
The fix is to key on the subscriber instead, and to verify the request actually came through the gateway before trusting the header that names them:
func (m *RapidAPIAuth) ServeHTTP(w http.ResponseWriter, r *http.Request) {
secret := r.Header.Get("X-RapidAPI-Proxy-Secret")
if secret == "" {
m.next.ServeHTTP(w, r) // not a gateway request, carry on
return
}
if subtle.ConstantTimeCompare([]byte(secret), []byte(m.expected)) != 1 {
writeError(w, http.StatusForbidden, "forbidden", "...")
return
}
user := r.Header.Get("X-RapidAPI-User")
// ... attribute the request to `user`, and rate limit on that
}
A limit you are not allowed to set
I wanted paid plans at 30 requests per hour, matching what the box can actually do.
RapidAPI does not let you. Paid tiers are locked at 1,000 requests per hour and the provider has no field to lower it. So the platform advertises a number my hardware cannot serve.
The resolution is not technical, it is honesty. The server enforces 120 per hour
Better a customer reads that before subscribing than discovers it at request 121.
Endpoints do not exist until you declare them
Last one, and it cost me real time.
The origin server has twelve endpoints. I defined four in RapidAPI's dashboard and assumed the gateway was a dumb proxy that would forward anything.
$ curl -X POST "https://...p.rapidapi.com/ocr-pdf" -H "X-RapidAPI-Key: $KEY" ...
{"message":"Endpoint '/ocr-pdf' does not exist"}
What made this hard to catch: an undeclared path returns 401 without a key, exactly like a declared one. Authentication runs before path resolution. So the obvious probe tells you nothing, and you need a valid key to discover that eight of your twelve endpoints are unreachable.
I had already written documentation for all twelve. Had I shipped it, every subscriber would have hit a 404 on two thirds of the product.
You can find it on RapidAPI where 12 enpoinds live.
Free tier is 50 requests a month, no card. It is on RapidAPI as Papyrio PDF Tools, and the reference is at here.
Top comments (1)
The PDF/A result matches what I see from the conversion side. The PDF/A step I run subsets and embeds the fonts, adds an sRGB profile as the output intent and cleans up the colour spaces. Nothing in it resamples an image, so it has no way to make an image-heavy file smaller.
Handing back the original when the output is not smaller means the tool can no longer make a file worse, whatever the next converter does.
Did PDF/A stay in the toolkit as its own option for people who need the archival format, or did it go when compress moved to Ghostscript?