Our document ingest path used to run OCR inside the app pods. The library wanted glibc. The app images were Alpine. Missing language packs got pulled from github.com at parse time. That last part is the one that kept me up: a PDF upload could trigger an outbound fetch from a pod that otherwise had no business talking to the public internet.
So we moved parsing into its own service, baked the traineddata into the image, and closed egress entirely. Including DNS.
Why the app was the wrong place
Three separate failures stacked in one process:
- The native liteparse binary didn't load on Alpine. You discover that in the target environment, not in a glibc laptop build.
- Tessdata that wasn't already on disk downloaded at runtime. OCR "just worked" until the network policy, the rate limit, or the broken mirror made it not work.
- The parse library accepted path-shaped options (
tessdataPath,imageOutputDir,ocrServerUrl). A caller that could set those could point the parser at another filesystem or another host.
None of those are model problems. They're packaging and trust-boundary problems that show up the first time you treat document upload as untrusted input, which it is.
One HTTP service, three caller knobs
The service is a small Bun HTTP app: multipart upload in, pages of text out. Callers may pass only:
z.strictObject({
ocrEnabled: z.boolean().optional(),
ocrLanguage: z.string().regex(/^[a-z]{3}(\+[a-z]{3})*$/).optional(),
maxPages: z.number().int().positive().optional(),
});
strictObject rejects every other field by name. tessdataPath is not an option; it comes from LITEPARSE_TESSDATA_PATH and nowhere else. Languages are checked against the .traineddata files present at startup. Ask for fra when only eng and deu are baked in and you get 400, not a download.
The image pins deu and eng from tesseract-ocr/tessdata_best at a fixed commit SHA with sha256 verification. Build time is when we fetch. Runtime is when we refuse to.
Auth is a shared bearer token, sixteen characters minimum. Health checks stay unauthenticated so Kubernetes can probe without holding the secret. There is no /metrics scrape path, because nothing scrapes it yet and an open metrics port would be another thing to reason about.
Closed egress is the point
In the cluster the service sits in its own namespace. The NetworkPolicy sets egress: [] — no outbound traffic, DNS included. Ingress is an allowlist of caller namespaces from stack config. Empty list means deny-all until the first consumer is declared.
That shape only works because the container never needs the network. Puppeteer, sitting next to it in our stack, cannot do the same: it renders user HTML with external subresources. OCR can, once the models live in the image.
App pods call http://liteparse-service…:3000 over the cluster network with the shared token. They never embed the native binary. Alpine stays Alpine. The parse blast radius is one dedicated Deployment with a memory tier and no way out.
What we'd do again
Pin traineddata in the image and fail closed on unknown languages. Whitelist request options so a multipart field cannot redirect filesystem or network paths. Put the NetworkPolicy next to the Deployment in the same PR, with ingress empty by default so a forgotten consumer config doesn't open the service to the whole cluster.
If your OCR still downloads language packs on first use, you don't have an offline parser. You have a deferred curl with a PDF as the trigger.
Top comments (0)