DEV Community

Cover image for Three things I learned shipping ONNX to Fly.io
Digital Craft Workshop
Digital Craft Workshop

Posted on Originally published at Medium

Three things I learned shipping ONNX to Fly.io

ld-linux-x86-64.so.2: No such file or directory. That error cost me 20 minutes. Then OOM killed my Fly machine within 30 seconds of first inference. Then Pinecone's SDK threw "Must pass in at least 1 recordID" on a call where I clearly passed an array of IDs. Three errors in one afternoon, all misleading.

I was shipping a small Node feature for grownote, my Substack growth dashboard. The feature: a button that pulls memories from my mem0 (Pinecone-backed) and bakes them into AI-generated engagement drafts. To query Pinecone semantically I needed the same embedding model that wrote the index: sentence-transformers/all-MiniLM-L6-v2. The Node-friendly route is @xenova/transformers, an ONNX runtime port of the same model.

Three things had to work: getting the ONNX native binary running on Linux, fitting the model in machine memory, and making the Pinecone client behave like its docs say. None of them did on the first try.


1. Alpine vs glibc: the silent native binary trap

My Dockerfile was the standard node:22-alpine recipe. Multi-stage build, deps + build + runtime, ~180MB image. It's been good for grownote since launch.

@xenova/transformers installed cleanly and the build passed. The container came up, but the first call to the embedding pipeline returned this from Fly logs:

Error loading shared library ld-linux-x86-64.so.2: 
No such file or directory 
(needed by /app/node_modules/onnxruntime-node/bin/napi-v3/linux/x64/libonnxruntime.so.1.14.0)
Enter fullscreen mode Exit fullscreen mode

Fly logs showing the ld-linux-x86-64.so.2 shared library error

The ld-linux error that really means Alpine has no glibc | Generated with Claude

The lie in the error message: ld-linux-x86-64.so.2 is the glibc dynamic linker. The file is missing because Alpine doesn't have glibc. Alpine uses musl libc, which uses a different linker (ld-musl-x86_64.so.1).

@xenova/transformers ships prebuilt binaries via onnxruntime-node. Those binaries are linked against glibc. They don't run on musl, and there's no onnxruntime-node build for musl in the official package.

You can fix this two ways:

Option A: switch to a glibc base image. Change ARG NODE_VERSION=22-alpine to node:22-bookworm-slim. Image grows ~85MB (Debian is heavier than Alpine). Otherwise drop-in compatible. ONNX prebuilt binaries load without complaint.

Option B: keep Alpine and build ONNX from source against musl. Theoretically possible. Practically not worth the time unless you have hard reasons to keep Alpine (image size for very large fleets, security policy). For one app, this is hours of yak-shaving.

I went with Option A. The Dockerfile change was one line:

ARG NODE_VERSION=22-bookworm-slim
Enter fullscreen mode Exit fullscreen mode

Rebuilt, pushed, the container came up clean.

Lesson: Native Node modules that ship prebuilt binaries are usually glibc-only. Alpine works for pure-JS packages. Once your dep tree includes a compiled module like ONNX or sqlite3, Debian slim is the safer base. If a package fails at runtime instead of install time, you want to catch it in 30 seconds via docker run locally, not after deploy.


2. 512 MB Fly machine: the OOM with misleading symptoms

The next deploy started fine. Health check passed. I clicked "Sync from mem0" in the dashboard. Frontend showed a spinner. After 30 seconds: 502 from Fly's edge proxy.

Browser DevTools showed an empty response body. Frontend logged a generic "Uncaught Error: An unexpected response was received from the server."

I checked Fly logs:

[125.024082] Out of memory: Killed process 634 (node) 
total-vm:23813904kB, anon-rss:391536kB, file-rss:184kB
INFO Process appears to have been OOM killed!
Enter fullscreen mode Exit fullscreen mode

Fly OOM kill log with virtual and resident memory numbers

23.8 GB virtual is noise; 391 MB resident is what got it killed | Generated with Claude

The numbers: 23.8 GB virtual memory, 391 MB resident. The kernel killed the process because it was using ~76% of the 512 MB machine RAM, and its heuristic for "this process is about to thrash and bring down everything else" fired.

The 23 GB virtual is misleading. ONNX Runtime mmaps the model file. mmap reserves virtual address space without committing physical memory upfront, and pages get pulled in only as you read them. So the virtual number is meaningless. The resident number (391 MB) is what mattered. On a 512 MB machine, with Node + Next.js + Drizzle + Pinecone client all running, 391 MB just for the embedding pipeline pushed total RSS past the limit.

Two options:

Option A: bump machine memory. Edit fly.toml:

[[vm]]
  cpu_kind = "shared"
  cpus = 1
  memory = "1024mb"
Enter fullscreen mode Exit fullscreen mode

Cost delta: ~$2.50/month at full uptime, less if you have auto_stop_machines = "stop" and min_machines_running = 0 (idle machines don't bill). For grownote, which I use evenings, real cost delta is ~$0.40/month.

Option B: externalize the embedding. Use HuggingFace Inference API for the model. POST text, receive vector. No local model. It adds an HF API token as a new secret and 200-500ms of HTTP roundtrip per query, but no native deps and no memory overhead.

I tried Option B first because it felt cleaner. Then I realized HF Inference API for a personal app is yet another vendor dependency. If HF goes down, my engagement queue can't sync mem0. The whole point of self-hosted Pinecone with a local-friendly model was to own the stack.

I went with Option A. One-line config change, 1 GB machine, problem disappears. Fly auto-stop means I'm paying ~$0.40/mo extra for the few hours per day the app actually serves traffic.

Lesson: When local model loading OOMs on a small VM, don't reflexively reach for an API. Check the cost of bumping memory first. Auto-stopped machines are nearly free at idle. Bumping memory is usually cheaper than adding an external API, which means one more token to manage and one more service that can go down.

Wiring memory into your own Claude Code setup is its own rabbit hole; I packed the hooks, the config, and the mistakes that bite first into a free email series, The Claude Code Memory Starter.


3. Pinecone v7 SDK: the misleading validator

With Alpine swapped and machine bumped, the embedding worked. Now the actual feature: list all my mem0 records, translate Czech ones to English, re-embed, upsert. Pinecone has explicit APIs for all four steps.

The list call worked:

const page = await index.listPaginated({});
const ids = page.vectors.map(v => v.id);
Enter fullscreen mode Exit fullscreen mode

The fetch call did not:

const fetched = await index.fetch(ids);
// PineconeArgumentError: Must pass in at least 1 recordID.
Enter fullscreen mode Exit fullscreen mode

I logged ids to confirm: an array of 10 valid UUIDs. The validator was clearly checking something, but not what the error message said.

index.fetch with a bare array failing versus the object form working

index.fetch(ids) throws; the validator wants { ids }, not a bare array | Generated with Claude

I dug into the fetch validator in node_modules/@pinecone-database/pinecone:

const validator = (options) => {
    if (!options.ids || options.ids.length === 0) {
        throw new PineconeArgumentError('Must pass in at least 1 recordID.');
    }
};
Enter fullscreen mode Exit fullscreen mode

It checks options.ids, not the argument I handed in. My bare array arrived as options itself, so options.ids was undefined and the guard fired, even though the array held 10 entries.

The signature that works:

const fetched = await index.fetch({ ids });
Enter fullscreen mode Exit fullscreen mode

Object-form, with ids as a property. That's the v7 FetchOptions shape ({ ids: string[] }); the error message just never names the property it actually wants.

The clearer case is upsert, which Pinecone v7 explicitly moved to object-form. Pass the old array and you get the same kind of validator complaint:

// v6 / examples in older docs:
await index.upsert([{ id, values, metadata }]);

// v7:
await index.upsert({ records: [{ id, values, metadata }] });
Enter fullscreen mode Exit fullscreen mode

Same unhelpful message ("Must pass in at least 1 record to upsert") on a non-empty array.

This cost me 30 minutes between the two endpoints. The fix was trivial. The error message was actively unhelpful.

Lesson: When an SDK error says "must pass in at least 1 X" on a call where you clearly passed multiple X's, the validator is probably reading a property off an options object you didn't pass. Try wrapping the array in an object with a sensibly named property (ids, records, vectors). It's a 5-second test that works disturbingly often.

The Pinecone v7 release notes do mention the object-form move. I found them after I'd already debugged it. Read the migration guide first, even when the major version bump feels minor.


The payoff: same embedding model, two languages

The mem0 server that wrote the Pinecone index is a Python MCP server using the mem0 library, configured to use huggingface/all-MiniLM-L6-v2 for embeddings. The Node app reading the index uses @xenova/transformers with Xenova/all-MiniLM-L6-v2, an ONNX export of the same model.

I'd half-expected a vector dim mismatch or subtle drift between the Python and Node implementations. There was none. The ONNX export is deterministic. Vectors written by the Python pipeline and queried by the Node pipeline match within float precision.

The only thing I had to ensure: the same pooling: 'mean' and normalize: true settings on both sides. Otherwise the model is the model, regardless of which runtime you call it from.

This is genuinely useful: you can write data in one language ecosystem (Python's mature ML tooling) and read it in another (Node's web ecosystem) using the same embedding space. As long as both runtimes load the same model weights and apply the same post-processing, the vectors are interchangeable.

Python writes and Node reads the same Pinecone index using one embedding model

One embedding model, written from Python and read from Node | Generated with Claude


The rules I kept

Here's what I'm carrying into the next build:

Default to Debian for any Node app that might add native deps later. The Alpine image-size win is real for pure-JS apps. The moment you add a native dep, Alpine becomes a tax. Debian slim just works.

Probe small VM memory limits with the actual workload before assuming "it'll fit." A 512 MB machine on Fly comfortably runs Next.js + Drizzle + a few API integrations. It does not run Next.js + ONNX + a model in RAM at the same time. The OOM kill happens fast and the error is generic. Test on the target hardware, not your 32 GB Mac.

When SDK errors don't make sense, check argument shape before checking your data. "Must pass at least 1 X" on a non-empty array is almost always an arguments signature mismatch, not a data problem. Try the object-wrapped form first.

The actual feature now works. Sync runs in 8-12 seconds, costs ~$0.02 per click, and returns a 5-section markdown block of my recent decisions and project state that feeds into AI-drafted Substack engagement comments. The build was supposed to take an evening. It took two evenings plus a memory bump. Worth it for the next year of automated drafts that actually reference what I'm working on.

If you're shipping ML-adjacent Node code to Fly: bookworm-slim base and at least 1 GB of memory. Use Pinecone v7 in object-form. You'll skip three afternoons of confused debugging.

If you want the memory setup behind grownote's mem0 sync, it is a free email series, The Claude Code Memory Starter.


External Sources


I build small tools and kits for solo creators. You can find them here: https://danielrusnok.gumroad.com

Top comments (0)