A free VIN decode proxy often shares one worker pool across every DecodeVinValues call. One pathological path -- a hung upstream, a retry storm on a bad VIN, or a client that floods the same edge -- can exhaust concurrency and leave healthy submissions waiting behind the pile-up. Per-request deadlines and hedged fetches (covered elsewhere) limit how long a single attempt lives. This post is about bulkheading: partitioning concurrency so one bad VIN path cannot starve the rest of the pool.
The goal is narrow: isolate in-flight decode slots by tenant, route, or failure class so a saturated bulkhead rejects or queues locally instead of consuming every shared worker. Do not confuse this with circuit breakers that open on error rates, idempotency keys, or jittered retries. Here the job is: cap how many concurrent NHTSA-bound calls one noisy path may hold.
Where pool starvation comes from
Starvation appears when every decode shares one unbounded (or globally capped) queue:
- A single API key or browser session opens dozens of parallel DecodeVinValues calls
- A stuck upstream keeps workers occupied until client timeouts pile retries on top
- Batch import and interactive form traffic share the same pool with no separate ceilings
- A "bad VIN" hot path retries forever while interactive users wait for free slots
Each case turns free-quota and worker time into a tragedy of the commons. Bulkheads give each path its own concurrency budget so exhaustion stays local.
Cap concurrency per bulkhead key
Assign each decode a bulkhead key (tenant id, route name, or interactive vs batch). Acquire a slot before calling the proxy; release it in finally. When the bulkhead is full, fail fast or queue with a short local timeout -- do not steal slots from other keys.
export type Bulkhead = {
key: string;
limit: number;
inFlight: number;
waiters: Array<() => void>;
};
export function createBulkhead(key: string, limit: number): Bulkhead {
return { key, limit, inFlight: 0, waiters: [] };
}
export async function withBulkhead<T>(
bh: Bulkhead,
run: () => Promise<T>,
): Promise<T> {
if (bh.inFlight >= bh.limit) {
await new Promise<void>((resolve, reject) => {
const timer = setTimeout(() => {
const i = bh.waiters.indexOf(wake);
if (i >= 0) bh.waiters.splice(i, 1);
reject(new Error("BULKHEAD_TIMEOUT"));
}, 2_000);
const wake = () => {
clearTimeout(timer);
resolve();
};
bh.waiters.push(wake);
});
}
bh.inFlight += 1;
try {
return await run();
} finally {
bh.inFlight -= 1;
const next = bh.waiters.shift();
if (next) next();
}
}
export function bulkheadKey(opts: {
tenantId?: string;
route: "interactive" | "batch";
}): string {
return `${opts.route}:${opts.tenantId ?? "anon"}`;
}
Interactive form traffic can use a higher per-tenant limit than batch imports. Anonymous clients get a small shared ceiling so one scraper cannot monopolize DecodeVinValues.
Forbidden "fixes"
Product pressure often asks for:
- Raising the global worker limit instead of partitioning noisy paths
- Letting batch jobs borrow interactive slots when the batch bulkhead is full
- Retrying
BULKHEAD_TIMEOUTimmediately without backoff (which re-saturates the same key) - Sharing one bulkhead across all tenants "for simplicity"
- Treating bulkhead rejection as an NHTSA outage and opening a global breaker
Refuse those. A full bulkhead is a local capacity signal, not proof that vPIC is down. Keep breaker logic on upstream error rates; keep bulkheads on concurrency ownership.
export type DecodePool = Map<string, Bulkhead>;
export function getBulkhead(
pool: DecodePool,
key: string,
limits: { interactive: number; batch: number },
): Bulkhead {
let bh = pool.get(key);
if (!bh) {
const route = key.startsWith("batch:") ? "batch" : "interactive";
bh = createBulkhead(
key,
route === "batch" ? limits.batch : limits.interactive,
);
pool.set(key, bh);
}
return bh;
}
export async function decodeBehindBulkhead<T>(
pool: DecodePool,
opts: { tenantId?: string; route: "interactive" | "batch" },
run: () => Promise<T>,
): Promise<T> {
const key = bulkheadKey(opts);
const bh = getBulkhead(pool, key, { interactive: 8, batch: 2 });
return withBulkhead(bh, run);
}
Separate interactive and batch ceilings. Never let a saturated batch key drain the interactive pool, and never map bulkhead rejection into a fabricated "NHTSA unavailable" banner when other bulkheads still succeed.
UI and API copy that stays honest
Prefer:
- "Decode busy for this session -- try again shortly" on
BULKHEAD_TIMEOUT - Metrics labeled by bulkhead key, not a single global "queue depth"
- A short footnote in ops docs: local concurrency cap, not an upstream grade
Avoid:
- "NHTSA is down" when only one tenant bulkhead is full
- Silently queueing batch work on interactive workers
- Grey spinners that look like a decode in flight after you already rejected the slot
Quick checks
import assert from "node:assert/strict";
const bh = createBulkhead("interactive:t1", 1);
let released = false;
const first = withBulkhead(bh, async () => {
await new Promise((r) => setTimeout(r, 50));
released = true;
return "ok";
});
await assert.rejects(
withBulkhead(bh, async () => "nope"),
/BULKHEAD_TIMEOUT/,
);
assert.equal(released, false);
assert.equal(await first, "ok");
assert.equal(bh.inFlight, 0);
assert.equal(
bulkheadKey({ tenantId: "a", route: "batch" }),
"batch:a",
);
const pool: DecodePool = new Map();
const interactive = getBulkhead(pool, "interactive:x", {
interactive: 8,
batch: 2,
});
const batch = getBulkhead(pool, "batch:x", { interactive: 8, batch: 2 });
assert.equal(interactive.limit, 8);
assert.equal(batch.limit, 2);
assert.notEqual(interactive, batch);
Review rule: bulkhead modules must fail closed on local saturation, never steal slots across keys, and must not relabel rejection as a global NHTSA outage.
Takeaway
Bulkheading VIN decode workers keeps one bad path from starving the shared pool. Cap concurrency per tenant and route, fail or wait locally when a bulkhead is full, and keep that signal separate from breakers, hedges, and deadlines. Your free VIN proxy stays fair when interactive decodes cannot be crowded out by a single noisy batch or retry storm -- and when rejection copy admits local capacity, not invented upstream failure.
I maintain VIN Lookup, a free VIN decode based on NHTSA data.
Top comments (0)