Nakodo (nakodo.app) searches for creators on YouTube, Instagram and TikTok, which means it leans on several outside sources with very different failure modes. One is metered and bills us per request. Two are free but rate limited, and start refusing if you push.
I wrote earlier about the YouTube quota, where the daily allowance is 10,000 units handed down by Google and the only interesting decisions are how to divide it. This post is the opposite problem: the sources where we choose the number, and where I had the wrong idea about what the number is for.
A cap is either a budget or a guard, and ours were pretending to be both
The first version was a constant per source. Pick a number that covers a busy day, with some headroom, and refuse past it.
The flaw shows up the first time the product grows. A constant sized for last month's traffic is a ceiling on this month's: real users doing real work hit it, searches stop halfway through the afternoon, and the only fix is a human noticing and editing a number in a deploy. Meanwhile the thing the cap was supposed to protect against, a runaway loop or a wave of bot sign-ups burning through paid requests at machine speed, is not actually prevented by any number large enough to let normal growth through.
The two jobs are not the same job:
- A budget bounds the money. It belongs where money is controlled.
- A guard catches an anomaly. It belongs in the code, and it has to be relative to normal, because "normal" moves.
Once they are separate, the budget goes where it should: the metered provider has its own monthly spending limit in its own dashboard, and that is what actually bounds the bill. Our cap stopped trying to be the wallet and became a guard.
The whole guard is twelve lines
// Spike guards: daily caps that follow recent use instead of a fixed number.
// Today's cap is `surge` times the busiest of the last week's days, and never
// below `floor`. Steady growth raises it on its own, so more users never hit
// it; only a sudden jump (a job stuck in a loop, a wave of bot sign-ups) is
// held until tomorrow. A day that reached its cap counts at the cap, so real
// demand that keeps coming lifts the cap by `surge` a day.
export type SpikeGuard = { floor: number; surge: number };
export const SPIKE_LOOKBACK_DAYS = 7;
// `busiest` is the most used on one day in the last week, today not included.
export function spikeCap(guard: SpikeGuard, busiest: number): number {
return Math.max(guard.floor, Math.ceil(guard.surge * busiest));
}
That is it. No I/O, no clock, no database handle, so the tests are three assertions about arithmetic:
const guard = { floor: 1000, surge: 3 };
test("a quiet week keeps the floor", () => {
assert.equal(spikeCap(guard, 0), 1000);
assert.equal(spikeCap(guard, 333), 1000);
});
test("a busy week lifts the cap to a multiple of its busiest day", () => {
assert.equal(spikeCap(guard, 400), 1200);
assert.equal(spikeCap(guard, 5000), 15000);
});
Four properties fall out of the shape:
- Today is excluded from the lookback. Today's own use cannot raise today's ceiling, or the guard unwinds itself the moment it is approached.
- The busiest day, not the mean. A mean over a week lets a quiet weekend pull the ceiling down onto Monday morning.
- A day that hit the cap counts at the cap. That is the self-raising part: genuine demand that keeps arriving lifts the ceiling by the surge factor each day, so a real growth spurt costs you one afternoon rather than a support ticket.
- The floor is the budget sentence. It is the one number chosen in money rather than in traffic. For our metered source the floor is sized at roughly five dollars a day, and that is the comment next to it.
The surge factor is a judgement about what an anomaly looks like for your workload. Three times the busiest day in a week is loose enough that no normal pattern touches it, and tight enough that something looping flat out gets caught on the same day.
Reserving a request has to be one statement
The guard is only useful if two concurrent workers cannot both spend the last request. The counter lives in one row per source per UTC day, and the reservation is a single conditional upsert:
const rows = await db
.insert(providerDays)
.values({ provider, day: today(), requestsUsed: n })
.onConflictDoUpdate({
target: [providerDays.provider, providerDays.day],
set: { requestsUsed: sql`${providerDays.requestsUsed} + ${n}` },
setWhere: sql`${providerDays.requestsUsed} + ${n} <= ${limit} and ${providerDays.exhaustedAt} is null`,
})
.returning({ used: providerDays.requestsUsed });
if (rows.length === 0 || rows[0].used > limit) {
throw new SourceLimitError(`Today's ${provider} requests are used up`, provider, nextUtcMidnight());
}
The setWhere is the Drizzle spelling of the WHERE clause on ON CONFLICT DO UPDATE, and it is doing the real work. If adding n would cross the cap, the update matches nothing, RETURNING gives zero rows, and that is the refusal. No read-then-write, no advisory lock, no transaction to get wrong. The second condition in the same clause handles a different kind of stop, below.
The rows[0].used > limit check next to the empty-rows check is belt and braces for the insert path: the row did not exist, so the conflict clause never ran, and a large n could land above the cap on its own.
Two ways to be out of requests
The cap is our opinion. The provider has its own, and it does not negotiate: out of credit, rate limited for the day, or just angry. That gets its own column rather than a fake increment of the counter:
export async function markExhausted(provider: Provider): Promise<void> { /* sets exhaustedAt for today */ }
Setting exhaustedAt makes every later reservation fail through the same setWhere, because exhausted_at is null is part of it. One statement, two reasons, one refusal path. And the admin run log reports used, cap and exhausted separately, because "we stopped" and "they stopped us" need different reactions from a human.
Hitting a cap is not an error
The last piece is what the error means to the caller:
export class SourceLimitError extends Error {
constructor(message: string, readonly provider: Provider, readonly until: Date) { ... }
}
The until is why the class exists. Our job queue is one Postgres table, and the distinction it cares about most is defer versus fail. A failed job burns an attempt and eventually gives up. A deferred job keeps its attempts and comes back later. A source being used up until midnight is the textbook deferral: nothing is wrong with the job, the world is simply not ready for it, so the run pushes the whole stage past until and spends the rest of the invocation on work that can proceed.
That is also why the fixed caps for the free-but-rate-limited sources still exist alongside the guard. Their numbers are not protecting a budget and not catching a spike. They are sized to keep us comfortably under the point where the provider starts refusing, because a deferral we chose is cheaper than a refusal we have to interpret.
Go and look at the other half
Caps that users see are a different design problem, because a user-facing limit has to be a promise rather than a guard. Every number on nakodo.app/pricing comes out of one PlanLimits object in the repository: campaigns at a time, creators contacted, keywords per campaign, how often keywords are searched again, and the search queue priority that decides whose work runs first when the shared YouTube quota is busy. The comparison table on that page is generated from those fields, so the marketing copy cannot drift from the enforcement.
The prices are in your own currency on that page too, which is a separate story I have told before for a sibling app, and the free tier needs no card if you want to watch the limits behave from the inside.
Top comments (0)