DEV Community

Daniel Pertu
Daniel Pertu

Posted on

10,000 units a day, one API key, and three caps that are not the same number

Nakodo searches YouTube for creators who fit a brand, so the YouTube Data API is not a nice-to-have integration, it is the product's floor. And that API has a hard daily budget: 10,000 units, resetting at midnight Pacific time.

The unit costs are the whole story:

  • search.list costs 100 units
  • channels.list, videos.list, playlistItems.list, commentThreads.list cost 1 unit each

So a project gets a hundred searches a day. Not a hundred per user. A hundred, total, across everybody. Everything else is effectively free by comparison.

That one fact shapes a surprising amount of the code. This post is about the budgeting layer, not about what we search for or how results are judged, which is the part I am keeping to myself.

Three caps, deliberately different

export const DAILY_QUOTA = 10_000;
// Stop short of the real limit; failed requests also cost quota.
export const QUOTA_CAP = 9_500;
// Searches stop here so 1-unit enrichment and comment calls always have room.
export const SEARCH_CAP = 8_000;

export const COST = { search: 100, list: 1 } as const;
Enter fullscreen mode Exit fullscreen mode

QUOTA_CAP at 9,500 is headroom. Failed requests still cost quota, so our counter and Google's counter drift apart in the direction that hurts, and running up to 10,000 means the last few calls of the day fail for everybody at once.

SEARCH_CAP at 8,000 is the more interesting one, and it comes from a bad afternoon. A search returns raw video rows that are worth nothing until the cheap 1-unit calls turn them into channels, upload histories and comment samples. If searches are allowed to run until the budget is gone, you spend 100 units at a time on new material and then cannot afford the 1-unit calls that make any of it usable. You end the day with a large pile of half-finished work and no complete results. So the expensive call is only allowed to eat 80% of the budget and the cheap calls always have 1,500 units waiting for them.

There is a third cap, and this one is a product decision rather than an engineering one. Per-account daily searches live in the plan table:

searchesPerDay: 8,   // Free
searchesPerDay: 30,  // Pro
searchesPerDay: 60,  // Business
Enter fullscreen mode Exit fullscreen mode

It is on the pricing page as "Up to 8 YouTube searches a day", and the row underneath it, "Search queue", is the same constraint seen from the other end: when there is a backlog, paid plans are served first. Open that page and scroll to the comparison table; those two rows are the shared 10,000 units divided up in public. I would rather sell a visible limit than silently slow people down.

The day key is a Pacific date string

Google resets at midnight Pacific. If you key your counter on a UTC date, your "day" ends at 4pm or 5pm Pacific depending on the season, and then you get an unexplained second reset in the evening.

export function pacificDay(date = new Date()): string {
  return new Intl.DateTimeFormat("en-CA", {
    timeZone: "America/Los_Angeles",
    year: "numeric",
    month: "2-digit",
    day: "2-digit",
  }).format(date);
}
Enter fullscreen mode Exit fullscreen mode

en-CA is a small trick worth keeping: it formats as 2026-10-02, so you get an ISO-shaped date for a named timezone with no date library and no manual offset arithmetic.

Reserving units is one statement, not a read and a write

The counter lives in a quota_days table, one row per Pacific day. The naive version of this is read the row, add the cost, write it back. With scheduled runs that overlap and several jobs in flight inside each one, that loses increments and sails past the cap.

Instead, reserve first and let Postgres decide, with Drizzle's setWhere on the conflict clause:

export async function reserveUnits(units: number, cap: number = QUOTA_CAP): Promise<boolean> {
  const day = pacificDay();
  const rows = await db
    .insert(quotaDays)
    .values({ day, unitsUsed: units })
    .onConflictDoUpdate({
      target: quotaDays.day,
      set: { unitsUsed: sql`${quotaDays.unitsUsed} + ${units}`, updatedAt: new Date() },
      setWhere: sql`${quotaDays.unitsUsed} + ${units} <= ${cap} and ${quotaDays.exhaustedAt} is null`,
    })
    .returning({ unitsUsed: quotaDays.unitsUsed });
  return rows.length > 0;
}
Enter fullscreen mode Exit fullscreen mode

Three things fall out of this shape:

  1. The increment and the cap check are the same statement, so two concurrent callers cannot both be told yes for the last 100 units.
  2. An empty returning array is the refusal. When setWhere does not hold, no row is written and nothing comes back, so rows.length > 0 is the answer. No exception, no second query.
  3. exhaustedAt is null in the same predicate means one flag can shut the whole thing down.

Every call goes through one function that reserves before it fetches:

async function call<T>(endpoint: string, params: Params, cost: number, cap: number, meter: QuotaMeter) {
  if (!(await reserveUnits(cost, cap))) {
    throw new QuotaError(`Daily YouTube quota budget reached (${cap} units)`, false);
  }
  meter.units += cost;
  // ...fetch...
}
Enter fullscreen mode Exit fullscreen mode

Reserving before the request, rather than counting after a success, is the honest model: Google charges for the attempt. The cap is a parameter because that is how the two caps are enforced; the search call passes SEARCH_CAP, everything else passes QUOTA_CAP. One line per call site, no branching.

When Google disagrees with our counter, Google wins

Our arithmetic can be wrong. The cost table could be out of date, or something else could be sharing the project's quota. So the error path trusts the API over itself:

const reason: string | undefined = body?.error?.errors?.[0]?.reason;
if (reason === "quotaExceeded" || reason === "dailyLimitExceeded") {
  await markQuotaExhausted();
  throw new QuotaError("YouTube says the daily quota is used up", true);
}
Enter fullscreen mode Exit fullscreen mode

markQuotaExhausted() stamps exhaustedAt on today's row, which the setWhere above already checks, so one 403 stops every further reservation for the day without a single extra read anywhere else. Note the boolean on QuotaError: exhausted: false means we stopped ourselves at our own cap, exhausted: true means Google stopped us. The first is a budgeting decision we can relax; the second is final.

The reset time walks forward an hour at a time, because of daylight saving

Deferred work needs a time to wake up at, which means computing the next Pacific midnight. The arithmetic version of that is wrong twice a year.

export function nextQuotaReset(now = new Date()): Date {
  const today = pacificDay(now);
  let t = new Date(now.getTime() + 60 * 60 * 1000);
  while (pacificDay(t) === today) t = new Date(t.getTime() + 60 * 60 * 1000);
  return new Date(t.getTime() + 5 * 60 * 1000);
}
Enter fullscreen mode Exit fullscreen mode

It steps forward in one-hour jumps until the Pacific date string changes, then adds five minutes of slack. It is not clever and it does at most 25 iterations of string formatting, which is free compared to one API call. The DST case is the test I care about:

test("reset works across the DST change", () => {
  const now = new Date("2026-11-01T06:00:00Z"); // Oct 31, 23:00 PDT
  const reset = nextQuotaReset(now);
  assert.equal(pacificDay(reset), "2026-11-01");
  assert.ok(reset.getTime() - now.getTime() <= 2.1 * 60 * 60 * 1000);
});
Enter fullscreen mode Exit fullscreen mode

On the night the clocks go back, 23:00 is an hour and a bit from the next Pacific day, not 60 minutes. An addHours(24) or a hardcoded offset gets this wrong and either wastes a run or wakes up before the reset and burns a failed request.

Running out of quota is not an error

The last piece is what happens to work in flight when the budget goes. It is not a failure, because nothing is wrong: the work simply cannot happen until tomorrow. So quota-starved jobs are put back without counting an attempt against them.

export async function deferJobs(ids: string[], until: Date): Promise<void> {
  await db.update(jobs)
    .set({ status: "pending", lockedAt: null, runAfter: until, attempts: sql`greatest(${jobs.attempts} - 1, 0)` })
    .where(and(inArray(jobs.id, ids), eq(jobs.status, "running")));
}
Enter fullscreen mode Exit fullscreen mode

The attempts - 1 is the point. Claiming a job increments its attempt counter, and jobs give up after three, so without the decrement two quiet days of quota pressure would retire a job that never actually ran. greatest(..., 0) keeps it from going negative when a job is deferred before it was ever properly tried.

And only the stages that spend units go to sleep:

const USES_YOUTUBE = new Set<JobType>(["enrich", "comments", "search"]);
Enter fullscreen mode Exit fullscreen mode

Everything else carries on. When the quota is gone, we stop fetching from YouTube and keep working through what is already stored, so the app is still doing something useful for the rest of the day, and the message in the UI says so rather than showing a spinner that will not finish until tomorrow.

The summary I would give my past self

A third-party rate limit is not an edge case to handle at the call site, it is a resource to budget. Give it a counter keyed the way the provider resets it, reserve from that counter in one atomic statement before you spend, keep separate caps for the expensive and the cheap calls, trust the provider's "no" over your own arithmetic, and make running out a deferral rather than a failure.

If you want to see where this surfaces for a user, the pricing comparison is the same budget written as plan limits, and the free plan takes no card if you want to watch eight searches a day being spent on your own keywords.

Top comments (0)