Ten to thirty seconds. That is how long a streamed AI call leaves a quota unenforced if the counter goes up when the stream closes, which is the obvious place to put it and the place I nearly put it.
I have wired a usage quota in front of four things now: Claude API calls, video renders, outbound email batches, and one cached lookup that barely deserved one. Every one came down to one small decision. Does the counter go up before the work, or after it comes back?
I wrote the underlying bug up in March, back when it was still a surprise to me. The postmortem is published, so this is what came after it.
The short version: for metered work, increment before the work happens. What decides it is duration. A ten-second call needs the counter in front of it; so does a ten-second render.
Only the write can move
Every metered feature has the same skeleton. Read the counter, compare it against the plan limit, do the work, write the counter back. The write is the only one of the four I can slide.
checkPlanLimits() is the boring part. It looks up the tenant and reads the usage row for the current calendar month, keyed by a plain string like "2026-03". A new month creates a new row on the first call and the old one stays behind as history, so nothing has to be reset on a schedule.
Put the write last and you get the version that reads as fair, because only successful work gets charged.
const limits = await checkPlanLimits(tenant.id);
if (limits.aiUsage.used >= limits.aiUsage.limit) return tooManyRequests();
const response = await anthropic.messages.create({ ... });
await incrementAiUsage(tenant.id); // count after
Put the write before the work and you get the version I ship.
const limits = await checkPlanLimits(tenant.id);
if (limits.aiUsage.used >= limits.aiUsage.limit) return tooManyRequests();
await incrementAiUsage(tenant.id); // count first
const response = await anthropic.messages.create({ ... });
Two lines swapped. In Drippery, the drip email tool I build, those two lines sit in the route handler behind AI email generation, capped at ten prompts a month on the Starter plan.
tooManyRequests() returns a 429 with a short JSON body, and the client turns that into a counter in the UI reading five of ten used this month. Both versions return the same 429 and render an identical usage counter. The only difference is which side of the model call the write lands on.
Counting last opens a gap between the read and the write. Anything arriving inside that gap reads a stale number and starts work it should not have been allowed to start. How wide the gap gets depends entirely on how long the work runs.
For the non-streaming call that builds a whole email sequence, the work runs five to fifteen seconds. For a streamed generation, the response object stays open until the last token lands, and the natural place to put the increment is the point where the stream closes.
for await (const event of stream) {
if (event.type === 'content_block_delta' && event.delta.type === 'text_delta') {
const data = JSON.stringify({ text: event.delta.text });
controller.enqueue(encoder.encode(`data: ${data}\n\n`));
}
}
controller.close();
// counting here puts the write 10-30 seconds after the read
Ten to thirty seconds is a long time to leave a paid endpoint unguarded. It is also long enough that nobody has to be malicious to walk through it. A user who clicks generate and then, seeing nothing happen, clicks again has already done it.
Most rate-limiting advice was written for CRUD endpoints. A form submission finishes in forty milliseconds, so hitting the gap takes a deliberately timed second request rather than an impatient second click. A model call finishes in twenty seconds. That leaves enough time for a second click to land inside the gap without anyone trying.
The race window is the whole span between the check and the increment. | Generated with Claude
The cheaper mistake is the one with a ceiling
Counting first has an obvious cost. When the model call fails, the user has spent a prompt and received nothing back. I sat on that one for a while before shipping it.
Then I put numbers on both sides of it. Anthropic API errors show up in my logs at well under one percent of calls. When one lands, a Starter user has nine prompts left for the month instead of ten, and if they write to me about it I add one back by hand.
I have never automated that refund. It happens rarely enough that a hand-written reply is cheaper than the code to avoid it. People also seem to like getting an actual reply.
The error on the other side has no ceiling. Someone firing requests faster than the work completes gets as many paid generations as they can queue inside the window, and I pay Anthropic for every one of them.
A generated email runs a few thousand tokens, so one call costs me somewhere around four cents at the tier I use. Ten of those a month come to forty cents against a nine-dollar subscription, which is why I could price the quota generously in the first place. Without a working quota, there is no ceiling except how fast a user can click.
I keep running into this shape of decision, where the only honest question is which mistake I would rather absorb. I wrote up the way I think about it as a free email series, Good-Enough Engineering.
Card processors settled this argument decades ago. The authorization hold goes on the card before the warehouse picks the item, and it is captured or released once the outcome is known. Same bet here. Stripe releases the hold automatically; I do it by editing a row by hand.
Both orders make a mistake; only one of them has a ceiling. | Generated with Claude
The second race is still in there
Moving the increment forward closes the wide race. A smaller one is still in there, inside the increment function.
export async function incrementAiUsage(tenantId: string): Promise<void> {
const month = getCurrentMonth();
const existing = await db.select().from(aiUsage).where(
and(eq(aiUsage.tenantId, tenantId), eq(aiUsage.month, month))
);
if (existing.length > 0) {
await db.update(aiUsage)
.set({ promptCount: existing[0].promptCount + 1, updatedAt: new Date() })
.where(and(eq(aiUsage.tenantId, tenantId), eq(aiUsage.month, month)));
} else {
await db.insert(aiUsage).values({ tenantId, month, promptCount: 1 });
}
}
Two concurrent calls can both select a count of four. Both then write five, and the tenant gets a free prompt out of the arithmetic.
Postgres will collapse the read and the write into one statement whenever I decide it matters. The unique index on tenant and month is already in the schema, which is the only thing ON CONFLICT needs to key on.
INSERT INTO drippery_ai_usage (tenant_id, month, prompt_count)
VALUES ($1, $2, 1)
ON CONFLICT (tenant_id, month)
DO UPDATE SET prompt_count = drippery_ai_usage.prompt_count + 1;
That statement buys less than it sounds like. It removes the lost update between the SELECT and the UPDATE. It does not cap the counter, because the quota comparison is a separate read further up the route, so prompt_count will still go from ten to eleven if the handler asks it to. The March piece said the rewrite left "no race window at all", which was too strong.
The other option was to wrap the select and the update in a transaction with SELECT ... FOR UPDATE, which holds the row until the transaction commits. That closes the lost update, but it still needs a retry for the first call of a new month, when there is no row to lock yet and both inserts race the unique index. It is also more code than the single statement above, so if I am going to touch this at all I would rather touch it once.
I have not shipped it either way. The window here is one database round trip rather than thirty seconds of streaming, and Drippery serves a few dozen AI requests a day, so two of them colliding inside a couple of milliseconds has never appeared in the logs. I ran those numbers here too and got the opposite answer.
Two reads of four both write five, and one prompt goes unbilled. | Generated with Claude
Three of my four quotas count first
The rule is not universal, and the four quotas I run do not all land the same way.
Claude API calls. Count first. The work takes seconds, and it costs money on someone else's meter. Once the request is out I cannot un-spend it.
Video renders. Count first. The job is queued long before the GPU picks it up, so the counter goes up at enqueue. A render sits in a queue for minutes before it starts and then occupies a single graphics card for several more. Counting at completion would leave a window measured in minutes, and the queue is allowed to get long.
Outbound email batches. Count first, at enqueue. Sent email is the least reversible thing on this list. A batch that goes out twice is a deliverability problem I get to explain to a mail provider, and that one I cannot absorb quietly.
A cached lookup. Count after, and honestly it barely matters. The work is a local read that finishes in single-digit milliseconds. Charging a user for a lookup that errored would be the more annoying of the two mistakes, so this is the one case where counting after is the correct call.
A four-cent API call and a render on hardware I already own leave the same window open if they both take thirty seconds. Which is why the cache and the render queue land on opposite sides despite costing me nothing but electricity.
Only the cache finishes fast enough that counting after stays safe. | Generated with Claude
The render queue is the one that surprised me. I had it counting on completion for a while, on the reasoning that a render that crashes halfway through a scene should not cost anybody anything. Then I watched the queue back up and realised I was measuring the wrong span. The window is queue wait plus render time, which on a busy night is most of an hour.
The Drippery code has not changed since I wrote that first piece. The decision takes me about thirty seconds now instead of an evening. Any time I put a quota in front of something that runs longer than a database write, the counter goes up before the slow thing starts, and I move on to the interesting half of the feature.
External Sources
- Anthropic streaming docs — the delta types the streaming loop has to guard on
- PostgreSQL INSERT reference — ON CONFLICT DO UPDATE semantics
- Stripe on authorization holds — the hold-then-capture pattern
- MDN on Server-Sent Events — the streaming transport itself
Nearby: getting Claude to return strict JSON without drift, from the tool that publishes these posts, and the $0/month background job architecture behind the email side of Drippery.
I build Drippery in public and write up the decisions as they happen. If the trade-off in this piece is the part you want more of, Good-Enough Engineering is a free five-email series on making that call.
I build small tools and kits for solo creators. You can find them here: https://danielrusnok.gumroad.com




Top comments (0)