CogniPrep is a practice platform for the aptitude tests and game based assessments that employers run on graduate applicants. Part of it is a lifecycle email cron: an hourly job that finds people who started something and did not finish it, and sends them one relevant email.
The job was losing emails, and the reason was a loop that looked completely ordinary.
The loop
for (const candidate of candidates) {
await resend.emails.send(buildEmail(candidate));
}
Sequential. One at a time. Nothing concurrent about it, no Promise.all, no fan out. It is the shape you write when you are being careful.
Our email provider allows 10 requests per second per account, and returns a 429 when you exceed it. A single send returns in well under 100ms. So that loop runs at roughly 13 iterations per second, and everything past the tenth send inside any given second comes back rate_limit_exceeded.
The failure mode is the nasty kind. Nothing throws. The API responds, the response carries an error field, the loop moves on. From the outside the job "ran", and every recipient past the cap in a given second was never emailed.
Sequential does not imply slow enough. It only implies one at a time, and one at a time can still be far too fast.
Two defences, at different layers
The first is a paced schedule inside the shared client. Every send claims a slot on a monotonically increasing timestamp, and the slot is claimed before any await, so two concurrent callers cannot both read the same Date.now() and pick the same slot:
const SENDS_PER_SECOND = 8; // deliberately under the cap
const SEND_SPACING_MS = Math.ceil(1000 / SENDS_PER_SECOND);
async function awaitSendSlot() {
const now = Date.now();
const slot = Math.max(now, nextSendSlotAt);
nextSendSlotAt = slot + SEND_SPACING_MS; // reserved synchronously
if (slot > now) await sleep(slot - now);
}
A 429 that slips through anyway is retried with exponential backoff, and the retry pushes the shared schedule back too, so every other caller in the process backs off with it rather than marching into the same closed window.
That is enough for transactional mail. It is not enough for a campaign, because the schedule is per process and a serverless platform runs many processes. Eight per second in each of four concurrent instances is 32 per second.
The second defence is to stop sending one email per request. The provider accepts up to 100 fully rendered emails in one batch request. A 500 recipient campaign then costs 5 requests instead of 500, and the rate limit stops being the constraint at all rather than being something you tiptoe around.
The interesting part is failure, not throughput
Batching is easy. Batching while keeping a per recipient send log honest is where the design lives.
Our send log is not a separate table. Each campaign mints one row per user, keyed uniquely on (user_id, campaign). That row is what makes the job idempotent: if the same person matches the audience window on two consecutive runs, the insert conflicts, the mint returns null, and they are skipped. The mint is the send log.
Which creates a problem the per recipient loop did not have. If you mint 100 rows and then fire one batch request, and that request fails, you have just recorded 100 emails that were never sent, permanently.
Three rules fix it:
- Mint one chunk at a time, not the whole campaign up front. Anything minted but not yet emailed is exposed if the function dies, so chunking bounds that exposure to a single batch of 100 instead of the entire audience.
- Send with permissive batch validation, and reconcile by index. One malformed recipient does not sink the other 99. The provider accepts the rest and reports rejects by their position in the payload, so the caller can discard exactly the rows that did not go out and leave the other 99 marked as sent:
const rejected = new Map(result.failures.map((f) => [f.index, f.message]));
for (const [index, pending] of chunk.entries()) {
if (rejected.has(index)) await discardUnsentCode(pending.codeId);
else sent++;
}
- Send an idempotency key derived from the batch itself. The key is built from the first minted row id in the chunk, which is unique per batch by construction. If the whole request is retried after a 429, the provider recognises the key and does not deliver the batch twice.
A whole request failure rolls every row in it back. A per recipient rejection rolls back exactly one. Neither leaves a user permanently marked as emailed when no email arrived, which is the only invariant that actually matters here.
Ordering is a feature, not an accident
The campaigns run sequentially in priority order, and each audience query excludes anyone who has been minted a code recently. That one property does a lot of work: a freshly minted higher priority code removes that user from every later campaign in the same run, so nobody gets two emails at once, and the higher value email wins.
The lowest priority campaign is the one that sells nothing. It sits last precisely because anyone who also qualifies for a conversion email should get that one instead. There is no second rule enforcing it. The ordering is the rule.
The clock
The route has a 300 second maximum duration. The job stops on its own 240 second budget, and it checks the deadline between chunks, never inside one.
That distinction matters for the same reason as rule 1 above. Being killed mid chunk means rows minted for emails that never went out. Stopping between chunks means the leftover candidates were never minted, so the next hourly run simply picks them up. The audience windows are wide enough (one to two days for most campaigns) that deferring an hour costs nothing.
See it
The lifecycle emails are documented in plain language on cogniprep.app/privacy. Search the page for "In the days after you sign up" and you will find the exact behaviour this job implements, including the fact that it is a small number of emails about CogniPrep only, and the provider it goes through. The same page lists which timestamp each email is anchored to, which is the audience selection above described from the other side.
If you want the receipt rather than the description, sign up and look at the footer of whatever arrives. Every marketing email carries a signed unsubscribe URL, and the token in it is per user, so one click works without a login.
The takeaway
If you have a loop that calls a rate limited API once per item, work out its actual iteration rate before you trust it. Ours was 13 per second against a cap of 10, written in the most conservative style available, and it quietly dropped mail for weeks.
Top comments (0)