Try this:
curl -s https://munchable.app/api/cron/indexnow
# {"error":"unauthorized"}
That is one of our two scheduled endpoints, answering a stranger. It is a 401, not a 404, and it is not quietly doing the work either. This post is about the two guards behind that response, and about the afternoon our scheduler reported success every fifteen minutes while nothing happened for three hours.
Two guards, because a scheduled route is a public route
A cron job on a serverless platform is an HTTP endpoint with a timer pointed at it. Which means it is reachable by anyone who guesses the path, and it can be invoked twice at once. Every one of our scheduled routes starts with the same two lines of defence.
The first is authentication, and the only interesting decision in it is what to do when the secret is not configured at all:
export function cronAuthStatus(authHeader: string | null, secret: string | undefined): CronAuth {
if (!secret || secret.length < 16) return 'unset';
const presented = authHeader?.startsWith('Bearer ') ? authHeader.slice('Bearer '.length).trim() : '';
if (!presented) return 'unauthorized';
const a = Buffer.from(presented, 'utf8');
const b = Buffer.from(secret, 'utf8');
if (a.length !== b.length) return 'unauthorized';
return timingSafeEqual(a, b) ? 'ok' : 'unauthorized';
}
Three states, not two. A wrong token is a 401. A missing configuration is a 503 and the job refuses to run. The tempting alternative, running unauthenticated when no secret is set, is how a job that spends money ends up invocable by a search crawler in a preview environment nobody remembered deploying.
The comparison is constant time, with a length check first because the comparison function throws on mismatched lengths. Yes, that leaks the length of the secret. A secret whose length is its security is not a secret.
The second guard, and the day it went wrong
The other guard is an overlap lock: a Redis key called cron:<name>, set with NX, so a slow run still going when the next tick fires skips instead of doubling. Our curation runner also shares that lock with its command line equivalent, so an operator running it by hand cannot collide with the schedule.
The original version set the key's TTL to the job's whole worst-case runtime. Three hours on the scheduled runner, six from the CLI. The reasoning was airtight: the lock must outlive the run, so the TTL has to be at least as long as the longest run.
It does outlive the run. The cost landed somewhere else.
A run that is cancelled or killed never reaches its finally, so it never deletes the key. The key then sat there for its full three hours. The runner, on a fifteen-minute schedule at the time, kept firing into it and getting:
{ "skipped": true, "reason": "locked" }
Which is an HTTP 200. Which the platform records as a successful run.
So the dashboard was green every quarter of an hour, twelve consecutive times, while absolutely nothing was being curated. We watched that happen on 14 September 2026 after two runs had been cancelled by hand.
The fix is to stop calling it a lock
The lock is not a timeout on the work. It is a lease, and the running job renews it:
/** How long a held lock stays valid without being renewed. */
export const LOCK_LEASE_S = 300;
/** Renewed comfortably inside the lease, so one failed renewal is survivable. */
export const RENEW_EVERY_MS = 90_000;
Five minutes of validity, renewed every ninety seconds. Three renewals per lease, so a single failed renewal is a non-event. A job that is alive holds the lease for as long as it likes, because it keeps saying so. A job that dies loses it within five minutes instead of within its worst-case runtime.
const renew = setInterval(() => {
void client.expire(key, lease).catch(() => {
// A blip is survivable: the next tick renews, and the lease is three
// ticks long. Losing the lock entirely only risks an overlap, which is
// the lesser failure against curation stopping for hours.
});
}, RENEW_EVERY_MS);
renew.unref?.();
Two details in there that are worth stealing.
unref(), so a pending renewal timer can never be the reason a worker process stays alive after the job is finished. An interval that outlives its purpose is a process that will not exit, and on a platform that bills by duration that is a bill.
And the comment in the catch, which is the actual engineering content of the whole change. Losing the lease risks an overlap. Holding a dead lease stops the work entirely. Those are both bad, they are not equally bad, and the code says which one it prefers and why. A silent catch that does not name the failure it is choosing is a decision somebody will have to re-derive at the worst possible time.
The finally still deletes the key on the happy path. If that delete fails, the lease releases it in five minutes rather than three hours, which is the difference between a blip and an outage.
The signature that invited the bug
/**
* `leaseSeconds` is exposed for tests. Callers should not pass it: a caller
* that thinks it is declaring "how long my job may run" is the bug this
* signature used to invite.
*/
export async function withCronLock<T>(
name: string,
fn: () => Promise<T>,
opts: { leaseSeconds?: number } = {},
): Promise<CronLockResult<T>>
This is my favourite part of the change, because the original bug was not in a line of logic. It was in a parameter name. A ttlSeconds next to a job meant every caller reasoned about "how long my job runs", and every caller reached the same wrong conclusion independently. The parameter was an invitation, and the fix included un-inviting it.
No Redis means no run
const client = redis;
if (!client) return { ran: false, reason: 'lock_unavailable' };
If the lock store is unreachable, the job does not run. Not "runs without a lock, just this once". These jobs write curated data and make paid model calls, and an unguarded double run of that is worse than a missed run of it.
Which is the same shape as the 503 on a missing secret. Both say: the guard is not optional decoration around the work, it is part of the work.
The reporting lesson, which cost us more than the code
The code change here is twenty lines. The lesson is bigger: a skipped run must not be indistinguishable from a successful one.
Your scheduler only knows HTTP status. It cannot tell "did the thing" from "politely declined to do the thing", and it will paint both green. So the distinction has to exist in something you actually watch: a summary that counts what was done rather than reporting that an invocation happened, and an alert on consecutive skips rather than on errors. Twelve green ticks in a row told us everything was fine, and they were technically correct every single time.
Check the live endpoints
-
curl -s https://munchable.app/api/cron/indexnowgives you the 401 above. Our other scheduled route, the monthly contributor-reward settlement, answers the same way. - munchable.app/sitemap.xml is what the hourly job feeds. It is the visible output of the thing that was silently not happening.
- The search-submission side of that job has its own story: our IndexNow key sits in the repo on purpose. And for a different flavour of a job that failed silently every single run, the bug was a Date object.
Top comments (0)