CogniPrep has a health check at /api/health. It has no authentication, no rate limit and no secret query parameter. That is deliberate: an uptime monitor that needs credentials is a thing you turn off during an incident, which is the exact moment you wanted it.
Open it right now:
https://cogniprep.app/api/health
Today it says this:
{
"status": "healthy",
"timestamp": "2026-09-28T08:04:00.782Z",
"checks": {
"database": { "status": "healthy", "responseTime": 67 },
"redis": { "status": "configured" },
"environment": { "status": "healthy" },
"monitoring": { "status": "configured" }
}
}
Four statuses and one number. That is the whole payload, and the interesting part is everything that is not in it.
A public endpoint is a reconnaissance surface
The useful debugging information in a health check is exactly the information you would want if you were probing someone else's stack. Two places in this route had to be written twice because the first version was helpful to the wrong audience.
The first is the database check. A failed query has an error message attached, and postgres driver errors are chatty: host names, role names, pooler internals, sometimes a truncated statement.
} catch (error) {
// Log the detail for us; return only the status to the caller. Database error
// messages can disclose host names, roles and driver internals.
logError('[health] Database check failed', error);
checks.database = { status: 'unhealthy' };
isHealthy = false;
}
The detail still exists. It goes to the logger, which goes to Sentry, which is where the person fixing it is looking anyway. The caller gets one word.
The second is the environment check, and it is the one I would have got wrong by default:
const requiredEnvVars = [
'DATABASE_URL',
'STRIPE_SECRET_KEY',
'STRIPE_WEBHOOK_SECRET',
'CRON_SECRET',
];
const missingEnvVars = requiredEnvVars.filter((v) => !process.env[v]);
// Report only the count. Naming the missing variables tells an anonymous caller
// exactly which part of our configuration is absent, e.g. that
// STRIPE_WEBHOOK_SECRET is unset, which implies webhook verification is broken.
checks.environment = {
status: missingEnvVars.length === 0 ? 'healthy' : 'unhealthy',
missingCount: missingEnvVars.length > 0 ? missingEnvVars.length : undefined,
};
{"missingCount": 1} tells an operator to go and look at the deploy. {"missing": ["STRIPE_WEBHOOK_SECRET"]} tells a stranger that our webhook signature verification is currently unenforced, which is an invitation to post a fake checkout.session.completed at us. The names go to the log line, not the response body.
Same reasoning for Redis and Sentry: the response says configured or not_configured, never a URL, never a project, never a DSN.
Degraded is not down
The database check times itself and grades the result:
const dbResponseTime = Date.now() - dbStart;
checks.database = {
status: dbResponseTime > 1000 ? 'degraded' : 'healthy',
responseTime: dbResponseTime,
};
degraded still returns HTTP 200. Only a real failure, a thrown query or a missing required variable, returns 503.
That split exists because a monitor that pages on slowness gets muted, and a muted monitor does not page on outages either. A one second query on a pooled Postgres connection that has just cold started is not an outage, it is a Monday. The number is published so a human can watch it drift; the status code is reserved for "a thing is actually broken".
responseTime is the one piece of internal telemetry the endpoint does hand out. It is a duration with no identifiers attached, and it is the single most useful thing to have in a screenshot when someone reports that the site feels slow.
The five second timer that has to be cleared
The database check is a race, because a hanging connection should fail the check rather than hang the function until the platform kills it:
let dbTimeout: ReturnType<typeof setTimeout> | undefined;
try {
// The timer is cleared in `finally` - otherwise a fast query still leaves a
// pending 5s timeout holding the serverless function open.
await Promise.race([
db.execute('SELECT 1'),
new Promise((_, reject) => {
dbTimeout = setTimeout(() => reject(new Error('Database timeout')), 5000);
}),
]);
...
} finally {
if (dbTimeout) clearTimeout(dbTimeout);
}
Promise.race settles as soon as the first promise settles, but it does not cancel the loser. When the query comes back in 67ms, that setTimeout is still armed. On a serverless runtime the invocation does not finish while a timer is pending, so a 67ms health check bills like a five second one, every single time the monitor calls it. At one call a minute that is a quiet, permanent, entirely self-inflicted cost.
clearTimeout in a finally is four lines. It is also the difference between a health check that costs nothing and one that is the most expensive route in the app.
It must never be cached
The last trap is framework shaped. Next.js will happily prerender a route handler that looks static, and a cached health check is worse than no health check: it reports the state of the world at build time, with a timestamp that makes it look live.
export const dynamic = 'force-dynamic';
export const revalidate = 0;
and on the response:
headers: { 'Cache-Control': 'no-cache, no-store, must-revalidate' }
Both are needed. The route segment config stops the framework and the CDN caching it; the response header stops everything between us and the monitor doing the same.
See it
- cogniprep.app/api/health: the live response. Count the fields. There is no host, no version, no library, no region, no variable name.
- Open it in DevTools and look at the response headers:
cache-control: no-cache, no-store, must-revalidate, and nox-vercel-cache: HITno matter how many times you reload. - Reload it a few times and watch
checks.database.responseTimemove. That number is the check doing real work rather than reading a cached answer.
The general rule I would now apply to any unauthenticated status route: every field has to justify itself to an audience of strangers. Statuses, durations and counts pass. Names, hosts, versions and error strings go to the log.
Top comments (1)