I run a small SaaS called VoiceDash. It gives voice AI agencies a white label portal for their clients. Next.js on Vercel, Postgres behind Prisma, Stripe for billing.
A few weeks ago I wrote about how five code paths write the same subscription row on purpose. This is the follow-up nobody warns you about: what happens to all those carefully synced rows when you point the app at a different Stripe account.
Short version: every Stripe id in my database turned into a guess. And the first helper I wrote to deal with that had a catch block that was quietly deciding who got a free trial.
Ids are scoped to an account
This sounds obvious once you say it. A Stripe customer id like cus_... and a subscription id like sub_... only mean something inside the account that created them. Swap the secret key for one from another account and those same strings resolve to nothing.
My workspace table stores a stripeCustomerId. The subscription table stores a stripeSubscriptionId. Both were written against the old account. So the moment the new key goes live:
- Checkout gets passed a
customerthat does not exist, and fails with "No such customer". - The billing portal does the same thing.
- The plan change route tries to retrieve a subscription that is not there and throws.
Net effect: every existing workspace is unable to subscribe at all. Not a degraded experience. A wall.
The first fix, and why it was wrong
The obvious move is to check a stored id against the current account before using it, and if it is not there, clear it and treat the workspace as a new customer. So that is what I wrote:
try {
const customer = await getStripe().customers.retrieve(storedCustomerId);
if ((customer as any)?.deleted) throw new Error("customer deleted");
return storedCustomerId;
} catch {
await prisma.workspace
.update({ where: { id: workspaceId }, data: { stripeCustomerId: null } })
.catch(() => {});
return null;
}
It worked. I tested it against a stale id, the id got cleared, checkout created a fresh customer. Shipped.
The next morning I reread it and did not like the bare catch.
That block does not say "if the customer is missing". It says "if anything at all goes wrong". A network blip. A rate limit. A Stripe 5xx. A timeout on a slow cold start. Every one of those lands in the same branch as "this customer does not exist here", and that branch deletes the customer id.
Now follow the null downstream. In the checkout route:
// A trial is per Stripe account, so a customer we cannot see here has not
// used one and is offered the trial like any new signup.
const isReturningCustomer = !!customerId;
A workspace with no customer id is a first-timer. First-timers get a seven day trial. So one transient error during checkout would have detached a paying customer from their billing record and offered them a second free trial on a brand new customer object. The subscription lookup had the same shape: swallow the error, return null, fall through to a fresh checkout, and start a second subscription for someone who already had one.
None of that throws. None of it pages anyone. It just makes billing state wrong in a way that looks exactly like billing state being right.
Only a confirmed miss counts as a miss
The rewrite splits errors into two kinds: Stripe explicitly telling me the object is not in this account, and everything else.
function isResourceMissing(err: unknown): boolean {
const e = err as { code?: string; type?: string; statusCode?: number };
if (e?.code === "resource_missing") return true;
return e?.type === "StripeInvalidRequestError" && e?.statusCode === 404;
}
Then the resolver only clears the id on a confirmed miss:
let customer;
try {
customer = await getStripe().customers.retrieve(storedCustomerId);
} catch (err) {
if (!isResourceMissing(err)) {
console.error(
`[stripe] could not verify customer ${storedCustomerId} for workspace ${workspaceId}; keeping it as-is`,
(err as any)?.message
);
return storedCustomerId;
}
return clearStaleCustomer(workspaceId, storedCustomerId);
}
// A deleted customer still resolves but cannot be reused for checkout.
if ((customer as any)?.deleted) {
return clearStaleCustomer(workspaceId, storedCustomerId);
}
return storedCustomerId;
The interesting line is return storedCustomerId on an ambiguous failure. I am knowingly handing a possibly broken id to the caller. If Stripe really is down, the checkout call right after will fail too, and the user sees an error and tries again. That is the worst case, and it is the same worst case the app had before any of this code existed. It is a bad minute, not corrupted data.
The subscription lookup goes one step further and rethrows:
export async function getSubscriptionIfInAccount(subscriptionId: string) {
try {
return await getStripe().subscriptions.retrieve(subscriptionId);
} catch (err) {
if (isResourceMissing(err)) {
console.warn(
`[stripe] subscription ${subscriptionId} is not in the current Stripe account; ignoring it`
);
return null;
}
throw err;
}
}
Returning the stored id would not help here, because the caller wants the subscription object itself. So an ambiguous failure becomes a 500 on the plan change request instead of a silent second subscription.
The rule I took away: when a function's "not found" answer triggers a destructive or generous action, not found has to be proven, not inferred from failure. Clearing a column is destructive. Offering a trial is generous. Both were hanging off a catch that could not tell the difference between "gone" and "busy".
The other thing hiding in the same checkout
While I was in there I found why a checkout could complete in Stripe and still drop the customer back on the plan picker.
Newer Stripe API versions moved current_period_end off the subscription object and onto its items. Three places in my code (the onboarding page, the subscribe page and the verify route) read it off the subscription, got undefined, multiplied by 1000 and wrote new Date(NaN) into a non-nullable column. The upsert threw. The money side had already succeeded.
It is the same family of bug as the catch-all. A value that is technically present in the type system, an Invalid Date, standing in for "I do not know". The fix reads every place the answer could live and refuses to produce garbage:
const candidates: unknown[] = [
s.current_period_end,
s.items?.data?.[0]?.current_period_end,
s.trial_end,
];
const ts = candidates.find(
(v): v is number => typeof v === "number" && Number.isFinite(v) && v > 0
);
if (ts) return new Date(ts * 1000);
const days = billingCycle === "annual" ? 365 : 30;
return new Date(Date.now() + days * 24 * 60 * 60 * 1000);
The projected fallback is deliberately temporary. The read time reconcile from the earlier post rewrites the row from Stripe on a later page load, so a projected date gets corrected on its own. What it never does is write a value that makes the whole upsert fail.
Those three copies had also drifted apart from each other, which is part of why the bug was hard to see. They are one shared syncCheckoutSession function now.
Smaller things the move shook loose
-
A lazy Stripe client. Routes used to construct
new Stripe(...)at import time, which throws during a build in any environment without the key. Now there is onegetStripe()that builds the client on first use, withmaxNetworkRetries: 3and a 30 second timeout. Retries on the client are exactly what makes the "ambiguous failure" branch rarer, but they do not make it impossible, so the error classification still has to be right. -
Return URLs from the request. Checkout, portal and billing emails were building links from
NEXT_PUBLIC_APP_URL, and wherever that was unset they producedundefined/agency/subscribe. The helper now prefers the env var, then the request host, thenVERCEL_URL, and ignores a localhost value that leaked into a Vercel deployment. - The portal says what is true. A workspace whose customer lives in the old account now gets "No billing account found. Please subscribe to a plan first." instead of a raw Stripe error.
What I would tell myself before the move
Before you switch Stripe keys, grep for every column that stores a Stripe id and ask what happens when that id stops resolving. Then look at every catch around the retrieve calls and ask a meaner question: what does this branch do when the network hiccups instead?
In my case the answer was "gives away a free trial and forgets who paid". The fix was not more retries or a smarter fallback. It was being precise about which error means which thing, and letting the ambiguous ones fail loudly.
If you are building a Stripe-backed SaaS and want to see what the portal side of VoiceDash looks like, it lives at voice-dash.com.
Top comments (0)