The feature put a summary at the top of every long thread in a customer's inbox: three sentences, generated when the thread was opened and cached after that. It had been in production for five months, and it worked well enough that customers had started to mention it.
On a Tuesday in May the model provider had an overloaded afternoon and started answering a meaningful fraction of requests with 529. Our code caught the error, logged it and returned an empty summary, and the UI rendered an empty summary box. For four hours every customer who opened a long thread saw a grey rectangle with nothing in it. The on-call engineer saw an error rate graph that was up but not alarming, because the errors were being handled.
Nobody had decided what the feature should do when the model isn't there. The code had decided by default, and its default was to show the customer a broken feature and tell nobody.
The four things in this post all come from that afternoon, and none of them have anything to do with the model.
The switch
The first one is the most boring. The feature sits behind a flag, anyone on call can turn the flag off in under a minute without a deploy, and turning it off makes the feature disappear instead of degrade.
We did have a flag. It was the rollout flag from launch, which had been set to 100 percent afterwards and forgotten. Technically it could have been flipped, and flipping it would have hidden the feature, which is what we wanted. But nobody on call knew it existed, it wasn't in the runbook, and the on-call engineer didn't know hiding the feature was an option, because nobody had ever put it to them as one.
The kill switch is now its own flag, summaries.enabled. It's listed in the incident runbook under "things you can turn off", with a sentence next to it saying what the customer sees when it's off, and that sentence matters most. A kill switch is only useful if whoever holds it knows what happens when they use it. "The summary box does not render" is a perfectly good outcome, and it beats "the summary box renders empty" every time.
export async function threadSummary(threadId: string): Promise<Summary | null> {
if (!(await flags.enabled("summaries.enabled"))) return null;
// ... the rest
}
The UI treats null as "do not show the box".1
The fallback
The second is what the feature does when the model call fails and the switch is still on.
For a summary, the fallback I'd stand behind is the previous summary if there is one, and nothing if there isn't. A thread summarised yesterday that got two new messages today can show yesterday's summary with a small "may be out of date" marker. A thread that's never been summarised shows no box. Neither is a broken feature. Both are the feature being a bit less good, and that's all a fallback should be.
We also built the other option, a second provider. The summariser calls one model by default, and a different one from a different company when the first returns a 5xx or times out.2 It costs something: two API keys, two sets of rate limits, two bills, and running the eval against both models on every prompt change. For a feature customers have started to mention, that's worth paying. For an internal tool it wouldn't be.
So the order is: try the primary, if that fails try the secondary, if that fails return the cached previous summary, and if there's no cache return null. Each step was decided in daylight, in code, and not left to whatever the catch block happened to do at 3pm on a bad Tuesday.
The budget
The third is a limit, enforced in code, on how much the feature can spend in tokens and money per hour and per day.
Nobody thinks about this until the first surprising bill. Ours came from a different feature, a bulk export that called the model once per row, which a customer pointed at a table with 400,000 rows. It ran for six hours before anyone noticed, and that afternoon's bill was bigger than the feature's whole previous month.
The budget is a counter in Redis. Each response's token usage gets added to it, and it's checked before every call. When the hourly budget runs out, the feature falls back as if the model had failed and an alert fires. When the daily budget runs out, the switch flips off by itself and someone gets paged. A feature that's spent its daily budget by 11am is either far more popular than yesterday or being abused, and either way a human needs to look.
const budget = { hourly: 4_000_000, daily: 40_000_000 }; // tokens
async function withinBudget(): Promise<boolean> {
const [h, d] = await redis.mget(hourKey(), dayKey());
if (Number(d) >= budget.daily) { await flags.disable("summaries.enabled"); page("summaries daily budget"); return false; }
return Number(h) < budget.hourly;
}
The numbers come from the feature's real usage plus headroom, and we review them monthly. In normal running they don't save anything. They're there so an abnormal afternoon costs an abnormal afternoon's worth and not a quarter's.
The canary
The fourth would have turned four hours into twenty minutes. Every two minutes a synthetic request runs the whole feature path with a fixed input and checks the output.
It doesn't check whether the call succeeded. On the bad Tuesday the call did succeed, in the sense that it returned an error we handled. It checks whether the feature produced something that looks like a summary: not empty, under 400 characters, and mentioning a word from the fixed input. If the canary fails three times in a row it pages with the reason, which on the 529 afternoon would have been "provider returned 529", six minutes after the problem started.
It checks the fallback too. Once an hour the canary runs with the primary provider switched off on purpose and makes sure the secondary produced a summary. The one time that failed, the secondary provider's API key had expired. Otherwise we'd have found that out during the next primary outage, at the worst possible moment.
What this is not
None of this is prompt engineering or model quality.3 These four keep the feature working as a feature when the model, the provider or the usage pattern does something unexpected. They're the same four things you'd build around any external dependency: a way to turn it off, a way to degrade, a limit on what it can use up, and a probe that tells you when it's unwell.
My guess at why they get skipped for LLM features in particular is that the model feels like the hard part, so once the model works the feature feels done. But the model is a vendor API that sometimes returns 529, and it needs the same wrapping every other vendor API has earned.
Rolling out a model change behind the same switch
We built the switch, the fallback, the budget and the canary for outages. They turned out to be what we needed for something that happens much more often than an outage, which is changing the model.
A provider retires a model version. A newer one scores better on the eval set. A cheaper one scores nearly as well. Every one of those changes what customers see. Before this setup each was a deploy that either went fine or produced a Slack thread three days later saying the summaries "feel different". Nobody could say how, and the only thing to roll back to was the previous deploy.
Now a model change is a flag value with a percentage. summaries.model is a flag whose value is the model name, and a rollout means giving the new value to 5 percent of tenants, then 25, then everyone, over a week. The eval set gates the first step, so a model that scores below the current one doesn't get its 5 percent. The canary runs against both values the whole time, so a regression the eval set missed shows up as a canary failure in the 5 percent cohort before it reaches anyone else. The budget is per model, because a new model with a longer default output can double the token spend without any visible change.
Rolling back means setting the flag to the old value, which takes as long as the flag takes to propagate, under a minute. The previous deploy doesn't come into it.
The price is that the code has to handle two models at once, which mostly means the output parser can't rely on one model's formatting habits.4 And once a feature has two models it can have three, at no extra cost.
The runbook entry
All of the above is only worth as much as the person on call at 3am can find. Here's the entry in full, because a runbook that says "see the design doc" isn't a runbook.
Summaries. Flag summaries.enabled turns the feature off, the summary box disappears, customers see nothing broken. Flag summaries.model picks the model, current value in the flag UI. Provider errors are handled with a second provider, then a cached previous summary, then nothing. If the canary alert fires and the error is 5xx from the primary, do nothing for ten minutes and check that the fallback is producing summaries. If the fallback is also failing, turn the feature off. If the budget alert fires, look at the per tenant usage panel for one tenant doing something unusual, and if there is one, rate limit that tenant rather than turning the feature off for everyone. Re-enable by setting the flag back.
That's eight sentences. An on-call engineer who'd read it on the bad Tuesday would have turned the feature off in the first fifteen minutes and gone back to bed.
The thing that was actually hard
None of the four pieces took more than a day. The hard part was the decision under all of them: what should the customer see when the feature can't work? Nobody had ever asked. When we asked the product manager, it took a week to get an answer, because answering honestly meant admitting the feature was optional, and nobody had talked about it that way in the launch deck.
It is optional. Every LLM feature is, in the sense that the product worked before it existed, and on the day the model is unavailable the product has to work without it again. Writing that down in one sentence next to a flag name is the whole design, and the four pieces just implement it.
The checklist
Before an LLM feature goes to customers:
A kill switch, separate from the rollout flag and in the runbook, with a sentence saying what the customer sees when it's off.
A fallback chain decided in code: second provider, cached previous output, or nothing. null is a valid output and the UI knows what to do with it.
A budget in tokens per hour and per day, enforced before the call, that trips the switch and pages someone when it runs out.
A canary that checks the shape of the output instead of the status code, and runs the fallback on a schedule.
It's a day or two of work, and the grey rectangle doesn't come back.
Originally published at zeybek.dev.
-
That was a one line change in the component, and it should have been there from day one. ↩
-
Same prompt, same output format, and the eval set scores the second model about two points lower, which is fine for a fallback. ↩
-
The eval set takes care of quality, separately. ↩
-
The fallback provider had already forced that on us anyway. ↩
Top comments (0)