I'm a solo founder, and my product, an AI chat widget that qualifies leads on business websites, ran on a free-tier stack: Cloudflare Pages, Netlify Functions, Supabase, and Groq for the LLM. It worked well. Then Groq deprecated the Llama 3.3 70B model my whole product was built around.
Nothing dramatic broke at first. But I realized I had no plan for this, and that was the real problem.
What actually broke
My model name was hardcoded in a few places. The system prompts were tuned for that specific model's behavior. And the qualification flow depended on it returning clean structured output.
Swapping in a "similar" model wasn't a one-line change. Different models follow instructions differently, format JSON differently, and handle long system prompts differently. My prompts worked because I had unknowingly fitted them to one model's quirks.
Lesson 1: your prompts are coupled to your model, even if your code isn't.
Why I moved to OpenRouter
I wanted two things: no single provider that could pull the rug out from under me, and the ability to change models without redeploying half my app.
OpenRouter gives you one OpenAI-compatible API in front of many models and providers. For me, the migration was mostly a base URL and an env var:
// netlify/functions/chat.js
const MODELS = (process.env.LLM_MODELS || "").split(","); // primary first, then fallbacks
async function callLLM(messages) {
for (const model of MODELS) {
try {
const res = await fetch("https://openrouter.ai/api/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.OPENROUTER_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({ model, messages, temperature: 0.3 }),
});
if (!res.ok) throw new Error(`${model} failed: ${res.status}`);
const data = await res.json();
return data.choices[0].message.content;
} catch (err) {
console.error(err.message); // log it, then try the next model
}
}
throw new Error("All models failed");
}
Now the model list lives in an environment variable. If a model gets deprecated, I change a config value, not code.
Lesson 2: treat the model name as config, and always have a fallback.
The migration itself
[Describe what you actually did here: which model you moved to, how long it took, what broke in your prompts, and what you had to re-tune. Even one concrete example, like "the new model wrapped JSON in markdown fences and my parser choked", is worth more than a paragraph of general advice.]
Before switching fully, I ran a set of real conversations through both models and compared the lead qualification output. That test set is now the most valuable file in my repo. I'd build it earlier next time.
Lesson 3: keep a small regression set of real inputs. You can't judge a model swap by vibes.
Why I wouldn't run production on free models
This is the part I'd tell my past self. Free tiers and free models are great for building and validating. They are a risky foundation for the part of your product that customers rely on:
No stability promise. A free model can be rate limited harder, slowed down, or removed with little warning. That is what happened to me, and it's the business model working as designed.
Rate limits hit exactly when you succeed. Free limits are fine with 5 test users. A traffic spike or a customer's ad campaign is when you need capacity most.
No one owes you support. When something breaks at 2am, there's no SLA and no one to escalate to.
Data handling can differ. Free endpoints can have different logging or training terms than paid ones. If your users send customer conversations through your product, read those terms before you ship. [Verify the current terms for the specific model you use.]
For a product where each failed response is a lost lead for your customer, paying for inference is part of the cost of the product. The spend is small compared with the cost of an outage.
My rule now: free tier for prototypes and internal tools, paid and fallback-protected for anything a customer depends on.
What I'd do differently
Put the model name in config from day one.
Build a small test set of real conversations before launch.
Add a fallback model before I need one.
Pay for the inference layer as soon as paying customers depend on it.
If you want to see what I built with this, it's Zappiq AI, but the lessons above apply to any app built on an LLM API.
Has a model deprecation ever caught you off guard? I'd like to hear how you handled it.
Top comments (0)