DEV Community

Cover image for Shipping an LLM Feature Without Breaking Production
Diginatives LLC
Diginatives LLC

Posted on

Shipping an LLM Feature Without Breaking Production

Shipping an LLM Feature Without Breaking Production

The demo worked. Everyone clapped. Then you deployed it, and within a week something strange happened: a user pasted a 40-page document, the model returned a response in a format your parser had never seen, and an error page appeared where a helpful answer should have been.

Sound familiar? Language model features behave differently from normal code, and the gap between "works on my laptop" and "works for real users" is wider than usual. Here's the checklist I wish more teams used.

1. Treat the model like a flaky third-party service

Because it is one. Responses can be slow, rate-limited, or occasionally nonsense. So wrap every call with the basics:

  • A hard timeout (don't let a user stare at a spinner for a minute)
  • A small number of retries with backoff
  • A fallback path that still gives the user something useful

The fallback matters most. If the summary feature fails, show the original text. If the classifier fails, route the item to a human queue. A boring fallback beats a broken page every time.

2. Never trust the output shape

Even when you ask for JSON, you won't always get valid JSON. Validate before you use anything.

import json

REQUIRED = {"category", "confidence"}

def parse_reply(raw: str) -> dict | None:
    try:
        data = json.loads(raw)
    except json.JSONDecodeError:
        return None
    if not isinstance(data, dict) or not REQUIRED <= data.keys():
        return None
    return data

def classify(ticket_text: str) -> dict:
    raw = call_model(ticket_text)  # your model call, with timeout
    result = parse_reply(raw)
    if result is None:
        return {"category": "needs_review", "confidence": 0.0}
    return result
Enter fullscreen mode Exit fullscreen mode

Notice the last branch. When parsing fails, the code doesn't crash. It degrades gracefully and flags the item for review.

3. Build a tiny evaluation set early

You don't need a research lab. Collect 30 to 50 real examples with the answers you'd consider correct, and store them in a file. Run them every time you change the prompt, the model or the settings.

Why bother? Because prompt tweaks are sneaky. You fix one case and quietly break three others. Without a test set, you'll never notice until a user does.

Keep the set honest, too. Include the ugly inputs: typos, mixed languages, empty fields, absurdly long text.

4. Log enough to debug, not enough to leak

When something goes wrong, you'll want to know what went in and what came out. Log the prompt version, model name, latency, token counts and a request ID. For the actual text, be careful. User content can contain names, phone numbers and other private details. Redact or hash what you can, and set a retention limit.

Debuggability and privacy pull in opposite directions. Decide the trade-off deliberately, not by accident.

5. Put the feature behind a flag

Ship it dark. Turn it on for internal staff first, then 5% of users, then everyone. And keep a kill switch that disables the model path instantly without a redeploy.

You'll be glad of it the first time a provider has an outage or a prompt change goes sideways at 2 a.m.

6. Watch the bill

Model calls cost money per request, and costs scale with traffic and input length. Add a few cheap guards:

  • Cap input length and truncate sensibly
  • Cache repeated requests
  • Set a daily spending alert

It's far nicer to receive an alert than a surprise invoice.

7. Be honest in the interface

Tell users when text is machine-generated, and give them an easy way to correct or reject it. A "this isn't right" button doubles as free evaluation data. Over a few weeks, those corrections become your best test cases.

Who should own this?

Whether you're working solo or inside a software development company Lahore clients rely on, decide who owns the feature after launch. Prompts and models drift. Somebody needs to run the evaluation set monthly and review the logs. If nobody owns it, quality decays slowly and silently.

And if the feature involves heavier work, such as fine-tuning, retrieval pipelines or custom data handling, it may be worth bringing in an AI development company Lahore teams have worked with before, rather than learning every lesson the expensive way.

The short version

  • Timeouts, retries and fallbacks on every call
  • Validate output before using it
  • Keep a small evaluation set and run it often
  • Log carefully, protect user data
  • Feature flag with a kill switch
  • Cap input size and watch costs
  • Let users correct the output

None of this is glamorous. All of it is what separates a demo from a product.

What would you add to the list? Drop it in the comments.

Top comments (0)