DEV Community

Shanthanu C
Shanthanu C

Posted on

From AI Demo to Production: 6 Things We Add Before Shipping an LLM Feature

Getting an LLM feature to work in a demo takes an afternoon. Getting it to keep working when real users, real data and real traffic arrive is a different job.

Most AI pilots I see stall at the same point: the prototype calls a model API directly from a route handler, and everything "works" until the first timeout, malformed response or surprise bill.

Here are six things worth adding before you ship. Examples use Node.js/TypeScript, but the ideas apply to any stack (Python, .NET, Java).

  1. Put a timeout and retry policy around every model call

Model APIs are network calls. They will be slow or fail sometimes. Never let a request hang forever.

ts
import OpenAI from "openai";

const client = new OpenAI({
apiKey: process.env.OPENAI_API_KEY,
timeout: 20_000, // fail fast instead of hanging the request
maxRetries: 2, // retries on transient errors (429, 5xx)
});

Decide up front what the user sees when the call fails. A clear fallback message beats a spinner that never ends.

  1. Never trust the output: validate it

If your app expects JSON, parse and validate it. Models occasionally return extra text, missing fields or wrong types.

ts
import { z } from "zod";

const SummarySchema = z.object({
title: z.string().min(1).max(120),
bullets: z.array(z.string()).min(1).max(5),
sentiment: z.enum(["positive", "neutral", "negative"]),
});

export async function summarize(text: string) {
const res = await client.chat.completions.create({
model: process.env.LLM_MODEL!,
response_format: { type: "json_object" },
messages: [
{
role: "system",
content:
"Return JSON with keys: title, bullets (array of up to 5 strings), sentiment (positive|neutral|negative).",
},
{ role: "user", content: text },
],
});

const raw = res.choices[0]?.message?.content ?? "{}";
const parsed = SummarySchema.safeParse(JSON.parse(raw));

if (!parsed.success) {
// log it, retry once, or fall back, but never pass bad data downstream
throw new Error("Model returned invalid structure");
}
return parsed.data;
}

Treat model output like user input: untrusted until validated.

  1. Cache what you can

Many requests are repeated or near-identical. A simple cache cuts both latency and cost.

ts
import { createHash } from "node:crypto";

const cache = new Map(); // use Redis in production

function keyFor(input: string) {
return createHash("sha256").update(input).digest("hex");
}

export async function cachedSummarize(text: string) {
const key = keyFor(text);
if (cache.has(key)) return cache.get(key);

const result = await summarize(text);
cache.set(key, result);
return result;
}

Swap the Map for Redis with a TTL once you run more than one instance.

  1. Log the right things (and protect the wrong ones)

You can't improve what you can't see. For every call, log:

model name and prompt version
latency
token usage
validation pass/fail
an anonymous request ID

Do not log raw user content or personal data by default. Decide what you're allowed to store, mask it, and set a retention period.

  1. Set cost and rate guardrails

A bug or a bad actor can burn through your budget fast. Add:

per-user and per-IP rate limits
a maximum input length before the call is made
a max_tokens cap on responses
a monthly spend alert at your provider
ts
if (text.length > 8_000) {
throw new Error("Input too long");
}

  1. Version your prompts and test them

Prompts are code. Keep them in files, give them versions, and run a small regression set whenever you change one.

ts
// prompts/summarize.v2.ts
export const SUMMARIZE_PROMPT_V2 = ...;

A set of 20-30 representative inputs with expected properties (valid JSON, correct sentiment, no empty fields) catches most regressions before users do. Run it in CI next to your normal tests.

A quick pre-launch checklist

  • Timeouts and retries on every model call
  • Output validated against a schema
  • Cache for repeated inputs
  • Structured logs without sensitive data
  • Rate limits, input caps and spend alerts
  • Versioned prompts with a small test set
  • A user-facing fallback when the model is unavailable

Wrapping up

None of this is glamorous, but it's what separates an impressive demo from a feature your team can trust. Start with validation and timeouts, since those two prevent most of the pain.

I'm part of the team at CodvikScribe, where we build custom software, mobile apps and AI-powered products. If you're working on an AI feature and want a second pair of eyes, you can reach us through our contact page.

What's the biggest thing that broke when you took an LLM feature to production? Tell me in the comments.

Devto post ai pilot to production
MD
Web search

Top comments (0)