Building a Production-Ready AI Integration with Large Language Models
Integrating a large language model (LLM) into a customer-facing product is more than a simple API call. The code below reliable a set of practical steps that keep latency low, costs predictable, and output reliable.
1. Choose the right model and endpoint
Select a model that matches the required context length and token budget. For most SaaS features a 7-billion-parameter model offers a good balance between quality and cost. Use the provider’s hosted endpoint with a dedicated API key to isolate traffic.
import requests
API_URL = "https://api.provider.com/v1/completions"
HEADERS = {"Authorization": f"Bearer {YOUR_API_KEY}"}
def generate(prompt: str) -> str:
payload = {
"model": "gpt-7b",
"prompt": prompt,
"max_tokens": 200,
"temperature": 0.7,
}
response = requests.post(API_URL, json=payload, headers=HEADERS)
response.raise_for_status()
return response.json()["choices"][0]["text"]
2. Prompt engineering for consistency
Write a prompt that includes a short system instruction, the user query, and a clear request for format. Keep the instruction under 50 tokens to reduce latency.
You are a helpful assistant that answers technical questions about cloud infrastructure. Provide a concise answer in two sentences. If the question is outside the scope, reply with "I am not able to help with that."
3. Caching frequent queries
Many SaaS applications see repeated requests for the same information. Store the hash of the prompt and the model response in a fast key-value store such as Redis with a TTL of one hour.
import hashlib, redis
r = redis.Redis(host="localhost", port=6379, db=0)
def cached_generate(prompt: str) -> str:
key = hashlib.sha256(prompt.encode()).hexdigest()
cached = r.get(key)
if cached:
return cached.decode()
answer = generate(prompt)
r.setex(key, 3600, answer)
return answer
4. Latency monitoring and timeout handling
Wrap the API call in a timeout and record the duration. If the request exceeds 2 seconds, fall back to a static answer or a simplified rule-based response.
import time
def safe_generate(prompt: str) -> str:
start = time.time()
try:
answer = cached_generate(prompt)
except requests.exceptions.Timeout:
answer = "Please try again later."
elapsed = time.time() - start
# Log elapsed to monitoring system
print(f"LLM latency: {elapsed:.2f}s")
return answer
5. Cost tracking
Log token usage for each request. Most providers expose prompt_tokens and completion_tokens in the response. Aggregate these metrics daily to detect spikes.
def generate_with_metrics(prompt: str):
payload = {"model": "gpt-7b", "prompt": prompt, "max_tokens": 200}
resp = requests.post(API_URL, json=payload, headers=HEADERS, timeout=5)
data = resp.json()
usage = data.get("usage", {})
# Store usage metrics
print(f"Prompt tokens: {usage.get('prompt_tokens')}, Completion tokens: {usage.get('completion_tokens')}")
return data["choices"][0]["text"]
6. Security and data privacy
Never send raw user data to the LLM. Strip personally identifiable information and apply a whitelist of allowed fields before constructing the prompt.
def sanitize(user_input: str) -> str:
# Simple example: remove email addresses
import re
return re.sub(r"[\w\.-]+@[\w\.-]+", "[redacted]", user_input)
7. Deployment checklist
- Verify the API key is stored in a secret manager.
- Enable request retries with exponential backoff.
- Run load tests with a realistic mix of short and long prompts.
- Set up alerts for latency > 2 seconds or error rate > 1 %.
By following these steps you can move from a proof-of-concept script to a production-grade AI feature that respects latency budgets, cost constraints, and user privacy. need this built? developerz.ai #ai #devops
Top comments (0)