DEV Community

developerz.ai
developerz.ai

Posted on

Building a Production-Ready AI Integration with Large Language Models

Building a Production-Ready AI Integration with Large Language Models

Integrating a large language model (LLM) into a customer-facing product is more than a simple API call. The code below reliable a set of practical steps that keep latency low, costs predictable, and output reliable.

1. Choose the right model and endpoint

Select a model that matches the required context length and token budget. For most SaaS features a 7-billion-parameter model offers a good balance between quality and cost. Use the provider’s hosted endpoint with a dedicated API key to isolate traffic.

import requests

API_URL = "https://api.provider.com/v1/completions"
HEADERS = {"Authorization": f"Bearer {YOUR_API_KEY}"}

def generate(prompt: str) -> str:
    payload = {
        "model": "gpt-7b",
        "prompt": prompt,
        "max_tokens": 200,
        "temperature": 0.7,
    }
    response = requests.post(API_URL, json=payload, headers=HEADERS)
    response.raise_for_status()
    return response.json()["choices"][0]["text"]
Enter fullscreen mode Exit fullscreen mode

2. Prompt engineering for consistency

Write a prompt that includes a short system instruction, the user query, and a clear request for format. Keep the instruction under 50 tokens to reduce latency.

You are a helpful assistant that answers technical questions about cloud infrastructure. Provide a concise answer in two sentences. If the question is outside the scope, reply with "I am not able to help with that."
Enter fullscreen mode Exit fullscreen mode

3. Caching frequent queries

Many SaaS applications see repeated requests for the same information. Store the hash of the prompt and the model response in a fast key-value store such as Redis with a TTL of one hour.

import hashlib, redis
r = redis.Redis(host="localhost", port=6379, db=0)

def cached_generate(prompt: str) -> str:
    key = hashlib.sha256(prompt.encode()).hexdigest()
    cached = r.get(key)
    if cached:
        return cached.decode()
    answer = generate(prompt)
    r.setex(key, 3600, answer)
    return answer
Enter fullscreen mode Exit fullscreen mode

4. Latency monitoring and timeout handling

Wrap the API call in a timeout and record the duration. If the request exceeds 2 seconds, fall back to a static answer or a simplified rule-based response.

import time

def safe_generate(prompt: str) -> str:
    start = time.time()
    try:
        answer = cached_generate(prompt)
    except requests.exceptions.Timeout:
        answer = "Please try again later."
    elapsed = time.time() - start
    # Log elapsed to monitoring system
    print(f"LLM latency: {elapsed:.2f}s")
    return answer
Enter fullscreen mode Exit fullscreen mode

5. Cost tracking

Log token usage for each request. Most providers expose prompt_tokens and completion_tokens in the response. Aggregate these metrics daily to detect spikes.


def generate_with_metrics(prompt: str):
    payload = {"model": "gpt-7b", "prompt": prompt, "max_tokens": 200}
    resp = requests.post(API_URL, json=payload, headers=HEADERS, timeout=5)
    data = resp.json()
    usage = data.get("usage", {})
    # Store usage metrics
    print(f"Prompt tokens: {usage.get('prompt_tokens')}, Completion tokens: {usage.get('completion_tokens')}")
    return data["choices"][0]["text"]
Enter fullscreen mode Exit fullscreen mode

6. Security and data privacy

Never send raw user data to the LLM. Strip personally identifiable information and apply a whitelist of allowed fields before constructing the prompt.


def sanitize(user_input: str) -> str:
    # Simple example: remove email addresses
    import re
    return re.sub(r"[\w\.-]+@[\w\.-]+", "[redacted]", user_input)
Enter fullscreen mode Exit fullscreen mode

7. Deployment checklist

  • Verify the API key is stored in a secret manager.
  • Enable request retries with exponential backoff.
  • Run load tests with a realistic mix of short and long prompts.
  • Set up alerts for latency > 2 seconds or error rate > 1 %.

By following these steps you can move from a proof-of-concept script to a production-grade AI feature that respects latency budgets, cost constraints, and user privacy. need this built? developerz.ai #ai #devops

Top comments (0)