DEV Community

jareer nauman
jareer nauman

Posted on Originally published at keencraft.tech

Why Your AI Demo Is Fast and Production Is Slow (and the 4 Fixes That Worked)

The demo was fast. Production was not.

In a typical AI demo, one request comes back in about 400ms. It feels instant. Then real users arrive, and the same endpoint takes 3 to 4 seconds at p95 during peak hours, with occasional timeouts.

Nothing is broken. The code does exactly what it did in the demo. The demo just never asked it to do that work under real traffic.

After hardening a few AI products that hit this wall, including the backends behind the voice platforms I lead development on, I keep finding the same four causes:

Sequential calls that should run in parallel
A cache that is cold, or does not exist
A database that was never asked to serve real traffic
Retrieval that is fast in testing and slow (or wrong) at scale

Here is each one, with the code change that fixed it. Examples are Python with FastAPI, Redis, and Postgres, but the patterns carry to any stack.

Fix 1: Stop waiting in line

The most common demo pattern is a chain of awaits. Each one is fine alone. Together they add up.

python
# Before: three independent calls, run one after another
async def build_context(user_id: str, query: str):
    profile = await crm.get_contact(user_id)      # ~250ms
    history = await db.get_recent_calls(user_id)  # ~180ms
    docs = await retriever.search(query)          # ~300ms
    return profile, history, docs                 # total: ~730ms

Enter fullscreen mode Exit fullscreen mode

None of those calls depends on the others. So run them together, and put a timeout on the whole group so one slow service cannot hold every request hostage.

python
import asyncio

# After: the same calls, concurrently, with a ceiling
async def build_context(user_id: str, query: str):
    profile, history, docs = await asyncio.wait_for(
        asyncio.gather(
            crm.get_contact(user_id),
            db.get_recent_calls(user_id),
            retriever.search(query),
        ),
        timeout=1.5,
    )
    return profile, history, docs                 # total: ~300ms, the slowest call
Enter fullscreen mode Exit fullscreen mode

The latency drops from the sum of the calls to the slowest single call. In a voice agent, where every extra 100ms is a pause the caller hears, this one change often matters more than switching models.

One caution: parallel calls multiply load on your downstream services. If three endpoints each fan out to the CRM, check its rate limits before you ship.

Fix 2: Warm the cache (or build one)

In a demo you hit the same few records over and over, so everything feels cached even when nothing is. In production, thousands of users each ask for different data, and every request goes all the way to the source.

A read-through cache in Redis handles the data that changes rarely but gets read constantly: business hours, service lists, agent configs, calendar availability windows.

python
import json
import redis.asyncio as redis

cache = redis.Redis(host="localhost", port=6379, decode_responses=True)

async def get_agent_config(workspace_id: str) -> dict:
    key = f"agent_config:{workspace_id}"
    cached = await cache.get(key)
    if cached:
        return json.loads(cached)

    config = await db.fetch_agent_config(workspace_id)
    await cache.set(key, json.dumps(config), ex=300)  # 5-minute TTL
    return config

async def update_agent_config(workspace_id: str, data: dict):
    await db.save_agent_config(workspace_id, data)
    await cache.delete(f"agent_config:{workspace_id}")  # invalidate on write
Enter fullscreen mode Exit fullscreen mode

Two details matter more than the cache itself. First, invalidate on write, so an owner who changes their hours does not see the agent quote the old ones for five minutes. Second, key everything by workspace or tenant. A cache key without the tenant in it is how one client's config ends up in another client's call.

After launch, warm the cache for your busiest tenants on deploy, so the first callers after a release do not pay the cold-start penalty.

Fix 3: Ask the database what it is actually doing

With 200 demo rows, a full table scan takes a millisecond. With 2 million rows, the same query takes seconds, and it runs on every call. The query did not get worse. The table got bigger.

Start with EXPLAIN ANALYZE on your slowest endpoint's queries:

sql

EXPLAIN ANALYZE
SELECT id, caller_phone, outcome, created_at
FROM calls
WHERE workspace_id = 'ws_123'
ORDER BY created_at DESC
LIMIT 20;

Enter fullscreen mode Exit fullscreen mode

-- Seq Scan on calls ... (actual time=0.02..1843.207 rows=...)

A Seq Scan on a large table in a hot path is the signal. A composite index that matches both the filter and the sort fixes this query shape:

sql

CREATE INDEX CONCURRENTLY idx_calls_workspace_created
ON calls (workspace_id, created_at DESC);
Enter fullscreen mode Exit fullscreen mode

CONCURRENTLY lets you add it to a live table without locking writes.

The second database problem is connections. A demo opens one. Production opens one per request, and Postgres starts refusing them at peak. Use a pool and size it deliberately:

python

import asyncpg

pool = await asyncpg.create_pool(
    dsn=DATABASE_URL,
    min_size=5,
    max_size=20,   # stay well under Postgres max_connections, across all workers
    command_timeout=5,
)

async def get_recent_calls(workspace_id: str):
    async with pool.acquire() as conn:
        return await conn.fetch(
            "SELECT id, caller_phone, outcome, created_at FROM calls "
            "WHERE workspace_id = $1 ORDER BY created_at DESC LIMIT 20",
            workspace_id,
        )

Enter fullscreen mode Exit fullscreen mode

Remember the multiplier: 4 workers with a pool of 20 each is 80 connections. If you run more services against the same database, put PgBouncer in front of it.

Fix 4: Retrieval that stays fast and correct

The demo knowledge base had 30 documents. Production has 30,000, across dozens of tenants, and two things go wrong at once: vector search gets slower, and it starts returning the wrong tenant's content.

With pgvector, the fix for both is the same query shape: filter by tenant first, and give the vector column an approximate index so it stops scanning everything.

sql
CREATE INDEX idx_chunks_embedding
ON doc_chunks USING hnsw (embedding vector_cosine_ops);

CREATE INDEX idx_chunks_workspace ON doc_chunks (workspace_id);
python

async def search(workspace_id: str, query_embedding: list[float], k: int = 5):
    async with pool.acquire() as conn:
        return await conn.fetch(
            """
            SELECT id, content, embedding <=> $2 AS distance
            FROM doc_chunks
            WHERE workspace_id = $1
            ORDER BY embedding <=> $2
            LIMIT $3
            """,
            workspace_id, query_embedding, k,
        )

Enter fullscreen mode Exit fullscreen mode

The WHERE workspace_id line is not an optimization. It is a correctness rule. Never rely on a prompt instruction like "only use documents for location A." If the data can reach the model, eventually it will.

Two more things keep retrieval fast in practice. Cache embeddings for repeated questions (callers ask "what are your hours" constantly), and keep chunks small enough that you send the model 3 to 5 relevant passages, not a wall of text it has to read on every turn.

Find out before your users do

Every fix above is easy once you know where the time goes. The hard part is finding out before launch day. A short Locust script against a staging copy with production-sized data does that:

python

from locust import HttpUser, task, between

class Caller(HttpUser):
    wait_time = between(1, 3)

    @task
    def build_context(self):
        self.client.post("/agent/context", json={
            "user_id": "test_user",
            "query": "what time do you open tomorrow",
        })
Enter fullscreen mode Exit fullscreen mode

Run it at your expected peak, then double it, and watch p95 latency rather than the average. Averages hide the callers who waited four seconds.

As a rough picture of what to expect: on an endpoint like the one above, these four fixes together commonly take p95 from around 3 seconds to under 800ms at a couple of hundred concurrent users. Your numbers will depend on your data and providers, but the direction is reliable. None of it requires a new model or a bigger server, just asking the system to do real work before real users do.

I build production voice agents and backends at KeenCraft, and lead development on VoiceCake, Rawk.ai, and Dynaris. Happy to answer questions about any of these fixes in the comments.

Top comments (0)