DEV Community

Cover image for Fast in the Demo, Slow in Production: Fixing Latency with Postgres, Redis and Parallel Pipelines
jareer nauman
jareer nauman

Posted on

Fast in the Demo, Slow in Production: Fixing Latency with Postgres, Redis and Parallel Pipelines

The demo works. One user, a warm laptop, and a presenter who already knows which button to press.

Then the first week of real traffic hits, and the same product waits, times out, or returns a stale answer. People call that a scaling problem. It's usually a path problem: the request was designed for an audience of one.

I lead development on a few production platforms, including Rawk.ai, a voice agent builder where we cut end-to-end response time from roughly 900ms to roughly 320ms. None of that came from a smarter prompt or a bigger server. It came from fixing the path. This post covers where the time actually goes and the fixes that hold up under load.

Load is a tail, not an average

A product can look healthy on average and still lose the people who hit the slow path. Voice callers hang up. Checkout users retry and double-charge themselves. Staff refresh a CRM and act on yesterday's data.

The number that matters is the slow tail (p95, p99), and whether the answer is still correct when it finally arrives. Three things to watch:

  • Speed: time to first useful response, not time until a spinner disappears.
  • Correctness: a fast stale cache is a bug with better graphs.
  • Cost: a path that calls a model or vendor API on every keystroke gets expensive before it gets slow.

If you can't name the request you're timing, you're not ready to optimize. "The app feels slow" is a complaint. "Time to first audio on a live call" is a job.

What the Rawk.ai latency pass actually changed

Callers don't forgive a 900ms gap after every sentence. They talk over the agent, the agent talks over them, and the call collapses.

The drop to ~320ms came from the pipeline:

  • Streaming speech-to-text instead of waiting for a final transcript
  • Streaming the LLM output instead of buffering a full paragraph
  • Starting audio on the first clause
  • Prefetching tool results before the agent needed them
  • Colocating services so no turn crossed extra regions

The voice-specific version of that budget is in how to build an AI voice agent that actually books appointments. But the same failure modes show up in products that never place a phone call.

Where the time actually goes

Most slow products are a chain of reasonable steps that were never allowed to overlap.

1. Sequential I/O

The page waits for the user, then the permissions, then the list, then the counts. Each call is fine on its own. The sum is the demo that died.

If the calls don't depend on each other, run them concurrently:

import asyncio

# Before: four round trips, one after another
user = await get_user(user_id)
perms = await get_permissions(user_id)
items = await get_items(workspace_id)
counts = await get_counts(workspace_id)

# After: independent calls run at the same time
user, perms, items, counts = await asyncio.gather(
    get_user(user_id),
    get_permissions(user_id),
    get_items(workspace_id),
    get_counts(workspace_id),
)
Enter fullscreen mode Exit fullscreen mode

The request now takes about as long as the slowest call instead of the sum of all four. Before you do this, map the dependencies honestly. If get_items needs perms, keep that pair sequential and parallelize the rest.

2. A cache that's missing, or lying

PostgreSQL stays the system of record. Redis sits in front of reads that are hot, identical across requests, and safe to serve twice.

The design work is the boundary: which keys exist, how long they live, and which writes delete them. A cache without an invalidation rule is how a paid invoice keeps showing "new lead" in the UI staff trusts.

import json
from redis.asyncio import Redis

redis = Redis()
HOURS_TTL = 300  # seconds; a safety net, not the invalidation strategy

async def get_location_hours(location_id: str) -> dict:
    key = f"location:{location_id}:hours"
    cached = await redis.get(key)
    if cached:
        return json.loads(cached)

    hours = await db.fetch_location_hours(location_id)
    await redis.set(key, json.dumps(hours), ex=HOURS_TTL)
    return hours

async def update_location_hours(location_id: str, hours: dict) -> None:
    await db.update_location_hours(location_id, hours)
    # The write owns the invalidation. Delete, don't wait for the TTL.
    await redis.delete(f"location:{location_id}:hours")
Enter fullscreen mode Exit fullscreen mode

The TTL is there to limit damage if an invalidation is ever missed. The real rule is that every write path for that data deletes the key.

3. Queries that scan because nobody indexed the access pattern

Postgres will answer. It'll just answer late, and later every week as the table grows. Check the query the slow request actually runs:

EXPLAIN ANALYZE
SELECT id, name, stage, created_at
FROM leads
WHERE workspace_id = 'ws_123' AND stage = 'new'
ORDER BY created_at DESC
LIMIT 50;
Enter fullscreen mode Exit fullscreen mode

If the plan shows a Seq Scan on a large table, the index doesn't match how you read the data. Build one that does, matching the filter columns and the sort:

CREATE INDEX CONCURRENTLY idx_leads_workspace_stage_created
ON leads (workspace_id, stage, created_at DESC);
Enter fullscreen mode Exit fullscreen mode

Run EXPLAIN ANALYZE again and confirm the plan switched to an index scan. CONCURRENTLY avoids locking writes on a live table while the index builds.

4. Retrieval on every request

Embeddings, vector search, and a model call for a question the last hundred users already asked. Pinecone or pgvector belongs in the path only when the answer actually depends on search. If the fact is already a column, read the column.

When repeated questions are common, cache the answer, but scope the key so one tenant can never see another's result:

import hashlib

def answer_cache_key(workspace_id: str, question: str) -> str:
    normalized = " ".join(question.lower().split())
    digest = hashlib.sha256(normalized.encode()).hexdigest()
    return f"answer:{workspace_id}:{digest}"
Enter fullscreen mode Exit fullscreen mode

The workspace_id in the key is the important part. A shared answer cache across tenants is a data leak waiting for the right question.

5. Cross-region hops

Each vendor defaulted to a different cloud, so every request crosses regions. That's tens of milliseconds per hop, on every turn, and no amount of prompt or query tuning buys it back. Colocate the services that talk on every request. Leave the rare admin job wherever it is.

Audit the live path before proposing a rebuild

When a product already has users, the instinct is often to rewrite it. Usually that's wrong. Measure first, then change the smallest layer that moves the tail:

  1. Name the one request users feel. Time it on production-like data, including the slow tail.
  2. List every downstream call on that request, in order, with what each returns.
  3. Mark which results are identical across users and which must never be shared.
  4. Cache only the identical, safe reads, with an explicit delete on write.
  5. Colocate the services that talk on every request.
  6. Re-measure the same request. If the tail didn't move, the cache wasn't the bottleneck.

A caching tier or one fixed pipeline is often the whole job. A rewrite is what you propose after that pass can't hit the bar.

And write it down: an architecture overview and the API behavior your team will need when the person who tuned it isn't in the room. Undocumented speed is a demo you'll fail to repeat.


I work on production voice AI and backend systems at KeenCraft. This post was originally published on the KeenCraft blog.

Top comments (0)