DEV Community

Cover image for Your AI Demo Works. Will Your LLM Architecture Survive 15 Seconds?
Gaurav Talesara
Gaurav Talesara

Posted on

Your AI Demo Works. Will Your LLM Architecture Survive 15 Seconds?

Imagine you are building an AI customer support assistant for an e-commerce company.

A customer opens the chat and asks:

"Where is my order? It was supposed to arrive yesterday."

Your application retrieves the order status, checks the shipping information, sends the relevant details to an LLM, and returns a helpful answer.

During development, everything works.

Then a major sale begins. Thousands of customers start using the assistant. The LLM provider becomes slower, some requests receive HTTP 429 responses, and a few requests time out.

Now customers are waiting for answers about their orders. Support tickets begin to pile up.

The model is still capable of generating an answer. Your application is struggling to deliver it reliably.

This is where production LLM architecture matters.

1. Understand the business workflow first

Let's define a simple target: the application should return a response within 15 seconds.

That is an illustrative service objective for this example, not a universal recommendation. The appropriate deadline depends on the user experience and business requirements.

A typical request might follow this path:

Customer
   |
   v
Support Chat API
   |
   v
Authentication + Order Lookup
   |
   v
LLM Gateway
   |
   +------> Primary LLM Provider
   |
   +------> Approved Fallback
   |
   v
Validate Answer
   |
   v
Return to Customer
Enter fullscreen mode Exit fullscreen mode

Notice that the LLM is not responsible for looking up the order by itself.

Your backend should retrieve authoritative order and shipping information from the relevant systems. The model can explain those facts in natural language, but it should not invent an order status or guess a delivery date.

This separation is the first important design decision.

2. What can go wrong within 15 seconds?

Suppose the request takes this much time:

Operation Time
Authenticate customer 200 ms
Retrieve order and shipment 800 ms
Build prompt and context 200 ms
Wait for LLM response 8 seconds
Validate and return answer 300 ms
Total 9.5 seconds

This request succeeds within the 15-second budget.

Now imagine the LLM takes 12 seconds and returns a rate-limit error. Your application retries, and the second request takes another 8 seconds.

The customer has now waited more than 20 seconds.

The lesson is simple: setting a timeout on each API call is not the same as setting a deadline for the entire business operation.

3. Enforce one end-to-end deadline

Start with a deadline for the complete request.

const REQUEST_BUDGET_MS = 15_000;

function createDeadline() {
  return Date.now() + REQUEST_BUDGET_MS;
}

function remainingTimeMs(deadline: number) {
  return Math.max(0, deadline - Date.now());
}
Enter fullscreen mode Exit fullscreen mode

When the order lookup finishes, calculate the remaining time before calling the model.

const deadline = createDeadline();

const order = await getOrderForCustomer(
  customerId,
  orderId
);

const remaining = remainingTimeMs(deadline);

if (remaining <= 0) {
  throw new Error("Request deadline exceeded");
}

const answer = await generateSupportAnswer({
  order,
  timeoutMs: remaining
});
Enter fullscreen mode Exit fullscreen mode

The generateSupportAnswer function must pass timeoutMs to the actual SDK or HTTP client. The snippet illustrates the design, not a complete implementation.

In a real application, you should also account for authentication, database timeouts, cancellation, and any other work performed before or after the model call.

Important: a timeout means your application stopped waiting. It does not always mean the provider stopped processing the request.

That distinction becomes especially important when a request can trigger a real-world action.

4. Handle rate limits without creating a retry storm

During the sale, your provider starts returning HTTP 429 responses.

A tempting implementation is to retry every failed request immediately.

That can make the situation worse. If 1,000 requests fail and each retries three times, the provider could receive up to 3,000 additional attempts from that batch.

Instead, use a bounded retry policy.

function retryDelayMs(attempt: number): number {
  const baseMs = 250;
  const maxMs = 2_000;

  const cap = Math.min(
    maxMs,
    baseMs * 2 ** attempt
  );

  return Math.random() * cap;
}
Enter fullscreen mode Exit fullscreen mode

This function calculates a randomized backoff delay. It is only one part of a retry policy.

A production implementation should:

  • Classify the error before retrying.
  • Respect the provider's Retry-After guidance when available.
  • Limit the number of attempts.
  • Check the remaining deadline before sleeping or retrying.
  • Avoid retrying invalid requests or authentication failures.

For example, if only 700 milliseconds remain and the next retry would require a longer delay plus a model response, the application should stop trying.

It is better to return a controlled response than to keep a customer waiting for a deadline the application has already missed.

5. Design a useful fallback

What should happen when the primary LLM provider is unavailable?

For this support assistant, there are two reasonable paths.

Path A: Use an approved fallback model.

This can work when the alternative model has been evaluated for the task and meets your latency, quality, privacy, and cost requirements.

Path B: Return a deterministic response.

If no suitable model is available, the backend can return a simple message based on the verified order data:

"Your order is currently marked as in transit. We cannot generate a detailed explanation right now. You can check the latest tracking status here."

This response may be less conversational, but it is based on real data and still helps the customer.

The application should never fabricate a delivery date just to make the conversation feel complete.

Fallbacks should be designed around the business outcome, not around the assumption that another model will always work.

6. Keep order access outside the model's authority

This is a critical security boundary.

Suppose a customer asks:

"Show me the status of order #84521."

The backend must verify that the authenticated customer is authorized to access that order before including any order information in the model context.

A safe request flow is:

Authenticated Customer
         |
         v
Authorize Order Access
         |
         v
Fetch Verified Order Data
         |
         v
Send Minimum Required Context
         |
         v
Generate Answer
         |
         v
Validate and Return
Enter fullscreen mode Exit fullscreen mode

Do not let the model decide whether the customer has permission to view an order. That decision belongs to your application's authorization layer.

Likewise, if the assistant can cancel orders or issue refunds, those actions should use explicit backend authorization, business rules, and appropriate confirmation steps. The LLM can help interpret the customer's intent, but it should not bypass the controls that protect the transaction.

7. Control the cost of each support interaction

During a large sale, request volume can increase sharply.

Suppose the assistant sends the customer's entire conversation history and a large amount of order documentation with every request. Token usage can grow even when the actual question is simple.

For "Where is my order?", the model probably needs only the relevant order status, tracking information, and a small amount of conversational context.

Set limits for:

  • Input size and conversation history.
  • Retrieved context.
  • Maximum generated tokens.
  • Requests per user or tenant.
  • Overall spending and usage alerts.

Track the cost per successfully resolved support interaction, not just the average cost of an LLM request.

A request that is cheap but repeatedly fails may still increase total support costs if customers abandon the assistant and contact a human agent.

8. Monitor the customer outcome

A dashboard showing successful HTTP responses is not enough.

Imagine that the model returns HTTP 200 for 99% of requests, but many answers are too slow or fail to address the customer's question.

The infrastructure may appear healthy while the support experience is deteriorating.

Track the full request lifecycle.

Metric Why it matters
End-to-end latency Measures the customer's wait
Model latency Shows how much time the provider consumes
HTTP 429 rate Identifies rate-limit pressure
Retry count Helps reveal repeated recovery attempts
Fallback rate Shows how often the primary path fails
Token usage and cost Tracks spending per interaction
Answer validation failures Finds malformed or unusable responses
Resolution rate Measures whether the assistant actually helps

For answer quality, evaluate whether the response matches the verified order data. For business impact, measure whether the customer receives a useful answer or still needs a support agent.

Avoid logging sensitive order details or personal information unnecessarily. Use request IDs for tracing, and apply appropriate access controls and retention rules to logs.

9. Test failure before the sale begins

Do not test only the happy path.

Run controlled tests where:

  • The LLM takes 12 seconds to respond.
  • The provider returns HTTP 429.
  • The provider returns HTTP 503.
  • The request exceeds its deadline.
  • The model returns malformed JSON or unsupported claims.
  • The primary provider becomes unavailable.

Check that the application stops retrying when the budget is exhausted, falls back only when allowed, and never exposes another customer's order information.

If the assistant can initiate transactions, test the case where the provider times out after the backend may already have completed an action. A retry must not accidentally cancel an order twice or create duplicate refunds.

10. Production-readiness checklist

Before releasing this assistant to real customers, I would want these controls in place:

  • [ ] A deadline for the complete request.
  • [ ] Bounded retries with appropriate backoff.
  • [ ] Explicit handling for rate limits and provider outages.
  • [ ] A tested fallback or deterministic response.
  • [ ] Authorization checks before retrieving customer data.
  • [ ] Limits on context size, generated tokens, and spending.
  • [ ] Monitoring for latency, failures, retries, cost, and answer quality.
  • [ ] Tests for timeouts and duplicate side effects.
  • [ ] A clear response when the AI service is unavailable.

You do not need to build a complex microservice architecture immediately. A well-structured module inside your existing backend can enforce many of these controls. Extract it into a separate gateway when operational needs justify the added complexity.

Final thought

The real test of an AI support assistant is not whether it can answer a question when everything is working.

It is whether the business can still serve the customer when the model is slow, the provider is rate-limiting requests, or the request budget is nearly exhausted.

The LLM generates the language. Your architecture determines whether the feature is reliable enough to use.

If your primary LLM provider became unavailable during your busiest hour, would your application still be able to help the customer?

Top comments (0)