"Design a URL shortener" isn't going away. But a new question keeps showing up in system design loops: design a ChatGPT-style chat service.
It looks like a chat app. It isn't. A normal chat app moves small messages between people. An LLM chat service runs an expensive computation for every message, streams the answer one token at a time, and rations a scarce resource: GPUs.
Most candidates draw a chat app and put a box labeled "LLM" on the right. Interviewers are looking for the parts that box hides. Let's open it up.
If you're preparing for AI-focused rounds, Grokking the AI System Design Interview goes deeper on questions like this one.
Clarify the scope
Questions worth asking up front:
- Is this a consumer product, an API for developers, or both?
- Do we host the models, or call a third-party provider?
- Do users need conversation history across devices?
- Free and paid tiers? Different limits per tier?
- Any tools, file uploads, or web search? (Usually out of scope for 45 minutes.)
For this walkthrough: a consumer chat app, our own hosted models, history synced across devices, free and paid tiers, text only.
Requirements
Functional:
- Users send a message and see the reply stream in token by token.
- Conversations are saved and listed per user.
- Users can stop a response mid-stream or regenerate it.
- Paid users get access to a larger model and higher limits.
Non-functional:
- Time to first token under about one second. Users judge speed by when text starts appearing, not when it ends.
- Survive spikes without collapsing. When GPUs are full, degrade gracefully.
- Never lose a saved conversation.
- Control cost. GPUs are the biggest line item by far.
Back-of-the-envelope estimates
Assumptions: 10 million daily active users, 10 messages each per day, an average reply of 500 output tokens, and about 2,000 tokens of context sent with each request.
- Requests: 100 million per day. That's about 1,200 per second on average. Plan for roughly 3x at peak, so 3,500 per second.
- Output tokens: 100M x 500 = 50 billion per day. That's about 580,000 tokens per second on average, and around 1.7 million at peak.
- Open streams: If a model generates about 50 tokens per second per user, a 500-token reply takes 10 seconds. At 3,500 new requests per second, you have about 35,000 streams open at any moment.
- GPUs: Suppose one model replica, with batching, produces 2,500 output tokens per second. Then peak needs about 700 replicas. The real number depends heavily on model size and hardware, but the conclusion holds: this is a fleet, not a server.
- Storage: About 200 million messages per day at roughly 2 KB each is around 400 GB per day. Large, but ordinary for a distributed database.
The punchline: storage and request counts are normal web-scale numbers. GPU throughput is the bottleneck. The design should revolve around it.
High-level design
The request path:
- API gateway. Authenticates the user and applies token-based rate limits.
- Chat service. Saves the user's message, loads conversation history, and builds the prompt.
- Inference router. Picks a model pool based on the user's tier and current load. Queues the request if GPUs are busy.
- Model servers. Run the model with continuous batching and stream tokens back.
- Streaming back to the client. Tokens flow back through the chat service to the user over server-sent events.
- Persistence and metering. When the reply finishes, the chat service saves it and emits a usage event to a queue for billing and analytics.
Let's go deeper on the parts interviewers probe most.
Deep dive 1: Streaming
Users expect text to appear as it's generated. Two common options:
- Server-sent events (SSE): One HTTP response that stays open while the server pushes chunks. Simple, works through most proxies, and fits a one-way stream well.
- WebSockets: Two-way and more flexible, but more connection state to manage.
SSE is usually enough. The client sends a normal POST, and the reply streams back.
The hard part is disconnects. A user on a train loses signal at token 300. Options:
- Keep generating on the server, and write tokens to a short-lived buffer keyed by message ID.
- When the client reconnects, it asks for the message from the last token it received.
- When generation finishes, the full reply goes to the database either way.
This also makes "open the same chat on your laptop" work. The second device reads from the same buffer.
Deep dive 2: Rate limiting by tokens, not requests
A request-per-minute limit is unfair here. One request might ask for a haiku. Another might paste in a whole book.
Limit by tokens instead, with a token bucket per user:
- Before the request runs, reserve an estimate: input tokens plus the maximum output tokens.
- After it finishes, refund the difference between the estimate and what was really used.
Think of it like a hotel holding a deposit on your card. It reserves the worst case, then settles the real bill at checkout.
Tiers become simple: paid users get a bigger bucket and a faster refill.
Deep dive 3: Managing context
Models can only read a limited number of tokens at once. Long conversations eventually exceed that.
Common strategies:
- Sliding window: Send only the most recent messages that fit.
- Summarization: Periodically compress older turns into a short summary, and send the summary plus recent messages.
- Retrieval: Store older messages as embeddings and pull in only the relevant ones.
Context also drives cost. Every token you send is a token the GPU must process. A related trick is prefix caching. The system prompt and earlier turns are the same on every request in a conversation. Model servers can cache the computed state for that shared prefix and skip recomputing it. The router helps by sending follow-up messages in a conversation to the same replica when it can.
Deep dive 4: Scheduling scarce GPUs
This is where strong candidates separate themselves.
- Continuous batching. A GPU is most efficient when it processes many requests together. Modern inference servers add new requests to a running batch as old ones finish, instead of waiting for a whole batch to complete.
- Queues with priorities. When the fleet is full, requests wait in a queue. Paid users get a higher priority. Free users wait a bit longer.
- Autoscaling is slow. Starting a new model replica means loading many gigabytes of weights. That can take minutes. You can't scale up in the middle of a spike, so keep headroom and scale on predictions, like daily traffic patterns.
- Graceful degradation. If the queue gets too long, route free users to a smaller, faster model, shorten max output, or show a "high demand" message. A slower answer beats an error.
Deep dive 5: Data model
Chat history is write-heavy, append-only, and read by conversation. That's a good fit for a wide-column or key-value store.
-
conversations: partition keyuser_id, sorted byupdated_at. Powers the sidebar list. -
messages: partition keyconversation_id, sorted bymessage_id(a time-ordered ID). Powers loading a chat.
Store the model name and token counts on each assistant message. You'll need them for billing, debugging, and evaluating model quality later.
Failure modes worth naming
- A GPU node dies mid-stream. The router retries the request on another replica. The client sees a short pause, then the stream restarts. Mark the partial reply so the UI can replace it cleanly.
- Duplicate sends. Users double-click. Attach a client-generated message ID and ignore repeats.
- Billing drift. Usage events go through a durable queue, so a crash in the billing service doesn't lose records. Consumers deduplicate by message ID.
- Safety. Run moderation checks on input and output. Keep them fast, or run the output check on chunks as they stream, so they don't destroy time to first token.
Key Takeaways
- An LLM chat service is a chat app wrapped around a scarce, expensive compute resource. Design around GPU throughput.
- Optimize time to first token. Stream with SSE and make streams resumable.
- Rate limit by tokens with a reserve-then-refund bucket, not by request count.
- Manage context with windows, summaries, or retrieval. Use prefix caching to cut repeat work.
- Use continuous batching, priority queues, and graceful degradation. Autoscaling alone is too slow for spikes.
- Store conversations in a store partitioned by user and conversation, with model and token counts on every reply.
FAQ
What if we call a third-party model API instead of hosting our own?
The GPU fleet becomes someone else's problem, but the same ideas still apply. You still need token-based rate limits, streaming, retries with backoff, and fallback to another model or provider when one is slow or down.
SSE or WebSockets?
SSE is simpler and fits a one-way token stream. Pick WebSockets if you need real two-way features, like live collaboration or voice.
How do I estimate GPU capacity in an interview?
State an assumption for tokens per second per replica, then divide peak token demand by it. Interviewers care that you identify GPU throughput as the bottleneck, not the exact number.
Should conversation history live in SQL?
It can at small scale. At hundreds of millions of messages per day, an append-heavy, partition-by-conversation store is a more natural fit. Keep account and billing data in a relational database.
How deep should I go on model internals?
Not very. Know that batching improves throughput, that longer context costs more, and that caching shared prefixes saves work. You're being tested on systems thinking, not on machine learning research. If classic building blocks like caching, queues, and partitioning still feel shaky, shore those up first with Grokking the System Design Interview.
Have you been asked an AI system design question in an interview yet? What was the prompt?


Top comments (0)