Most AI architecture diagrams are grey rectangles with "LLM" written on one of them. You can read them only if the person who drew them is in the room.
Below are ten architectures that show up in almost every AI project, drawn so that every component has a defined role: a model, a vector database, an agent, an MCP server, a guardrail. They are vendor-neutral on purpose. Replace the generic blocks with whatever you actually run.
Each diagram is interactive. Follow the link under it to pan, zoom and export it as PNG or SVG. No account needed.
Disclosure: we drew these with DuctTape.io, the diagram editor we build. The architectures themselves are common patterns and not tied to the tool.
1. RAG architecture
A RAG system has two paths. The ingestion path loads documents, splits them into chunks, turns each chunk into a vector with an embedding model and stores it in a vector database. The query path takes the user's question, retrieves the most similar chunks and sends them to the LLM together with the question. The diagram shows both paths meeting at the vector database. Use it as the starting point for any assistant that has to answer from your own documents.
When to use it: You need answers grounded in your own documents and can live with semantic search only.
2. Advanced RAG with hybrid search and reranking
Basic RAG misses exact terms such as product codes or names. Hybrid search runs a vector query and a keyword query (BM25) side by side. A reranker scores the merged candidates against the question and passes only the top few to the answer model. A small model can rewrite the question first, and tracing records what was retrieved for each answer. This is the usual next step when a basic RAG setup returns the wrong passages.
When to use it: Basic RAG returns the wrong passages, or users search for exact terms such as product codes and names.
3. AI agent with MCP servers
The Model Context Protocol gives an agent a standard way to discover and call tools. In this architecture the agent talks to one MCP server per backend system: code, tickets and data. Each server holds its own credentials and exposes only what the agent may do, for example read-only access to the database. The LLM decides which tool to call; the conversation state lives in short-term memory. Swap in your own servers to document what your agent can touch.
When to use it: Your agent needs to act on several systems and you want one place per system that defines what it may do.
4. Multi-agent system: orchestrator and workers
In the orchestrator-worker pattern one agent breaks a task into steps and hands each step to a specialised worker. Here a research agent searches the web, a coding agent runs code in a sandbox, and a review agent checks the result before a human approves it. The pattern pays off when sub-tasks need different tools or different prompts. It also costs more tokens than a single agent, so start with one agent and split only where it helps.
When to use it: Sub-tasks need different tools or prompts. If one agent can do the job, stay with one agent.
5. LLM gateway with model fallback
An LLM gateway is the single entry point for every model call in your company. Applications, batch jobs and agents send requests to the gateway instead of to a provider. The gateway answers repeated questions from a semantic cache, routes simple tasks to a small model, and falls back to a second model when the primary one fails or is rate limited. Because every call passes through it, cost and latency are traced in one place.
When to use it: More than one team or application calls models, and you want caching, fallback and cost tracking in one place.
6. Guardrails and human-in-the-loop
Guardrails sit on both sides of the model. Input guardrails reject prompt injection and off-topic requests; a PII filter redacts personal data before it reaches the model. Output guardrails check the draft answer against policy. Actions that cannot be undone, such as a refund or an outgoing e-mail, wait for a human approval. Every decision is written to an audit log. The diagram is a template for systems that must be explainable to compliance.
When to use it: The system handles personal data or can trigger actions that cannot be undone.
7. LLM observability and evaluation pipeline
You cannot improve an LLM application you cannot see. This architecture records every model call as a trace and stores it. Online evals score a sample of live traffic and raise alerts. Interesting traces and user feedback are curated into an evaluation dataset. Offline evals run that dataset in CI whenever a prompt or model changes, so a regression is caught before release. Prompts live in a registry, not in the code.
When to use it: You change prompts or models regularly and want to know whether a change made things better or worse.
8. Customer support chatbot architecture
A support chatbot needs three things beyond a model: knowledge, actions and an exit. Knowledge comes from a vector index of the help center. Actions are tools, here an order lookup against the shop and CRM API. The exit is the handoff to a human agent when the customer asks for it or the bot is unsure. Conversation state is kept in a session store so the customer does not have to repeat themselves.
When to use it: Support questions are repetitive, answers exist in a help center, and a human must stay reachable.
9. Voice agent architecture
A voice agent is a pipeline that has to finish in well under a second. Audio from the phone line is streamed to a speech-to-text model. The agent receives the transcript, calls the LLM and, if needed, tools such as a booking system. The reply is streamed through a text-to-speech model back to the caller. Call context is held in short-term memory. The diagram marks the two streaming edges where latency matters most.
When to use it: The conversation happens on the phone and latency decides whether it feels natural.
10. Document processing pipeline with LLMs
Invoices, contracts and forms arrive as PDFs and scans. A queue decouples the inbox from the workers. Each worker runs OCR and parsing, then asks an LLM to extract the fields as JSON. A validation step checks the result against a schema. Valid records go to the database; low-confidence results go to a review queue where a person corrects them. The pattern scales by adding workers and keeps humans on the difficult cases only.
When to use it: Documents arrive in volume and the result has to land in a database as structured records.
How to use these
- Pick the diagram closest to your system.
- Redraw it with your own components. The generic "LLM" becomes the model you call, the "vector DB" becomes the store you run.
- Share the link in the design review, the README or the customer call, instead of a screenshot that is outdated the next day.
All ten are collected at theducttape.io/examples.
Which architecture is missing from this list? Tell us in the comments and we will draw it.










Top comments (0)