The problem
Health records, payslips, identity documents, personal e-mails, private notes: this is the data a large language model (LLM) would be most useful on, and the data that should never be sent to a third-party API.
Sending it to a hosted model means losing control of where it is stored, who can read it, and how long it is kept.
The alternative is to run everything on one machine: the documents, the search index and the model. Nothing leaves the host.
That solves privacy. It leaves three questions open: what is stored and exposed, what each request costs, and where the bottleneck is. Observability answers them.
This article follows the data in four steps: ingestion, the LangGraph workflow, Kubernetes, and observability applied to security, cost and scaling.
Source code: github.com/jmiliamine/rag-observability-k8s-lab
The system at a glance

Architecture: a folder of notes, an index in PostgreSQL, a local model, and the observability stack
Four components:
- Data lake. A folder of Markdown and text files on the host. It is the single source of truth for the documents.
- Vector index. PostgreSQL with the pgvector extension. It stores each piece of text with its embedding and runs the similarity search.
- Models. Two local models served by Ollama: one generates the embeddings, one generates the answers. No API key, no external call.
- Observability stack. An OpenTelemetry Collector, Prometheus for metrics, Tempo for traces, Loki for logs, Grafana for dashboards.
The pattern is RAG, retrieval-augmented generation: for each question, the service retrieves the most relevant pieces of the documents and passes them to the model as context. The model answers from that context, not from its training data.
A RAG system has two flows. The write flow puts documents into the index. The read flow answers questions from the index. The next two sections cover them in that order.
Step 1. A note enters the system
Take a concrete case: notes written during a conference, saved as a Markdown file.

The path of a note: folder, ingest job, index, answer
The file is written to the data lake. Saved by hand or copied by a sync tool: any process that writes a text file into the folder works.
The ingest job reads the folder. It runs every night, and can be triggered on demand:
task app:ingest
The job splits each file into chunks. A chunk is a few paragraphs: small enough to match a precise question, large enough to keep its meaning.
Each chunk is embedded. The embedding model converts a chunk into a vector, a list of numbers that represents its meaning. Two chunks about the same topic produce vectors that are close to each other. The vectors are written to PostgreSQL.
The index is swapped atomically. The job builds the new index next to the current one and switches in a single transaction. The API serves the previous index until the new one is complete. A failed ingest leaves the served index untouched.
Two access rules apply:
- The ingest job is the only workload that reads the data lake, and it mounts it read-only.
- The API never reads the files. It queries the index with a database role limited to
SELECT.
The note is now searchable. The next step is the read flow.
Step 2. A question is asked: the LangGraph workflow

A question in the terminal, then the dashboard and the trace
Answering a question is not a single model call. It is a sequence of decisions: is this a follow-up question, did the search return anything useful, should the query be retried, is the answer supported by the documents.
LangGraph models this as a graph. Each step is a node, a function that reads a shared state and returns an update to it. Edges connect the nodes. A conditional edge chooses the next node from the content of the state.

The LangGraph workflow: condense, retrieve, rewrite, generate, grade, fallback
The workflow has six nodes:
| Node | What it does | Calls the model |
|---|---|---|
condense |
Turns a follow-up question into a standalone question, using the previous turns of the conversation | Only on follow-up questions |
retrieve |
Embeds the question and runs the similarity search in PostgreSQL. Keeps the chunks above a minimum score | Embedding model |
rewrite |
Rewrites the question as a short search query when no chunk passed the score | Yes |
generate |
Sends the question and the relevant chunks to the model and streams the answer | Yes |
grade |
Computes a groundedness score: the share of the words of the answer that are present in the retrieved chunks | No |
fallback |
Returns a fixed "I don't know" answer | No |
The decision happens after retrieve, on a conditional edge with three outcomes:
-
At least one relevant chunk: go to
generate, thengrade. -
No relevant chunk, first attempt: go to
rewrite, then back toretrievewith the new query. This retry happens once. -
No relevant chunk after the retry: go to
fallback.
Three examples show the three paths.
A direct question. "What is an error budget?" goes through condense (nothing to resolve on a first question, no model call), retrieve, generate, grade. One embedding call, one generation call.
A follow-up question. "And how is it spent?" has no subject on its own. condense uses the previous turns to rewrite it as "How is an error budget spent?" before the search. The conversation history is stored in PostgreSQL by a LangGraph checkpointer, one checkpoint per answered turn, limited to the last three turns.
A question the notes cannot answer. retrieve finds no chunk above the score, rewrite produces a new query, retrieve runs again and still finds nothing, and fallback answers "I don't know". The generation model is never asked to answer without context. On private data, a confident wrong answer is worse than no answer.
The graph structure has a direct benefit for observability: each node becomes one span in a trace. The path a question took, and the time spent in each node, can be read without adding any logging by hand. Step 4 comes back to this.
The application is now described end to end. The next question is where it runs.
Step 3. Where it runs: Kubernetes on one machine
Why Kubernetes for something that runs on one machine? Because the rules that keep the data safe become declarative objects instead of habits: a NetworkPolicy, a read-only volume, a restricted pod, a scheduled job. The same manifests also run unchanged on a real cluster later.
The cluster is created with k3d, which runs k3s inside Docker: one control-plane node and one worker node, about 500 MB of memory each.
It is split into three namespaces:
-
obslab: the application (API, PostgreSQL, scheduled jobs). -
observability: the OpenTelemetry Collector, Tempo and Loki. -
monitoring: Prometheus and Grafana.
Each need from the previous steps maps to a standard Kubernetes object:
| Need | Kubernetes object |
|---|---|
| Serve the API with no downtime during a rollout | Deployment with 2 replicas, PodDisruptionBudget |
| Keep the index and the conversations on disk | StatefulSet for PostgreSQL with a persistent volume |
| Index the notes every night, purge old conversations | Two CronJobs |
| Give the ingest job access to the notes, and only to it | PersistentVolume and PersistentVolumeClaim, mounted read-only |
| Expose the API and the UIs under one entry point | Gateway API: one Gateway, one HTTPRoute per service |
| Restrict who can talk to the database | NetworkPolicy |
| Store generated passwords | Secrets, created at install time |
| Declare the alert rules | PrometheusRule (custom resource of the Prometheus Operator) |
A request therefore enters through the Gateway, is routed by an HTTPRoute to the API Service, and reaches one of the two API pods, where the LangGraph workflow runs.
Ollama runs on the host, next to the cluster. Pods reach it through an ExternalName Service, so they call it by a cluster DNS name like any other backend.
The platform components are installed with Helm, each chart pinned to an exact version. The application is deployed with Kustomize: a base and two overlays, one with the real models and one with fake models for tests. Nothing is applied by hand, so the whole stack can be deleted and recreated with one command.
The system now runs, privately. What remains is to see what it does.
Step 4. Seeing what happens: observability
A local model is still a black box. A slow or wrong answer gives no indication of its cause: retrieval, context size, model speed, or the database.
Observability makes the system explain itself through three signals:
- Metrics: numeric values over time, such as latency or token counts.
- Traces: the path of one request through every step, with the duration of each.
- Logs: timestamped events, linked to the trace that produced them.
The service emits all three with OpenTelemetry and follows its semantic conventions for generative AI, a standard set of names for model calls, token usage and operation durations. Standard names mean the dashboards and alerts are not tied to one model or one provider.

The trace of one question: one span per node of the graph
The trace above is the direct-question path from Step 2. The root span is the HTTP request. Under it, one span for the workflow, then one span per LangGraph node: condense (a few microseconds, since a first question needs no rewriting), retrieve with the embedding call and the SQL query, generate with the model call, grade. Almost all of the 8 seconds are spent in generate. A question that took the retry path would show rewrite and a second retrieve.
The pods send their telemetry to the OpenTelemetry Collector over OTLP. The Collector adds Kubernetes metadata to every signal: namespace, pod name, deployment name. A slow trace can then be tied to the exact pod that served it, and a log line to the trace that produced it.
The three questions from the introduction can now be answered one by one.
Scaling: where the time goes

The dashboard: request rate, latency per node, tokens, answer quality
The dashboard breaks a request down:
-
Latency per node.
retrievetakes tens of milliseconds.generatetakes seconds. The model is the bottleneck, not the database. - Time to first chunk. The delay before the first part of the answer is returned.
- Tokens per second. The generation throughput of the model on the current hardware.
These numbers show what to change first: a smaller context, a smaller model, or more hardware.
On the Kubernetes side, three settings control how the API behaves under load:
- Requests and limits. Every container declares CPU and memory requests, and a memory limit. There is no CPU limit: CPU throttling adds latency without protecting anything.
- Probes. The liveness probe restarts a pod that is stuck. The readiness probe removes a pod from the Service when it cannot answer, for example when the index is not reachable. It does not depend on Ollama: a slow model must not take every replica out of rotation.
- Replicas. Two API replicas allow a rolling update with no interruption. Adding replicas does not make answers faster, since the model is the bottleneck. The latency per node makes that visible before any time is spent scaling the wrong component.
Four alerts are defined as service level objectives (SLOs), targets the service is expected to meet:
- Error rate.
- p95 latency.
- Fallback rate: how often the workflow ends in
fallback. - Groundedness: the score computed by
grade.
The last two measure quality. A RAG service can respond fast, return HTTP 200, and still be useless because retrieval stopped finding the right chunks. Latency and error rate do not detect that. Quality indicators do.
Security: what is stored, what is exposed
Telemetry is a common place for sensitive data to leak. The simplest thing to attach to a trace is the prompt and the answer, which is exactly the content that must stay private.
The setup applies these rules:
-
No content in telemetry by default. Spans record the structure of a request: nodes, durations, token counts, model name. Prompts and answers are recorded only when
OBSLAB_CAPTURE_CONTENT=trueis set explicitly. -
No port exposed on the network. Every published port is bound to
127.0.0.1. A unit test fails if one is not. - Restricted pods. The application namespace enforces the Kubernetes restricted Pod Security Standard: containers run as non-root, with a read-only root file system and no privilege escalation. A pod that does not comply is rejected by the API server.
- Least privilege on the database. The API uses a read-only role for the index. A NetworkPolicy allows only the application workloads to connect to PostgreSQL.
- Retention. Conversations are deleted after seven days by a CronJob.
Cost: tokens, time and cardinality
A local model has no invoice. It still consumes CPU or GPU time and memory, and the unit that drives that consumption is the token.
The dashboard tracks input and output tokens per request. Input size is the one to watch: it grows with the number of chunks sent to generate, and latency grows with it. A follow-up question costs one extra short model call, in condense.
One panel converts usage into money: the API-equivalent cost. With the per-token price of a hosted API set in two variables, it shows what the same traffic would cost there. It gives a factual basis to compare local and hosted inference.
The telemetry has a cost of its own. In Prometheus, every distinct combination of label values is a separate time series. A label that changes at every pod restart, such as a pod UID or a start time, multiplies the number of series without adding information. This is the cardinality problem. The Collector drops those attributes from metrics before export, and an end-to-end test checks that they do not come back.
Run it, and current limits
Requirements: Docker, the command-line tools listed in the README, and about 10 GB of memory.
git clone https://github.com/jmiliamine/rag-observability-k8s-lab
cd rag-observability-k8s-lab
task doctor
task setup
task up
The first start takes 15 to 20 minutes, mostly image downloads. Then:
task ask Q="What is an error budget?"
The repository ships with sample notes, so the question above returns an answer before any personal document is added.
Current limits:
-
Text only. The ingest job reads
.mdand.txtfiles. PDF files and scanned documents need a text extraction step that is not implemented. - No built-in sync. Moving notes from a phone to the data lake is left to an external sync tool.
- Batch indexing. A new file is searchable after the next ingest run, not immediately.
Summary
- Private data stays private when the documents, the index and the model run on the same host.
- A RAG workflow is a set of decisions. Modelling it as a LangGraph graph makes each decision a node that can be traced and measured.
- Kubernetes turns the protection rules into objects that can be reviewed and tested: NetworkPolicy, read-only volumes, Pod Security Standards, Secrets.
- Telemetry must be designed not to carry the content it is meant to protect.
- Token counts and latency per node show where the time goes and what to optimize.
- Quality indicators such as fallback rate and groundedness detect failures that latency and error rate miss.
Top comments (0)