<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Nerav Doshi</title>
    <description>The latest articles on DEV Community by Nerav Doshi (@agenticdevops).</description>
    <link>https://dev.to/agenticdevops</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3916785%2F423b2322-f2d4-4fee-8576-b0537c2866f0.png</url>
      <title>DEV Community: Nerav Doshi</title>
      <link>https://dev.to/agenticdevops</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/agenticdevops"/>
    <language>en</language>
    <item>
      <title>Reconnected Claude Code Over SSE — and Found the Distance Scores Aren't as Stable as I Thought</title>
      <dc:creator>Nerav Doshi</dc:creator>
      <pubDate>Fri, 02 Oct 2026 20:54:31 +0000</pubDate>
      <link>https://dev.to/agenticdevops/reconnected-claude-code-over-sse-and-found-the-distance-scores-arent-as-stable-as-i-thought-4fh2</link>
      <guid>https://dev.to/agenticdevops/reconnected-claude-code-over-sse-and-found-the-distance-scores-arent-as-stable-as-i-thought-4fh2</guid>
      <description>&lt;p&gt;Loose end from Entry 14: switching &lt;code&gt;mcp_search_server.py&lt;/code&gt; from stdio to SSE broke the Entry 08 Claude Code connection, which was registered expecting a spawned process, not a network server. Fixing it turned out to be the easy part:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp remove today-i-ran-notes
claude mcp add &lt;span class="nt"&gt;--transport&lt;/span&gt; sse today-i-ran-notes http://localhost:8090/sse
claude mcp list
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;✔ Connected&lt;/code&gt;, server still running from Entry 14. Clean.&lt;/p&gt;

&lt;p&gt;Re-ran Entry 08's original test — ask it to use the tool, search for the oc pod-status question. It worked, and gave a solid narrative answer identifying Entry 02's wrong command. But I wanted the actual distance score this time, not just the summary, so I asked for the raw tool output directly. That's where this got more interesting than "confirmed the reconnect works."&lt;/p&gt;

&lt;p&gt;Two calls came back, for two slightly different phrasings of basically the same question:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Query&lt;/th&gt;
&lt;th&gt;Distance to &lt;code&gt;02-oc-cli-mentor-system-prompt.md&lt;/code&gt;
&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;"how to check pod status with oc"&lt;/td&gt;
&lt;td&gt;430.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"oc get pods command output troubleshooting"&lt;/td&gt;
&lt;td&gt;375.8&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Same target document, same intent, a 54-point swing just from rewording. For comparison, the gap between the &lt;em&gt;best&lt;/em&gt; and &lt;em&gt;worst&lt;/em&gt; result within a single query was 442.6 → 461.1 — about 18 points. The variance from paraphrasing the question was three times larger than the variance between a genuinely-relevant result and a marginal one in the same result set.&lt;/p&gt;

&lt;p&gt;That's a sharper version of the open question from Entry 13, and it resolves cleanly once you see both numbers side by side. Entry 13 guessed the embedding model might be rewarding lexical/syntactic overlap over pure meaning — a broken, command-shaped query scoring tighter than a clean one. This data points at something more basic underneath that: distance isn't calibrated for "is this relevant" at all, full stop. It's only meaningful for ranking results within one specific query's wording. Compare distances &lt;em&gt;across&lt;/em&gt; different phrasings of the same question, and the number tells you more about word choice than about relevance.&lt;/p&gt;

&lt;p&gt;There's also a reproducibility result worth stating plainly, since it cuts the other way: Entry 14's n8n test and this one both queried the &lt;em&gt;same&lt;/em&gt; exact string — "how do I check pod status with oc" — through different clients, and both landed on 437.7. That's not in tension with today's finding. It's the other half of it: identical query text reliably gives identical distance, regardless of client or transport. Reword the query, even slightly, and the number moves more than you'd expect from meaning alone.&lt;/p&gt;

&lt;p&gt;One more thing worth being honest about, since it affects how much to trust the narrative answer from the first run: Claude Code's accurate quote of the wrong &lt;code&gt;oc&lt;/code&gt; command — the one from Entry 02 — didn't actually come from the MCP tool. The tool's results are truncated to ~200 characters per snippet, and that quote wasn't in any of them. The session separately grepped and read the full file directly off disk. So the earlier clean-looking answer wasn't a pure test of the RAG pipeline; it was the RAG pipeline plus Claude Code's own filesystem access filling the gap. Worth knowing before holding that result up as proof the retrieval system alone is doing the work — some of it was something else entirely.&lt;/p&gt;

&lt;p&gt;Collection size matters here too, and it's worth naming rather than letting the clean distance numbers imply otherwise: every query in this entry returned the exact same five results, just reordered. The index doesn't have much in it yet. These rankings are real, but they're ranking a very short list.&lt;/p&gt;

</description>
      <category>ollama</category>
      <category>mcp</category>
      <category>rag</category>
      <category>chromadb</category>
    </item>
    <item>
      <title>Wired n8n to Ollama and the MCP Tool — and Hit Almost Every Wall on the Way</title>
      <dc:creator>Nerav Doshi</dc:creator>
      <pubDate>Wed, 30 Sep 2026 01:09:19 +0000</pubDate>
      <link>https://dev.to/agenticdevops/wired-n8n-to-ollama-and-the-mcp-tool-and-hit-almost-every-wall-on-the-way-30id</link>
      <guid>https://dev.to/agenticdevops/wired-n8n-to-ollama-and-the-mcp-tool-and-hit-almost-every-wall-on-the-way-30id</guid>
      <description>&lt;p&gt;n8n running as a container is a decent stand-in for "how would a real automation platform actually call this stuff". Everything up to now has been me calling things directly. Two integrations, same workflow: Ollama first, since it's simpler, then the MCP tool from Entry 08.&lt;/p&gt;

&lt;p&gt;Ollama went cleanly. Container running, one HTTP Request node pointed at &lt;code&gt;http://host.docker.internal:11434/api/generate&lt;/code&gt; — &lt;code&gt;host.docker.internal&lt;/code&gt; matters here specifically, it's how a container reaches something running on the host machine itself, not &lt;code&gt;localhost&lt;/code&gt;, which inside the container just points back at the container. Real response came back, &lt;code&gt;done: true&lt;/code&gt;, no drama.&lt;/p&gt;

&lt;p&gt;Except for the response itself, which is worth keeping verbatim: &lt;em&gt;"To check the status of a Pod in Kubernetes using &lt;code&gt;oc&lt;/code&gt;, you can use the &lt;code&gt;kubectl&lt;/code&gt; command-line tool."&lt;/em&gt; Asked about &lt;code&gt;oc&lt;/code&gt;. Answered with &lt;code&gt;kubectl&lt;/code&gt;, immediately, in the same sentence that acknowledged the question was about &lt;code&gt;oc&lt;/code&gt;. Same unconstrained-model bias from Entries 03 and 10, just a cleaner, funnier example of it this time — the model contradicts itself in one breath.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fpipelineandprompts.com%2Fimages%2Fdiagrams%2Fn8n-workflow-canvas.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fpipelineandprompts.com%2Fimages%2Fdiagrams%2Fn8n-workflow-canvas.png" alt="n8n workflow: Manual Trigger node connected to an HTTP Request node calling Ollama" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;MCP was the actual project. I'd built the Entry 08 server over stdio — the transport where a client spawns your script as a subprocess and talks to it over stdin/stdout. n8n has a community node, &lt;code&gt;n8n-nodes-mcp&lt;/code&gt;, that supports exactly that. Installed it, configured a credential, tried to execute:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Failed to execute operation: The file or directory does not exist
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Worth stopping on this one rather than just patching around it, because the real cause is bigger than a wrong path. n8n runs inside its own container — a completely separate filesystem from the Mac. Pointing the STDIO credential at &lt;code&gt;~/mcp-env/bin/python3&lt;/code&gt; was asking the container to find a file that, from its perspective, doesn't exist anywhere. And even mounting that path in wouldn't have actually fixed it: that venv's &lt;code&gt;python3&lt;/code&gt; is a macOS binary, and the n8n container runs Linux. A Linux container can't execute a macOS binary no matter where you mount it — that's not a path problem, that's a "these two things are architecturally incompatible" problem.&lt;/p&gt;

&lt;p&gt;So: switched the server from stdio to SSE — a real network transport instead of a spawned subprocess — which sidesteps the whole cross-filesystem, cross-OS mess entirely. The container doesn't need to touch anything on the Mac's disk; it just makes a network request to a server the Mac is already running.&lt;/p&gt;

&lt;p&gt;First attempt at that:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;mcp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;transport&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sse&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;host&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;0.0.0.0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;port&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;8090&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TypeError: FastMCP.run() got an unexpected keyword argument 'host'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Turns out &lt;code&gt;host&lt;/code&gt; and &lt;code&gt;port&lt;/code&gt; belong on the &lt;code&gt;FastMCP()&lt;/code&gt; constructor in this SDK version, not on &lt;code&gt;.run()&lt;/code&gt; — an easy mistake given how many things about this exact package have already moved around mid-series (the FastMCP → MCPServer rename from Entry 08 wasn't even the last surprise it had). Fixed, confirmed the server actually starts:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Uvicorn running on http://0.0.0.0:8090
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Didn't assume the SSE path was &lt;code&gt;/sse&lt;/code&gt; — checked directly instead:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;curl -N http://localhost:8090/sse
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;event: endpoint
data: /messages/?session_id=f2000338ea514f89bcd2c695d56f49f7
: ping - 2026-09-28 20:28:31.952790+00:00
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Real session, real pings. Confirmed.&lt;/p&gt;

&lt;p&gt;Then n8n's side: &lt;code&gt;Failed to execute operation: The service refused the connection&lt;/code&gt;. One more thing worth checking before assuming it was a config typo — this machine only has Podman installed, not Docker Desktop, even though the &lt;code&gt;docker&lt;/code&gt; command works (Podman provides Docker-API compatibility). Whether &lt;code&gt;host.docker.internal&lt;/code&gt; resolves the same way under Podman wasn't something I wanted to guess at, so tested it directly from inside the container:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;podman exec -it n8n node -e "require('http').get('http://host.docker.internal:8090/sse', r =&amp;gt; console.log('status:', r.statusCode))"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;status: 200
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It actually worked fine — Podman handles that hostname the same way Docker does here. So the earlier "connection refused" was something else, most likely a typo or leftover default in the node's endpoint field rather than a real networking gap. Fixed the field, re-ran.&lt;/p&gt;

&lt;p&gt;Real result:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;type: text
text: [02-oc-cli-mentor-system-prompt.md] (distance: 437.7) ...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fpipelineandprompts.com%2Fimages%2Fdiagrams%2Fmcp-execute-tool-result.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fpipelineandprompts.com%2Fimages%2Fdiagrams%2Fmcp-execute-tool-result.png" alt="n8n MCP Client node output showing a real search_notes result with distance 437.7" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;437.7 — the exact same distance Claude Code got for this same query back in Entry 08. Different client, different transport, same underlying Chroma store, same number. That's about as clean a confirmation as this series has produced that the whole thing is actually wired together correctly, not just superficially working.&lt;/p&gt;

&lt;p&gt;One correction on my own assumption before closing this out: I'd been telling myself we needed to switch to n8n's separate, official MCP Client Tool node once stdio was ruled out. Turned out unnecessary — the same community node that failed on STDIO also supports an SSE credential type directly. Same node, just a different saved credential, confirmed by opening the credential's own configuration panel directly rather than assuming from the node's label:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fpipelineandprompts.com%2Fimages%2Fdiagrams%2Fmcp-credential-sse-vs-stdio.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fpipelineandprompts.com%2Fimages%2Fdiagrams%2Fmcp-credential-sse-vs-stdio.png" alt="n8n credential dropdown showing both the failed STDIO account and the working SSE API credential on the same MCP Client node" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Looking back at the whole thing: none of the individual failures here were exotic. A wrong path, a renamed API, an untested assumption about hostname resolution, a mislabeled credential. What made this entry different from something like Entry 08 is that every layer — the model, the container runtime, the SDK, the automation tool — was a separate thing that could be individually right and still add up to a broken chain. Wiring four systems together doesn't multiply the failure points, it compounds them: each fix had to survive contact with the next layer before I actually knew it worked.&lt;/p&gt;

&lt;p&gt;The one number that made all of it worth doing was that 437.7. Not because it's a good number — it's just a distance — but because it was the same 437.7 Claude Code got back in Entry 08, through a completely different path. That's the actual proof this series has been chasing since the RAG entries started: the same underlying system, reachable correctly from more than one direction, giving the same answer either way. Everything else in this entry was just the cost of getting there.&lt;/p&gt;

</description>
      <category>ollama</category>
      <category>n8n</category>
      <category>mcp</category>
      <category>docker</category>
    </item>
    <item>
      <title>Kubernetes Cost Attribution: Namespace vs. Cost-Center</title>
      <dc:creator>Nerav Doshi</dc:creator>
      <pubDate>Sun, 27 Sep 2026 13:42:38 +0000</pubDate>
      <link>https://dev.to/agenticdevops/kubernetes-cost-attribution-namespace-vs-cost-center-2mm7</link>
      <guid>https://dev.to/agenticdevops/kubernetes-cost-attribution-namespace-vs-cost-center-2mm7</guid>
      <description>&lt;p&gt;&lt;em&gt;Pipeline &amp;amp; Prompts | Byte size guides on DevOps, Cloud and AI&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;☁️ Cloud Without the Chaos #5&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;⚡ Byte Size Summary&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why native Kubernetes cost attribution (namespace/project-level, via CMMO) and organizational cost-center attribution are different problems, and why one mechanism can't do both jobs&lt;/li&gt;
&lt;li&gt;How cost detection (fast, alert-based) and cost attribution (detailed, reconciled) are separate concerns that don't need — and shouldn't need — the same mechanism&lt;/li&gt;
&lt;li&gt;What actually happens to attribution when the organizational-label layer goes away, and why that's not the same as losing all visibility&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  The Story
&lt;/h2&gt;

&lt;p&gt;Cloud cost estimates tend to look straightforward when a workload is first being planned. Size the application, estimate average utilization, account for expected growth, enable autoscaling, and build the procurement number.&lt;/p&gt;

&lt;p&gt;The problem starts when the workload doesn't behave like the estimate.&lt;/p&gt;

&lt;p&gt;The procurement number for this engagement looked solid on paper. The application was homegrown, autoscaling was enabled, and the infrastructure had been sized against what appeared to be a reasonable average-utilization estimate. Nothing raised a concern during review, and there were no issues at go-live.&lt;/p&gt;

&lt;p&gt;Then the cloud spend alerts started firing.&lt;/p&gt;

&lt;p&gt;When we looked at the usage pattern by time of day, the reason became obvious. The heaviest workload wasn't happening during normal business hours. It was happening in the evenings, overnight, and during weekends — exactly the periods that had been smoothed into an average during the original sizing exercise.&lt;/p&gt;

&lt;p&gt;Nothing unusual was happening with the application itself. API calls, logging, request processing, and the application's normal request/response activity were all behaving as expected. The autoscaler was doing what it was designed to do: responding to increased demand by adding capacity.&lt;/p&gt;

&lt;p&gt;The problem was that the demand was occurring at a different time, and at a different shape, than the original estimate assumed.&lt;/p&gt;

&lt;p&gt;The resulting cloud bill was approximately 20-30% higher than the original projection. And it stayed there. The workload's actual demand curve simply didn't match the curve used to build the estimate.&lt;/p&gt;

&lt;p&gt;That gap between what we estimate and what the workload actually does is where the cost-attribution problem starts — the third of the five dimensions worth &lt;a href="https://pipelineandprompts.com/posts/hybrid-cloud-architecture-on-prem-vs-cloud-tradeoffs/" rel="noopener noreferrer"&gt;placing deliberately&lt;/a&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;The person who feels this first is usually whoever owns the cloud budget. The person who has to explain it is often whoever performed the original sizing.&lt;/p&gt;

&lt;p&gt;That leads to a fairly simple question: which workload is actually driving the cost?&lt;/p&gt;

&lt;p&gt;On a shared Kubernetes cluster running applications for multiple teams, knowing that the cluster cost a certain amount this month doesn't answer that question. Was the increase caused by one team's production application with genuinely spiky traffic? Was it a development or test workload that wasn't scaled down? Was a logging-heavy application consuming more resources than expected? Or was part of the cost associated with shared cluster infrastructure that shouldn't be assigned to an individual application team?&lt;/p&gt;

&lt;p&gt;A cluster-level cost number doesn't provide that level of visibility.&lt;/p&gt;

&lt;p&gt;Without granular attribution, the available responses tend to be broad ones: reduce the autoscaling ceiling, change the overall budget, or absorb the additional cost and try to improve the estimate next time. None of those approaches really solve the underlying problem.&lt;/p&gt;

&lt;p&gt;Before we can optimize the cost, we need to understand where the cost came from.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Existing Approaches Fall Short
&lt;/h2&gt;

&lt;p&gt;The first instinct is usually to look at the cloud provider's native cost-management capabilities. That's a reasonable place to start.&lt;/p&gt;

&lt;p&gt;Cloud providers can associate costs with resources using tags, resource groups, subscriptions, accounts, and other infrastructure-level dimensions. This works well when the infrastructure resource and the accountability boundary are essentially the same thing.&lt;/p&gt;

&lt;p&gt;Kubernetes changes that relationship. The cloud provider may be billing for virtual machines or node pools, while the application team is thinking in terms of namespaces, deployments, pods, and services. The infrastructure being billed and the workload consuming that infrastructure aren't necessarily the same object.&lt;/p&gt;

&lt;p&gt;If three application teams share the same Kubernetes cluster, the cloud provider sees the infrastructure. The application teams see their workloads. Both views are correct. Neither view, by itself, answers the complete cost question.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Architecture
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Starting with native Kubernetes cost attribution
&lt;/h3&gt;

&lt;p&gt;The next layer is understanding how much of the shared infrastructure each workload actually consumes. In an OpenShift environment, Red Hat Cost Management Metrics Operator (CMMO) provides this capability by using Prometheus/Thanos usage data from the cluster.&lt;/p&gt;

&lt;p&gt;CMMO allows cost to be distributed at the project level, where an OpenShift project corresponds to a Kubernetes namespace. This is a significant improvement over looking only at the underlying infrastructure. Instead of asking &lt;em&gt;how much did this cluster cost?&lt;/em&gt; we can start asking &lt;em&gt;how much of that infrastructure consumption belongs to each project?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;For many environments, that may be sufficient. In this case, it wasn't. The technical structure of the cluster didn't map cleanly to the organizational structure. A department could own several application namespaces. A namespace didn't necessarily map to exactly one department or cost center.&lt;/p&gt;

&lt;p&gt;Namespace-level attribution correctly answers &lt;em&gt;which workload or project consumed the resources?&lt;/em&gt; It doesn't necessarily answer &lt;em&gt;which department or cost center should own that cost?&lt;/em&gt; That second question required another attribution dimension.&lt;/p&gt;

&lt;h3&gt;
  
  
  Adding organizational context
&lt;/h3&gt;

&lt;p&gt;This is where labels become useful. Rather than replacing namespace-level attribution, we add organizational metadata to the workload. Deployment objects can carry labels representing information such as department, cost center, business unit, application owner, or environment.&lt;/p&gt;

&lt;p&gt;An admission policy ensures the required attribution information is present before a deployment reaches production or is scheduled onto a node — see Implementation, Step 2, for the specific mechanisms this can be built on.&lt;/p&gt;

&lt;p&gt;This creates two complementary dimensions. The namespace tells us where the workload lives. The label tells us who owns it from an organizational perspective. That distinction becomes particularly useful when the organization and the Kubernetes structure don't line up one-to-one. A department that owns five application namespaces can be shown per-namespace, or the label lets those five namespaces be viewed together as a single organizational cost center.&lt;/p&gt;

&lt;p&gt;The label isn't replacing native Kubernetes cost attribution. It's adding context to it.&lt;/p&gt;

&lt;h3&gt;
  
  
  The two data sources
&lt;/h3&gt;

&lt;p&gt;The implementation has two primary data sources.&lt;/p&gt;

&lt;p&gt;The first is cluster usage data. CMMO reads Prometheus/Thanos usage data and uploads it to Red Hat's cost-management service on an approximately six-hour cycle. This provides the resource-consumption side of the equation: who used what?&lt;/p&gt;

&lt;p&gt;The second is cloud billing data. For Azure, a native Azure Cost Export provides actual-cost data. In this implementation, the daily CSV export lands in a storage account in the same resource group as the cluster, and a service principal provides the access required for the cost-management platform to read that data. This provides the billing side: what did Azure actually charge?&lt;/p&gt;

&lt;p&gt;The cost-management system correlates Azure VM instance IDs with OpenShift nodes and uses the usage information to distribute infrastructure costs across projects. The native attribution boundary is therefore the OpenShift project or namespace. The organizational label provides an additional dimension on top of that.&lt;/p&gt;

&lt;h3&gt;
  
  
  Two pipelines, not one
&lt;/h3&gt;

&lt;p&gt;Cost attribution isn't a real-time process. There are two separate data pipelines that need to stay active:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Cluster usage:&lt;/strong&gt; Prometheus/Thanos → CMMO → cost-management platform, on a roughly six-hour upload cycle&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Azure billing:&lt;/strong&gt; Azure Cost Export → storage account → service-principal-scoped access → cost-management platform, on a daily export cycle&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;As a result, the attributed cost view can take up to 24 hours to reflect the current state of the environment. The dashboard isn't a live query into the cluster — it's a view of the environment after the usage and billing data have passed through their respective collection and processing cycles.&lt;/p&gt;

&lt;h3&gt;
  
  
  Cost detection and cost attribution are different problems
&lt;/h3&gt;

&lt;p&gt;The mechanism that tells us we have a cost problem doesn't need to be the same mechanism that tells us who caused it.&lt;/p&gt;

&lt;p&gt;In this implementation, native Azure Cost Management spend alerts provide the faster detection mechanism — that's what identified the original overrun. The alerting path operates directly against Azure billing information and isn't dependent on the CMMO or Cost Export processing cycle.&lt;/p&gt;

&lt;p&gt;Spend alerting asks &lt;em&gt;has spending crossed the threshold?&lt;/em&gt; Cost attribution asks &lt;em&gt;where did that spending come from?&lt;/em&gt; The first needs to be fast. The second needs to be detailed. Trying to make one mechanism do both jobs creates unnecessary complexity.&lt;/p&gt;

&lt;h3&gt;
  
  
  What happens when something goes wrong
&lt;/h3&gt;

&lt;p&gt;Two failure points are worth naming.&lt;/p&gt;

&lt;p&gt;The first is the cloud billing integration. If the service principal used to access the Azure Cost Export is given broader permissions than necessary, the integration could expose cost information beyond the intended scope.&lt;/p&gt;

&lt;p&gt;The second is the workload-label enforcement mechanism. If the admission policy is bypassed or disabled, workloads may be deployed without the required organizational labels. That doesn't break namespace-level attribution — the native project/namespace cost information keeps working. What's lost is the additional department or cost-center dimension for workloads that don't carry the required metadata.&lt;/p&gt;

&lt;p&gt;That's another reason to treat organizational labels as an additional attribution layer, not the foundation of the entire cost model — the foundation is CMMO's native namespace attribution, which keeps functioning whether or not the label layer does.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fpipelineandprompts.com%2Fimages%2Fdiagrams%2Fcost-profile-demand-shape-cloud-elasticity-architecture.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fpipelineandprompts.com%2Fimages%2Fdiagrams%2Fcost-profile-demand-shape-cloud-elasticity-architecture.png" alt="Architecture Diagram" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The diagram earns its place here by making one thing visible that prose can't: two independent pipelines, on two different cadences, converging on one dashboard — with the label/admission-policy layer drawn as an overlay on native namespace attribution rather than something the whole model depends on. That's the design decision worth seeing, not just reading.&lt;/p&gt;

&lt;p&gt;The pattern itself isn't Azure- or ARO-specific: native platform-level usage attribution, paired with an organizational label layer for the cases namespace boundaries and org charts don't agree. Amazon EKS (Elastic Kubernetes Service) has its own usage/billing correlation through Cost and Usage Reports and Kubecost-style tooling; Azure AKS (Kubernetes Service) and Google GKE (Kubernetes Engine) have their own native cost management surfaces. The mechanics below — the operator, the export cadence, the specific IAM roles — are the ARO implementation of that pattern, because that's the engagement this comes from. Swap the platform-specific pieces and the same two-pipeline shape holds.&lt;/p&gt;




&lt;h2&gt;
  
  
  Implementation
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Prerequisites
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Azure Red Hat OpenShift (ARO) cluster with &lt;code&gt;cluster-admin&lt;/code&gt; &lt;code&gt;oc&lt;/code&gt; access&lt;/li&gt;
&lt;li&gt;Azure subscription hosting the cluster, &lt;code&gt;az&lt;/code&gt; CLI logged in with rights on the cluster's resource group&lt;/li&gt;
&lt;li&gt;Red Hat Hybrid Cloud Console access with the Cloud Administrator role (or equivalent cost-management write access)&lt;/li&gt;
&lt;li&gt;Kubernetes 1.30+ (or an OpenShift version that ships it) for native &lt;code&gt;ValidatingAdmissionPolicy&lt;/code&gt; support — see Step 2 for the OPA Gatekeeper / Kyverno alternatives if you're not on a version that has it&lt;/li&gt;
&lt;li&gt;Everything below (storage account, export, service principal) is scoped to the single resource group that holds the ARO cluster — not the whole subscription&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Step 1 — Install the Cost Management Metrics Operator
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Verify the operator is available, then install via OperatorHub/Software Catalog&lt;/span&gt;
&lt;span class="c"&gt;# or apply directly:&lt;/span&gt;
&lt;span class="nb"&gt;cat&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt;&lt;span class="no"&gt;EOF&lt;/span&gt;&lt;span class="sh"&gt; | oc apply -f -
apiVersion: operators.coreos.com/v1
kind: OperatorGroup
metadata:
  name: costmanagement-metrics-operator
  namespace: costmanagement-metrics-operator
spec:
  targetNamespaces:
    - costmanagement-metrics-operator
&lt;/span&gt;&lt;span class="no"&gt;EOF

&lt;/span&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt;&lt;span class="no"&gt;EOF&lt;/span&gt;&lt;span class="sh"&gt; | oc apply -f -
apiVersion: operators.coreos.com/v1alpha1
kind: Subscription
metadata:
  name: costmanagement-metrics-operator
  namespace: costmanagement-metrics-operator
spec:
  channel: stable
  installPlanApproval: Automatic
  name: costmanagement-metrics-operator
  source: redhat-operators
  sourceNamespace: openshift-marketplace
&lt;/span&gt;&lt;span class="no"&gt;EOF
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then create the Hybrid Console service account (&lt;strong&gt;Settings → Identity &amp;amp; Access Management → Service Accounts&lt;/strong&gt;, added to a group with the Cloud Administrator role), store its &lt;code&gt;client_id&lt;/code&gt;/&lt;code&gt;client_secret&lt;/code&gt; as a cluster Secret, and apply a &lt;code&gt;CostManagementMetricsConfig&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;oc create namespace costmanagement-metrics-operator &lt;span class="nt"&gt;--dry-run&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;client &lt;span class="nt"&gt;-o&lt;/span&gt; yaml | oc apply &lt;span class="nt"&gt;-f&lt;/span&gt; -

&lt;span class="nb"&gt;cat&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt;&lt;span class="no"&gt;EOF&lt;/span&gt;&lt;span class="sh"&gt; | oc apply -f -
apiVersion: v1
kind: Secret
metadata:
  name: service-account-auth-secret
  namespace: costmanagement-metrics-operator
type: Opaque
stringData:
  client_id: "&amp;lt;CLIENT_ID&amp;gt;"
  client_secret: "&amp;lt;CLIENT_SECRET&amp;gt;"
&lt;/span&gt;&lt;span class="no"&gt;EOF

&lt;/span&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt;&lt;span class="no"&gt;EOF&lt;/span&gt;&lt;span class="sh"&gt; | oc apply -f -
apiVersion: costmanagement-metrics-cfg.openshift.io/v1beta1
kind: CostManagementMetricsConfig
metadata:
  name: costmanagementmetricscfg
  namespace: costmanagement-metrics-operator
spec:
  authentication:
    type: service-account
    secret_name: service-account-auth-secret
  packaging:
    max_reports_to_store: 30
    max_size_MB: 100
  prometheus_config:
    collect_previous_data: true
    context_timeout: 120
  source:
    check_cycle: 1440
    create_source: true
    name: aro-prod-cost
  upload:
    upload_cycle: 360
    upload_toggle: true
&lt;/span&gt;&lt;span class="no"&gt;EOF
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This engagement used the service-account auth type rather than the deprecated basic/token mode — worth calling out, since the token mode is what most quick-start examples default to. This is what produces native, project/namespace-level cost distribution from Prometheus/Thanos usage data — the baseline every other layer sits on top of.&lt;/p&gt;

&lt;p&gt;Rollback here is a clean uninstall: delete the &lt;code&gt;CostManagementMetricsConfig&lt;/code&gt;, then the &lt;code&gt;Subscription&lt;/code&gt; and &lt;code&gt;OperatorGroup&lt;/code&gt;, then the namespace. CMMO doesn't write anything back to the cluster's application workloads — removing it stops future uploads but doesn't touch anything already reported to the Hybrid Console.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2 — Enforce the cost-center label at admission time
&lt;/h3&gt;

&lt;p&gt;The requirement itself is simple to state: no &lt;code&gt;Deployment&lt;/code&gt; reaches production or gets scheduled onto a node without a &lt;code&gt;cost-center&lt;/code&gt; label present. This engagement enforced it with Kubernetes' own native admission policy — no external operator installed on the cluster.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What was actually used — native &lt;code&gt;ValidatingAdmissionPolicy&lt;/code&gt;.&lt;/strong&gt; CEL-based, GA from Kubernetes 1.30; confirm your OpenShift version ships it before relying on this path:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;admissionregistration.k8s.io/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ValidatingAdmissionPolicy&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;require-cost-center-policy"&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;failurePolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Fail&lt;/span&gt;
  &lt;span class="na"&gt;matchConstraints&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;resourceRules&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;apiGroups&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;apps"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
        &lt;span class="na"&gt;apiVersions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;v1"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
        &lt;span class="na"&gt;operations&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CREATE"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;UPDATE"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
        &lt;span class="na"&gt;resources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;deployments"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
  &lt;span class="na"&gt;validations&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;expression&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;has(object.metadata.labels)&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;&amp;amp;&amp;amp;&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;'cost-center'&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;in&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;object.metadata.labels"&lt;/span&gt;
      &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Deployments&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;must&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;include&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;'cost-center'&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;label."&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;admissionregistration.k8s.io/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ValidatingAdmissionPolicyBinding&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;require-cost-center-binding"&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;policyName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;require-cost-center-policy"&lt;/span&gt;
  &lt;span class="na"&gt;validationActions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Deny"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No new controller to install or upgrade, at the cost of being the newest and least battle-tested option on OpenShift specifically.&lt;/p&gt;

&lt;p&gt;If your cluster isn't on a version that ships native &lt;code&gt;ValidatingAdmissionPolicy&lt;/code&gt;, or you're already standardized on a policy engine, the same requirement maps onto either of these — shown here for readers on a different cluster, not what this engagement ran:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Option — OPA Gatekeeper&lt;/strong&gt;, via a constraint on &lt;code&gt;K8sRequiredLabels&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;constraints.gatekeeper.sh/v1beta1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;K8sRequiredLabels&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;require-cost-center-label&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;match&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;kinds&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;apiGroups&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;apps"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
        &lt;span class="na"&gt;kinds&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Deployment"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
  &lt;span class="na"&gt;parameters&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cost-center"&lt;/span&gt;
        &lt;span class="na"&gt;allowedRegex&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;^[a-z0-9-]+$"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This assumes the underlying &lt;code&gt;ConstraintTemplate&lt;/code&gt; for &lt;code&gt;K8sRequiredLabels&lt;/code&gt; is already installed from Gatekeeper's constraint template library — the constraint above references it, it doesn't define it. Pin to Gatekeeper v3.x.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Option — Kyverno&lt;/strong&gt;, via a &lt;code&gt;ClusterPolicy&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;kyverno.io/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ClusterPolicy&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;require-cost-center-label&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;validationFailureAction&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Enforce&lt;/span&gt;
  &lt;span class="na"&gt;background&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
  &lt;span class="na"&gt;rules&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;check-cost-center&lt;/span&gt;
      &lt;span class="na"&gt;match&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;any&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;resources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="na"&gt;kinds&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;Deployment&lt;/span&gt;
      &lt;span class="na"&gt;validate&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;The&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;'cost-center'&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;label&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;is&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;mandatory&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;on&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;all&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Deployments."&lt;/span&gt;
        &lt;span class="na"&gt;pattern&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="na"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="na"&gt;cost-center&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;?*"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;validationFailureAction: Enforce&lt;/code&gt; is what actually blocks the deploy — set to &lt;code&gt;Audit&lt;/code&gt; first if you want visibility before you start rejecting anything.&lt;/p&gt;

&lt;p&gt;No deployment artifact reaches a node without this label, regardless of which mechanism enforces it. Without it, Cost Management still gives you namespace-level attribution from CMMO — you just lose the department-level cut for any namespace that doesn't map cleanly to one cost center.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3 — Export Azure billing data and grant read access
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Storage account in the same resource group as the cluster.&lt;/span&gt;
&lt;span class="c"&gt;# public-network-access is disabled — reached only through a private&lt;/span&gt;
&lt;span class="c"&gt;# endpoint, not a public IP. Where a private endpoint isn't an option,&lt;/span&gt;
&lt;span class="c"&gt;# VNet rules plus IP firewall restrictions are the documented fallback.&lt;/span&gt;
az storage account create &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--name&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$CM_STORAGE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--resource-group&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$ARO_RG&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--location&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$ARO_LOCATION&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--sku&lt;/span&gt; Standard_LRS &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--kind&lt;/span&gt; StorageV2 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--min-tls-version&lt;/span&gt; TLS1_2 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--public-network-access&lt;/span&gt; Disabled

&lt;span class="c"&gt;# Service principal for Red Hat's read access — needs BOTH roles&lt;/span&gt;
&lt;span class="nv"&gt;SP_JSON&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;az ad sp create-for-rbac &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--name&lt;/span&gt; sp-rh-cost-management &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--role&lt;/span&gt; &lt;span class="s2"&gt;"Storage Blob Data Reader"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--scopes&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$EXPORT_SCOPE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--json-auth&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;

az role assignment create &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--assignee&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$SP_JSON&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; | jq &lt;span class="nt"&gt;-r&lt;/span&gt; .clientId&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--role&lt;/span&gt; &lt;span class="s2"&gt;"Cost Management Reader"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--scope&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$EXPORT_SCOPE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Store the returned &lt;code&gt;clientSecret&lt;/code&gt; immediately — it is not retrievable again after this command returns. Put it in Azure Key Vault if secrets are managed centrally, or in the CI/CD pipeline's own secure secret store (GitHub Actions Secrets, an Azure DevOps variable group) — never in a pipeline variable that lands in shell history or a plaintext log.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# CLI export creation needs the costmanagement extension and an existing container&lt;/span&gt;
az extension add &lt;span class="nt"&gt;--name&lt;/span&gt; costmanagement

az storage container create &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--account-name&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$CM_STORAGE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--name&lt;/span&gt; costexport &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--auth-mode&lt;/span&gt; login

&lt;span class="nv"&gt;STORAGE_ACCOUNT_ID&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"/subscriptions/&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;SUBSCRIPTION_ID&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;/resourceGroups/&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;ARO_RG&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;/providers/Microsoft.Storage/storageAccounts/&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;CM_STORAGE&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;

&lt;span class="c"&gt;# Recurrence start must be today or future — Azure CLI requirement&lt;/span&gt;
&lt;span class="nv"&gt;EXPORT_FROM&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;date&lt;/span&gt; &lt;span class="nt"&gt;-u&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'+1 day'&lt;/span&gt; +%Y-%m-%dT00:00:00Z&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="nv"&gt;EXPORT_TO&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;date&lt;/span&gt; &lt;span class="nt"&gt;-u&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'+2 years'&lt;/span&gt; +%Y-%m-%dT00:00:00Z&lt;span class="si"&gt;)&lt;/span&gt;

&lt;span class="c"&gt;# Daily actual-cost export, scoped to the resource group&lt;/span&gt;
az costmanagement &lt;span class="nb"&gt;export &lt;/span&gt;create &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--name&lt;/span&gt; rh-cost-export-daily &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--scope&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$EXPORT_SCOPE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--type&lt;/span&gt; ActualCost &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--timeframe&lt;/span&gt; MonthToDate &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--storage-account-id&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$STORAGE_ACCOUNT_ID&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--storage-container&lt;/span&gt; costexport &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--storage-directory&lt;/span&gt; daily &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--recurrence&lt;/span&gt; Daily &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--recurrence-period&lt;/span&gt; &lt;span class="nv"&gt;from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$EXPORT_FROM&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nv"&gt;to&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$EXPORT_TO&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--schedule-status&lt;/span&gt; Active
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If Red Hat Cost Management rejects the export schema, fall back to creating it through the Azure Portal instead — &lt;strong&gt;Cost Management + Billing → Cost export → Daily export&lt;/strong&gt;, using the "Cost and usage details (actual)" template. That's the path the guide points to when the CLI-created export doesn't validate; worth knowing before spending an afternoon debugging a schema error the CLI doesn't explain.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;Cost Management Reader&lt;/code&gt; role matters as much as &lt;code&gt;Storage Blob Data Reader&lt;/code&gt; — the read permission alone doesn't get you a functioning Azure integration in the Hybrid Console. Rollback is clean on this step in isolation: removing the export or the role assignment stops new billing data from landing, but doesn't touch what CMMO already reported from the cluster side.&lt;/p&gt;




&lt;h2&gt;
  
  
  Security Considerations
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Service principal scope on the Azure billing integration.&lt;/strong&gt; The service principal reading the daily cost export is, by definition, reading billing data that spans every namespace and every department on the cluster. If it's scoped more broadly than &lt;code&gt;Storage Blob Data Reader&lt;/code&gt; + &lt;code&gt;Cost Management Reader&lt;/code&gt; on the single resource group — subscription-level access, for instance — someone with access to the cost dashboard can infer more than cost: traffic volume and scaling patterns for departments other than their own. Confirmed as a real concern on this engagement; the mitigation is scoping strictly to the resource group holding the cluster, not the subscription.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Service principal credential storage.&lt;/strong&gt; The &lt;code&gt;az ad sp create-for-rbac&lt;/code&gt; command that provisions Red Hat's read access returns a client secret that Azure will not show again. Store it immediately in Azure Key Vault if secrets are managed centrally, or in the CI/CD pipeline's own &lt;a href="https://pipelineandprompts.com/posts/secrets-management-multi-cloud-pipelines/" rel="noopener noreferrer"&gt;secure secret store&lt;/a&gt; — GitHub Actions Secrets or an Azure DevOps variable group — so it never lands in shell history or a plaintext log.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;CMMO authentication mode.&lt;/strong&gt; This engagement used the Hybrid Console service-account auth type (&lt;code&gt;client_id&lt;/code&gt;/&lt;code&gt;client_secret&lt;/code&gt; stored as a cluster Secret) rather than the deprecated basic/token mode — worth stating explicitly, since token auth is what most quick-start examples default to and is the weaker of the two options for the cluster-to-console leg of this pipeline.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Read access on the cost dashboard itself is a separate control from the pipeline feeding it.&lt;/strong&gt; Scoping the service principal correctly limits what the &lt;em&gt;data pipeline&lt;/em&gt; can reach. It says nothing on its own about who can &lt;em&gt;view&lt;/em&gt; the resulting cost report. This engagement handled that separately, through Red Hat Hybrid Cloud Console User Access Groups — only authorized finance and department leads had access to aggregated cost metrics, distinct from whoever manages the underlying Azure integration.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The storage account itself.&lt;/strong&gt; The exported billing CSV sitting in the storage account is a sensitive, cross-department asset independent of anyone going through the Hybrid Console properly — someone with direct storage account access bypasses the Console's own access model entirely. This engagement secured it with Azure Private Endpoints, so the raw billing CSVs are never reachable over a public-facing IP; virtual network rules and IP firewall restrictions are the fallback where a private endpoint isn't an option.&lt;/p&gt;




&lt;h2&gt;
  
  
  Tradeoffs
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What you gain / what you give up — granularity vs. currency.&lt;/strong&gt; Two-layer attribution (native namespace/project plus label-based cost-center) gives you real per-application, per-department visibility that a flat cluster bill never could. What you give up is real-time accuracy: the ~24-hour cycle across CMMO's 6-hour uploads and Azure's daily cost export means the attributed-cost view is always behind. If you're making a same-day budget decision, you're making it against the faster but less granular native Azure Cost Management alert, not the project/label-attributed dashboard.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What you gain / what you give up — deployment friction vs. cost-center accuracy.&lt;/strong&gt; Blocking unlabeled deployments at admission time is what makes department-level attribution work for namespaces that don't map cleanly to one cost center. What it costs is friction: every deployment pipeline now has a hard dependency on getting the label right, and a missing or wrong label doesn't just skew a report, it blocks a deploy. That's a deliberate tradeoff, not an accident, but it needs to be communicated to every team shipping to the cluster before it surfaces as a support ticket.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Operability tradeoff — two independent pipelines instead of one, with a graceful (not flat) failure mode.&lt;/strong&gt; Running CMMO (cluster-side, usage-based) and the Azure Cost Export (billing-side, dollar-based) as two separate integrations that both have to stay active is more moving parts than a single source of truth would be. The upside of that separation: when the label layer goes down, attribution doesn't collapse to an even split across departments — it falls back to CMMO's native namespace/project view, which keeps working independently of the label/admission-policy layer. You lose the department-level cut, not all visibility.&lt;/p&gt;




&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;Cloud Without the Chaos #6 picks up the framework's fifth dimension directly: blast radius, tested honestly against a real failure instead of assumed on paper. This article's own two failure points — an over-scoped billing integration, and a bypassed label policy that degrades gracefully rather than catastrophically — are a small preview of that larger question.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;There's no companion repository for this one. This documents a cost-attribution architecture and a set of operational decisions specific to one ARO engagement, not a general-purpose tool — the Gatekeeper/Kyverno alternatives in Step 2 are illustrative, not what was actually run here.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Written by Pipeline &amp;amp; Prompts | Byte size guides on DevOps, Cloud and AI&lt;/em&gt;&lt;/p&gt;

</description>
      <category>costattribution</category>
      <category>aro</category>
      <category>openshift</category>
      <category>cloudwithoutthechaos</category>
    </item>
    <item>
      <title>Built a Hand-Rolled Agent Loop and Found a Weird Retrieval Bias</title>
      <dc:creator>Nerav Doshi</dc:creator>
      <pubDate>Thu, 24 Sep 2026 22:17:53 +0000</pubDate>
      <link>https://dev.to/agenticdevops/built-a-hand-rolled-agent-loop-and-found-a-weird-retrieval-bias-11c4</link>
      <guid>https://dev.to/agenticdevops/built-a-hand-rolled-agent-loop-and-found-a-weird-retrieval-bias-11c4</guid>
      <description>&lt;p&gt;Every "agent" framework boils down to the same three-step loop underneath: look at the world, decide what to do, do it. I wanted to build that by hand instead of reaching for a framework, mostly because I've never actually seen the seams of one up close, and frameworks are good at hiding exactly the parts I wanted to look at.&lt;/p&gt;

&lt;p&gt;Started with a completely fake skeleton, hardcoded task and all, just to prove the shape works before adding anything real:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;observe&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;task: find out how to check pod status with oc&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;decide&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;observation&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;search_notes: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;observation&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;act&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Would call: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;obs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;observe&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;plan&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;decide&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;act&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;plan&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Fine, obviously — it's just string formatting at this point. The real test was replacing &lt;code&gt;decide()&lt;/code&gt; with an actual model call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;ollama&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;decide&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;observation&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;You are an agent with one tool available: search_notes(query).
It searches a local knowledge base and returns relevant notes.

Given this task: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;observation&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

Respond with ONLY the exact search query you&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;d pass to search_notes — nothing else, no explanation.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ollama&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;llama3.2:1b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;response&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;First two runs gave me the identical output, which made me assume it was deterministic. It wasn't — a third run proved that wrong immediately. Three runs, three different behaviors:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A full &lt;code&gt;oc&lt;/code&gt; command wrapped in backticks — the model ignored the "respond with only a query" instruction entirely and answered the underlying question instead. Worth noting: it's the exact same flawed command from Entry 02 (&lt;code&gt;--all-namespaces&lt;/code&gt; instead of scoping to the actual namespace).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;find -f 'oc status pod*' -A -v&lt;/code&gt; — not a real command in any tool, just a strange mashup of &lt;code&gt;find&lt;/code&gt; syntax and &lt;code&gt;oc&lt;/code&gt; concepts.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;How to check pod status with oc&lt;/code&gt; — the one time it actually did what was asked.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That instability is itself the finding. A "decide" step built on a small local model with no real constraint around it isn't reliable even for one fixed task run back to back — which is a pretty direct argument for why Entry 02's system-prompt approach exists in the first place.&lt;/p&gt;

&lt;p&gt;Wired &lt;code&gt;act()&lt;/code&gt; to the real MCP tool from Entry 08 next, to see what each of those three outputs actually retrieved when fed to real search — not through a full MCP client this time, just calling the underlying function directly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;sys&lt;/span&gt;
&lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;insert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;mcp_search_server&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;search_notes&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;act&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;search_notes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;act&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;`oc get pods --all-namespaces -o jsonpath=&lt;/span&gt;&lt;span class="se"&gt;\'&lt;/span&gt;&lt;span class="s"&gt;{.items[*].metadata.name}&lt;/span&gt;&lt;span class="se"&gt;\'&lt;/span&gt;&lt;span class="s"&gt;`&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;---&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;act&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;find -f &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;oc status pod*&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt; -A -v&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;---&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;act&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;How to check pod status with oc&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Worth noting: &lt;code&gt;search_notes&lt;/code&gt; is wrapped in &lt;code&gt;@mcp.tool()&lt;/code&gt;, and I genuinely wasn't sure that would still work as a plain callable outside a real MCP session. It did — no errors, clean results, all three queries went straight through to the Chroma index.&lt;/p&gt;

&lt;p&gt;And this is where it got genuinely strange. All three queries surfaced the same top document — &lt;code&gt;02-oc-cli-mentor-system-prompt.md&lt;/code&gt;, which makes sense, it's the only stored content actually about &lt;code&gt;oc&lt;/code&gt;. But the distances ran backwards from what I expected:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Decide() output&lt;/th&gt;
&lt;th&gt;Distance to correct match&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Garbled backtick command&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;324.3&lt;/strong&gt; (best)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nonsensical &lt;code&gt;find&lt;/code&gt;/&lt;code&gt;oc&lt;/code&gt; hybrid&lt;/td&gt;
&lt;td&gt;378.7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Clean, correct query&lt;/td&gt;
&lt;td&gt;430.1 (worst)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The most broken query retrieved the correct document with the tightest confidence. The one time the model actually did what I asked — produced a clean, sensible search query — retrieved the &lt;em&gt;same&lt;/em&gt; correct document, but with noticeably weaker confidence than either of the broken attempts.&lt;/p&gt;

&lt;p&gt;I don't have a confirmed reason for this, and I'd rather say that plainly than invent one. My best guess: Entry 02's actual content contains real &lt;code&gt;oc&lt;/code&gt; command syntax, and the garbled query text — being made of similar command-like tokens — may be matching on surface-level lexical overlap rather than pure semantic meaning. That's a testable claim, not a settled one, and it's worth checking directly before trusting it.&lt;/p&gt;

&lt;p&gt;So the loop works end to end — a real model decision, feeding a real search tool, returning real results. But this entry didn't end where I expected. The interesting result wasn't "the agent works," it's that the search layer built across Entries 05-08 might be rewarding queries that look like the stored content syntactically, more than queries that actually match its meaning. If that's true, it's a real weakness in the retrieval system itself, not just in this one flaky decide() step — worth its own entry to actually test properly.&lt;/p&gt;

</description>
      <category>ollama</category>
      <category>mcp</category>
      <category>agents</category>
      <category>rag</category>
    </item>
    <item>
      <title>What EU Data Sovereignty Actually Requires Beyond Hosting</title>
      <dc:creator>Nerav Doshi</dc:creator>
      <pubDate>Mon, 21 Sep 2026 22:01:03 +0000</pubDate>
      <link>https://dev.to/agenticdevops/data-sovereignty-in-practice-what-its-in-an-eu-data-centre-actually-covers-and-doesnt-i2o</link>
      <guid>https://dev.to/agenticdevops/data-sovereignty-in-practice-what-its-in-an-eu-data-centre-actually-covers-and-doesnt-i2o</guid>
      <description>&lt;p&gt;&lt;em&gt;Pipeline &amp;amp; Prompts | Byte size guides on DevOps, Cloud and AI&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;☁️ Cloud Without the Chaos #4&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;⚡ Byte Size Summary&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why "the data is in an EU data centre" doesn't answer a sovereignty question. Support access, logs, and metrics are separate surfaces, and each needs its own boundary&lt;/li&gt;
&lt;li&gt;What happened when a blanket EU-only routing fix broke a 2-hour P1 SLA, and the tier we built instead&lt;/li&gt;
&lt;li&gt;Why sovereignty isn't something you automate. You can only shrink the surface area where a human has to touch the boundary&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  The Story
&lt;/h2&gt;

&lt;p&gt;The data residency box was already checked. Managed Kubernetes clusters, multi-region, customer workload data pinned to EU regions from day one. That part of the compliance conversation had been closed for months.&lt;/p&gt;

&lt;p&gt;Then a support ticket came in from an EU customer's compliance team, and it wasn't about where the data lived. It was about who could see it while fixing a problem.&lt;/p&gt;

&lt;p&gt;Our support model at the time was 24x7 global follow-the-sun. A P1 ticket routed to whichever region had bench depth at that hour, which meant a customer's cluster could get worked by an engineer sitting in Asia at 3am their time. That engineer had platform-only access: control plane nodes and management-layer components, no customer workload data, no application logs, no application namespaces. Scoped correctly, by design, and had been for a long time.&lt;/p&gt;

&lt;p&gt;The customer objected anyway. Not because anything had actually been accessed improperly — nothing had. The objection was on principle, GDPR-era: personal data protections don't stop being a concern just because the access is read-only and platform-scoped. Their compliance team wasn't comfortable with a data processing agreement that allowed any support engineer, anywhere, any kind of access path into infrastructure holding EU personal data, even platform-adjacent access with no visibility into the data itself.&lt;/p&gt;

&lt;p&gt;That's the sentence that reframes the whole problem: "it's in an EU data centre" answers a data residency question. It says nothing about who can touch the systems around that data, from where, under what access model, logged how. Sovereignty isn't one control — it's one of five dimensions worth &lt;a href="https://dev.to/posts/hybrid-cloud-architecture-on-prem-vs-cloud-tradeoffs/"&gt;placing deliberately&lt;/a&gt;, and it's a set of surfaces in its own right: storage location, support access, logs, metrics, key management. Each one needs its own answer.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;Platform engineers and SREs running managed Kubernetes clusters for customers under multiple compliance regimes at once feel this pain first. GDPR, SOC 2, and industry-specific overlays on top — not SOC 2 alone. GDPR puts a human body in a specific relationship to EU personal data. SOC 2 wants an audit trail: who touched what, when. Industry-specific requirements, financial services and healthcare especially, stack their own access-logging and retention rules on top of both.&lt;/p&gt;

&lt;p&gt;What breaks without an explicit answer to "who can access this, from where, under what conditions" is trust, not uptime. The systems keep running. The compliance team stops trusting the DPA. And once one EU customer's legal team flags a sovereignty gap, it doesn't stay contained to that account. It becomes a question every other EU prospect's procurement team asks in the next sales cycle.&lt;/p&gt;

&lt;p&gt;The cost shows up as engineering time spent redesigning access models after the fact, a support organization now split along a line that didn't exist before, and a rollback path that isn't a policy flip. It's staffing and contract changes.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Existing Approaches Fall Short
&lt;/h2&gt;

&lt;p&gt;The obvious first move is blanket EU-only ticket routing: if the customer is in the EU, route the ticket to EU-based support staff, full stop. We tried this. It broke the 2-hour P1 production SLA, because EU bench depth wasn't built for follow-the-sun coverage. It was built assuming tickets could route globally. A P1 that hits outside EU business hours with no EU engineer awake to take it either blows the SLA or forces someone to page a much smaller on-call rotation than the global model ever needed.&lt;/p&gt;

&lt;p&gt;The other common move is to treat this as a policy document problem: write a stricter DPA, keep the access model the same. That doesn't survive contact with an auditor or a customer's own legal review. A written policy that says "support engineers only access what they need" isn't a control, it's a description of intent. GDPR and SOC 2 auditors want the boundary enforced in the access model itself: who has a credential, what that credential can reach, how long it's valid, what's logged when it's used. A policy without RBAC and logging behind it fails the first real audit.&lt;/p&gt;

&lt;p&gt;Neither actually answers what the customer raised. Blanket geographic routing solves for where the person sits, not what they can access and for how long. A stricter policy document solves for looking compliant, not for being enforceable.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Architecture
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F32cd9vlvlewkibg54jzd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F32cd9vlvlewkibg54jzd.png" alt="Architecture Diagram" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The diagram shows the EU sovereign support tier as a distinct zone from the global support pool, with the RBAC boundary (read-only default, short-lived-token break-glass) drawn explicitly between them, and the EU-boundary-contained data plane (logs, metrics, KMS keys) called out separately from the cluster control plane. What it needs to prove, not just show: two independent boundaries have to hold — an access boundary and a data boundary — not one.&lt;/p&gt;

&lt;p&gt;The architecture that replaced blanket routing is sized per hosting environment, because not every managed cluster offering draws from the same support staffing pool. Different clusters can sit on different managed platforms with different regional engineer depth, so a single "EU tier" designed around one platform's bench doesn't automatically map onto another. Each hosting environment got its own EU sovereign support tier, sized to its own bench.&lt;/p&gt;

&lt;p&gt;Staffing is business-hours-plus-pager-crew, not full 24x7 in-region rotation. Building a true 24x7 EU-only rotation for each platform would have meant hiring at a scale the customer base didn't justify yet. The pager crew covers the overnight P1 gap without requiring a full parallel global team inside the EU boundary.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Control plane.&lt;/strong&gt; RBAC default is read-only for the EU-scoped support tier. Elevation beyond read-only goes through an automated break-glass path issuing short-lived tokens — no standing write access, no permanent elevated role sitting unused between incidents. Every break-glass elevation is itself an auditable event.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Data plane.&lt;/strong&gt; Debug logs, metrics, and &lt;a href="https://dev.to/posts/secrets-management-multi-cloud-pipelines/"&gt;KMS keys&lt;/a&gt; used for this support tier stay entirely inside the EU boundary. This is the piece "the data is in an EU data centre" misses: it's not enough that the customer's application data stays in-region if the diagnostic exhaust generated while fixing their problem doesn't.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Knowledge gap.&lt;/strong&gt; A smaller EU-scoped expert pool doesn't have the same depth of tribal knowledge as the full global team. An internal &lt;a href="https://dev.to/posts/ai-in-the-stack-02-rag-runbooks/"&gt;RAG system&lt;/a&gt; built from internal runbooks and past incident resolutions gives the EU pool a way to close knowledge gaps without routing questions outside the boundary.&lt;/p&gt;

&lt;p&gt;Blast radius is the useful lens here. The failure mode the old model exposed wasn't a security incident — nothing was ever improperly accessed. It was an unbounded set of humans who &lt;em&gt;could&lt;/em&gt; touch the boundary, with no way to prove the boundary held other than trusting intent. The new model bounds that set explicitly: fewer people, less standing access, shorter-lived elevation, every elevation logged.&lt;/p&gt;




&lt;h2&gt;
  
  
  Implementation
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Prerequisites
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Separate RBAC groups per hosting environment for the EU sovereign support tier, distinct from the global support RBAC groups&lt;/li&gt;
&lt;li&gt;Break-glass tooling capable of issuing short-lived tokens tied to an incident ticket ID, not a standing credential&lt;/li&gt;
&lt;li&gt;EU-region log and metrics storage isolated from the global aggregation pipeline — an infrastructure decision, not just an access-policy one&lt;/li&gt;
&lt;li&gt;An internal knowledge base with enough historical runbook and incident content to make a RAG system useful to a smaller support pool&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Step 1 — Scope the RBAC boundary per hosting environment
&lt;/h3&gt;

&lt;p&gt;Default role for the EU support tier is read-only against control plane resources. No default write access, no default access beyond the control plane nodes and management-layer components already established under the original platform-only scoping.&lt;/p&gt;

&lt;p&gt;Rollback consideration: removing this RBAC group is reversible on its own, but doing so without also reverting the log and metrics isolation in Step 3 leaves the boundary half-enforced. Access reverts to global scope while the data still lives in an EU-only pipeline nobody outside the EU tier can query during an incident.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2 — Wire break-glass elevation to short-lived tokens
&lt;/h3&gt;

&lt;p&gt;Elevation above read-only requires an incident ticket ID as input and issues a token scoped to a defined TTL, not a standing elevated role. The token expires; there's no cleanup step where someone has to remember to revoke it.&lt;/p&gt;

&lt;p&gt;Rollback consideration: if break-glass tooling fails or is unavailable during an incident, the fallback can't quietly become "give someone standing write access." That recreates the exact problem this architecture exists to avoid. The fallback has to be a manual, logged, time-boxed exception process, not a bypass.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3 — Isolate EU support logs, metrics, and KMS keys
&lt;/h3&gt;

&lt;p&gt;Diagnostic data generated by EU-tier support activity — debug logs, metrics scraped during troubleshooting, keys used for that access — needs to land in EU-region storage, not the global aggregation pipeline the rest of support uses.&lt;/p&gt;

&lt;p&gt;Rollback consideration: this is the expensive one to unwind. Once EU-only log storage exists and EU support workflows depend on querying it, reverting means either migrating that data's access model back to global or maintaining two parallel pipelines indefinitely. Not a flag flip.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4 — Stand up the internal RAG system for the EU pool
&lt;/h3&gt;

&lt;p&gt;Index internal runbooks, past incident write-ups, and platform-specific troubleshooting docs. Scope retrieval to what the EU support tier actually needs. This closes the tribal-knowledge gap between a smaller regional pool and the full global team without routing questions outside the boundary.&lt;/p&gt;

&lt;p&gt;Rollback consideration: this piece is closer to a convenience than a compliance control. Removing it degrades EU-tier effectiveness but doesn't reopen the sovereignty gap on its own — it isn't carrying regulated data across the boundary.&lt;/p&gt;




&lt;h2&gt;
  
  
  Security Considerations
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Break-glass is the attack surface, not the read-only default.&lt;/strong&gt; A read-only RBAC default is easy to reason about. The real risk sits in the elevation path: if the break-glass tooling has a weak authentication step, or short-lived tokens get logged in plaintext somewhere they shouldn't, an attacker doesn't need to compromise the sovereignty boundary directly. They need to compromise the mechanism that grants exceptions to it. Every break-glass elevation needs its own audit trail, separate from the general access logs, because it's the highest-value target in this design.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Log and metrics isolation can leak through the RAG system if scoped wrong.&lt;/strong&gt; The internal RAG system exists to close a knowledge gap for the EU support pool. But if it's indexed against a shared internal knowledge base that includes non-EU incident write-ups referencing customer-identifying details from other regions, the RAG system itself becomes an unintended cross-boundary data path. Scoping what gets indexed, and confirming incident write-ups used as retrieval material are scrubbed of cross-customer identifying detail, matters as much as scoping the RBAC.&lt;/p&gt;




&lt;h2&gt;
  
  
  Tradeoffs
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What you gain:&lt;/strong&gt; an enforceable, auditable answer to "who can access what, from where" that survives a GDPR or SOC 2 audit — an access model with logs behind it, not a policy document. You also gain a defensible sales conversation: the next EU prospect's compliance team gets a specific, technical answer instead of "trust our intent."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What you give up:&lt;/strong&gt; SLA resilience during off-hours EU incidents. Business-hours-plus-pager-crew staffing is a real tradeoff against the old follow-the-sun model's bench depth. The EU tier will always have thinner overnight coverage than the global pool did, and that's a permanent operational cost, not a startup pain that goes away with scale.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Operability tradeoff:&lt;/strong&gt; two support models now exist where one did before, and every team supporting these clusters has to maintain both. That's ongoing operational overhead, not a one-time migration cost. Every future support tooling change has to account for the EU tier's separate RBAC scoping and separate log pipeline, or it risks quietly breaking the boundary the whole architecture exists to hold.&lt;/p&gt;




&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Cloud Without the Chaos #5 — &lt;a href="https://dev.to/posts/cost-profile-demand-shape-cloud-elasticity/"&gt;Kubernetes Cost Attribution: Namespace vs. Cost-Center&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Dimension 3 of the original five-dimension framework, walked through with a real before/after cost model.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Written by Pipeline &amp;amp; Prompts | Byte size guides on DevOps, Cloud and AI&lt;/em&gt;&lt;/p&gt;

</description>
      <category>datasovereignty</category>
      <category>kubernetes</category>
      <category>compliance</category>
      <category>cloudwithoutthechaos</category>
    </item>
    <item>
      <title>Found the Real Kubernetes Memory Ceiling — It Wasn't Double</title>
      <dc:creator>Nerav Doshi</dc:creator>
      <pubDate>Mon, 21 Sep 2026 22:01:01 +0000</pubDate>
      <link>https://dev.to/agenticdevops/found-the-real-kubernetes-memory-ceiling-it-wasnt-double-53b0</link>
      <guid>https://dev.to/agenticdevops/found-the-real-kubernetes-memory-ceiling-it-wasnt-double-53b0</guid>
      <description>&lt;p&gt;Last entry left a real question hanging: the exact Podman number got killed under &lt;code&gt;kind&lt;/code&gt;, twice. Did that mean the gap between the two runtimes was huge — needing something close to double the memory — or was the real ceiling only a little higher than 1663Mi, and I'd just clipped it? Only one way to find out.&lt;/p&gt;

&lt;p&gt;Redeployed with the limit bumped to 2048Mi — a deliberately big jump, so the answer would be unambiguous either way:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl delete pod ollama-sized
kubectl apply &lt;span class="nt"&gt;-f&lt;/span&gt; ollama-pod-2gb.yaml
kubectl &lt;span class="nb"&gt;exec&lt;/span&gt; &lt;span class="nt"&gt;-it&lt;/span&gt; ollama-2gb &lt;span class="nt"&gt;--&lt;/span&gt; ollama pull llama3.2:1b
kubectl &lt;span class="nb"&gt;exec&lt;/span&gt; &lt;span class="nt"&gt;-it&lt;/span&gt; ollama-2gb &lt;span class="nt"&gt;--&lt;/span&gt; ollama run llama3.2:1b
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Asked it a real question, checked the outcome:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl top pod ollama-2gb
kubectl describe pod ollama-2gb | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-A5&lt;/span&gt; &lt;span class="s2"&gt;"Last State&lt;/span&gt;&lt;span class="se"&gt;\|&lt;/span&gt;&lt;span class="s2"&gt;Restart Count"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No crash. &lt;code&gt;Restart Count: 0&lt;/code&gt;, and &lt;code&gt;kubectl top&lt;/code&gt; showed real usage of &lt;strong&gt;1775Mi — about 1.73GB.&lt;/strong&gt; Which answers the question, and not the way I expected: that's only about 6.7% higher than Podman's 1.663GB. Not anywhere near double. The 2Gi limit worked, but it turns out to have been way more generous than actually necessary — the real ceiling sits much closer to 1663Mi than the size of my jump would suggest.&lt;/p&gt;

&lt;p&gt;So that reframes last entry's open question. This wasn't a big structural gap between how Podman and Kubernetes handle memory. It was a narrow miss. Setting the limit at exactly a steady-state number, with zero margin for whatever memory spike happens during model loading before things settle down, was always going to be right on the edge of failing. Give it something like 10% headroom above the measured number, and it holds fine.&lt;/p&gt;

&lt;p&gt;Which is really the actual lesson from this whole stretch of entries, more than "runtimes differ": a memory snapshot taken after a model is already loaded and idle isn't a safe number to use as a hard limit on its own. It doesn't account for the peak during loading. Going forward, my rule is measured steady-state number plus 10-15%, used as the real limit — not the raw number itself. 1663Mi measured should've meant something like 1875-1900Mi requested, and Entry 11's whole detour probably never happens.&lt;/p&gt;

</description>
      <category>ollama</category>
      <category>kubernetes</category>
      <category>kind</category>
      <category>resourceplanning</category>
    </item>
    <item>
      <title>The Same Memory Number That Worked in Podman Got OOMKilled in Kubernetes</title>
      <dc:creator>Nerav Doshi</dc:creator>
      <pubDate>Fri, 18 Sep 2026 19:49:52 +0000</pubDate>
      <link>https://dev.to/agenticdevops/the-same-memory-number-that-worked-in-podman-got-oomkilled-in-kubernetes-1g2k</link>
      <guid>https://dev.to/agenticdevops/the-same-memory-number-that-worked-in-podman-got-oomkilled-in-kubernetes-1g2k</guid>
      <description>&lt;p&gt;&lt;code&gt;OOMKilled&lt;/code&gt; is what Kubernetes says when a container tries to use more memory than its limit allows and the kernel steps in and kills it. Going into this one, my assumption was straightforward: Entry 09 measured 1.663GB under Podman, so setting a Kubernetes limit to that exact number should be enough. Same model, same host, same measurement. That assumption was wrong, and finding out why turned into the most interesting result in the series so far.&lt;/p&gt;

&lt;p&gt;Deployed the same image to a fresh &lt;code&gt;kind&lt;/code&gt; cluster with &lt;code&gt;requests.memory&lt;/code&gt; and &lt;code&gt;limits.memory&lt;/code&gt; both set to the precise Podman figure — 1663Mi. Pulled the model, tried to load it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl &lt;span class="nb"&gt;exec&lt;/span&gt; &lt;span class="nt"&gt;-it&lt;/span&gt; ollama-sized &lt;span class="nt"&gt;--&lt;/span&gt; ollama pull llama3.2:1b
kubectl &lt;span class="nb"&gt;exec&lt;/span&gt; &lt;span class="nt"&gt;-it&lt;/span&gt; ollama-sized &lt;span class="nt"&gt;--&lt;/span&gt; ollama run llama3.2:1b
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Pull worked. Loading gave me &lt;code&gt;command terminated with exit code 137&lt;/code&gt; — no real error message, just a dead session. &lt;code&gt;kubectl describe pod&lt;/code&gt; filled in what actually happened:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Last State:     Terminated
  Reason:       OOMKilled
  Exit Code:    137
Restart Count:  2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Killed twice, for exceeding the exact number that had worked cleanly under Podman on the same machine, for the same model.&lt;/p&gt;

&lt;p&gt;I don't have a confirmed explanation yet, and I'd rather say that plainly than make one up. A few real possibilities: the Podman number was a snapshot taken after the model had already loaded and settled, not a peak — memory during actual loading (KV cache setup, context buffers) might spike higher than that steady-state figure ever captured. Or cgroup memory accounting just differs between runtimes — Kubernetes' limit on containerd typically counts page cache against you, Podman's VM might not count it the same way. Or the two VMs underneath — Podman's &lt;code&gt;applehv&lt;/code&gt; machine, &lt;code&gt;kind&lt;/code&gt;'s node — have different baseline overhead that was never separately measured. Any of these could be true. None of them are confirmed.&lt;/p&gt;

&lt;p&gt;What is confirmed: a real, measured number from one container runtime doesn't automatically carry over to another, even on identical hardware running the identical model. "It worked under Podman" turned out to be necessary, not sufficient. The honest next move isn't guessing a bigger number and hoping — it's actually finding the real ceiling under &lt;code&gt;kind&lt;/code&gt; and seeing how far off 1663Mi was.&lt;/p&gt;

</description>
      <category>ollama</category>
      <category>kubernetes</category>
      <category>kind</category>
      <category>resourceplanning</category>
    </item>
    <item>
      <title>Deployed Ollama to Local Kubernetes With No Resource Limits</title>
      <dc:creator>Nerav Doshi</dc:creator>
      <pubDate>Fri, 18 Sep 2026 19:49:50 +0000</pubDate>
      <link>https://dev.to/agenticdevops/deployed-ollama-to-local-kubernetes-with-no-resource-limits-1eio</link>
      <guid>https://dev.to/agenticdevops/deployed-ollama-to-local-kubernetes-with-no-resource-limits-1eio</guid>
      <description>&lt;p&gt;Before testing whether last entry's real Podman number (1.663GB) holds up as an actual Kubernetes resource limit, I wanted a baseline: what happens with nothing set at all. &lt;code&gt;kind&lt;/code&gt; runs a full local Kubernetes cluster inside containers acting as nodes — Kubernetes on top of the same containers everything else in this series has been using.&lt;/p&gt;

&lt;p&gt;Spun up a fresh cluster and deployed Ollama with no &lt;code&gt;resources&lt;/code&gt; section in the pod spec whatsoever:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kind create cluster &lt;span class="nt"&gt;--name&lt;/span&gt; today-i-ran
kubectl apply &lt;span class="nt"&gt;-f&lt;/span&gt; ollama-pod-naive.yaml
kubectl &lt;span class="nb"&gt;exec&lt;/span&gt; &lt;span class="nt"&gt;-it&lt;/span&gt; ollama-naive &lt;span class="nt"&gt;--&lt;/span&gt; ollama pull llama3.2:1b
kubectl &lt;span class="nb"&gt;exec&lt;/span&gt; &lt;span class="nt"&gt;-it&lt;/span&gt; ollama-naive &lt;span class="nt"&gt;--&lt;/span&gt; ollama run llama3.2:1b
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Asked it a real question — "What does BGP do in OpenShift Networking" — to force a full response, then checked the actual resource configuration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl describe pod ollama-naive | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-A5&lt;/span&gt; &lt;span class="s2"&gt;"Limits&lt;/span&gt;&lt;span class="se"&gt;\|&lt;/span&gt;&lt;span class="s2"&gt;Requests"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Pulled and loaded with zero constraints. &lt;code&gt;QoS Class: BestEffort&lt;/code&gt;, which is Kubernetes plainly telling you no requests or limits were set. One real annoyance worth flagging: the image pull took &lt;strong&gt;4 minutes 26 seconds&lt;/strong&gt; on this fresh node, against maybe 30 seconds under Podman in the last entry. A brand new &lt;code&gt;kind&lt;/code&gt; cluster has nothing cached — the first pull anywhere is always going to be the slow one.&lt;/p&gt;

&lt;p&gt;Also worth being honest about: the model's actual answer to the OpenShift networking question wasn't good. It invented a fake expansion for "OpenShift Networking" and made up BGP/BGPsec behavior that doesn't match how OpenShift actually works. A 1B local model sounding confident on something specialized like this is a different risk than the more mechanical &lt;code&gt;oc&lt;/code&gt; questions from earlier entries — worth remembering not to trust this stuff just because it sounds sure of itself.&lt;/p&gt;

&lt;p&gt;With no limit, Kubernetes will let a pod take whatever it wants — which is exactly why this baseline matters. Without it, there'd be no way to tell whether the Podman number from last entry was actually a safe ceiling or not. Spoiler for the next one: it wasn't.&lt;/p&gt;

</description>
      <category>ollama</category>
      <category>kubernetes</category>
      <category>kind</category>
      <category>resourceplanning</category>
    </item>
    <item>
      <title>Containerized Ollama and Found the Real Memory Overhead</title>
      <dc:creator>Nerav Doshi</dc:creator>
      <pubDate>Wed, 09 Sep 2026 00:47:46 +0000</pubDate>
      <link>https://dev.to/agenticdevops/containerized-ollama-and-found-the-real-memory-overhead-j87</link>
      <guid>https://dev.to/agenticdevops/containerized-ollama-and-found-the-real-memory-overhead-j87</guid>
      <description>&lt;p&gt;On Linux, containers run natively — straight on the host kernel. On macOS, they can't, because containers need a Linux kernel underneath and your Mac isn't running one. So Podman and Docker Desktop quietly spin up a small Linux VM in the background and put every container inside that instead. Which means every container on a Mac is sharing a fixed slice of memory carved out for that VM — not your Mac's actual RAM — and that distinction is about to matter a lot.&lt;/p&gt;

&lt;p&gt;Started Ollama as a container, exposed on a different port so it wouldn't collide with the native app already running:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;podman run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;--name&lt;/span&gt; ollama-container &lt;span class="nt"&gt;-p&lt;/span&gt; 11435:11434 ollama/ollama
podman &lt;span class="nb"&gt;exec&lt;/span&gt; &lt;span class="nt"&gt;-it&lt;/span&gt; ollama-container ollama pull llama3.2:1b
podman &lt;span class="nb"&gt;exec&lt;/span&gt; &lt;span class="nt"&gt;-it&lt;/span&gt; ollama-container ollama run llama3.2:1b
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Pull went fine, after one transient network hiccup. Loading the model didn't:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Error: 500 Internal Server Error: model requires more system memory (1.3 GiB) than is available (620.8 MiB)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;podman machine list&lt;/code&gt; explained why immediately — the VM backing every container on this machine was set to &lt;strong&gt;2GiB total&lt;/strong&gt;, for the OS, the runtime, and every container combined. After overhead, there was only ~620MB actually free. Nowhere close to enough.&lt;/p&gt;

&lt;p&gt;Fixed by resizing the VM (has to be stopped first):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;podman machine stop
podman machine &lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="nt"&gt;--memory&lt;/span&gt; 4096
podman machine start
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That restart also killed the container, which then refused to &lt;code&gt;exec&lt;/code&gt; into ("container state improper") until I explicitly restarted it too:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;podman start ollama-container
podman &lt;span class="nb"&gt;exec&lt;/span&gt; &lt;span class="nt"&gt;-it&lt;/span&gt; ollama-container ollama run llama3.2:1b
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Worked that time. Asked it "what is kubernetes?" to force a real response, then pulled the memory number with Podman's version of &lt;code&gt;docker stats&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;podman stats ollama-container &lt;span class="nt"&gt;--no-stream&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Entry 01 (bare metal, native app)&lt;/th&gt;
&lt;th&gt;Containerized&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Process memory&lt;/td&gt;
&lt;td&gt;~1.24 GB RSS&lt;/td&gt;
&lt;td&gt;1.663 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Processor&lt;/td&gt;
&lt;td&gt;100% GPU (Metal)&lt;/td&gt;
&lt;td&gt;CPU only — no Metal passthrough in a container&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;About &lt;strong&gt;34% more memory&lt;/strong&gt; for the exact same model. Some of that's Podman/Ollama server overhead, some is probably the lack of GPU acceleration pushing more work onto the CPU side — that second part's an inference from the numbers, not something I directly measured, so I'm not going to state it more confidently than that. Either way, "same model, same memory" turned out to be wrong. Containerizing this isn't memory-neutral.&lt;/p&gt;

&lt;p&gt;Which means the resource request I'd actually set for this model in a container should be closer to 1.663GB than the 1.24GB bare-metal number from Entry 01 — sizing off the bare-metal figure alone would under-provision it. And separately: Podman's VM has its own memory ceiling, completely independent of how much RAM your Mac actually has. That's the first thing worth checking before assuming a model is just too big to run.&lt;/p&gt;

</description>
      <category>ollama</category>
      <category>podman</category>
      <category>containers</category>
      <category>resourceplanning</category>
    </item>
    <item>
      <title>Wired the Local MCP Server into Claude Code</title>
      <dc:creator>Nerav Doshi</dc:creator>
      <pubDate>Wed, 02 Sep 2026 19:58:27 +0000</pubDate>
      <link>https://dev.to/agenticdevops/wired-the-local-mcp-server-into-claude-code-3e1f</link>
      <guid>https://dev.to/agenticdevops/wired-the-local-mcp-server-into-claude-code-3e1f</guid>
      <description>&lt;p&gt;Everything up to this point — building the Chroma index, chunking, querying — was done by hand, one Python command at a time. An MCP server is what turns that into an actual tool: a small program exposing specific capabilities to an AI client, so it can call your stuff directly instead of you running queries yourself every time. What I actually wanted to prove here wasn't that the server starts. It's that a real client can reach it and get something useful back.&lt;/p&gt;

&lt;p&gt;After the &lt;code&gt;Connection closed&lt;/code&gt; error from last time — Claude Code was launching a bare &lt;code&gt;python3&lt;/code&gt;, which resolved to system Python outside the venv, missing every package the server needed — I re-registered it pointing straight at the venv's Python:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp remove today-i-ran-notes
claude mcp add today-i-ran-notes &lt;span class="nt"&gt;--&lt;/span&gt; ~/mcp-env/bin/python3 mcp_search_server.py
claude mcp list
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Got a &lt;code&gt;✔ Connected&lt;/code&gt; back. Then, in an actual Claude Code session:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Use the today-i-ran-notes server to search for how to check pod status with oc&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The tool got called — twice. Once for a broad search, then a follow-up pulling more detail from the closest match. Claude Code's own response is worth quoting in full, because it's a better result than a clean success would have been:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The search returned results, but none of them directly cover checking pod status with oc. The closest match is a note about building an "oc CLI Mentor" system prompt... but doesn't appear to include specific pod-status commands.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's exactly right. Entry 02 is about building a constrained system prompt, not a list of pod-status commands — the tool didn't force a match that wasn't there. It said so, plainly, then answered the actual question from its own general knowledge and was clear about the difference between "your notes don't cover this" and "here's the answer anyway."&lt;/p&gt;

&lt;p&gt;One more thing worth flagging so nobody gets confused reading this later: when I asked it to save the new commands for future reference, it saved them to Claude Code's own built-in memory feature — not to the Chroma index this whole arc has been building. Two separate systems, easy to mix up. The note lives in Claude Code's memory now, not in &lt;code&gt;chroma_db&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;So — the full loop actually works. A local embeddings pipeline, exposed as a real tool, correctly called by a real client, with honest reporting when the answer genuinely isn't in the index instead of a confident wrong guess. That last part is the result I actually care about here. A search tool that says "not in here" is worth a lot more than one that always returns something and leaves you to assume it's relevant.&lt;/p&gt;

</description>
      <category>ollama</category>
      <category>rag</category>
      <category>mcp</category>
      <category>chromadb</category>
    </item>
    <item>
      <title>What is DevOps? A Plain English Guide</title>
      <dc:creator>Nerav Doshi</dc:creator>
      <pubDate>Wed, 02 Sep 2026 15:51:45 +0000</pubDate>
      <link>https://dev.to/agenticdevops/what-is-devops-a-plain-english-guide-2102</link>
      <guid>https://dev.to/agenticdevops/what-is-devops-a-plain-english-guide-2102</guid>
      <description>&lt;h2&gt;
  
  
  Ever Wondered How Netflix Never Seems to Go Down?
&lt;/h2&gt;

&lt;p&gt;Think about this for a second. Netflix has over 260 million subscribers worldwide. People are watching shows in Tokyo, London, Lagos, and New York — all at the same time. And yet, when was the last time Netflix crashed on you?&lt;/p&gt;

&lt;p&gt;Now think about your favourite food delivery app. You open it, order food, track your driver in real time, and get a notification the moment your burger arrives. All of that happens in seconds.&lt;/p&gt;

&lt;p&gt;Behind all of this is a way of working called DevOps. And by the end of this article, you'll understand exactly what it is — no jargon, no complicated diagrams, just plain English.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Old Way (And Why It Was a Nightmare)
&lt;/h2&gt;

&lt;p&gt;To understand DevOps, we first need to understand the problem it solved.&lt;/p&gt;

&lt;p&gt;Imagine a software company in the early 2000s. They had two completely separate teams:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Developers&lt;/strong&gt; — the people who wrote the code and built new features&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Operations team&lt;/strong&gt; — the people who managed the servers and kept everything running&lt;/p&gt;

&lt;p&gt;These two teams barely talked to each other. Developers would spend months building new features, then hand over a massive pile of code to the operations team and say "here you go, make it work."&lt;/p&gt;

&lt;p&gt;The operations team would panic. They hadn't been involved in building it, had no idea what it did, and now they had to deploy it to millions of users without breaking anything.&lt;/p&gt;

&lt;p&gt;The result? Deployments took weeks. Bugs slipped through. Systems crashed. Customers complained. And the two teams blamed each other.&lt;/p&gt;

&lt;p&gt;Sound stressful? It was.&lt;/p&gt;




&lt;h2&gt;
  
  
  So What is DevOps?
&lt;/h2&gt;

&lt;p&gt;DevOps is simply the practice of bringing developers and operations teams together to build, test, and release software faster and more reliably.&lt;/p&gt;

&lt;p&gt;The name itself is a combination of &lt;strong&gt;Dev&lt;/strong&gt; (Development) and &lt;strong&gt;Ops&lt;/strong&gt; (Operations). Instead of two teams working in silos, they work as one team with shared goals, shared tools, and shared responsibility.&lt;/p&gt;

&lt;p&gt;Think of it like a restaurant kitchen.&lt;/p&gt;

&lt;p&gt;In a badly run kitchen, the chefs cook the food and just slide it through a hatch to the waiters. The waiters don't know what's in the dish, the chefs don't know what the customers are saying, and when something goes wrong, everyone points fingers.&lt;/p&gt;

&lt;p&gt;In a well run kitchen — like the ones you see at a great restaurant — the chefs and waiters communicate constantly. They know the menu inside out, they get feedback from customers quickly, and they work as one team to give people a great experience.&lt;/p&gt;

&lt;p&gt;DevOps is that well run kitchen, but for software.&lt;/p&gt;




&lt;h2&gt;
  
  
  A Real World Example: Amazon
&lt;/h2&gt;

&lt;p&gt;Amazon deploys new code to its website thousands of times per day.&lt;/p&gt;

&lt;p&gt;That means engineers are constantly making small improvements — fixing a bug here, improving the checkout experience there, tweaking a recommendation — and those changes go live almost instantly.&lt;/p&gt;

&lt;p&gt;How? Because Amazon uses DevOps practices. Small changes are automatically tested, automatically checked for problems, and automatically deployed without anyone having to manually press a button.&lt;/p&gt;

&lt;p&gt;In the old way of working, those same changes might have taken weeks to go live, gone through five teams, and required a late night deployment session that everyone dreaded.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Three Big Ideas Behind DevOps
&lt;/h2&gt;

&lt;p&gt;You don't need to memorise these, but it helps to know the thinking behind DevOps.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Work in Small Steps
&lt;/h3&gt;

&lt;p&gt;Instead of building for six months and releasing everything at once (terrifying), DevOps teams release small changes frequently. If something breaks, it's easy to find and fix because the change was tiny.&lt;/p&gt;

&lt;p&gt;Uber does this constantly. Every few weeks, the Uber app gets tiny updates — a new button here, a faster map there. You barely notice, but the team is constantly improving without disrupting your experience.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Automate the Boring Stuff
&lt;/h3&gt;

&lt;p&gt;Testing code manually, deploying to servers manually, checking for errors manually — all of this is slow and humans make mistakes. DevOps teams automate these tasks so they happen instantly and consistently every single time.&lt;/p&gt;

&lt;p&gt;Think of it like a car factory. Cars aren't built by hand anymore — robots do the repetitive work faster and with fewer errors. DevOps applies the same thinking to software.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Get Feedback Fast
&lt;/h3&gt;

&lt;p&gt;When something breaks, DevOps teams know about it within seconds, not days. Monitoring tools watch the system constantly and send alerts the moment something looks wrong.&lt;/p&gt;

&lt;p&gt;Netflix actually has a famous practice where they intentionally break parts of their own system during working hours to make sure their team can fix things quickly. They call it Chaos Engineering. It sounds mad, but it means they're never caught off guard.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Does a DevOps Engineer Actually Do?
&lt;/h2&gt;

&lt;p&gt;A DevOps engineer is the person who builds and maintains the systems that help developers work faster and more safely. They work on things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Setting up automated testing so bugs are caught before they reach users&lt;/li&gt;
&lt;li&gt;Building pipelines that automatically deploy code (we'll cover this in a future article)&lt;/li&gt;
&lt;li&gt;Managing cloud infrastructure on platforms like AWS or Azure&lt;/li&gt;
&lt;li&gt;Monitoring systems and making sure everything is running smoothly&lt;/li&gt;
&lt;li&gt;Writing scripts to automate repetitive tasks&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It's one of the most in-demand roles in tech right now, and the skills involved are exactly what this blog is here to help you build.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Should You Care About DevOps?
&lt;/h2&gt;

&lt;p&gt;Whether you're a developer, a system admin, a project manager, or someone just getting into tech — DevOps matters because it's how modern software is built.&lt;/p&gt;

&lt;p&gt;Every major tech company in the world uses DevOps practices. Banks use it to deploy new banking features. Airlines use it to update booking systems. Hospitals use it to improve patient management software. It's not just for Silicon Valley startups — it's everywhere.&lt;/p&gt;

&lt;p&gt;Learning DevOps opens doors. And the best part is, you don't need to know everything at once. We'll take it one byte at a time.&lt;/p&gt;




&lt;h2&gt;
  
  
  Quick Recap
&lt;/h2&gt;

&lt;p&gt;Here's everything we covered in plain English:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;DevOps&lt;/strong&gt; = Developers and Operations working together instead of in separate silos&lt;/li&gt;
&lt;li&gt;It solves the old problem of slow, painful, risky software releases&lt;/li&gt;
&lt;li&gt;The core ideas are: &lt;strong&gt;small changes, automation, and fast feedback&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Companies like Amazon, Netflix, and Uber use DevOps to deploy changes thousands of times a day&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;DevOps engineer&lt;/strong&gt; builds the tools and systems that make all of this possible&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What's Next?
&lt;/h2&gt;

&lt;p&gt;In the next article we're going to look at &lt;strong&gt;&lt;a href="https://dev.to/posts/linux-basics-for-devops/"&gt;Linux — The Operating System That Runs the Internet&lt;/a&gt;&lt;/strong&gt; — the OS that powers most of the internet and why every DevOps engineer needs to know the basics.&lt;/p&gt;

&lt;p&gt;It's going to be short, practical, and you'll be typing your first Linux commands before the end of the article. See you there.&lt;/p&gt;




</description>
      <category>devops</category>
      <category>beginners</category>
      <category>cloud</category>
      <category>careerswitch</category>
    </item>
    <item>
      <title>Direct Connect and ExpressRoute: Fixing Asymmetric BGP Routing</title>
      <dc:creator>Nerav Doshi</dc:creator>
      <pubDate>Tue, 01 Sep 2026 01:57:01 +0000</pubDate>
      <link>https://dev.to/agenticdevops/direct-connect-and-expressroute-fixing-asymmetric-bgp-routing-3c67</link>
      <guid>https://dev.to/agenticdevops/direct-connect-and-expressroute-fixing-asymmetric-bgp-routing-3c67</guid>
      <description>&lt;h2&gt;
  
  
  The Story
&lt;/h2&gt;

&lt;p&gt;Back in &lt;a href="https://pipelineandprompts.com/posts/hybrid-cloud-architecture-on-prem-vs-cloud-tradeoffs/" rel="noopener noreferrer"&gt;Article 1&lt;/a&gt;, I said we'd get back to this: once you've decided &lt;em&gt;what&lt;/em&gt; goes in the cloud and &lt;a href="https://pipelineandprompts.com/posts/managed-vs-self-hosted-handing-over-keys/" rel="noopener noreferrer"&gt;&lt;em&gt;who&lt;/em&gt; manages it&lt;/a&gt;, there's a third question that decides whether any of it actually works — how does your data get there?&lt;/p&gt;

&lt;p&gt;A telecom customer I worked with was pushing sustained real-time Kafka streams past 850 Mbps between on-prem and the cloud, with big unpredictable spikes on top, over an AWS Site-to-Site VPN. A single tunnel is rated up to 1.25 Gbps. On paper they had headroom.&lt;/p&gt;

&lt;p&gt;They still hit a wall. It wasn't AWS's fault.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;The wall was the on-prem VPN appliance's CPU — specifically, it couldn't keep up with IPsec encryption for every packet, made worse by two tunnels that weren't splitting traffic evenly. I confirmed this from CloudWatch tunnel metrics and CLI inspection on the appliance itself, not from guessing based on symptoms.&lt;/p&gt;

&lt;p&gt;This is the trap for any platform team running high-throughput streaming over a public-internet VPN: the bottleneck almost never shows up where the bandwidth numbers say it should. Think of a VPN over the public internet like a public road — cheap, open to everyone, fine most days. But there's a single-lane toll booth at the entrance where every car gets checked. On a normal day, minor delay. On a heavy day, that booth &lt;em&gt;is&lt;/em&gt; the road, no matter how wide the highway gets afterward. The toll booth here was IPsec — every packet individually encrypted and decrypted, with real CPU cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the Obvious Fixes Didn't Work
&lt;/h2&gt;

&lt;p&gt;Two fixes made sense on paper. Both failed, for different reasons.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scaling up&lt;/strong&gt; meant an emergency maintenance window to go from 4 vCPUs to 8, plus more RAM. No help — IPsec/IKE encryption for a given tunnel is bound to a single worker thread, and that thread doesn't spread across cores just because more exist. The new cores sat idle. The one thread doing the work stayed pegged at 98%.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scaling out&lt;/strong&gt; meant adding two more VPN tunnels, on the theory that more tunnels meant more hash buckets and distributed crypto load. Also failed. ECMP hashes traffic by flow — source/destination IP, source/destination port, protocol — and this Kafka stream was one sustained flow. It didn't matter how many tunnels existed; that flow kept landing on the same path. More lanes feeding the same toll booth. Retries got worse, not better.&lt;/p&gt;

&lt;p&gt;Direct Connect and ExpressRoute are the private road built just for you — no public traffic, no toll booth. They cost more and take time to build, but they behave predictably once they're in place. Most teams start on the public road because it's what's available on day one. This is what happens once your traffic outgrows it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Architecture
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqi68b9qdw30hrpca3bul.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqi68b9qdw30hrpca3bul.png" alt="Architecture Diagram" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The diagram lays out the "two private roads, no shared map" problem: the on-prem VPN appliance with its two IPsec tunnels to AWS; an HAProxy layer using 16 secondary IPs to feed ECMP; Direct Connect (AWS) and ExpressRoute (Azure) running as parallel dedicated circuits, each with dual-location redundancy; independent BGP autonomous systems per cloud with zero shared visibility into each other's routing; and the asymmetric-path failure itself — a request leaving via Direct Connect, its response coming back over ExpressRoute, hitting a stateful firewall that drops it as an unmatched session.&lt;/p&gt;

&lt;p&gt;Direct Connect took six weeks to provision. A single-provider option was ruled out almost immediately — this customer was already deliberately multi-cloud for capacity, cost, and reliability reasons that had nothing to do with this problem. The design became Direct Connect into AWS and ExpressRoute into Azure, each with redundant physical locations, driven by a real contractual requirement: 99.99% uptime.&lt;/p&gt;

&lt;p&gt;Here's where this specific engagement took an unexpected turn. In week two of running both circuits live, a new failure mode showed up: an asymmetric BGP routing loop that broke stateful firewalls. AWS and Azure each run independent BGP autonomous systems with no visibility into each other's routing decisions. Each cloud picked its own "best" path back to the same on-prem address block — unaware the other had picked differently. A request could leave over Direct Connect while the response came back over ExpressRoute, and to a stateful firewall expecting a matched pair, that looks like a session that was never opened. Dropped.&lt;/p&gt;

&lt;p&gt;This is exactly why the failure was subtle: the data plane — the actual Kafka traffic — looked completely healthy right up until the control plane's routing asymmetry collided with a stateful security device.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fixing It
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Prerequisites:&lt;/strong&gt; two cloud providers with existing dedicated-circuit relationships (Direct Connect for AWS, ExpressRoute for Azure), BGP peering already established on both circuits, a stateful firewall in the on-prem path, and TLS termination capability at the compute tier rather than just at a central gateway.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 1 — buy time while Direct Connect provisions&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Pulled non-critical workloads off the IPsec tunnels to free capacity for the critical stream&lt;/li&gt;
&lt;li&gt;Deployed local HAProxy bound to 16 separate secondary IPs on the on-prem appliance&lt;/li&gt;
&lt;li&gt;Why it worked: ECMP hashes partly on source/destination IP, so splitting one flow across 16 source IPs made it look like sixteen distinct flows — forcing real distribution across both tunnels instead of one path absorbing everything&lt;/li&gt;
&lt;li&gt;Not off-the-shelf load balancing — HAProxy's only job here was manufacturing enough distinct source IPs to break ECMP's per-flow hash&lt;/li&gt;
&lt;li&gt;Rollback: trivial. Pull HAProxy out of the path, nothing left to unwind&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Step 2 — fix the asymmetric routing&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Rejected the easy fix: loosening firewall statefulness to just allow asymmetric traffic — would have weakened a required security posture&lt;/li&gt;
&lt;li&gt;Real fix: BGP community tagging, MED tuning, and AS-path prepending on the Azure side, plus subnet-specific routing, to force one deterministic, symmetric path per address block&lt;/li&gt;
&lt;li&gt;Rollback: reversible by withdrawing the community tags and MED values — but only cleanly if you documented the pre-change route table first. Untangling which of several manual tweaks caused a new asymmetry after the fact is much harder than reverting to a known-good baseline&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Step 3 — get encryption out of the choke point&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Rather than reintroduce a central encryption bottleneck — recreating the exact problem that started this story — we moved TLS to the application layer, handled independently by each service instance. Spreading crypto across the compute tier kept continuous pod-to-pod encryption without a single-threaded chokepoint.&lt;/p&gt;

&lt;h2&gt;
  
  
  Security Considerations
&lt;/h2&gt;

&lt;p&gt;A private circuit isn't an encryption exemption. This was telecom data under SOC 2, which requires encryption in transit regardless of whether the path is public or private — "it's on a dedicated circuit" doesn't satisfy that requirement by itself.&lt;/p&gt;

&lt;p&gt;And the obvious fix for the BGP asymmetry — loosening firewall statefulness — was a security regression we explicitly rejected. It would have widened the firewall's tolerance for traffic patterns that look identical to spoofing and session-hijacking attempts. We took the slower, harder BGP-policy route instead of the fast one that degraded posture.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tradeoffs
&lt;/h2&gt;

&lt;p&gt;Dedicated circuits gave us predictable throughput and latency, and a real path to 99.99% uptime through dual-cloud, dual-location redundancy. In exchange: six weeks of lead time per circuit, and an entirely new failure mode — BGP asymmetry across independent cloud ASNs — that a single-provider VPN never has to deal with.&lt;/p&gt;

&lt;p&gt;Moving TLS to the application layer removed the single-threaded bottleneck entirely, but it cost us centralized visibility. One gateway doing crypto means one place to audit; per-service TLS means per-service certificate management and rotation discipline — more operational surface area, not less.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd Do Differently
&lt;/h2&gt;

&lt;p&gt;Design application traffic to use multiple distinguishable connections from the start — one giant flow is always a point of contention, VPN or dedicated circuit alike. And put BGP routing policy templates in place &lt;em&gt;before&lt;/em&gt; the first circuit goes live, not after an outage forces the question. Both failures in this story were, in hindsight, predictable the moment more than one path existed.&lt;/p&gt;

&lt;p&gt;Has any of this actually been tested? Partially. Automatic BGP failover from the primary Direct Connect circuit to the secondary has been proven for real — a production fiber cut triggered it, and it worked. Full failover down to the standby VPN has only run in a controlled drill, never a real outage. Worth naming rather than assuming away.&lt;/p&gt;

&lt;p&gt;And this design has a ceiling. It holds for two clouds and a handful of regions, but it doesn't scale by just repeating the pattern — add enough regions and individually-advertised address blocks and you hit a hard limit on how many routes a provider's edge router will accept. Past that limit it doesn't degrade gracefully; it silently drops excess routes or drops the BGP session entirely, taking every route with it. That blackholes active streams — the exact failure this architecture was built to prevent. The fix at that scale isn't more hand-tuned BGP policy per circuit. It's a managed transit gateway: a centralized hub that handles routing between clouds and regions so no circuit is negotiating its own path in isolation.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;Next in Cloud Without the Chaos: &lt;strong&gt;Article 4 — Data Sovereignty in Practice: What "It's in an EU Data Centre" Actually Covers (and Doesn't)&lt;/strong&gt; — going deeper on the compliance angle this article touched on with SOC 2.&lt;/p&gt;

</description>
      <category>directconnect</category>
      <category>expressroute</category>
      <category>bgprouting</category>
      <category>cloudwithoutthechaos</category>
    </item>
  </channel>
</rss>
