<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Jakson Tate</title>
    <description>The latest articles on DEV Community by Jakson Tate (@jaksontate).</description>
    <link>https://dev.to/jaksontate</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3844606%2F248b4fa0-86c4-40f6-9b8d-d410fdbb9e72.jpeg</url>
      <title>DEV Community: Jakson Tate</title>
      <link>https://dev.to/jaksontate</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/jaksontate"/>
    <language>en</language>
    <item>
      <title>How to Deploy Multi-Agent AI: Setup CrewAI on Ubuntu 24.04 Bare Metal GPU</title>
      <dc:creator>Jakson Tate</dc:creator>
      <pubDate>Thu, 17 Sep 2026 11:08:31 +0000</pubDate>
      <link>https://dev.to/jaksontate/how-to-deploy-multi-agent-ai-setup-crewai-on-ubuntu-2404-bare-metal-gpu-b92</link>
      <guid>https://dev.to/jaksontate/how-to-deploy-multi-agent-ai-setup-crewai-on-ubuntu-2404-bare-metal-gpu-b92</guid>
      <description>&lt;p&gt;Multi-agent AI frameworks have shifted from experimental prototypes to production-grade automation engines. But running CrewAI or AutoGen agents via cloud APIs introduces crippling costs: a single agent task can trigger 20 to 50 LLM calls due to autonomous thinking loops.&lt;/p&gt;

&lt;p&gt;By deploying your multi-agent AI framework on a Dedicated Bare Metal GPU server, your VRAM becomes a fixed cost. Whether your agent loops 10 times or 10,000 times, your infrastructure bill remains flat.&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 1: VRAM &amp;amp; KV Cache Math
&lt;/h2&gt;

&lt;p&gt;When multiple agents run concurrently, each agent consumes VRAM for model weights plus KV Cache (Key-Value Cache) for context memory:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A &lt;strong&gt;14B model&lt;/strong&gt; (like Qwen2.5 or Llama-3.3) takes ~15GB of VRAM in FP8 precision.&lt;/li&gt;
&lt;li&gt;Each concurrent agent context requires roughly &lt;strong&gt;1.5GB of KV cache&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Running a 4-agent crew concurrently requires unthrottled GPU access and high PCIe throughput to prevent Out-Of-Memory (OOM) crashes and latency spikes.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Phase 2: Server Hardening &amp;amp; UFW Setup
&lt;/h2&gt;

&lt;p&gt;Never expose your local LLM engine to the public internet. Leaving inference ports open invites GPU hijacking. Lock down your Ubuntu 24.04 firewall:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Update Ubuntu packages&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;apt update &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;sudo &lt;/span&gt;apt upgrade &lt;span class="nt"&gt;-y&lt;/span&gt;

&lt;span class="c"&gt;# Allow only SSH, HTTP, HTTPS&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;ufw default deny incoming
&lt;span class="nb"&gt;sudo &lt;/span&gt;ufw default allow outgoing
&lt;span class="nb"&gt;sudo &lt;/span&gt;ufw allow 22/tcp
&lt;span class="nb"&gt;sudo &lt;/span&gt;ufw allow 80/tcp
&lt;span class="nb"&gt;sudo &lt;/span&gt;ufw allow 443/tcp
&lt;span class="nb"&gt;sudo &lt;/span&gt;ufw &lt;span class="nb"&gt;enable&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Phase 3: Install Ollama &amp;amp; Pull the Model
&lt;/h2&gt;

&lt;p&gt;Install Ollama to host your sovereign inference engine locally:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;https://ollama.com/install.sh]&lt;span class="o"&gt;(&lt;/span&gt;https://ollama.com/install.sh&lt;span class="o"&gt;)&lt;/span&gt; | sh

&lt;span class="c"&gt;# Pull a 14B+ model for robust agent reasoning&lt;/span&gt;
ollama pull qwen2.5:14b
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Engineering Tip:&lt;/strong&gt; 8B models struggle with complex JSON structured outputs required for CrewAI task delegation. Use at least 14B models for worker agents and 32B/70B models for manager agents.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Phase 4: Deploying CrewAI with &lt;code&gt;uv&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;Avoid raw &lt;code&gt;pip&lt;/code&gt; to prevent Python environment conflicts. Use &lt;code&gt;uv&lt;/code&gt; from Astral for ultra-fast environment isolation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Install UV package manager&lt;/span&gt;
curl &lt;span class="nt"&gt;-LsSf&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;https://astral.sh/uv/install.sh]&lt;span class="o"&gt;(&lt;/span&gt;https://astral.sh/uv/install.sh&lt;span class="o"&gt;)&lt;/span&gt; | sh
&lt;span class="nb"&gt;source&lt;/span&gt; &lt;span class="nv"&gt;$HOME&lt;/span&gt;/.cargo/env

&lt;span class="c"&gt;# Initialize project scaffold&lt;/span&gt;
uv tool &lt;span class="nb"&gt;install &lt;/span&gt;crewai
crewai create crew servermo_agents
&lt;span class="nb"&gt;cd &lt;/span&gt;servermo_agents
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  The LiteLLM Trap
&lt;/h3&gt;

&lt;p&gt;CrewAI utilizes LiteLLM under the hood, which crashes if it doesn't detect an OpenAI API key—even when routing strictly to &lt;code&gt;localhost&lt;/code&gt;. Bypass this by creating a &lt;code&gt;.env&lt;/code&gt; file with dummy values:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OPENAI_API_KEY="NA"
OPENAI_API_BASE="http://localhost:11434/v1"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Phase 5: Production Crew Code (&lt;code&gt;crew.py&lt;/code&gt;)
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;crewai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Agent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Task&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Crew&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Process&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;LLM&lt;/span&gt;

&lt;span class="c1"&gt;# 1. Define Local GPU LLM
&lt;/span&gt;&lt;span class="n"&gt;bare_metal_llm&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;LLM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ollama/qwen2.5:14b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;http://localhost:11434&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.2&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# 2. Architect Agents (Disable infinite delegation loops)
&lt;/span&gt;&lt;span class="n"&gt;research_agent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Agent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;role&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Infrastructure Security Analyst&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;goal&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Discover vulnerabilities in cloud VM networking&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;backstory&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Elite SRE who trusts only bare metal servers.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;llm&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;bare_metal_llm&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;verbose&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;allow_delegation&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;writer_agent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Agent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;role&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;DevSecOps Technical Writer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;goal&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Draft an actionable security report based on findings&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;backstory&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Synthesizes complex security data into Markdown.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;llm&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;bare_metal_llm&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;verbose&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;allow_delegation&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# 3. Define Tasks
&lt;/span&gt;&lt;span class="n"&gt;research_task&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Task&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Analyze why multi-tenant Cloud VMs are less secure than Dedicated Bare Metal. List 3 key points.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;expected_output&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;3 technical bullet points regarding hypervisor vulnerabilities.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;agent&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;research_agent&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;write_task&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Task&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Write a 2-paragraph security advisory from the 3 bullet points.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;expected_output&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Formatted Markdown security advisory document.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;agent&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;writer_agent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;research_task&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# 4. Initialize and Run Crew
&lt;/span&gt;&lt;span class="n"&gt;production_crew&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Crew&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;agents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;research_agent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;writer_agent&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;tasks&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;research_task&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;write_task&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;process&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;Process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sequential&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;production_crew&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;kickoff&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;--- FINAL OUTPUT ---&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Deployment Model Comparison
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Deployment Model&lt;/th&gt;
&lt;th&gt;Compute Cost Model&lt;/th&gt;
&lt;th&gt;Latency Stability&lt;/th&gt;
&lt;th&gt;Data Sovereignty&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Public Cloud APIs&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Punishing per-token fees&lt;/td&gt;
&lt;td&gt;High network jitter&lt;/td&gt;
&lt;td&gt;Third-party privacy risk&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Shared Cloud VMs&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Low initial cost&lt;/td&gt;
&lt;td&gt;Noisy neighbors cause OOMs&lt;/td&gt;
&lt;td&gt;Shared hypervisor risk&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;ServerMO Bare Metal GPU&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100% Fixed monthly rate&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Unthrottled PCIe &amp;amp; VRAM&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100% Private network&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;p&gt;👉 &lt;strong&gt;Ready to deploy sovereign multi-agent AI on bare metal GPUs? Read the full deployment guide on ServerMO:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;a href="https://www.servermo.com/howto/deploy-crewai-ubuntu-bare-metal/" rel="noopener noreferrer"&gt;Setup CrewAI on Ubuntu 24.04 Bare Metal GPU | ServerMO&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>devops</category>
      <category>sre</category>
    </item>
    <item>
      <title>How to Build Production-Ready AI Agents on Bare Metal</title>
      <dc:creator>Jakson Tate</dc:creator>
      <pubDate>Thu, 17 Sep 2026 09:58:08 +0000</pubDate>
      <link>https://dev.to/jaksontate/how-to-build-production-ready-ai-agents-on-bare-metal-3p1j</link>
      <guid>https://dev.to/jaksontate/how-to-build-production-ready-ai-agents-on-bare-metal-3p1j</guid>
      <description>&lt;p&gt;An AI agent becomes serious the moment it touches something real—a customer database, an internal file system, or a payment API. Before that, it is merely a demo.&lt;/p&gt;

&lt;p&gt;Many engineering teams discover this gap after a prototype that dazzled stakeholders on Friday silently corrupts a database on Monday. A production agent is not just an LLM with a better prompt; it is a distributed software system. The model can think, but the surrounding architecture must dictate what that thinking is allowed to do. &lt;/p&gt;

&lt;p&gt;Here are the elite SRE practices required to govern and deploy autonomous agents securely on Bare Metal.&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 1: The 6 SRE Layers of Agentic Architecture
&lt;/h2&gt;

&lt;p&gt;At its core, every agent executes the same loop: Perceive, Reason, Act, Observe. Hand-rolling this loop takes an afternoon. Wrapping that loop in the engineering required to keep it safe under real traffic demands strict architecture across 6 critical layers:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Tools:&lt;/strong&gt; Typed function schemas, strict input validation, and safe error returns.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory:&lt;/strong&gt; Short-term session context and long-term vector stores.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retrieval (RAG):&lt;/strong&gt; Chunking, embeddings, and cross-encoder re-ranking.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Orchestration:&lt;/strong&gt; Routing, retries, bounded loops, and human-in-the-loop handoffs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Evals &amp;amp; Guardrails:&lt;/strong&gt; Automated testing frameworks to measure quality before deployment.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Observability:&lt;/strong&gt; Tracing execution paths, logging token costs, and capturing telemetry.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Phase 2: Strict Tool Schemas &amp;amp; Model Context Protocol (MCP)
&lt;/h2&gt;

&lt;p&gt;The most common cause of a flaky agent is a tool returning an unstructured stack trace that the LLM cannot parse. When an agent calls a tool, it generates a JSON string. Without strict constraints, LLMs inevitably hallucinate non-existent arguments or pass arrays instead of strings.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The "All Tools" Antipattern:&lt;/strong&gt; Exposing more than 8 tools simultaneously causes severe tool confusion. Narrow, job-specific tools (e.g., &lt;code&gt;get_invoice_by_id&lt;/code&gt;) drastically outperform broad wrappers (e.g., &lt;code&gt;query_database&lt;/code&gt;).&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;To fix this, utilize &lt;strong&gt;Pydantic&lt;/strong&gt; to enforce rigid type constraints, Enum bounds, and min/max values. Furthermore, adopting the &lt;strong&gt;Model Context Protocol (MCP)&lt;/strong&gt; standardizes how agents exchange context with external databases and APIs, ensuring the LLM only receives pre-validated, secure schemas.&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 3: Neurosymbolic Guardrails (&lt;code&gt;BeforeToolCallEvent&lt;/code&gt;)
&lt;/h2&gt;

&lt;p&gt;Writing &lt;em&gt;"CRITICAL: Never confirm bookings without payment verification"&lt;/em&gt; in a system prompt is not security—prompts are suggestions, not hard constraints. An LLM can and will hallucinate compliance under injection attacks.&lt;/p&gt;

&lt;p&gt;According to OWASP Top 10 for LLM Applications (LLM08: Insecure Plugin Design), runtime security risks cluster at the plugin execution layer. You must implement &lt;strong&gt;Neurosymbolic Guardrails&lt;/strong&gt; by combining neural reasoning with deterministic Python code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Deterministic execution interception hook
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;validate_action&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;BeforeToolCallEvent&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="c1"&gt;# Evaluate explicit business rules BEFORE execution
&lt;/span&gt;    &lt;span class="n"&gt;passed&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;violations&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;validate_rules&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;rules&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tool_use&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;input&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;passed&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="c1"&gt;# Cancel tool execution entirely. The LLM cannot override this.
&lt;/span&gt;        &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cancel_tool&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;BLOCKED: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;, &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;violations&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Phase 4: Anchored Summarization &amp;amp; Context Degradation
&lt;/h2&gt;

&lt;p&gt;LLM reasoning quality degrades significantly when context window capacity hits just &lt;strong&gt;25%&lt;/strong&gt;. Waiting until 100% capacity means the agent has already lost its core instructions.&lt;/p&gt;

&lt;p&gt;Do not use naive sliding windows that simply delete historical messages. Instead, use &lt;strong&gt;Anchored Iterative Summarization&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Anchor Block:&lt;/strong&gt; Keep a fixed block containing the original system prompt, core user goal, and active constraints.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Iterative Compression:&lt;/strong&gt; Run a secondary background process to summarize only intermediate completed steps.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Phase 5: Open Source Agent Orchestration &amp;amp; Observability
&lt;/h2&gt;

&lt;p&gt;Shipping an agent without tracing makes debugging hallucinations impossible. While cloud platforms offer proprietary suites, they lock you into their ecosystem. &lt;/p&gt;

&lt;p&gt;By implementing &lt;strong&gt;OpenTelemetry&lt;/strong&gt;, you can capture full execution traces, step latency, and token consumption—visualizing telemetry inside self-hosted Grafana dashboards to monitor drift continuously.&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 6: The Bare Metal Security Advantage
&lt;/h2&gt;

&lt;p&gt;Executing a high-performance agentic architecture requires immense compute power. Cloud API egress fees and inter-node network latency can rapidly bankrupt projects as they scale.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Deployment Model&lt;/th&gt;
&lt;th&gt;Compute Performance&lt;/th&gt;
&lt;th&gt;API Egress Costs&lt;/th&gt;
&lt;th&gt;Data Governance&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Public Cloud SaaS&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Shared / Throttled&lt;/td&gt;
&lt;td&gt;High API Egress Fees&lt;/td&gt;
&lt;td&gt;Third-party Privacy Risks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Shared Cloud VMs&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Noisy Neighbor Latency&lt;/td&gt;
&lt;td&gt;Variable Network Taxes&lt;/td&gt;
&lt;td&gt;Shared Hypervisor Constraints&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;ServerMO Bare Metal Dedicated&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Unshared High-Core CPUs / GPUs&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0 Egress Taxes&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100% On-Prem / Sandboxed&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;p&gt;👉 &lt;strong&gt;Ready to escape cloud API taxes and deploy production AI agents? Read the full tutorial on ServerMO:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;a href="https://www.servermo.com/howto/build-production-ready-ai-agents/" rel="noopener noreferrer"&gt;How to Build Production-Ready AI Agents on Bare Metal | ServerMO&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>devops</category>
      <category>sre</category>
    </item>
    <item>
      <title>Stop Wasting GPUs on Embeddings: The RAG FinOps Guide</title>
      <dc:creator>Jakson Tate</dc:creator>
      <pubDate>Thu, 17 Sep 2026 09:27:03 +0000</pubDate>
      <link>https://dev.to/jaksontate/stop-wasting-gpus-on-embeddings-the-rag-finops-guide-3nbp</link>
      <guid>https://dev.to/jaksontate/stop-wasting-gpus-on-embeddings-the-rag-finops-guide-3nbp</guid>
      <description>&lt;p&gt;In the rush to build Retrieval-Augmented Generation (RAG) pipelines, engineering teams make a massive architectural blunder: assuming that because Large Language Models (LLMs) require massive GPU clusters, the embedding models vectorizing text must run on those same GPUs. This forces teams to rent $3,000 NVIDIA cards just to host tiny 1GB encoder models.&lt;/p&gt;

&lt;p&gt;Embedding models perform simple forward passes and do not require autoregressive generation. Elite SREs strictly decouple their infrastructure layers—relying on optimized CPU inference to preserve GPU VRAM exclusively for generative models.&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 1: The Decoupled Architecture (GPU vs CPU)
&lt;/h2&gt;

&lt;p&gt;When evaluating hardware profiles for RAG workloads, separate your workload profiles:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Real-time Queries (High-Core CPU):&lt;/strong&gt; Feeding a single 15-token user query to an H100 GPU sits idle 95% of the time. Modern CPUs process single vectors in under 20ms, matching GPU latency while avoiding PCIe bus overhead.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Massive Batch Ingestion (Entry-Level GPU):&lt;/strong&gt; For re-indexing 10 million documents, a CPU bottlenecks. Utilizing an entry-level datacenter GPU (like an NVIDIA L4) handles up to 4,500 chunks/sec for bulk ingestion tasks.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Phase 2: ONNX Runtime CPU Embeddings &amp;amp; AVX-512
&lt;/h2&gt;

&lt;p&gt;To achieve real-time CPU speeds, bypass PyTorch bottlenecks by converting models to ONNX format and leveraging &lt;strong&gt;AVX-512&lt;/strong&gt; processor instructions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# SRE method for high-speed CPU embedding generation
&lt;/span&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;transformers&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;AutoTokenizer&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;optimum.onnxruntime&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;ORTModelForFeatureExtraction&lt;/span&gt;

&lt;span class="n"&gt;model_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;philipp-zettl/BAAI-bge-m3-ONNX&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;tokenizer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;AutoTokenizer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_pretrained&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Force AVX-512/VNNI hardware optimizations on CPU
&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ORTModelForFeatureExtraction&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_pretrained&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; 
    &lt;span class="n"&gt;provider&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CPUExecutionProvider&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;inputs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;tokenizer&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;High speed ONNX inference on ServerMO Bare Metal&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;padding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;truncation&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;return_tensors&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;embeddings&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;inputs&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="n"&gt;last_hidden_state&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Phase 3: Deploying HuggingFace TEI &amp;amp; QInt8 Quantization
&lt;/h2&gt;

&lt;p&gt;While Ollama is popular for serving local LLMs, benchmark telemetry reveals that serving embeddings via Ollama averages ~99ms per request.&lt;/p&gt;

&lt;p&gt;Deploying &lt;strong&gt;HuggingFace Text-Embeddings-Inference (TEI)&lt;/strong&gt; Docker containers yields sub-20ms latencies—delivering a 5x speed improvement specifically optimized for vector extraction.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Security Note:&lt;/strong&gt; Avoid passing &lt;code&gt;-e HF_TOKEN="hf_..."&lt;/code&gt; directly in Docker run commands to prevent writing plain-text API keys to system Bash history. Map credentials using isolated &lt;code&gt;.env&lt;/code&gt; files.&lt;br&gt;
&lt;/p&gt;
&lt;/blockquote&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Securely define credentials&lt;/span&gt;
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"HF_TOKEN=your_secure_huggingface_read_token"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; .hf_env
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;MODEL_DATA&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;$PWD&lt;/span&gt;/embedding_cache
&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="nv"&gt;$MODEL_DATA&lt;/span&gt;

&lt;span class="c"&gt;# Deploy CPU-optimized HuggingFace TEI&lt;/span&gt;
docker run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--name&lt;/span&gt; tei-embeddings &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--env-file&lt;/span&gt; .hf_env &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-p&lt;/span&gt; 8080:80 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-v&lt;/span&gt; &lt;span class="nv"&gt;$MODEL_DATA&lt;/span&gt;:/data &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--pull&lt;/span&gt; always ghcr.io/huggingface/text-embeddings-inference:cpu-1.5 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--model-id&lt;/span&gt; BAAI/bge-m3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Phase 4: Defeating Vector DB RAM Explosions
&lt;/h2&gt;

&lt;p&gt;In PostgreSQL (&lt;code&gt;pgvector&lt;/code&gt;), a single 3,072-dimension vector consumes 12.3 KB of RAM. Scaling to 1 million documents burns &lt;strong&gt;12.3 GB of RAM&lt;/strong&gt; purely for raw table storage.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Vector Model Dimension&lt;/th&gt;
&lt;th&gt;Storage per 1M Docs&lt;/th&gt;
&lt;th&gt;MTEB Accuracy Retention&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;3,072 Dims (Standard)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~12.3 GB RAM&lt;/td&gt;
&lt;td&gt;100% Baseline&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;256 Dims (Matryoshka Truncated)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~1.02 GB RAM&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;&amp;gt;98% Retained&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Using &lt;strong&gt;Matryoshka Representation Learning&lt;/strong&gt;, you can truncate vectors from 3072 down to 256 dimensions—slashing RAM footprint by &lt;strong&gt;6x&lt;/strong&gt; with less than 2% degradation in retrieval accuracy.&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 5: Escaping the FP16 CPU Trap with QInt8
&lt;/h2&gt;

&lt;p&gt;Running 16-bit float (FP16) models on standard CPUs without specialized AMX instructions causes the kernel to downcast and upcast numeric types mid-operation, degrading inference speed by 2x to 7x. &lt;/p&gt;

&lt;p&gt;Leveraging &lt;strong&gt;QInt8 quantization&lt;/strong&gt; eliminates execution penalties and speeds up CPU matrix operations by &lt;strong&gt;3x&lt;/strong&gt;.&lt;/p&gt;




&lt;p&gt;👉 &lt;strong&gt;Read the full technical tutorial on ServerMO:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;a href="https://www.servermo.com/howto/cpu-inference-embeddings-rag/" rel="noopener noreferrer"&gt;Stop Wasting GPUs on Embeddings: The RAG FinOps Guide | ServerMO&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>devops</category>
      <category>sre</category>
    </item>
    <item>
      <title>The Cybersecurity SaaS-pocalypse: Open-Source SIEM on Bare Metal</title>
      <dc:creator>Jakson Tate</dc:creator>
      <pubDate>Thu, 17 Sep 2026 08:57:44 +0000</pubDate>
      <link>https://dev.to/jaksontate/the-cybersecurity-saas-pocalypse-open-source-siem-on-bare-metal-5b8i</link>
      <guid>https://dev.to/jaksontate/the-cybersecurity-saas-pocalypse-open-source-siem-on-bare-metal-5b8i</guid>
      <description>&lt;p&gt;As organizations deploy fleets of AI agents, autonomous workflows, and LLMs, security telemetry has exploded by an order of magnitude. Suddenly, the traditional Software-as-a-Service (SaaS) business model for SIEMs has become a massive financial liability.&lt;/p&gt;

&lt;p&gt;The era of blindly forwarding terabytes of raw logs to a Cloud SIEM is over. Here is an SRE and FinOps breakdown of the mathematical reality behind the cloud SIEM pricing trap and how modern Security Operations Centers (SOCs) build Open-Source Security Data Lakes on Bare Metal NVMe servers.&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 1: The Per-GB Pricing Trap
&lt;/h2&gt;

&lt;p&gt;Commercial SIEM platforms utilize data-volume pricing models averaging roughly &lt;strong&gt;$1,800 per GB per day, annually&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;For a modest 100GB/day ingestion pipeline, base licensing costs ~$180,000 per year. The critical flaw appears during an active cyber attack: forensic log volume multiplies by 10x. Because of ingest-based pricing, your cloud SIEM bill explodes precisely when your security team is under attack.&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 2: The False Cloud Alternative
&lt;/h2&gt;

&lt;p&gt;When evaluating open-source alternatives like Wazuh ($0 licensing fees), many organizations make a fatal architectural mistake: deploying high-throughput log ingestion servers on AWS or Azure.&lt;/p&gt;

&lt;p&gt;Deploying a production Wazuh/OpenSearch environment handling 5,000+ endpoints on public clouds simply replaces software licensing bills with massive infrastructure charges:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Provisioned IOPS (io2) Fees:&lt;/strong&gt; Driven by relentless Elasticsearch write operations.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;NAT Gateway Egress Taxes:&lt;/strong&gt; Incurred by continuous cross-zone log transfers.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Phase 3: Building a Security Data Lake Architecture
&lt;/h2&gt;

&lt;p&gt;To escape the SIEM trap, decouple compute from storage using a &lt;strong&gt;Security Data Lake Architecture&lt;/strong&gt;: route 100% of raw logs to cold object storage (MinIO) for long-term retention, and stream only filtered, high-priority events to OpenSearch for real-time threat hunting.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;&lt;span class="c"&gt;# fluent-bit.conf : Forwarding Telemetry to an Open-Source Data Lake
&lt;/span&gt;[&lt;span class="n"&gt;SERVICE&lt;/span&gt;]
    &lt;span class="n"&gt;Flush&lt;/span&gt;        &lt;span class="m"&gt;1&lt;/span&gt;
    &lt;span class="n"&gt;Daemon&lt;/span&gt;       &lt;span class="n"&gt;Off&lt;/span&gt;
    &lt;span class="n"&gt;Log_Level&lt;/span&gt;    &lt;span class="n"&gt;info&lt;/span&gt;

[&lt;span class="n"&gt;INPUT&lt;/span&gt;]
    &lt;span class="n"&gt;Name&lt;/span&gt;         &lt;span class="n"&gt;tail&lt;/span&gt;
    &lt;span class="n"&gt;Path&lt;/span&gt;         /&lt;span class="n"&gt;var&lt;/span&gt;/&lt;span class="n"&gt;log&lt;/span&gt;/&lt;span class="n"&gt;ai_agents&lt;/span&gt;/*.&lt;span class="n"&gt;json&lt;/span&gt;
    &lt;span class="n"&gt;Tag&lt;/span&gt;          &lt;span class="n"&gt;ai_security&lt;/span&gt;.&lt;span class="n"&gt;logs&lt;/span&gt;

[&lt;span class="n"&gt;FILTER&lt;/span&gt;]
    &lt;span class="c"&gt;# SRE Best Practice: Drop debug noise BEFORE network transmission
&lt;/span&gt;    &lt;span class="n"&gt;Name&lt;/span&gt;         &lt;span class="n"&gt;grep&lt;/span&gt;
    &lt;span class="n"&gt;Match&lt;/span&gt;        *
    &lt;span class="n"&gt;Exclude&lt;/span&gt;      &lt;span class="n"&gt;level&lt;/span&gt; &lt;span class="n"&gt;debug&lt;/span&gt;

[&lt;span class="n"&gt;OUTPUT&lt;/span&gt;]
    &lt;span class="c"&gt;# Output 1: Send ALL logs to MinIO (Cold Lake) for cheap retention
&lt;/span&gt;    &lt;span class="n"&gt;Name&lt;/span&gt;         &lt;span class="n"&gt;s3&lt;/span&gt;
    &lt;span class="n"&gt;Match&lt;/span&gt;        *
    &lt;span class="n"&gt;Bucket&lt;/span&gt;       &lt;span class="n"&gt;ai&lt;/span&gt;-&lt;span class="n"&gt;threat&lt;/span&gt;-&lt;span class="n"&gt;telemetry&lt;/span&gt;-&lt;span class="n"&gt;archive&lt;/span&gt;
    &lt;span class="n"&gt;Endpoint&lt;/span&gt;     [&lt;span class="n"&gt;http&lt;/span&gt;://&lt;span class="n"&gt;minio&lt;/span&gt;.&lt;span class="n"&gt;internal&lt;/span&gt;.&lt;span class="n"&gt;lan&lt;/span&gt;:&lt;span class="m"&gt;9000&lt;/span&gt;](&lt;span class="n"&gt;http&lt;/span&gt;://&lt;span class="n"&gt;minio&lt;/span&gt;.&lt;span class="n"&gt;internal&lt;/span&gt;.&lt;span class="n"&gt;lan&lt;/span&gt;:&lt;span class="m"&gt;9000&lt;/span&gt;)
    &lt;span class="n"&gt;Store_Dir&lt;/span&gt;    /&lt;span class="n"&gt;tmp&lt;/span&gt;/&lt;span class="n"&gt;fluent&lt;/span&gt;-&lt;span class="n"&gt;bit&lt;/span&gt;/&lt;span class="n"&gt;s3&lt;/span&gt;

[&lt;span class="n"&gt;OUTPUT&lt;/span&gt;]
    &lt;span class="c"&gt;# Output 2: Send ONLY critical/filtered events to OpenSearch (Hot Index)
&lt;/span&gt;    &lt;span class="n"&gt;Name&lt;/span&gt;         &lt;span class="n"&gt;es&lt;/span&gt;
    &lt;span class="n"&gt;Match&lt;/span&gt;        &lt;span class="n"&gt;ai_security&lt;/span&gt;.&lt;span class="n"&gt;logs&lt;/span&gt;
    &lt;span class="n"&gt;Host&lt;/span&gt;         &lt;span class="n"&gt;bare&lt;/span&gt;-&lt;span class="n"&gt;metal&lt;/span&gt;-&lt;span class="n"&gt;opensearch&lt;/span&gt;.&lt;span class="n"&gt;internal&lt;/span&gt;.&lt;span class="n"&gt;lan&lt;/span&gt;
    &lt;span class="n"&gt;Port&lt;/span&gt;         &lt;span class="m"&gt;9200&lt;/span&gt;
    &lt;span class="n"&gt;Index&lt;/span&gt;        &lt;span class="n"&gt;ai&lt;/span&gt;-&lt;span class="n"&gt;threat&lt;/span&gt;-&lt;span class="n"&gt;telemetry&lt;/span&gt;
    &lt;span class="n"&gt;Type&lt;/span&gt;         &lt;span class="err"&gt;_&lt;/span&gt;&lt;span class="n"&gt;doc&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Phase 4: Repatriating to Bare Metal NVMe
&lt;/h2&gt;

&lt;p&gt;Security Operations Centers generate intensive, constant write workloads. Cloud Block Storage (EBS) throttles write-heavy operations unless expensive provisioned IOPS tiers are purchased.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The 96.67% Cost Savings Matrix:&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Commercial SaaS SIEM:&lt;/strong&gt; 100GB/day = ~$180,000/year.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ServerMO Bare Metal Pivot:&lt;/strong&gt; Hosting Wazuh + OpenSearch + MinIO on ServerMO Dedicated Bare Metal with Enterprise PCIe NVMe = ~$500/month ($6,000/year).
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Net Savings:&lt;/strong&gt; &lt;strong&gt;$174,000/year reduction in TCO (96.67% savings).&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;p&gt;By moving to dedicated bare metal, security teams gain unthrottled write IOPS for real-time log parsing, eliminate data egress taxes entirely, and maintain complete data sovereignty.&lt;/p&gt;




&lt;p&gt;👉 &lt;strong&gt;Ready to build an unbreakable security posture without the cloud tax? Read the full architectural breakdown on ServerMO:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;a href="https://www.servermo.com/blogs/open-source-siem-bare-metal/" rel="noopener noreferrer"&gt;Open-Source SIEM &amp;amp; Security Data Lakes on Bare Metal | ServerMO&lt;/a&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>devops</category>
      <category>linux</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Semantic Caching for LLMs on Ubuntu 24.04: Reduce API Costs</title>
      <dc:creator>Jakson Tate</dc:creator>
      <pubDate>Thu, 17 Sep 2026 07:49:10 +0000</pubDate>
      <link>https://dev.to/jaksontate/semantic-caching-for-llms-on-ubuntu-2404-reduce-api-costs-2m85</link>
      <guid>https://dev.to/jaksontate/semantic-caching-for-llms-on-ubuntu-2404-reduce-api-costs-2m85</guid>
      <description>&lt;p&gt;When you deploy a Generative AI application to production, you quickly discover a painful financial truth: &lt;strong&gt;inference costs scale violently&lt;/strong&gt;. You are charged for every single token. But if you analyze prompt logs, up to 40% of user queries express the exact same intent, just phrased differently.&lt;/p&gt;

&lt;p&gt;"How do I reset my password?" and "I forgot my login credentials, help" express identical intent. However, a traditional key-value store treats these as two entirely different requests, invoking the OpenAI API twice and charging full price both times.&lt;/p&gt;

&lt;p&gt;Here is an SRE and Data Science guide breaking down the mechanics of semantic caching for LLM applications on Ubuntu 24.04.&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 1: Lexical Caching vs. Semantic Caching
&lt;/h2&gt;

&lt;p&gt;To effectively reduce OpenAI API costs, you must understand the critical difference between lookup mechanisms:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Lexical Caching (e.g., standard Redis key-value store):&lt;/strong&gt; Relies on exact string matching. If a user types "Where is my order?", it is cached. If the next user types "Track my package", it triggers a Cache Miss because the characters do not match.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Semantic Caching:&lt;/strong&gt; Uses an Embedding Model to convert the prompt into a high-dimensional mathematical vector. In this vector space, "Where is my order?" and "Track my package" cluster tightly together. The cache retrieves the stored response instantly, dropping lookup latency from 3000ms down to 15ms.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Phase 2: Implementing the Vector Cache (Qdrant)
&lt;/h2&gt;

&lt;p&gt;To build an enterprise LLM prompt caching architecture on Ubuntu 24.04 LTS, deploy Qdrant using Docker with persistent storage:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 1. Install Qdrant via Docker on Ubuntu 24.04 (Persistent Storage)&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;apt update &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;sudo &lt;/span&gt;apt &lt;span class="nb"&gt;install &lt;/span&gt;docker.io &lt;span class="nt"&gt;-y&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;docker run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;--restart&lt;/span&gt; unless-stopped &lt;span class="nt"&gt;-p&lt;/span&gt; 6333:6333 &lt;span class="nt"&gt;-p&lt;/span&gt; 6334:6334 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-v&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;pwd&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;/qdrant_storage:/qdrant/storage:z &lt;span class="se"&gt;\&lt;/span&gt;
    qdrant/qdrant

&lt;span class="c"&gt;# 2. Setup your Python environment&lt;/span&gt;
python3 &lt;span class="nt"&gt;-m&lt;/span&gt; venv llm_cache_env
&lt;span class="nb"&gt;source &lt;/span&gt;llm_cache_env/bin/activate
pip &lt;span class="nb"&gt;install &lt;/span&gt;qdrant-client sentence-transformers openai
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Phase 3: Preventing Multi-Tenant Cache Poisoning
&lt;/h2&gt;

&lt;p&gt;If you implement a global semantic cache in a SaaS product, you are exposing your application to cross-tenant data leaks.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The Security Risk (Cross-Tenant Leakage):&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;strong&gt;User A&lt;/strong&gt; asks "Summarize my recent transactions." The LLM generates a response containing User A's private financial data, which gets cached. &lt;strong&gt;User B&lt;/strong&gt; logs in and asks "Give me a summary of my transactions." The semantic cache detects a 98% intent match and serves User A's private response to User B.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;The SRE Fix:&lt;/strong&gt; Inject a strict Namespace (&lt;code&gt;tenant_id&lt;/code&gt;) into the Qdrant Payload. Vector similarity searches must ALWAYS execute underneath a hard metadata filter:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;qdrant_client&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;QdrantClient&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;qdrant_client.http&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;models&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;sentence_transformers&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;SentenceTransformer&lt;/span&gt;

&lt;span class="c1"&gt;# Initialize Local Embedding Model (Free &amp;amp; Fast)
&lt;/span&gt;&lt;span class="n"&gt;encoder&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;SentenceTransformer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;all-MiniLM-L6-v2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;QdrantClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;host&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;localhost&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;port&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;6333&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;COLLECTION_NAME&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;semantic_cache_prod&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="c1"&gt;# Explicitly create collection with correct Vector Dimensions (384)
&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;collection_exists&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;collection_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;COLLECTION_NAME&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create_collection&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;collection_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;COLLECTION_NAME&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;vectors_config&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;models&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;VectorParams&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;384&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="c1"&gt;# Matches 'all-MiniLM-L6-v2' output dimension
&lt;/span&gt;            &lt;span class="n"&gt;distance&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;models&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Distance&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;COSINE&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;check_secure_cache&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tenant_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;threshold&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.90&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;query_vector&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;encoder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_query&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;tolist&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="c1"&gt;# Hard Metadata Filtering by tenant_id
&lt;/span&gt;    &lt;span class="n"&gt;hits&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;collection_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;COLLECTION_NAME&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;query_vector&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;query_vector&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;query_filter&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;models&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;must&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;models&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;FieldCondition&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tenant_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="n"&gt;match&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;models&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;MatchValue&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;tenant_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="p"&gt;)]&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;limit&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;score_threshold&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;threshold&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;hits&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; Secure Cache HIT! (Score: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;hits&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;hits&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;llm_response&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Phase 4: The Similarity Threshold Dilemma
&lt;/h2&gt;

&lt;p&gt;Configuring your Similarity Threshold is a delicate precision vs. recall tradeoff. Standard guides often suggest a blanket threshold of 0.80, which causes severe hallucinations.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Embedding-Close is NOT Meaning-Equal:&lt;/strong&gt; "What is the capital of France?" and "What is the capital of Germany?" cluster very close together in vector space (often yielding cosine similarity &amp;gt;0.85) due to identical sentence structures.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If your threshold is set to 0.80, the user asking about Germany receives the cached answer for France. Tune thresholds dynamically per route:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;0.95+ Threshold:&lt;/strong&gt; For strict factual, technical, or financial queries.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;0.88 Threshold:&lt;/strong&gt; For generic conversational FAQs.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Phase 5: Overcoming Vector DB IOPS Bottlenecks
&lt;/h2&gt;

&lt;p&gt;Vector Databases using HNSW algorithms are intensely demanding on system memory and Disk I/O. Running high-throughput vector similarity searches on public clouds forces reliance on cloud block storage with astronomical &lt;strong&gt;Provisioned IOPS (io2)&lt;/strong&gt; fees.&lt;/p&gt;

&lt;p&gt;Deploying your Python vector semantic cache stack on &lt;strong&gt;ServerMO Dedicated Bare Metal Servers&lt;/strong&gt; eliminates cloud storage taxes entirely:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Unmetered NVMe IOPS:&lt;/strong&gt; Direct-attached PCIe Enterprise NVMe drives deliver millions of raw IOPS at zero extra cost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sub-15ms Latency:&lt;/strong&gt; High DDR5 RAM capacity paired with dedicated hardware keeps embedding lookups locked under 15ms under heavy concurrent load.&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;👉 &lt;strong&gt;Read the full technical tutorial on ServerMO:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;a href="https://www.servermo.com/howto/semantic-caching-llms-ubuntu/" rel="noopener noreferrer"&gt;Semantic Caching for LLMs on Ubuntu 24.04: Reduce API Costs | ServerMO&lt;/a&gt;&lt;/p&gt;

</description>
      <category>openai</category>
      <category>python</category>
      <category>ubuntu</category>
      <category>devops</category>
    </item>
    <item>
      <title>The OpenTelemetry K8s Cost Trap (And How to Fix It)</title>
      <dc:creator>Jakson Tate</dc:creator>
      <pubDate>Thu, 03 Sep 2026 06:11:30 +0000</pubDate>
      <link>https://dev.to/jaksontate/the-opentelemetry-k8s-cost-trap-and-how-to-fix-it-4mm2</link>
      <guid>https://dev.to/jaksontate/the-opentelemetry-k8s-cost-trap-and-how-to-fix-it-4mm2</guid>
      <description>&lt;p&gt;You instrumented your applications correctly, adopted OpenTelemetry (OTel), and finally achieved end-to-end tracing across your microservices. The engineering team is thrilled. Then, Finance forwards you the monthly cloud invoice, and the CTO demands a meeting. Your observability costs have surged 10x in a single month.&lt;/p&gt;

&lt;p&gt;What happened? The truth about OpenTelemetry in Kubernetes is that while the OTel software is open-source and free, the infrastructure required to transport and store the telemetry is not. When you combine the explosive data volume of OTel auto-instrumentation with the rapid churn of Kubernetes autoscaling, you create a perfect financial storm.&lt;/p&gt;

&lt;p&gt;Here is an SRE/FinOps breakdown exposing hidden cloud network taxes, metric cardinality explosions, distributed trace sampling bugs, and how to fix them.&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 1: HPA and The Cardinality Explosion
&lt;/h2&gt;

&lt;p&gt;When investigating why observability SaaS bills (like Datadog or Splunk) explode, Kubernetes clusters often reveal a hidden culprit: the Horizontal Pod Autoscaler (HPA). As traffic spikes, your HPA rapidly spins up dozens of new pods, and then destroys them when traffic subsides.&lt;/p&gt;

&lt;p&gt;Vendors charge heavily for Custom Metrics based on "Cardinality" (the number of unique metric combinations). Every time HPA creates a new pod, it generates ephemeral labels like &lt;code&gt;k8s.pod.uid&lt;/code&gt; or dynamic &lt;code&gt;k8s.pod.name&lt;/code&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;🚨 &lt;strong&gt;The OTTL Context Blindspot:&lt;/strong&gt; These labels are Resource Attributes, not Datapoint Attributes. If you try to strip them using &lt;code&gt;context: datapoint&lt;/code&gt; in your OpenTelemetry Transformation Language (OTTL) configuration, it will silently fail.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  The SRE Fix
&lt;/h3&gt;

&lt;p&gt;Filter high-cardinality attributes at the OTel Collector using the &lt;code&gt;resource&lt;/code&gt; context before exporting metrics:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;processors&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;transform/metrics&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;error_mode&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ignore&lt;/span&gt;
    &lt;span class="na"&gt;metric_statements&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;context&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;resource&lt;/span&gt; &lt;span class="c1"&gt;# CRITICAL: You must use 'resource' context for K8s pod labels!&lt;/span&gt;
        &lt;span class="na"&gt;statements&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="c1"&gt;# Strip ephemeral pod identifiers before export to prevent billing spikes&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;delete_key(attributes, "k8s.pod.uid")&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;delete_key(attributes, "k8s.pod.name")&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Phase 2: The Stealth Egress Tax (NAT &amp;amp; Cross-AZ)
&lt;/h2&gt;

&lt;p&gt;The most devious OpenTelemetry Kubernetes cost trap doesn't come from your observability vendor. It comes directly from AWS, Azure, or GCP. Cloud providers charge heavily for data leaving their network or moving between boundaries:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Cross-AZ Penalty:&lt;/strong&gt; A best-practice OTel architecture uses Edge DaemonSets that forward data to a centralized Gateway Collector. If a DaemonSet in &lt;code&gt;us-east-1a&lt;/code&gt; sends 5TB of traces to a Gateway in &lt;code&gt;us-east-1b&lt;/code&gt;, cloud providers charge ~$0.01/GB in &lt;em&gt;both&lt;/em&gt; directions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The NAT Gateway Processing Fee:&lt;/strong&gt; If your Kubernetes cluster sits in a private subnet, sending telemetry to a public SaaS backend requires passing through a NAT Gateway. You pay $0.045/GB for NAT processing &lt;strong&gt;PLUS&lt;/strong&gt; $0.09/GB for Internet Egress.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The Reality:&lt;/strong&gt; You are paying &lt;strong&gt;~$0.135 per GB&lt;/strong&gt; just to move your own data, before your vendor even bills you for ingestion!&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 3: Fixing the Sampling Blindspot
&lt;/h2&gt;

&lt;p&gt;To survive egress taxes, you must aggressively reduce trace volume before it leaves your cluster. However, defaulting to Head-Based Sampling randomly drops traces at inception, destroying 90% of your critical P99 ERROR traces. You must use &lt;strong&gt;Tail-Based Sampling&lt;/strong&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;🚨 &lt;strong&gt;The Multi-Replica Topology Error:&lt;/strong&gt; Tail sampling requires the processor to evaluate the complete trace. If you run multiple OTel Gateway Pods, Span 1 might hit Gateway A, and Span 2 might hit Gateway B. The tail sampling logic gets confused and drops critical error traces.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  The SRE Cure
&lt;/h3&gt;

&lt;p&gt;Deploy a Load Balancing Exporter on your Edge Agents (DaemonSets) with &lt;code&gt;routing_key: "traceID"&lt;/code&gt;. This ensures all spans of the same trace reliably hit the exact same Gateway replica where the tail-sampling processor lives.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# 1. Edge Agent (DaemonSet) Configuration&lt;/span&gt;
&lt;span class="na"&gt;exporters&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;loadbalancing&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;routing_key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;traceID"&lt;/span&gt; &lt;span class="c1"&gt;# Ensure all spans for a trace reach the same Gateway replica&lt;/span&gt;
    &lt;span class="na"&gt;protocol&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;otlp&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;endpoint&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;http&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;&lt;span class="nv"&gt;//gateway-service.observability.svc.cluster.local&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;&lt;span class="nv"&gt;4317&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;&lt;span class="s"&gt;(http://gateway-service.observability.svc.cluster.local:4317)&lt;/span&gt;

&lt;span class="c1"&gt;# ----------------------------------------------------&lt;/span&gt;

&lt;span class="c1"&gt;# 2. Gateway Collector Configuration&lt;/span&gt;
&lt;span class="na"&gt;processors&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;tail_sampling&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;decision_wait&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;10s&lt;/span&gt; &lt;span class="c1"&gt;# Buffer time to wait for trace completion&lt;/span&gt;
    &lt;span class="na"&gt;num_traces&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;100000&lt;/span&gt; &lt;span class="c1"&gt;# Memory sizing (Monitor OOM kills!)&lt;/span&gt;
    &lt;span class="na"&gt;policies&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="c1"&gt;# Policy 1: Always keep 100% of Errors&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;keep-errors&lt;/span&gt;
        &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;status_code&lt;/span&gt;
        &lt;span class="na"&gt;status_code&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;status_codes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;ERROR&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
      &lt;span class="c1"&gt;# Policy 2: Sample only 5% of healthy normal traffic&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;sample-healthy&lt;/span&gt;
        &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;probabilistic&lt;/span&gt;
        &lt;span class="na"&gt;probabilistic&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;sampling_percentage&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;5&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Phase 4: Repatriating Observability to Bare Metal
&lt;/h2&gt;

&lt;p&gt;Even with aggressive Tail-Based sampling, high-throughput microservices will still generate terabytes of vital telemetry data. The public cloud billing model fundamentally penalizes you for deeply monitoring your own infrastructure.&lt;/p&gt;

&lt;p&gt;This is why Elite engineering organizations are repatriating heavy observability stacks (Prometheus, Grafana Loki, ClickHouse) to &lt;strong&gt;ServerMO Dedicated Bare Metal Servers&lt;/strong&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Zero Egress Fees:&lt;/strong&gt; ServerMO provides Unmetered or massive Flat-Rate Bandwidth. Stream 50TB+ of OpenTelemetry data daily with zero cross-AZ or NAT fees.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Free Write IOPS:&lt;/strong&gt; Observability is a 100% write-heavy workload. Cloud providers charge astronomical Provisioned IOPS (e.g., AWS io2) fees for heavy storage writes. ServerMO's direct-attached Enterprise PCIe NVMe drives deliver millions of write IOPS at zero extra cost.&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;👉 &lt;strong&gt;Read the full FinOps guide on ServerMO:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;a href="https://www.servermo.com/blogs/opentelemetry-kubernetes-cost-trap/" rel="noopener noreferrer"&gt;The OpenTelemetry Trap: How K8s Autoscaling Bankrupts Your Cloud Bill | ServerMO&lt;/a&gt;&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>opentelemetry</category>
      <category>finops</category>
      <category>devops</category>
    </item>
    <item>
      <title>Fix Kubernetes CPU Throttling &amp; CFS Quotas</title>
      <dc:creator>Jakson Tate</dc:creator>
      <pubDate>Thu, 03 Sep 2026 05:45:34 +0000</pubDate>
      <link>https://dev.to/jaksontate/fix-kubernetes-cpu-throttling-cfs-quotas-137g</link>
      <guid>https://dev.to/jaksontate/fix-kubernetes-cpu-throttling-cfs-quotas-137g</guid>
      <description>&lt;p&gt;The P99 latency alert fires while Grafana reports 20% CPU usage. Welcome to &lt;strong&gt;Kubernetes CPU throttling&lt;/strong&gt;, where Linux CFS Quotas forcibly freeze containers in microscopic 100ms windows despite low average usage.&lt;/p&gt;

&lt;p&gt;Here is the complete SRE blueprint to detect micro-freezes, tune runtimes, and bypass CFS limits using bare-metal static core pinning.&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 1: The Dashboard Illusion (100ms CFS Quota)
&lt;/h2&gt;

&lt;p&gt;Standard CPU metrics average usage over 1 to 5 minutes, hiding microscopic kernel freezes. Kubernetes CPU limits are actively enforced by the Linux kernel's Completely Fair Scheduler (CFS) in &lt;strong&gt;100-millisecond windows&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;100m&lt;/code&gt; Limit:&lt;/strong&gt; Grants 10ms of execution time per 100ms window.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;500m&lt;/code&gt; Limit:&lt;/strong&gt; Grants 50ms of execution time per 100ms window.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When a bursty, multi-threaded application consumes its 100ms allocation in the first 10ms, &lt;strong&gt;the kernel forcibly freezes the container for the remaining 90ms&lt;/strong&gt;. Your dashboard reports low average CPU usage, but your application was completely dead for 90% of that second.&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 2: Detecting Throttling via PromQL
&lt;/h2&gt;

&lt;p&gt;Stop tracking raw CPU usage (&lt;code&gt;container_cpu_usage_seconds_total&lt;/code&gt;). Instead, track the percentage of 100ms windows where the kernel parked your application:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# Throttling Ratio Query (%)
sum by (namespace, pod, container) (
  rate(container_cpu_cfs_throttled_periods_total{container!="", container!="POD"}[5m])
) 
/ 
sum by (namespace, pod, container) (
  rate(container_cpu_cfs_periods_total{container!="", container!="POD"}[5m])
) * 100
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;🚨 &lt;strong&gt;SRE Rule:&lt;/strong&gt; If this metric exceeds &lt;strong&gt;15–25%&lt;/strong&gt;, your application is suffering from artificial latency spikes caused by cgroup throttling.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Phase 3: Fixing Thread Amplification (JVM &amp;amp; Go)
&lt;/h2&gt;

&lt;p&gt;On a bare-metal node with 64 physical cores, Java and Go runtimes inspect the host OS, detect 64 cores, and spawn 64 Garbage Collection or worker threads. If your container limit is set to 2 vCPUs, all 64 threads wake up simultaneously and burn through your 100ms quota in milliseconds.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Java / JVM (JDK 11+)
&lt;/h3&gt;

&lt;p&gt;Rely on &lt;code&gt;UseContainerSupport&lt;/code&gt; so the JVM auto-detects container cgroup limits:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="k"&gt;ENV&lt;/span&gt;&lt;span class="s"&gt; JAVA_OPTS="-XX:+UseContainerSupport"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  2. Golang
&lt;/h3&gt;

&lt;p&gt;Import Uber's &lt;code&gt;automaxprocs&lt;/code&gt; library to auto-tune &lt;code&gt;GOMAXPROCS&lt;/code&gt; to cgroup limits rather than host core counts:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="s"&gt;"go.uber.org/automaxprocs"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Phase 4: The 2x P99 Right-Sizing Rule
&lt;/h2&gt;

&lt;p&gt;Removing CPU limits entirely leaves nodes vulnerable to CPU Exhaustion DoS attacks and Node Starvation risks. Instead of removing limits, right-size using telemetry:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Measure &lt;strong&gt;P50 (median) CPU usage&lt;/strong&gt; over 7 days $\rightarrow$ Set as &lt;code&gt;requests.cpu&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Measure &lt;strong&gt;P99 (peak burst) CPU usage&lt;/strong&gt; over 7 days $\rightarrow$ Multiply by 2 $\rightarrow$ Set as &lt;code&gt;limits.cpu&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Phase 5: Eradicating Throttling on Bare Metal
&lt;/h2&gt;

&lt;p&gt;Public Cloud VMs suffer from "Double Throttling"—K8s CFS limits combined with Hypervisor Steal Time from noisy neighbors.&lt;/p&gt;

&lt;p&gt;On &lt;strong&gt;ServerMO Dedicated Bare Metal&lt;/strong&gt;, enable Kubelet's &lt;code&gt;cpuManagerPolicy: static&lt;/code&gt;. By deploying pods in the &lt;strong&gt;Guaranteed QoS class&lt;/strong&gt; (&lt;code&gt;requests&lt;/code&gt; equal &lt;code&gt;limits&lt;/code&gt; using integer CPU values), Kubernetes completely bypasses the CFS 100ms quota system and pins the container directly to dedicated physical CPU cores.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Pod&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;high-performance-api&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;api&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;enterprise-api:v1.2&lt;/span&gt;
    &lt;span class="na"&gt;resources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;limits&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;cpu&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;4"&lt;/span&gt;        &lt;span class="c1"&gt;# Integer value required&lt;/span&gt;
        &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;8Gi"&lt;/span&gt;
      &lt;span class="na"&gt;requests&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;cpu&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;4"&lt;/span&gt;        &lt;span class="c1"&gt;# Must equal limits for Guaranteed QoS&lt;/span&gt;
        &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;8Gi"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;p&gt;👉 &lt;strong&gt;Read the full SRE guide on ServerMO:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;a href="https://www.servermo.com/howto/fix-kubernetes-cpu-throttling/" rel="noopener noreferrer"&gt;Stop Kubernetes CPU Throttling | ServerMO&lt;/a&gt;&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>devops</category>
      <category>sre</category>
      <category>linux</category>
    </item>
    <item>
      <title>What is ECC RAM? True ECC vs DDR5 On-Die &amp; Server Needs</title>
      <dc:creator>Jakson Tate</dc:creator>
      <pubDate>Thu, 03 Sep 2026 05:07:11 +0000</pubDate>
      <link>https://dev.to/jaksontate/what-is-ecc-ram-true-ecc-vs-ddr5-on-die-server-needs-49l7</link>
      <guid>https://dev.to/jaksontate/what-is-ecc-ram-true-ecc-vs-ddr5-on-die-server-needs-49l7</guid>
      <description>&lt;p&gt;Data center engineers share a universal fear: &lt;strong&gt;silent data corruption&lt;/strong&gt;. You can have the fastest NVMe storage and the most powerful CPUs in the world, but if the volatile workspace linking them—your Random Access Memory (RAM)—is flawed, your entire enterprise infrastructure is built on sand.&lt;/p&gt;

&lt;p&gt;Here is a deep dive into cosmic ray bit-flips, 80-bit DDR5 bus math, and why consumer DDR5 "On-Die ECC" fails server workloads.&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 1: What Causes RAM Bit-Flips?
&lt;/h2&gt;

&lt;p&gt;Volatile memory cells hold binary data using tiny electrical charges that can be disrupted by external forces:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cosmic Rays (Single-Event Upsets):&lt;/strong&gt; Secondary neutrons from cosmic ray collisions constantly strike silicon chips, altering transistor states and flipping binary &lt;code&gt;0&lt;/code&gt;s to &lt;code&gt;1&lt;/code&gt;s.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Electromagnetic Interference (EMI):&lt;/strong&gt; Signal degradation inside server chassis caused by power supply fluctuations or unshielded components.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cell Aging &amp;amp; Thermal Stress:&lt;/strong&gt; High temperatures in 24/7 server environments degrade silicon's ability to retain electrical charge accurately.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Phase 2: How Does ECC RAM Work? (72-Bit vs 80-Bit Bus Math)
&lt;/h2&gt;

&lt;p&gt;Standard non-ECC RAM uses a 64-bit data path and blindly trusts incoming data. True ECC RAM adds dedicated parity pathways.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DDR3/DDR4 True ECC Architecture:
[ 64-bit Data Path ] + [ 8-bit Parity (9th Chip) ] = 72-bit Bus Width

DDR5 True Enterprise ECC Architecture:
[ Channel A: 32-bit Data + 8-bit Parity ] + [ Channel B: 32-bit Data + 8-bit Parity ] = 80-bit Bus Width
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;ECC uses &lt;strong&gt;Hamming codes&lt;/strong&gt; to provide &lt;strong&gt;SECDED&lt;/strong&gt; (Single Error Correction, Double Error Detection). When the CPU reads data, the memory controller recalculates the checksum. Single-bit errors are corrected on the fly, while double-bit errors safely halt the system before corrupted data reaches persistent storage.&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 3: The DDR5 "On-Die" Illusion vs True Side-Band ECC
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;🚨 &lt;strong&gt;SRE HARDWARE WARNING: The DDR5 Myth&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
A dangerous misconception claims that consumer DDR5 has built-in ECC, making server-grade RAM obsolete. &lt;strong&gt;This is false.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;As DDR5 memory cells shrank to achieve higher clock speeds, internal silicon bit-flips increased. Manufacturers added &lt;strong&gt;On-Die ECC&lt;/strong&gt; to consumer DDR5 purely to improve silicon manufacturing yields.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;On-Die ECC:&lt;/strong&gt; Fixes errors &lt;em&gt;internally&lt;/em&gt; inside the memory chip. Data in transit across the motherboard bus to the CPU remains &lt;strong&gt;100% unprotected&lt;/strong&gt;. It does not log errors to the OS or Baseboard Management Controller (BMC).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;True Side-Band ECC (80-bit RDIMM):&lt;/strong&gt; Provides end-to-end protection for data in transit and logs single-bit errors for predictive failure analysis.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Phase 4: The ZFS Reality (Debunking the Scrub of Death)
&lt;/h2&gt;

&lt;p&gt;A popular myth dictates that bad RAM causes ZFS to execute a "Scrub of Death" that overwrites good disk data with corrupted data. This is false—ZFS checksums protect existing data on disk.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The True Danger: Dirty Memory Buffers&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
When new data enters RAM (within Dirty Data Buffers or Transaction Groups / TXG) &lt;em&gt;before&lt;/em&gt; ZFS generates its checksum, an uncorrected bit-flip corrupts the payload in memory. ZFS then calculates a valid checksum for the corrupted payload and writes it permanently to storage.&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 5: Gaming vs Enterprise Performance
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Gamers:&lt;/strong&gt; Non-ECC is superior. Overclocked profiles (Intel XMP / AMD EXPO) reach 6000MT/s+ CL30, yielding a 15–20% FPS advantage over JEDEC baseline speeds.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enterprise Servers:&lt;/strong&gt; Absolute stability is mandatory. Enterprise ECC RDIMMs strictly adhere to baseline JEDEC standards (e.g., 4800MT/s CL40) to guarantee 24/7 reliability.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  ECC RAM &amp;amp; Server Reliability FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is DDR5 On-Die ECC the same as True Server ECC?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
No. On-Die ECC only fixes internal silicon errors inside individual memory chips. It does not protect data traveling over the motherboard bus to the CPU. Enterprise servers require True Side-Band ECC (80-bit bus) for end-to-end protection.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do I need ECC RAM for ZFS?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Yes. While the "Scrub of Death" myth is false, non-ECC RAM can flip bits in dirty memory buffers before ZFS generates checksums, causing ZFS to write permanently corrupted data to storage.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why don't gaming PCs use ECC RAM?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Enterprise ECC RAM runs at baseline JEDEC speeds and does not support consumer XMP/EXPO overclocking, resulting in lower clock speeds and higher latencies compared to consumer RAM.&lt;/p&gt;




&lt;p&gt;👉 &lt;strong&gt;Read the full guide on ServerMO:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;a href="https://www.servermo.com/blogs/ecc-ram-ddr5-vs-true-ecc/" rel="noopener noreferrer"&gt;What is ECC RAM? True ECC vs DDR5 On-Die | ServerMO&lt;/a&gt;&lt;/p&gt;

</description>
      <category>linux</category>
      <category>devops</category>
      <category>hardware</category>
      <category>sre</category>
    </item>
    <item>
      <title>Model Context Protocol: Setup MCP Server on Bare Metal</title>
      <dc:creator>Jakson Tate</dc:creator>
      <pubDate>Fri, 21 Aug 2026 06:30:05 +0000</pubDate>
      <link>https://dev.to/jaksontate/model-context-protocol-setup-mcp-server-on-bare-metal-1hdj</link>
      <guid>https://dev.to/jaksontate/model-context-protocol-setup-mcp-server-on-bare-metal-1hdj</guid>
      <description>&lt;p&gt;Integrating Large Language Models (LLMs) with enterprise databases using custom REST APIs requires constant glue-code maintenance. The &lt;strong&gt;Model Context Protocol (MCP)&lt;/strong&gt; provides a universal, standardized interface for dynamic tool discovery via JSON-RPC.&lt;/p&gt;

&lt;p&gt;Here is how to deploy a secure FastMCP server on Bare Metal infrastructure.&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 1: Environment Setup with &lt;code&gt;uv&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;Bypass standard &lt;code&gt;pip&lt;/code&gt; bloat and use &lt;code&gt;uv&lt;/code&gt;—the Rust-based Python package manager—for ultra-fast dependency isolation.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;⚠️ &lt;strong&gt;SRE WARNING: The FastMCP Import Anomaly&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Do not install &lt;code&gt;mcp[cli]&lt;/code&gt; when writing custom Python code using &lt;code&gt;from fastmcp import FastMCP&lt;/code&gt;. Namespace collisions cause an immediate &lt;code&gt;ModuleNotFoundError&lt;/code&gt;. Install &lt;code&gt;fastmcp&lt;/code&gt; directly.&lt;br&gt;
&lt;/p&gt;
&lt;/blockquote&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 1. Install 'uv' globally&lt;/span&gt;
curl &lt;span class="nt"&gt;-LsSf&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;https://astral.sh/uv/install.sh]&lt;span class="o"&gt;(&lt;/span&gt;https://astral.sh/uv/install.sh&lt;span class="o"&gt;)&lt;/span&gt; | sh
&lt;span class="nb"&gt;source&lt;/span&gt; &lt;span class="nv"&gt;$HOME&lt;/span&gt;/.local/bin/env

&lt;span class="c"&gt;# 2. Setup project directory&lt;/span&gt;
&lt;span class="nb"&gt;mkdir&lt;/span&gt; ~/enterprise-mcp &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd&lt;/span&gt; ~/enterprise-mcp
uv init
uv venv
&lt;span class="nb"&gt;source&lt;/span&gt; .venv/bin/activate

&lt;span class="c"&gt;# 3. SRE FIX: Install FastMCP package explicitly&lt;/span&gt;
uv add fastmcp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Phase 2: Architecting the Secure FastMCP Server
&lt;/h2&gt;

&lt;p&gt;MCP uses Standard I/O (&lt;code&gt;stdio&lt;/code&gt;) to stream JSON-RPC messages to AI clients.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;🚨 &lt;strong&gt;SECURITY ALERT: The print() Crash Trap&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Never use standard &lt;code&gt;print()&lt;/code&gt; statements in your MCP server code. Outputting text to &lt;code&gt;stdout&lt;/code&gt; injects raw string data into the JSON stream, corrupting the protocol and instantly crashing Claude Desktop. Force all logging exclusively to &lt;code&gt;sys.stderr&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Create &lt;code&gt;server.py&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;sys&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;logging&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;fastmcp&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;FastMCP&lt;/span&gt;

&lt;span class="c1"&gt;# SRE FIX: Force all logging to stderr to protect the JSON-RPC stdio stream
&lt;/span&gt;&lt;span class="n"&gt;logging&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;basicConfig&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;level&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;logging&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;INFO&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stderr&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;format&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;%(levelname)s: %(message)s&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;logger&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;logging&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getLogger&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;__name__&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;mcp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;FastMCP&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Enterprise-Data-Gateway&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nd"&gt;@mcp.tool&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_server_status&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Returns operational status of the Bare Metal server.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;logger&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Tool called: get_server_status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ServerMO Bare Metal Node 01: All systems operational. 0% Packet Loss.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;logger&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Starting MCP Server on stdio transport...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;mcp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;transport&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;stdio&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Phase 3: Hardening Against Path Traversal
&lt;/h2&gt;

&lt;p&gt;Granting AI agents file system tools requires strict sandboxing to prevent prompt injections from exfiltrating system secrets like &lt;code&gt;/etc/passwd&lt;/code&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;🚨 &lt;strong&gt;SECURITY ALERT: The Sibling Directory Bypass&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
String checks like &lt;code&gt;requested_path.startswith(str(ALLOWED_DIR))&lt;/code&gt; create critical vulnerabilities! If your allowed directory is &lt;code&gt;/home/data&lt;/code&gt;, requesting &lt;code&gt;/home/data-secret/pass.txt&lt;/code&gt; passes string matching. Use &lt;code&gt;Path.is_relative_to()&lt;/code&gt; instead.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Add file sandboxing to &lt;code&gt;server.py&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;

&lt;span class="n"&gt;ALLOWED_DIR&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/home/ubuntu/enterprise-mcp/data&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;resolve&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="nd"&gt;@mcp.tool&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;read_secure_file&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;filename&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Reads a text file strictly from the allowed data sandbox.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;requested_path&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ALLOWED_DIR&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;filename&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;resolve&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="c1"&gt;# SRE FIX: Prevent Sibling Directory Path Traversal Bypass
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;requested_path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;is_relative_to&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ALLOWED_DIR&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;logger&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Security Violation: Attempted path traversal to &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;requested_path&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ERROR: Access Denied. Path traversal detected.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;requested_path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;is_file&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ERROR: File &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;filename&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt; not found in sandbox.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;requested_path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;logger&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Read error: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ERROR: Could not read file due to permissions or locking.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Phase 4: Remote Stdio over SSH (Claude Desktop Integration)
&lt;/h2&gt;

&lt;p&gt;To connect local Claude Desktop to this remote MCP server without exposing public ports:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;⚠️ &lt;strong&gt;SRE WARNING: The MOTD &amp;amp; Working Directory Trap&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Running standard SSH commands causes two critical failures:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Ubuntu MOTD banners ("Welcome to Ubuntu") corrupt the JSON-RPC stream.&lt;/li&gt;
&lt;li&gt;SSH defaults to &lt;code&gt;/home/ubuntu&lt;/code&gt;, completely blinding &lt;code&gt;uv&lt;/code&gt; to your virtual environment.&lt;/li&gt;
&lt;/ol&gt;
&lt;/blockquote&gt;

&lt;p&gt;Pass &lt;code&gt;-q -T&lt;/code&gt; to SSH to kill login banners, and pass &lt;code&gt;--directory&lt;/code&gt; and &lt;code&gt;--quiet&lt;/code&gt; to &lt;code&gt;uv&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Edit &lt;code&gt;claude_desktop_config.json&lt;/code&gt; on your local laptop:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"enterprise-bare-metal"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ssh"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"args"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"-q"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"-T"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"-i"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"/path/to/private_key.pem"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"ubuntu@YOUR_REMOTE_SERVER_IP"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"/home/ubuntu/.local/bin/uv"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"--directory"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"/home/ubuntu/enterprise-mcp"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"run"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"--quiet"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"server.py"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Restart Claude Desktop to connect your local AI agent directly over an encrypted SSH pipe.&lt;/p&gt;




&lt;h2&gt;
  
  
  FastMCP &amp;amp; Model Context Protocol FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;MCP Protocol vs REST API: Which is better for AI Agents?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
REST APIs require custom integration code and static endpoint management for every tool. MCP standardizes client-server interactions, allowing AI agents to dynamically discover and execute tools via JSON-RPC.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why did my MCP Server crash Claude Desktop?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Standard &lt;code&gt;print()&lt;/code&gt; statements or SSH MOTD banners write raw text to &lt;code&gt;stdout&lt;/code&gt;, corrupting the JSON-RPC stream. Route Python logs to &lt;code&gt;sys.stderr&lt;/code&gt; and use &lt;code&gt;ssh -q -T&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why is &lt;code&gt;.startswith()&lt;/code&gt; dangerous for directory path checks?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;code&gt;string.startswith("/data")&lt;/code&gt; matches &lt;code&gt;/data-secret/file.txt&lt;/code&gt;. Always use &lt;code&gt;path.is_relative_to(ALLOWED_DIR)&lt;/code&gt; for filesystem sandboxing.&lt;/p&gt;




&lt;p&gt;👉 &lt;strong&gt;Read the full guide on ServerMO:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;a href="https://www.servermo.com/howto/setup-mcp-server-bare-metal/" rel="noopener noreferrer"&gt;Setup MCP Server on Bare Metal | ServerMO&lt;/a&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>ai</category>
      <category>devops</category>
      <category>security</category>
    </item>
    <item>
      <title>10 Linux Server Disasters &amp; Open-Source SRE Cures</title>
      <dc:creator>Jakson Tate</dc:creator>
      <pubDate>Thu, 20 Aug 2026 11:58:47 +0000</pubDate>
      <link>https://dev.to/jaksontate/10-linux-server-disasters-open-source-sre-cures-13dj</link>
      <guid>https://dev.to/jaksontate/10-linux-server-disasters-open-source-sre-cures-13dj</guid>
      <description>&lt;p&gt;When a production server goes down at 2 AM, standard beginner advice fails. Running &lt;code&gt;df -h&lt;/code&gt;, staring blindly at &lt;code&gt;top&lt;/code&gt;, or arbitrarily executing &lt;code&gt;systemctl restart&lt;/code&gt; without understanding the root cause leads to prolonged downtime and potential data loss.&lt;/p&gt;

&lt;p&gt;Here are 10 critical Linux server disasters and the open-source Site Reliability Engineering (SRE) techniques required to fix them permanently.&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 1: Surviving the OOM-Killer Blindspot
&lt;/h2&gt;

&lt;p&gt;When your server runs out of RAM, the Linux kernel invokes the OOM Killer to terminate processes.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;⚠️ &lt;strong&gt;SRE WARNING: The &lt;code&gt;overcommit_memory=2&lt;/code&gt; Trap&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Running &lt;code&gt;echo "vm.overcommit_memory = 2" &amp;gt; /etc/sysctl.conf&lt;/code&gt; forces strict memory checks. Databases like MySQL or PostgreSQL request large virtual memory blocks on boot. Under strict overcommit settings, they will throw "Cannot allocate memory" and refuse to start, even when physical RAM is free.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;The SRE Fix:&lt;/strong&gt; Shield critical services via Systemd, not kernel sysctl settings:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl edit mysql

&lt;span class="c"&gt;# Add the following lines to grant OOM immunity:&lt;/span&gt;
&lt;span class="o"&gt;[&lt;/span&gt;Service]
&lt;span class="nv"&gt;OOMScoreAdjust&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nt"&gt;-1000&lt;/span&gt;

&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl daemon-reload
&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl restart mysql
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Phase 2: High CPU Diagnosis &amp;amp; The &lt;code&gt;iowait&lt;/code&gt; Trap
&lt;/h2&gt;

&lt;p&gt;If &lt;code&gt;htop&lt;/code&gt; shows low CPU utilization but system load average is over 50, your CPU is stalled waiting for storage I/O (&lt;strong&gt;iowait&lt;/strong&gt;).&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Diagnosis:&lt;/strong&gt; Install &lt;code&gt;sysstat&lt;/code&gt; and run &lt;code&gt;iostat -xz 1&lt;/code&gt;. Check &lt;code&gt;%util&lt;/code&gt; and &lt;code&gt;await&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cure:&lt;/strong&gt; Adding CPU cores won't resolve disk I/O bottlenecks. Migrate I/O-intensive workloads to &lt;strong&gt;ServerMO Bare Metal&lt;/strong&gt; with enterprise direct-attached NVMe storage.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Phase 3: Fixing the Silent Disk Full Error (Inodes)
&lt;/h2&gt;

&lt;p&gt;An application throws &lt;code&gt;No space left on device&lt;/code&gt;, but &lt;code&gt;df -h&lt;/code&gt; shows the disk is 50% free. This indicates &lt;strong&gt;Inode Exhaustion&lt;/strong&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;🚨 &lt;strong&gt;CRITICAL SRE ALERT: The Server-Crashing &lt;code&gt;find&lt;/code&gt; Loop&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Avoid running nested &lt;code&gt;find&lt;/code&gt; loops on choked production machines. It induces massive I/O load and hangs the server.&lt;br&gt;
&lt;/p&gt;
&lt;/blockquote&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 1. Verify Inode usage&lt;/span&gt;
&lt;span class="nb"&gt;df&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt;

&lt;span class="c"&gt;# 2. Find top Inode consumers safely without hanging the machine&lt;/span&gt;
&lt;span class="nb"&gt;sudo du&lt;/span&gt; &lt;span class="nt"&gt;--inodes&lt;/span&gt; &lt;span class="nt"&gt;-xS&lt;/span&gt; / | &lt;span class="nb"&gt;sort&lt;/span&gt; &lt;span class="nt"&gt;-rh&lt;/span&gt; | &lt;span class="nb"&gt;head&lt;/span&gt; &lt;span class="nt"&gt;-20&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Phase 4: Resolving Database Choking
&lt;/h2&gt;

&lt;p&gt;If slow database queries bottleneck your app, enable &lt;code&gt;slow_query_log&lt;/code&gt; in MySQL/MariaDB or &lt;code&gt;log_min_duration_statement&lt;/code&gt; in PostgreSQL. Prefix captured queries with &lt;code&gt;EXPLAIN&lt;/code&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The SRE Fix:&lt;/strong&gt; If &lt;code&gt;EXPLAIN&lt;/code&gt; returns &lt;code&gt;type: ALL&lt;/code&gt; (MySQL) or &lt;code&gt;Seq Scan&lt;/code&gt; (PostgreSQL), the database engine is reading millions of rows manually. Create targeted indexes on columns used in &lt;code&gt;WHERE&lt;/code&gt;, &lt;code&gt;JOIN&lt;/code&gt;, or &lt;code&gt;ORDER BY&lt;/code&gt; clauses.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Phase 5: Network Botnets &amp;amp; CrowdSec
&lt;/h2&gt;

&lt;p&gt;Fail2Ban only analyzes local logs, making it ineffective against distributed multi-IP botnets.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Open-Source Fix:&lt;/strong&gt; Deploy &lt;strong&gt;CrowdSec&lt;/strong&gt;, an AI-driven, collaborative Intrusion Prevention System (IPS). Threat intelligence is shared globally across nodes to block malicious IPs before they reach your server.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Phase 6: SSL &amp;amp; Reverse Proxy Simplification
&lt;/h2&gt;

&lt;p&gt;Managing complex Nginx server blocks and Certbot cron jobs creates operational fragility.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Open-Source Fix:&lt;/strong&gt; Switch to &lt;strong&gt;Caddy Server&lt;/strong&gt; or &lt;strong&gt;Nginx Proxy Manager&lt;/strong&gt;. Caddy provisions TLS certificates automatically, supports HTTP/3 (QUIC) natively, and simplifies configurations into concise Caddyfiles.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Phase 7: Container Sprawl &amp;amp; Zombie Networks
&lt;/h2&gt;

&lt;p&gt;Orphaned Docker containers, untagged images, and dangling networks create IP conflicts and waste disk space.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Open-Source Fix:&lt;/strong&gt; Use &lt;strong&gt;Portainer&lt;/strong&gt; for web dashboard management, or run &lt;strong&gt;&lt;code&gt;ctop&lt;/code&gt;&lt;/strong&gt; in the CLI for real-time container metrics.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Phase 8: Configuration Drift &amp;amp; Spaghetti Servers
&lt;/h2&gt;

&lt;p&gt;Manually editing server configs creates unrepeatable "snowflake" servers.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Open-Source Fix:&lt;/strong&gt; Implement &lt;strong&gt;Ansible&lt;/strong&gt;. Define infrastructure as declarative YAML playbooks for automated, auditable deployments across your server fleet.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Phase 9: Zero-Trust Backups
&lt;/h2&gt;

&lt;p&gt;Standard &lt;code&gt;rsync&lt;/code&gt; or &lt;code&gt;tar&lt;/code&gt; scripts leave backup storage exposed to ransomware if root credentials are compromised.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Open-Source Fix:&lt;/strong&gt; Use &lt;strong&gt;Restic&lt;/strong&gt; or &lt;strong&gt;BorgBackup&lt;/strong&gt;. They generate encrypted, deduplicated, and append-only backups that prevent clients from deleting historical snapshots.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Phase 10: The Datadog Escape Plan
&lt;/h2&gt;

&lt;p&gt;Proprietary SaaS monitoring tools incur high per-host licensing costs.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The FinOps Stack:&lt;/strong&gt; Deploy &lt;strong&gt;Prometheus&lt;/strong&gt; (metrics) + &lt;strong&gt;Grafana Loki&lt;/strong&gt; (log aggregation) + &lt;strong&gt;Grafana&lt;/strong&gt; (visualization). Running this stack on dedicated Bare Metal ensures high log-ingestion performance without SaaS licensing fees.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  💬 Linux SRE Troubleshooting FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why does my server say "No space left on device" when &lt;code&gt;df -h&lt;/code&gt; shows space available?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
You have run out of index nodes (inodes). Check &lt;code&gt;df -i&lt;/code&gt;. Millions of tiny files (like session files) consume inodes regardless of remaining disk space.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do I prevent MySQL or PostgreSQL from being killed by the OOM Killer?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Set &lt;code&gt;OOMScoreAdjust=-1000&lt;/code&gt; inside a Systemd override file (&lt;code&gt;sudo systemctl edit mysql&lt;/code&gt;) to make the service immune to kernel OOM termination.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is the difference between Fail2Ban and CrowdSec?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Fail2Ban analyzes local logs on a single server. CrowdSec uses a collaborative network that shares IP blocklists across all users worldwide in real time.&lt;/p&gt;




&lt;p&gt;👉 &lt;strong&gt;Read the full guide on ServerMO:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;a href="https://www.servermo.com/blogs/linux-server-troubleshooting-tools/" rel="noopener noreferrer"&gt;10 Linux Server Disasters &amp;amp; Open-Source SRE Cures | ServerMO&lt;/a&gt;&lt;/p&gt;

</description>
      <category>linux</category>
      <category>devops</category>
      <category>sre</category>
      <category>sysadmin</category>
    </item>
    <item>
      <title>Setup JupyterLab &amp; PyTorch on Ubuntu 24.04: Remote GPU Server</title>
      <dc:creator>Jakson Tate</dc:creator>
      <pubDate>Thu, 20 Aug 2026 07:20:00 +0000</pubDate>
      <link>https://dev.to/jaksontate/setup-jupyterlab-pytorch-on-ubuntu-2404-remote-gpu-server-525h</link>
      <guid>https://dev.to/jaksontate/setup-jupyterlab-pytorch-on-ubuntu-2404-remote-gpu-server-525h</guid>
      <description>&lt;p&gt;Ditch your local laptop. Build a persistent Deep Learning lab on Bare Metal. Master Secure SSH Tunneling, avoid dependency hell, and connect directly via VSCode.&lt;/p&gt;




&lt;h2&gt;
  
  
  The End of Localhost AI
&lt;/h2&gt;

&lt;p&gt;Running serious Deep Learning models or fine-tuning Large Language Models (LLMs) on a local Mac or standard desktop is no longer viable. The VRAM requirements for modern AI mandate deploying workloads on a &lt;strong&gt;Remote GPU Server&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;However, transitioning to a headless Ubuntu server often results in "Dependency Hell" and security vulnerabilities. Modern SREs deploy &lt;strong&gt;JupyterLab&lt;/strong&gt; to provide a full browser-based IDE.&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 1: NVIDIA Drivers &amp;amp; The Miniconda Architecture
&lt;/h2&gt;

&lt;p&gt;Verify your server recognizes the NVIDIA hardware via &lt;code&gt;nvidia-smi&lt;/code&gt;. (If drivers are missing, execute &lt;code&gt;sudo ubuntu-drivers autoinstall&lt;/code&gt; and reboot).&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;⚠️ &lt;strong&gt;SRE WARNING: Anaconda Bloatware &amp;amp; Conda Crashes&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Anaconda installs gigabytes of unnecessary libraries. Use &lt;strong&gt;Miniconda&lt;/strong&gt; instead. Furthermore, running &lt;code&gt;conda activate&lt;/code&gt; immediately after &lt;code&gt;conda init&lt;/code&gt; triggers a &lt;code&gt;CommandNotFoundError&lt;/code&gt;. You must refresh your shell context using &lt;code&gt;source ~/.bashrc&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;
&lt;/blockquote&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 1. Verify NVIDIA Driver&lt;/span&gt;
nvidia-smi

&lt;span class="c"&gt;# 2. Install Miniconda (Lightweight Environment Manager)&lt;/span&gt;
wget &lt;span class="o"&gt;[&lt;/span&gt;https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh]&lt;span class="o"&gt;(&lt;/span&gt;https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="nt"&gt;-O&lt;/span&gt; miniconda.sh
bash miniconda.sh &lt;span class="nt"&gt;-b&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="nv"&gt;$HOME&lt;/span&gt;/miniconda
&lt;span class="nb"&gt;eval&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nv"&gt;$HOME&lt;/span&gt;/miniconda/bin/conda shell.bash hook&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
conda init

&lt;span class="c"&gt;# 3. SRE FIX: Refresh shell to prevent CommandNotFoundError&lt;/span&gt;
&lt;span class="nb"&gt;source&lt;/span&gt; ~/.bashrc

&lt;span class="c"&gt;# 4. Create an isolated environment for Python 3.11&lt;/span&gt;
conda create &lt;span class="nt"&gt;-n&lt;/span&gt; ai_lab &lt;span class="nv"&gt;python&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;3.11 &lt;span class="nt"&gt;-y&lt;/span&gt;
conda activate ai_lab
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Phase 2: Install PyTorch &amp;amp; JupyterLab
&lt;/h2&gt;

&lt;p&gt;Modern PyTorch 2.x ships with pre-compiled CUDA binaries, eliminating the need to install system-level CUDA toolkits manually.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 1. Install PyTorch with CUDA 12.1 Support&lt;/span&gt;
pip &lt;span class="nb"&gt;install &lt;/span&gt;torch torchvision torchaudio &lt;span class="nt"&gt;--index-url&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;https://download.pytorch.org/whl/cu121]&lt;span class="o"&gt;(&lt;/span&gt;https://download.pytorch.org/whl/cu121&lt;span class="o"&gt;)&lt;/span&gt;

&lt;span class="c"&gt;# 2. Install JupyterLab and IPykernel&lt;/span&gt;
pip &lt;span class="nb"&gt;install &lt;/span&gt;jupyterlab ipykernel

&lt;span class="c"&gt;# 3. Register your environment as a Jupyter Kernel&lt;/span&gt;
python &lt;span class="nt"&gt;-m&lt;/span&gt; ipykernel &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;--user&lt;/span&gt; &lt;span class="nt"&gt;--name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;ai_lab &lt;span class="nt"&gt;--display-name&lt;/span&gt; &lt;span class="s2"&gt;"PyTorch (GPU)"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Phase 3: Persistent Headless Execution
&lt;/h2&gt;

&lt;p&gt;Avoid training job crashes caused by dropped SSH sessions by setting a persistent hashed password and running JupyterLab in &lt;code&gt;tmux&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 1. Generate config and set a persistent password&lt;/span&gt;
jupyter server &lt;span class="nt"&gt;--generate-config&lt;/span&gt;
jupyter server password

&lt;span class="c"&gt;# 2. Start a persistent tmux session&lt;/span&gt;
tmux new &lt;span class="nt"&gt;-s&lt;/span&gt; jupyter_session

&lt;span class="c"&gt;# 3. Launch JupyterLab bound strictly to localhost&lt;/span&gt;
jupyter lab &lt;span class="nt"&gt;--no-browser&lt;/span&gt; &lt;span class="nt"&gt;--port&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;8888 &lt;span class="nt"&gt;--ip&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;127.0.0.1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Detach from tmux: Press &lt;code&gt;Ctrl+B&lt;/code&gt;, then &lt;code&gt;D&lt;/code&gt;.&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 4: Browser Access via SSH Tunnel (Zero Open Ports)
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;🚨 &lt;strong&gt;SECURITY ALERT: The Exposed Port Vulnerability&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Running &lt;code&gt;sudo ufw allow 8888&lt;/code&gt; exposes Jupyter directly to the internet. Automated botnets scan port 8888 to hijack GPUs for crypto-mining. Keep UFW closed and bridge connections securely via SSH Tunneling.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Run this command &lt;strong&gt;on your Local Laptop&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Forward Local Port 8888 to Remote Port 8888&lt;/span&gt;
ssh &lt;span class="nt"&gt;-N&lt;/span&gt; &lt;span class="nt"&gt;-L&lt;/span&gt; 8888:127.0.0.1:8888 your_username@YOUR_REMOTE_SERVER_IP
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now navigate to &lt;code&gt;http://localhost:8888&lt;/code&gt; in your local browser and enter your password.&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 5: Modern IDE Approach (VSCode Remote-SSH)
&lt;/h2&gt;

&lt;p&gt;For full local extension support (Pylance, GitHub Copilot) alongside remote execution:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Install the official &lt;strong&gt;Remote - SSH&lt;/strong&gt; extension in local VSCode.&lt;/li&gt;
&lt;li&gt;Press &lt;code&gt;F1&lt;/code&gt; -&amp;gt; &lt;code&gt;Remote-SSH: Connect to Host...&lt;/code&gt; -&amp;gt; Enter &lt;code&gt;ssh username@YOUR_SERVER_IP&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Open a &lt;code&gt;.ipynb&lt;/code&gt; notebook file on the server.&lt;/li&gt;
&lt;li&gt;Select Kernel -&amp;gt; Python Environments -&amp;gt; Select &lt;code&gt;ai_lab&lt;/code&gt;.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Verify GPU availability in VSCode Notebook
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;PyTorch Version: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;__version__&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CUDA Available: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cuda&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;is_available&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cuda&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;is_available&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Hardware Detected: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cuda&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_device_name&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;VRAM Allocated: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cuda&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;memory_allocated&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="mf"&gt;1e9&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; GB&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  💬 JupyterLab &amp;amp; Remote GPU FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;JupyterLab vs Jupyter Notebook: Which is better?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
JupyterLab is a complete browser-based IDE offering terminal access, file managers, and split views, making it superior to the legacy single-document Notebook interface for remote GPU workflows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why shouldn't I open Port 8888 on my Ubuntu Firewall?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Opening port 8888 exposes Jupyter to automated botnet scans that hijack GPU resources for crypto-mining. Access the server strictly via SSH Tunneling or VSCode Remote-SSH.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why did my 'conda activate' command crash on Ubuntu 24.04?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Running &lt;code&gt;conda init&lt;/code&gt; modifies &lt;code&gt;.bashrc&lt;/code&gt; but does not reload your active shell context automatically. You must run &lt;code&gt;source ~/.bashrc&lt;/code&gt; before running &lt;code&gt;conda activate&lt;/code&gt;.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Read the full tutorial on ServerMO:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;a href="https://www.servermo.com/howto/setup-jupyterlab-pytorch-ubuntu-gpu/" rel="noopener noreferrer"&gt;Setup JupyterLab &amp;amp; PyTorch on Ubuntu 24.04: Remote GPU Server | ServerMO&lt;/a&gt;&lt;/p&gt;

</description>
      <category>pytorch</category>
      <category>jupyterlab</category>
      <category>ubuntu</category>
      <category>gpu</category>
    </item>
    <item>
      <title>Deploy Apache CloudStack on Ubuntu 24.04: Build a Private AWS</title>
      <dc:creator>Jakson Tate</dc:creator>
      <pubDate>Thu, 20 Aug 2026 06:43:30 +0000</pubDate>
      <link>https://dev.to/jaksontate/deploy-apache-cloudstack-on-ubuntu-2404-build-a-private-aws-430f</link>
      <guid>https://dev.to/jaksontate/deploy-apache-cloudstack-on-ubuntu-2404-build-a-private-aws-430f</guid>
      <description>&lt;p&gt;The ultimate VMware Escape Plan. Master cloud repatriation, secure KVM hypervisors, and slash enterprise IT costs natively on ServerMO Bare Metal.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Era of Cloud Repatriation
&lt;/h2&gt;

&lt;p&gt;Organizations face two massive shifts: hyper-inflated public cloud bills (AWS/Azure) and exorbitant VMware license renewals post-Broadcom acquisition.&lt;/p&gt;

&lt;p&gt;Enterprises are executing aggressive cloud repatriation strategies. When comparing CloudStack vs OpenStack vs OpenNebula, CTOs quickly realize that OpenStack requires a dedicated DevOps team to maintain its modular dependency hell. Conversely, &lt;strong&gt;Apache CloudStack&lt;/strong&gt; is a monolithic, turnkey platform that gives you a "Private AWS" out-of-the-box.&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 1: Network Bridge Configuration (Netplan)
&lt;/h2&gt;

&lt;p&gt;To provide VPC routing, isolated guest networks, and floating IPs, CloudStack requires total control over a Linux Bridge. You must strip the IP address from your physical Network Interface Card (NIC) and assign it to a bridge.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;⚠️ &lt;strong&gt;SRE ARCHITECTURE WARNING: Physical NIC DHCP&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
You MUST disable DHCP on your physical ethernet interface (e.g., &lt;code&gt;eth0&lt;/code&gt;). If the physical NIC and the bridge both try to claim an IP address, your server will experience a catastrophic routing loop and disconnect from the network.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Edit &lt;code&gt;/etc/netplan/01-netcfg.yaml&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;network&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;version&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;2&lt;/span&gt;
  &lt;span class="na"&gt;renderer&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;networkd&lt;/span&gt;
  &lt;span class="na"&gt;ethernets&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;eth0&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;dhcp4&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
      &lt;span class="na"&gt;dhcp6&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
  &lt;span class="na"&gt;bridges&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;cloudbr0&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;addresses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;192.168.1.10/24&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt; &lt;span class="c1"&gt;# Your Bare Metal Server IP&lt;/span&gt;
      &lt;span class="na"&gt;routes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;to&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;default&lt;/span&gt;
          &lt;span class="na"&gt;via&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;192.168.1.1&lt;/span&gt;
      &lt;span class="na"&gt;nameservers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;addresses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;8.8.8.8&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;1.1.1.1&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
      &lt;span class="na"&gt;interfaces&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;eth0&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
      &lt;span class="na"&gt;parameters&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;stp&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
        &lt;span class="na"&gt;forward-delay&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;0&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Apply the configuration safely:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;netplan try
&lt;span class="nb"&gt;sudo &lt;/span&gt;netplan apply
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Phase 2: Hardening the KVM Hypervisor
&lt;/h2&gt;

&lt;p&gt;CloudStack uses KVM as its primary open-source hypervisor. Installing KVM on Ubuntu 24.04 introduces systemic challenges, specifically regarding &lt;code&gt;libvirtd&lt;/code&gt; socket activation and AppArmor blocking API calls.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;🚨 &lt;strong&gt;CRITICAL SRE ALERT: The Ubuntu 24.04 Libvirt Trap&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Ubuntu 24.04 uses socket-based activation for libvirt. CloudStack requires legacy TCP listening. If you do not mask these sockets and disable AppArmor for libvirt, your Virtual Machines will silently fail to start during deployment.&lt;br&gt;
&lt;/p&gt;
&lt;/blockquote&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Install KVM &amp;amp; CloudStack Agent&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;apt update &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;sudo &lt;/span&gt;apt &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-y&lt;/span&gt; qemu-kvm cloudstack-agent

&lt;span class="c"&gt;# Mask socket listeners to force legacy mode&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl mask libvirtd.socket libvirtd-ro.socket libvirtd-admin.socket libvirtd-tls.socket libvirtd-tcp.socket

&lt;span class="c"&gt;# Ensure disable directory exists and disable AppArmor for Libvirt&lt;/span&gt;
&lt;span class="nb"&gt;sudo mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; /etc/apparmor.d/disable/
&lt;span class="nb"&gt;sudo ln&lt;/span&gt; &lt;span class="nt"&gt;-s&lt;/span&gt; /etc/apparmor.d/usr.sbin.libvirtd /etc/apparmor.d/disable/
&lt;span class="nb"&gt;sudo &lt;/span&gt;apparmor_parser &lt;span class="nt"&gt;-R&lt;/span&gt; /etc/apparmor.d/usr.sbin.libvirtd

&lt;span class="c"&gt;# Configure Libvirt TCP Listening&lt;/span&gt;
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'listen_tls=0'&lt;/span&gt; | &lt;span class="nb"&gt;sudo tee&lt;/span&gt; &lt;span class="nt"&gt;-a&lt;/span&gt; /etc/libvirt/libvirtd.conf
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'listen_tcp=1'&lt;/span&gt; | &lt;span class="nb"&gt;sudo tee&lt;/span&gt; &lt;span class="nt"&gt;-a&lt;/span&gt; /etc/libvirt/libvirtd.conf
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'tcp_port="16509"'&lt;/span&gt; | &lt;span class="nb"&gt;sudo tee&lt;/span&gt; &lt;span class="nt"&gt;-a&lt;/span&gt; /etc/libvirt/libvirtd.conf
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'auth_tcp="none"'&lt;/span&gt; | &lt;span class="nb"&gt;sudo tee&lt;/span&gt; &lt;span class="nt"&gt;-a&lt;/span&gt; /etc/libvirt/libvirtd.conf

&lt;span class="c"&gt;# Restart the hypervisor engine&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl restart libvirtd
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Phase 3: Secure Storage &amp;amp; Database Setup
&lt;/h2&gt;

&lt;p&gt;CloudStack relies on MySQL for state management and NFS for Primary (VM Disks) and Secondary (ISOs/Snapshots) storage.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;🚨 &lt;strong&gt;SECURITY ALERT: Defeating NFS Vulnerabilities&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Never expose an NFS share with full &lt;code&gt;chmod 777&lt;/code&gt; permissions. Restrict your &lt;code&gt;/etc/exports&lt;/code&gt; strictly to your Management VLAN subnet, and use UFW to drop external port 2049 requests.&lt;br&gt;
&lt;/p&gt;
&lt;/blockquote&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 1. Install MySQL &amp;amp; NFS&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;apt &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-y&lt;/span&gt; mysql-server nfs-kernel-server

&lt;span class="c"&gt;# 2. Secure NFS Exports (Replace 192.168.1.0/24 with your subnet)&lt;/span&gt;
&lt;span class="nb"&gt;sudo mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; /export/primary /export/secondary
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"/export/primary 192.168.1.0/24(rw,async,no_root_squash,no_subtree_check)"&lt;/span&gt; | &lt;span class="nb"&gt;sudo tee&lt;/span&gt; &lt;span class="nt"&gt;-a&lt;/span&gt; /etc/exports
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"/export/secondary 192.168.1.0/24(rw,async,no_root_squash,no_subtree_check)"&lt;/span&gt; | &lt;span class="nb"&gt;sudo tee&lt;/span&gt; &lt;span class="nt"&gt;-a&lt;/span&gt; /etc/exports
&lt;span class="nb"&gt;sudo &lt;/span&gt;exportfs &lt;span class="nt"&gt;-ra&lt;/span&gt;

&lt;span class="c"&gt;# 3. Optimize MySQL for CloudStack&lt;/span&gt;
&lt;span class="nb"&gt;cat&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="no"&gt;EOF&lt;/span&gt;&lt;span class="sh"&gt;' | sudo tee /etc/mysql/mysql.conf.d/cloudstack.cnf
[mysqld]
server-id=1
innodb_rollback_on_timeout=1
innodb_lock_wait_timeout=600
max_connections=1000
log-bin=mysql-bin
binlog-format = 'ROW'
&lt;/span&gt;&lt;span class="no"&gt;EOF

&lt;/span&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl restart mysql nfs-kernel-server
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Phase 4: Deploying the Management Server
&lt;/h2&gt;

&lt;p&gt;The Management Server serves the UI, orchestrates KVM hosts, and provisions SystemVMs (Virtual Routers and Console Proxies).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Add the official ShapeBlue CloudStack Repository&lt;/span&gt;
&lt;span class="nb"&gt;sudo mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; /etc/apt/keyrings
wget &lt;span class="nt"&gt;-O-&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;http://packages.shapeblue.com/release.asc]&lt;span class="o"&gt;(&lt;/span&gt;http://packages.shapeblue.com/release.asc&lt;span class="o"&gt;)&lt;/span&gt; | gpg &lt;span class="nt"&gt;--dearmor&lt;/span&gt; | &lt;span class="nb"&gt;sudo tee&lt;/span&gt; /etc/apt/keyrings/cloudstack.gpg &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /dev/null
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"deb [signed-by=/etc/apt/keyrings/cloudstack.gpg] [http://packages.shapeblue.com/cloudstack/upstream/debian/4.20](http://packages.shapeblue.com/cloudstack/upstream/debian/4.20) /"&lt;/span&gt; | &lt;span class="nb"&gt;sudo tee&lt;/span&gt; /etc/apt/sources.list.d/cloudstack.list

&lt;span class="nb"&gt;sudo &lt;/span&gt;apt update
&lt;span class="nb"&gt;sudo &lt;/span&gt;apt &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-y&lt;/span&gt; cloudstack-management bzip2

&lt;span class="c"&gt;# Initialize the Database Schema&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;cloudstack-setup-databases cloud:P@ssw0rd123@localhost &lt;span class="nt"&gt;--deploy-as&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;root:YourMySQLRootPass

&lt;span class="c"&gt;# Launch the Management Server&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;cloudstack-setup-management

&lt;span class="c"&gt;# Seed the KVM SystemVM Template to Secondary Storage&lt;/span&gt;
wget &lt;span class="o"&gt;[&lt;/span&gt;http://download.cloudstack.org/systemvm/4.20/systemvmtemplate-4.20.1-x86_64-kvm.qcow2.bz2]&lt;span class="o"&gt;(&lt;/span&gt;http://download.cloudstack.org/systemvm/4.20/systemvmtemplate-4.20.1-x86_64-kvm.qcow2.bz2&lt;span class="o"&gt;)&lt;/span&gt;
bzip2 &lt;span class="nt"&gt;-d&lt;/span&gt; systemvmtemplate-4.20.1-x86_64-kvm.qcow2.bz2

&lt;span class="nb"&gt;sudo&lt;/span&gt; /usr/share/cloudstack-common/scripts/storage/secondary/cloud-install-sys-tmplt &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;-m&lt;/span&gt; /export/secondary &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;-f&lt;/span&gt; systemvmtemplate-4.20.1-x86_64-kvm.qcow2 &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;-h&lt;/span&gt; kvm &lt;span class="nt"&gt;-F&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Access your Private Cloud Dashboard at &lt;code&gt;http://YOUR_SERVER_IP:8080/client&lt;/code&gt; (Default login: &lt;code&gt;admin&lt;/code&gt; / &lt;code&gt;password&lt;/code&gt;).&lt;/p&gt;




&lt;h2&gt;
  
  
  Phase 5: The VMware Escape Plan
&lt;/h2&gt;

&lt;p&gt;CloudStack includes native &lt;code&gt;virt-v2v&lt;/code&gt; migration integration to move existing virtual machines off VMware ESXi without third-party licenses:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Install virt-v2v and nbdkit on KVM hosts for native VMware migration&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;apt &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-y&lt;/span&gt; virt-v2v nbdkit
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You can map your vCenter credentials inside CloudStack to automatically convert VMDK disk formats to QCOW2 and inject KVM virtio drivers on the fly. Pairing this with &lt;strong&gt;ServerMO Bare Metal Servers&lt;/strong&gt; slashes hypervisor licensing costs to zero.&lt;/p&gt;




&lt;h2&gt;
  
  
  💬 Apache CloudStack &amp;amp; Private Cloud FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why is Apache CloudStack the best VMware alternative post-Broadcom?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
CloudStack is a turnkey IaaS platform providing a VMware-like UI, VPC networking, and native migration tools (&lt;code&gt;virt-v2v&lt;/code&gt;) to move workloads to open-source KVM seamlessly without vendor lock-in.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;CloudStack vs OpenStack vs OpenNebula: Which is better?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
OpenStack requires a dedicated DevOps team to maintain modular complexity. OpenNebula lacks deep enterprise features. Apache CloudStack deploys out-of-the-box like a monolithic AWS clone with low operational overhead.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why does my CloudStack SystemVM stay in the "Starting" state?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
This is typically an NFS export permission or routing issue. Verify that Secondary NFS storage includes &lt;code&gt;no_root_squash&lt;/code&gt; and that your &lt;code&gt;cloudbr0&lt;/code&gt; bridge has internet access.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Read the full guide on ServerMO:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;a href="https://www.servermo.com/howto/deploy-apache-cloudstack-ubuntu-24-04/" rel="noopener noreferrer"&gt;Deploy Apache CloudStack on Ubuntu 24.04: Build a Private AWS | ServerMO&lt;/a&gt;&lt;/p&gt;

</description>
      <category>devops</category>
      <category>ubuntu</category>
      <category>cloud</category>
      <category>kvm</category>
    </item>
  </channel>
</rss>
