<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Shehroz Ali</title>
    <description>The latest articles on DEV Community by Shehroz Ali (@sherredev).</description>
    <link>https://dev.to/sherredev</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4161507%2F0b9dd1d8-12e1-4fd1-a9d1-898bfe24fc32.jpg</url>
      <title>DEV Community: Shehroz Ali</title>
      <link>https://dev.to/sherredev</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sherredev"/>
    <language>en</language>
    <item>
      <title>Achieving 99.99% uptime on single Pod rollout using NEG-based Network Load Balancing</title>
      <dc:creator>Shehroz Ali</dc:creator>
      <pubDate>Tue, 29 Sep 2026 15:03:52 +0000</pubDate>
      <link>https://dev.to/sherredev/achieving-9999-uptime-on-single-pod-rollout-using-neg-based-network-load-balancing-k7d</link>
      <guid>https://dev.to/sherredev/achieving-9999-uptime-on-single-pod-rollout-using-neg-based-network-load-balancing-k7d</guid>
      <description>&lt;p&gt;If you’re using an external passthrough Network Load Balancer sitting in front of your Kuberentes cluster, facts are most of the times its not aware of your Service Pod’s container health which results in users receiving Connection Reset (RST) or invalid response.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvse9lgx0eo2ynk606yjt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvse9lgx0eo2ynk606yjt.png" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This becomes increasingly more critical during peak hours of your traffic or especially during times when RPS is extremely high and you can’t afford to drop any network packet or cause connection reset as it can long-term hurt you business or revenue.&lt;/p&gt;

&lt;p&gt;In our company, we’ve been noticing for a long time that some of our uptime monitoring checks starts failing randomly especially during node upgrades, rollouts, pod restarts or any upgrade/maintainance work going on. At first, we thought this could be a random network blip over the wire, but sooner it became clear that this was happnening specifically whenever a Kubernetes node goes upgrading or workloads migrate to existing/new nodes.&lt;/p&gt;

&lt;p&gt;We started digging down and after a while it got clear to us that this was happening in our single-replica deployment meaning 1 proxy (NGNIX) pod on each node sitting behind 1 global external Network Load Balancer.&lt;/p&gt;

&lt;p&gt;Let me get into straight&lt;/p&gt;

&lt;p&gt;In a traditional NLB setup, the Load Balancer picks any random VM from the list (target-pool based) and forwards the TCP packet to that VM. Once the packet is forwarded to the VM, LB did its job and moves on to forwarding next packet. That’s all. So in short, the LB is not aware of the running Service Pod health, if Pod is Not Ready/Terminating LB will still end up forwarding traffic to that VM.&lt;/p&gt;

&lt;p&gt;That’s what causing connection lost and reset.&lt;/p&gt;

&lt;p&gt;Google Cloud offers another type of more modern Network Load Balancer and its called Regional service-based Network Load Balancer. So instead of maintaing a fixed target-pool of instances, it maintains NEGs (Network Endpoint Groups) which contains the VM IP inside. The interesting thing is the NEG list gets updated by a dedicated NEG-controller running inside GKE control plane which watches Kubernetes EndpointSlice specifically for the app defined in Service selector label for Load Balancer.&lt;/p&gt;

&lt;p&gt;To enable regional external NLB, add this in your manifest and apply (immutable):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Service&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;app&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;proxy&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;tls-proxy&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;externalTrafficPolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Local&lt;/span&gt;
  &lt;span class="na"&gt;loadBalancerClass&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;networking.gke.io/l4-regional-external&lt;/span&gt; &lt;span class="c1"&gt;# add this&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now let me walk you down through a scenario:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Pod goes in Terminating state&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F96ilx2psbiavfvrl4mzs.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F96ilx2psbiavfvrl4mzs.png" width="800" height="345"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;EndpointSlice gets updated marking that Pod Not Ready and Cilium reponds with 503 to mark node as unhealthy&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb52rbif0nyiybekdas3x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb52rbif0nyiybekdas3x.png" width="800" height="392"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;NEG controller watches EndpointSlice and updates the NEGs removing the VM IP of the terminating Pod&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2kkw8t3tpjcn3ppeo4yr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2kkw8t3tpjcn3ppeo4yr.png" width="800" height="367"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;LB send new connections to other nodes running health replica&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3kx4lfji1kfz9qkih033.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3kx4lfji1kfz9qkih033.png" width="800" height="370"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You might will think we’re done here and no more connection reset right? The answer is no!&lt;/p&gt;

&lt;p&gt;You just solved the problem for new connections routing to other nodes running healthy replica but what about existing connections with that VM with terminating Pod inside?&lt;/p&gt;

&lt;p&gt;Now either could be your situation:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Pod goes in Terminating state or shutting down during node upgrade, rollout, restart, etc. (we can fix)&lt;/li&gt;
&lt;li&gt;Pod dies immidiately, crashed, killed, server error (your luck!)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If a Pod is in terminating state and you want to gracefully handle all existing connections without dropping any packets you need to tweak and playaround your proxy settings and LB connection draining timeout.&lt;/p&gt;

&lt;p&gt;In our case since we’re using NGNIX as our proxy layer for TLS and backend routing, we tweaked these settings and observed close to 99.99% uptime without connection loss.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Added sleep 40s in preStop to let NGNIX stay alive and keep serving exisiting connections&lt;/li&gt;
&lt;li&gt;Set Connection Draining duration of regional LB to 30s to allow existing connections to finish before cutting off and removing from NEGs&lt;/li&gt;
&lt;li&gt;Added keepalive_time to 15s (or your choice) to gracefully close existing connections after 15s and let user/client open fresh TCP connection with other VMs running healthy replica&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff2l1v8ynof5wn6ogurnb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff2l1v8ynof5wn6ogurnb.png" width="800" height="374"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;During 40s NGNIX is still alive and serving so readiness probe keeps passing which means any packet entererd in VM until draining Cilium will forward it to NGNIX pod as EndpointSlice is still serving: true for terminating pod. Also we’ve externalTrafficPolicy: Local set which means no internal network hop allowed so the terminating NGNIX Pod stays serving and Cilium uses it as a fallback&lt;/li&gt;
&lt;li&gt;Make sure to increase terminationGracePeriodSeconds to sleep + N seconds so you allow enough time for shutdown before SIGKILL&lt;/li&gt;
&lt;li&gt;All existing connections gracefully closes within 40s sleep duration and no packet lost!&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcyq80iizrpvxj44k61tx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcyq80iizrpvxj44k61tx.png" width="800" height="375"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;New connections gets routed to other nodes running healthy replica of your LoadBalancer Service Pod&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk6fqowofol2jwg1r5o8m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk6fqowofol2jwg1r5o8m.png" width="799" height="381"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Once the new Pod is up and Ready, NEG controller will attach the same VM to NEGs again and new traffic starts coming.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft0fbytjdxbys37j7zyj3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft0fbytjdxbys37j7zyj3.png" width="800" height="374"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>platformengineering</category>
      <category>kubernetes</category>
      <category>googlecloudplatform</category>
    </item>
    <item>
      <title>Why the AI Era Demands a Shift from Product to Platform Engineering</title>
      <dc:creator>Shehroz Ali</dc:creator>
      <pubDate>Sat, 16 May 2026 14:33:01 +0000</pubDate>
      <link>https://dev.to/sherredev/why-the-ai-era-demands-a-shift-from-product-to-platform-engineering-2oj8</link>
      <guid>https://dev.to/sherredev/why-the-ai-era-demands-a-shift-from-product-to-platform-engineering-2oj8</guid>
      <description>&lt;p&gt;Right now, while a lot of software engineering is moving up the abstraction layer (with LLMs writing routine product code), the physical and architectural realities of &lt;em&gt;where&lt;/em&gt; and &lt;em&gt;how&lt;/em&gt; that code runs have never been more critical.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm5hjgzrt6xz6yk5cwfo4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm5hjgzrt6xz6yk5cwfo4.png" width="800" height="437"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  My Backstory — The Product Era (The Abstraction)
&lt;/h3&gt;

&lt;p&gt;I started where most people do — focusing on user value, feature delivery, and rapid iteration. But eventually, the magic wears off when you realize you’re constantly building &lt;em&gt;on top&lt;/em&gt; of systems you don’t fully control or understand. Everything is abstracted under the hood — application frameworks, libraries, methods, databases, memory management, etc. It kind of felt repetitive to me just calling functions/utils/methods to process data for input/output and return it back to customer.&lt;/p&gt;

&lt;p&gt;Especially after AI era, it didn’t feel like I’m actually solving any hardcore engineering problem, I sort of became more like a business/product person rather than calling myself an “engineer”. There was no engineering left, it was just prompting, gathering requirements, doing CRUD, reviewing, testing, merging code and that’s all. And all this because we programmers were already working at the most abstracted part of a software — very far from hardware.&lt;/p&gt;

&lt;p&gt;AI made me realize, the job of a programmer was essentially not the hardest part in the past let’s say 10 years because everything was already solved and abstracted. Our job was to call functions/methods, wait for data to get fetched, perform some math on it and return it back. There was no more thinking left to dig down hardware and question why and how to optimize it better. Everything felt repetitive to me, same APIs, same business requirements, same code, same principles, microservices patterns, I became more like a machine programmed to write code rather than a human to do some critical thinking and being curious.&lt;/p&gt;

&lt;p&gt;And to be honest, this is exactly what AI is so good at — writing application code, but miserable at building large-scale distributed systems that requires combination of human + systems thinking.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Revenge of the Hardware
&lt;/h3&gt;

&lt;p&gt;In the early days of SaaS, the mantra was “hardware is a commodity; code is king.” Cloud abstractions made us forget about bytes, disks, and network packets.&lt;/p&gt;

&lt;p&gt;The twist is AI and massive scale changed that. When you are provisioning clusters, managing GPUs, or optimizing Kubernetes nodes for high-throughput workloads, you suddenly care deeply about memory bandwidth, network latency, and compute density. Every single line of code matters alot.&lt;/p&gt;

&lt;p&gt;Platform engineering demands &lt;strong&gt;system thinking,&lt;/strong&gt; for e.g:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How does a failure in the database pool ripple through the ingress controller?&lt;/li&gt;
&lt;li&gt;How do we scale a cluster dynamically without cascading timeouts of millions of real-time users?&lt;/li&gt;
&lt;li&gt;If we suddenly spin up 20 more instances, can the underlying database handle 20x more concurrent connection pools?&lt;/li&gt;
&lt;li&gt;Will our NAT gateway handle the massive surge in outbound traffic?&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The AI Era
&lt;/h3&gt;

&lt;p&gt;AI boom shattered the “infrastructure is invisible” paradigm.&lt;/p&gt;

&lt;p&gt;Today running large-scale systems and cloud infrastructure today requires a deep understanding of low-level constraints. If you don’t respect the machine, the scale will break you.&lt;/p&gt;

&lt;p&gt;For the last decade, engineering glory was found at the top of the stack. It was about building slick UIs, optimizing user conversion funnels, and shipping features at lightning speed. But the AI era has completely inverted this dynamic.&lt;/p&gt;

&lt;h4&gt;
  
  
  The Death of Routine Product Code (Upstairs)
&lt;/h4&gt;

&lt;p&gt;At the top of the stack, abstractions have become so high that code is becoming heavily commoditized.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;With LLMs, generating a React component, scaffolding a CRUD API, or writing glue code for a product feature is fast, cheap, and increasingly automated.&lt;/li&gt;
&lt;li&gt;The “problems” at the top are becoming less about deep technical complexity and more about product design, prompt engineering, and stitching together pre-existing services. It’s a space where complexity is being managed &lt;em&gt;away&lt;/em&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  The Explosion of Physical Reality (Downstairs)
&lt;/h4&gt;

&lt;p&gt;Meanwhile, at the bottom of the stack, the problems have become intensely complex, fascinating, and deeply tied to physical constraints. You can’t “prompt engineer” your way out of a networking bottleneck, a noisy neighbor saturating the CPU cache, or a memory leak under massive concurrent loads.&lt;/p&gt;

&lt;p&gt;When a product app breaks, you check the logs and fix a null pointer. When a platform breaks at scale, you might be dealing with Linux kernel OOM (Out Of Memory) killers renegading through your pods, network packet drops at the NAT gateway, or storage IOPS throttling. It requires genuine detective work.&lt;/p&gt;

&lt;h3&gt;
  
  
  It’s Now More Fun
&lt;/h3&gt;

&lt;p&gt;AI models, LLMs, and massive data pipelines are the ultimate test of systems thinking and I’m really enjoying it learning, studying, building and practicing around.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How fast can data move from Cloud Storage to GPU memory (VRAM)?&lt;/li&gt;
&lt;li&gt;Are we bottlenecked by PCIe bandwidth?&lt;/li&gt;
&lt;li&gt;How do we orchestrate distributed training across multiple nodes without the network switches becoming a massive choking point?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Training or running inference on large models requires moving terrifying amounts of data. The bottle-neck isn’t how fast the code executes; it’s how fast data can cross the PCIe bus or the network switch.&lt;/p&gt;

&lt;p&gt;How do you dynamically provision Kubernetes nodes with massive GPU attachments right when a burst of heavy compute hits?&lt;/p&gt;

&lt;p&gt;I didn’t move to platform engineering to escape the AI revolution; I moved because that was always my core foundation — love for hardware and code, it’s a beautiful blend of systems, physics and humans.&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>platformengineering</category>
      <category>cloudnative</category>
      <category>claude</category>
    </item>
    <item>
      <title>A Weekend Project: Building a Sidecarless Observability Engine for K8s</title>
      <dc:creator>Shehroz Ali</dc:creator>
      <pubDate>Sun, 19 Apr 2026 12:28:04 +0000</pubDate>
      <link>https://dev.to/sherredev/a-weekend-project-building-a-sidecarless-observability-engine-for-k8s-mb4</link>
      <guid>https://dev.to/sherredev/a-weekend-project-building-a-sidecarless-observability-engine-for-k8s-mb4</guid>
      <description>&lt;p&gt;The best observability is invisible. It shouldn’t require SDKs, it shouldn’t require sidecars, and it definitely shouldn’t eat up half your cluster’s CPU.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why build another tool?
&lt;/h3&gt;

&lt;p&gt;In the world of Kubernetes, observability is often a choice between two extremes: &lt;strong&gt;too little&lt;/strong&gt; or &lt;strong&gt;too much.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;On one hand, you have standard application logs. They tell you what the app &lt;em&gt;thinks&lt;/em&gt; happened, but they don’t tell you what actually happened on the wire. On the other hand, you have massive Service Meshes like Istio or Linkerd. While they give you incredible insights, they come with a heavy “sidecar tax” — injecting a proxy into every pod, adding latency, and making your YAML files look like a phone book.&lt;/p&gt;

&lt;p&gt;I realized that all the information I needed every HTTP header, every JSON body, every latency spike was already floating through the air (or rather, the virtual wires) of my Kubernetes nodes. It was just a matter of reaching out and grabbing it.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Vision for KubeSocket
&lt;/h3&gt;

&lt;p&gt;I wanted to build something lean and simple, something which can just make the job done without any extra overhead or infrastructure configuration. The goal was:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;No extra sidecar container (zero maintenance)&lt;/li&gt;
&lt;li&gt;No SDKs and dependencies&lt;/li&gt;
&lt;li&gt;No dependency on any programming framework (zero configuration)&lt;/li&gt;
&lt;li&gt;Should be platform agnostic&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then, I asked myself; “How about we capture the internet traffic entering in the host machine before it touches anything?”&lt;/p&gt;

&lt;p&gt;This meant that something which can talk directly to host NIC (Network Interface Card), copy raw bytes (0s and 1s), parse ethernet headers and extract HTTP payload. Since HTTP payload starts at Layer 7, it means that we need to capture the entire raw Ethernet packet entering the host machine through NIC mounted on main motherboard before anyone touches it or it gets modified by any application/webserver, etc. Keeping it raw and simple. Pretty interesting.&lt;/p&gt;

&lt;p&gt;To visualise it, this is something I wanted:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx2y30omcj8afdgwa9uwv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx2y30omcj8afdgwa9uwv.png" width="781" height="441"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Ethernet packet flying from wire directly to RAM (OS kernel)&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Network Interface Card (NIC)
&lt;/h3&gt;

&lt;p&gt;A &lt;strong&gt;Network Interface Card (NIC)&lt;/strong&gt;, also known as a network adapter or LAN adapter, is a hardware component that allows a computer or other device to connect to a network and communicate with other devices.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5kfv71enylgladc26miv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5kfv71enylgladc26miv.png" width="800" height="486"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Credits: &lt;a href="https://levens.fr/523572/Card-10-100-1000-Mbps-Low-Profile-LAN-Adapter-For-Desktop-PC" rel="noopener noreferrer"&gt;https://levens.fr/523572/Card-10-100-1000-Mbps-Low-Profile-LAN-Adapter-For-Desktop-PC&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  The Physical Arrival (Layer 1 &amp;amp; 2)
&lt;/h4&gt;

&lt;p&gt;As electrical or optical signals arrive at the NIC, the hardware’s controller performs &lt;strong&gt;Frame Delimitation&lt;/strong&gt;. It identifies where a frame starts and ends, verifies the &lt;strong&gt;Checksum (FCS)&lt;/strong&gt; to ensure the data wasn’t corrupted in flight, and checks the &lt;strong&gt;Destination MAC address&lt;/strong&gt;.&lt;/p&gt;

&lt;h4&gt;
  
  
  The Ring Buffer
&lt;/h4&gt;

&lt;p&gt;The most important concept for a packet-sniffer developer is the &lt;strong&gt;RX (Receive) Ring Buffer&lt;/strong&gt;. Imagine a circular conveyor belt in your RAM.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The NIC&lt;/strong&gt; places a packet on a spot on the belt.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The NIC&lt;/strong&gt; sends a “signal” (an interrupt) to the CPU.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Kernel&lt;/strong&gt; (the Driver) walks over, picks up the packet from that spot, and moves it into the OS networking stack.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Belt&lt;/strong&gt; keeps spinning, ready for the next packet.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbikasnxpxr2gz98t0dcq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbikasnxpxr2gz98t0dcq.png" width="764" height="465"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;How an ethernet packet goes into OS kernel&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Keep Reading!
&lt;/h3&gt;

&lt;p&gt;That’s it, we have now understood how the electrical signals entering the host NIC gets converted into a packet and then how that packet flies through NIC to main system RAM making it accessible for Operating System (OS) kernel to process it for user.&lt;/p&gt;

&lt;p&gt;In next blog, I’ll talk about this raw ethernet packet, what’s inside it and how can we parse it to read the raw HTTP data.&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>cloudnative</category>
      <category>linux</category>
      <category>networking</category>
    </item>
    <item>
      <title>Inside AWS Nitro: The System Design Behind 100 Gbps Performance for EC2</title>
      <dc:creator>Shehroz Ali</dc:creator>
      <pubDate>Thu, 29 Jan 2026 22:34:05 +0000</pubDate>
      <link>https://dev.to/sherredev/inside-aws-nitro-the-system-design-behind-100-gbps-performance-for-ec2-d5f</link>
      <guid>https://dev.to/sherredev/inside-aws-nitro-the-system-design-behind-100-gbps-performance-for-ec2-d5f</guid>
      <description>&lt;p&gt;For years, virtualization came with a hidden cost: the ‘Hypervisor Tax.’ Every time your application sent a network packet, your CPU had to stop what it was doing, context-switch, and play the role of a traffic cop. It was a bottleneck that stood between your code and the raw power of the hardware.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flsfpqbcreas1pyv2nhv1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flsfpqbcreas1pyv2nhv1.png" width="800" height="437"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The Problem: The “Virtualization Tax”
&lt;/h3&gt;

&lt;p&gt;Before Nitro, AWS used a traditional hypervisor (specifically a customised version of &lt;strong&gt;Xen&lt;/strong&gt; ). In this setup, every time a virtual machine (VM) wanted to send a packet of data, it had to go through a “middleman.”&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Dom0 (The Privileged VM):&lt;/strong&gt; The physical server ran a specialized VM called “Domain 0” or &lt;strong&gt;Dom0&lt;/strong&gt;. This was a full-blown Linux environment that had direct access to the physical hardware (NICs, Disks).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context Switching:&lt;/strong&gt; When a customer VM (DomU) sent a network packet, the CPU had to “stop” what the customer was doing and “switch” to the Dom0 code to process that packet. This constant back-and-forth is known as a context switch.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Resource Stealing:&lt;/strong&gt; Dom0 required its own CPU cores and memory to function. On a large server, you might lose &lt;strong&gt;10% to 20%&lt;/strong&gt; of the physical hardware capacity just to run the management software. This is the &lt;strong&gt;Virtualization Tax&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyjxp0vfax38tyx6giwdl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyjxp0vfax38tyx6giwdl.png" width="800" height="453"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Overview of how Hypervisor talks to CPU for processing network packets&lt;/em&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Jitter:&lt;/strong&gt; Because the host CPU is busy “managing” the network for 50 different VMs, latency becomes inconsistent. A packet might be delayed because the CPU was busy processing a storage request for a different customer.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The Solution: AWS Nitro Offloading
&lt;/h3&gt;

&lt;p&gt;The Nitro System effectively “breaks apart” the hypervisor. It takes all the work that Dom0 used to do — networking, storage, and security — and moves it onto a separate piece of hardware: &lt;strong&gt;The Nitro Card.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hardware Separation:&lt;/strong&gt; The Nitro Card is a physical PCIe card with its own processor (ASIC) and memory. It is physically separate from the main motherboard where the customer’s CPU (Intel/AMD/Graviton) sits.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Forctx5il1eutwza48z0k.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Forctx5il1eutwza48z0k.png" width="686" height="379"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Credits: &lt;a href="https://dev.to/choonho/nitro-card-why-aws-is-best-46ph"&gt;https://dev.to/choonho/nitro-card-why-aws-is-best-46ph&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Zero-CPU Networking:&lt;/strong&gt; When a VM sends a packet, it goes directly to the Nitro Card via &lt;strong&gt;SR-IOV&lt;/strong&gt; (Single Root I/O Virtualization). The host CPU never has to “touch” the packet. It simply drops it into a memory queue, and the Nitro Card picks it up.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The Nitro Card for VPC
&lt;/h3&gt;

&lt;p&gt;These are independent System-on-Chips (SoCs) connected via the PCIe bus. They run their own operating systems and are responsible for specific tasks like networking, storage, and management.&lt;/p&gt;

&lt;h4&gt;
  
  
  SR-IOV Implementation
&lt;/h4&gt;

&lt;p&gt;It uses &lt;strong&gt;Single Root I/O Virtualization (SR-IOV)&lt;/strong&gt; to create “Virtual Functions” (VFs). This allows multiple VMs on the same host to have direct, high-speed paths to the hardware without going through a software switch.&lt;/p&gt;

&lt;h4&gt;
  
  
  The Magic: Direct Memory Access (DMA)
&lt;/h4&gt;

&lt;p&gt;There is a physical controller (the &lt;strong&gt;DMA Controller&lt;/strong&gt; ) on the Nitro Card. This circuit has the electrical authority to take control of the PCIe bus and move data from the system’s RAM directly into the card’s own memory.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj4pnwpu34p2k513amnwh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj4pnwpu34p2k513amnwh.png" width="679" height="389"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Hardware diagram for DMA controller inside NIC&lt;/em&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  1. Step One: The OS Kernel Writes to RAM (CPU Action)
&lt;/h4&gt;

&lt;p&gt;When an application inside your VM wants to send a packet (e.g., a “Hello World” message), it doesn’t know about hardware. It just hands the message to the &lt;strong&gt;OS Kernel&lt;/strong&gt;. The CPU takes that message and wraps it in standard network headers (TCP, IP, Ethernet). The CPU writes this completed packet into a specific “buffer” in the &lt;strong&gt;System RAM&lt;/strong&gt; (the slice of memory assigned to your VM).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsvz4ctkuveuh9leeh7of.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsvz4ctkuveuh9leeh7of.png" width="470" height="372"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  2. Step Two: The “Doorbell” (The Hand-off Signal)
&lt;/h4&gt;

&lt;p&gt;Once the packet is sitting in the RAM, the CPU needs to tell the Nitro Card, &lt;em&gt;“Hey, I’ve left a package for you in the lobby.”&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MMIO Write:&lt;/strong&gt; The CPU writes a tiny bit of data (a “doorbell” signal) directly to a specific address that is physically mapped to the &lt;strong&gt;Nitro Card&lt;/strong&gt; over the PCIe bus.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7xfweihki5oimnfv0dze.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7xfweihki5oimnfv0dze.png" width="800" height="442"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  3. Step Three: DMA Hardware Takeover (No CPU)
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;The Fetch:&lt;/strong&gt; The &lt;strong&gt;DMA Controller&lt;/strong&gt; inside the Nitro Card reads the memory address provided by the CPU and reaches across the &lt;strong&gt;PCIe Bus&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Direct Copy:&lt;/strong&gt; It copies the packet data from the &lt;strong&gt;System RAM&lt;/strong&gt; directly into the &lt;strong&gt;Nitro Card’s local memory&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fko9yva0v9b5wo34p4a9v.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fko9yva0v9b5wo34p4a9v.png" width="800" height="367"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The core innovation of the AWS Nitro System is the &lt;strong&gt;physical decoupling&lt;/strong&gt; of the networking data plane from the host CPU. By shifting virtualization tasks to custom hardware, AWS eliminates the “virtualization tax” and provides performance that is nearly indistinguishable from bare metal.&lt;/p&gt;

</description>
      <category>distributedsystems</category>
      <category>linux</category>
      <category>cloudcomputing</category>
      <category>aws</category>
    </item>
    <item>
      <title>Why AI is a Physics Problem: Heat, Size, and the Rise of Nvidia CUDA</title>
      <dc:creator>Shehroz Ali</dc:creator>
      <pubDate>Fri, 02 Jan 2026 15:41:59 +0000</pubDate>
      <link>https://dev.to/sherredev/why-ai-is-a-physics-problem-heat-size-and-the-rise-of-nvidia-cuda-4nfo</link>
      <guid>https://dev.to/sherredev/why-ai-is-a-physics-problem-heat-size-and-the-rise-of-nvidia-cuda-4nfo</guid>
      <description>&lt;p&gt;The CPU optimizes for &lt;strong&gt;Latency&lt;/strong&gt; through complex Branch Prediction and Out-of-Order execution; the GPU optimizes for &lt;strong&gt;Throughput&lt;/strong&gt; by stripping Control Logic to maximize ALU density. This blog explores that silicon-level trade-off and how CUDA bridges the gap.&lt;/p&gt;

&lt;p&gt;To understand the difference between a CPU and a GPU, it helps to think of them not just as “chips,” but as two different types of workers in a factory.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fphvm02ulobxb9i046rt9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fphvm02ulobxb9i046rt9.png" width="800" height="383"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  The Best Analogy: The Chef vs. The Assembly Line
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The CPU (The Master Chef):&lt;/strong&gt; A CPU is like a world-class chef. This chef is incredibly smart and can follow a complex recipe to make anything from a 5-course French dinner to a chocolate souffle. However, the chef is only one person (or a small team). They do things &lt;strong&gt;one at a time&lt;/strong&gt; but with great precision and logic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The GPU (The Burger Flippers):&lt;/strong&gt; A GPU is like a massive line of 1,000 junior cooks who only know how to flip burgers. They aren’t “smart” enough to cook a 5-course meal, but if you need to flip 1,000 burgers at the exact same time, they will finish the job much faster than the master chef ever could.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The Core
&lt;/h3&gt;

&lt;p&gt;A Core is a physical component on a processor which handle and computes instructions. Let’s talk about difference in CPU vs. GPU core and how they handle instructions.&lt;/p&gt;

&lt;h4&gt;
  
  
  CPU core
&lt;/h4&gt;

&lt;p&gt;A CPU core is like a collection of Control Unit, Arithmetic Logic Unit (ALU) and Cache, this combined makes a single core in a CPU.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1 Core = 1 ALU + 1 Control Unit + 1 Cache&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Because each core has its own “Control Unit” (the boss), it can work on a completely different task than the core next to it. So multiple cores in a CPU can work on different set of complex instructions simultaneously at a same time.&lt;/p&gt;

&lt;h4&gt;
  
  
  GPU core
&lt;/h4&gt;

&lt;p&gt;A GPU core doesn’t have its own Control Unit, but rather it only contains a basic ALU with a cache. Hundreds and thousands of cores in a GPU is controlled by 1 Control Unit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1 Core = 1 ALU + 1 Cache&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1000s Cores -&amp;gt; 1 Control Unit&lt;/strong&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  The difference
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;CPU:&lt;/strong&gt; 8 chefs, each with their own recipe book, cooking 8 different meals.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GPU:&lt;/strong&gt; 1 drill sergeant (Control Unit) screaming “JUMP!” at 1,000 soldiers (ALUs) at the same time. The soldiers don’t have their own brains; they just follow the one sergeant.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you gave a GPU a task where every core had to do something different (Core 1 adds, Core 2 subtracts, Core 3 divides), the GPU would break down.&lt;/p&gt;

&lt;p&gt;Because they &lt;strong&gt;share&lt;/strong&gt; a Control Unit, they all have to perform the &lt;strong&gt;same instruction&lt;/strong&gt; at the exact same time.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7kj0zir4xmxb0tc9349e.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7kj0zir4xmxb0tc9349e.png" width="800" height="379"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;CPU core vs. GPU core&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  GPUs are stupid at making decisions, but good at simple math
&lt;/h3&gt;

&lt;p&gt;If you have one giant complex math problem, the CPU will finish it first. If you have 10,000 tiny additions to do (like brightening every pixel in a photo), the GPU will finish the whole batch before the CPU even gets through the first hundred.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqwu4v005ej8oz4qkvz5r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqwu4v005ej8oz4qkvz5r.png" width="507" height="611"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;CPU has less workers but each of them are intelligent&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This is because in a CPU, every core has its own dedicated Control Unit. It’s like 8 people each having their own brain. They can each decide to do something different.&lt;/p&gt;

&lt;p&gt;In a GPU, &lt;strong&gt;one Control Unit&lt;/strong&gt; is shared by a group of, say, 32 or 64 ALUs. This creates a hardware limitation, so they have massive energy to do math but are dumb at making complex decisions due to shortage of Control Unit.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnq4gimq0cjais0dkm8zw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnq4gimq0cjais0dkm8zw.png" width="800" height="606"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;GPU has massive workers but each of them are dumb&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;So GPU has thousands of these tiny cores, but each of them are controlled by 1 Control Unit, and then all cores perform the same instruction on their own in parallel.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why can’t we make GPUs intelligent?
&lt;/h3&gt;

&lt;p&gt;It’s a great question — if the master chef (CPU) is so much better at everything, why not just hire 1,000 of them?&lt;/p&gt;

&lt;p&gt;The answer comes down to three cold, hard physical limits: &lt;strong&gt;Size&lt;/strong&gt; , &lt;strong&gt;Heat&lt;/strong&gt; , &lt;strong&gt;Memory and Cost.&lt;/strong&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  The Size Problem
&lt;/h4&gt;

&lt;p&gt;A “master” CPU core is physically massive compared to a “junior” GPU core.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;CPU Core:&lt;/strong&gt; Packed with complex features like branch prediction (guessing what you’ll do next) and huge “waiting rooms” for data (Cache).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpcfb053gcz3b8qefmjej.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpcfb053gcz3b8qefmjej.png" width="800" height="702"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;100x in physical size due to complex setup&lt;/em&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GPU Core:&lt;/strong&gt; Stripped down to just the math parts.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In the same physical space where you can fit &lt;strong&gt;one&lt;/strong&gt; high-end CPU core, you can often fit over &lt;strong&gt;500&lt;/strong&gt; GPU cores. To fit 1,000 “master” CPU cores, your computer chip would have to be the size of a dinner plate.&lt;/p&gt;

&lt;h4&gt;
  
  
  The Heat Problem
&lt;/h4&gt;

&lt;p&gt;Complex ALUs are “power hungry.” They run at very high speeds (4–5 GHz) and use a lot of electricity to power all that “smart” logic.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If you tried to pack 1,000 of those into a single chip, the amount of heat generated would be high enough to &lt;strong&gt;melt the silicon&lt;/strong&gt; instantly.&lt;/li&gt;
&lt;li&gt;GPUs stay cool(er) because their thousands of cores are much simpler and run at lower speeds (around 1–2 GHz). They are like 1,000 lightbulbs compared to 10 massive industrial spotlights.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fug5409jit8a1dujncczi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fug5409jit8a1dujncczi.png" width="799" height="292"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Visual comparison of ideal GPU vs. real GPU (size)&lt;/em&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  The Memory Problem
&lt;/h4&gt;

&lt;p&gt;Imagine having 1,000 master chefs in one kitchen. They would all be screaming for ingredients at the same time.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A CPU core needs a constant stream of complex instructions.&lt;/li&gt;
&lt;li&gt;Current memory technology (RAM) isn’t fast enough to feed 1,000 “smart” cores simultaneously. They would spend 99% of their time just sitting there waiting for data to arrive, making them a waste of space.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  The Cost Problem
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Electricity:&lt;/strong&gt; A “smart” CPU core uses significantly more power than a “dumb” GPU core because it’s constantly running complex “prediction” circuits.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Heat:&lt;/strong&gt; 1,000 smart cores would require an industrial-grade cooling system (like liquid nitrogen or massive fans).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Infrastructure:&lt;/strong&gt; To run a chip that complex, you would need a specialized motherboard and a power supply as big as a microwave.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Heterogeneous Computing: Nvidia CUDA
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Heterogeneous Computing&lt;/strong&gt; is the “teamwork” approach to computer design. Instead of trying to make one chip that is good at everything, it puts different types of specialized processors together on the same system (or even the same chip) to handle specific tasks.&lt;/p&gt;

&lt;p&gt;The hardest part of heterogeneous computing isn’t the hardware — it’s the &lt;strong&gt;Software&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A CPU and a GPU speak different languages.&lt;/li&gt;
&lt;li&gt;To make them work together, programmers have to write special code (using tools like &lt;strong&gt;OpenCL&lt;/strong&gt; or &lt;strong&gt;CUDA&lt;/strong&gt; ) to tell the computer: &lt;em&gt;“Send this math to the GPU, but keep the logic on the CPU.”&lt;/em&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftad9y9mgnc00tfnkmr13.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftad9y9mgnc00tfnkmr13.png" width="800" height="661"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;High-level architecture of Nvidia CUDA technology&lt;/em&gt;&lt;/p&gt;

</description>
      <category>gpuprogramming</category>
      <category>cuda</category>
      <category>artificialintelligen</category>
      <category>nvidia</category>
    </item>
    <item>
      <title>One OS Thread, Millions of Goroutines: The Magic of Go’s Scheduler Explained</title>
      <dc:creator>Shehroz Ali</dc:creator>
      <pubDate>Sat, 29 Nov 2025 13:11:45 +0000</pubDate>
      <link>https://dev.to/sherredev/one-os-thread-millions-of-goroutines-the-magic-of-gos-scheduler-explained-5gpm</link>
      <guid>https://dev.to/sherredev/one-os-thread-millions-of-goroutines-the-magic-of-gos-scheduler-explained-5gpm</guid>
      <description>&lt;p&gt;Go achieves high concurrency through its scheduler, which manages millions of goroutines on a small number of OS threads. Here’s how it works.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuc8fmfqou31cfza7lcqr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuc8fmfqou31cfza7lcqr.png" width="800" height="534"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Overview of Go’s scheduler model&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Go’s concurrency model lets you run millions of goroutines with just a few OS threads. This post explains how the scheduler multiplexes goroutines onto threads, how it handles blocking and non-blocking I/O, and how the G-M-P model enables efficient concurrency.&lt;/p&gt;

&lt;p&gt;Go uses goroutines, not threads. Important differences:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Threads (traditional):&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Managed by OS&lt;/li&gt;
&lt;li&gt;Heavy (~1–2MB stack each)&lt;/li&gt;
&lt;li&gt;Limited (typically 100s-1000s)&lt;/li&gt;
&lt;li&gt;Context switching is expensive (OS kernel involved)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Goroutines:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Managed by Go runtime&lt;/li&gt;
&lt;li&gt;Lightweight (~2KB stack initially)&lt;/li&gt;
&lt;li&gt;Can spawn millions&lt;/li&gt;
&lt;li&gt;Context switching is cheap (user-space)&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;
  
  
  GOMAXPROCS
&lt;/h3&gt;

&lt;p&gt;OS threads are limited by GOMAXPROCS.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="c"&gt;// Default: GOMAXPROCS = number of CPU cores&lt;/span&gt;
&lt;span class="c"&gt;// You can set it:&lt;/span&gt;
&lt;span class="n"&gt;runtime&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;GOMAXPROCS&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="m"&gt;4&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c"&gt;// Use 4 OS threads&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Example on an 8-core machine:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;8 OS threads&lt;/li&gt;
&lt;li&gt;Can handle thousands of concurrent goroutines&lt;/li&gt;
&lt;li&gt;Go scheduler multiplexes goroutines onto threads&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  How the OS thread handles multiple goroutines
&lt;/h3&gt;

&lt;p&gt;The OS thread doesn’t run all goroutines &lt;em&gt;(let’s say Gn in the diagram below)&lt;/em&gt; at the same time. The Go scheduler switches between them.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpmrjzu3a9bisih6vjb68.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpmrjzu3a9bisih6vjb68.png" width="800" height="522"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Overview of Go concurrency model&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Let’s zoom-in here and try to get a full clear picture how the Go runtime scheduler actually juggles between goroutines on a single OS thread.&lt;/p&gt;

&lt;p&gt;There can be two types of operation being performed by goroutine (Gn):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Non-blocking (HTTP/network)&lt;/li&gt;
&lt;li&gt;Blocking (disk I/O)&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;
  
  
  Non-blocking Task
&lt;/h4&gt;

&lt;p&gt;Let’s say the goroutine (G1) waiting to be scheduled needs to perform a HTTP call over the internet, here’s how the scheduler will work:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The goroutine (G1) initiates the HTTP call&lt;/li&gt;
&lt;li&gt;It yields control to the scheduler&lt;/li&gt;
&lt;li&gt;The scheduler switches to another goroutine (G2)&lt;/li&gt;
&lt;li&gt;The OS thread continues with other work&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp1wqqmuar0nielnm2n0o.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp1wqqmuar0nielnm2n0o.png" width="800" height="437"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Visual Timeline:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fclhvhcu6mjzgz9isnizt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fclhvhcu6mjzgz9isnizt.png" width="800" height="353"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code-level view:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="c"&gt;// G1's code&lt;/span&gt;
&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="n"&gt;handler1&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c"&gt;// G1 is running on OS Thread 1&lt;/span&gt;
    &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;http&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"https://api1.com"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="c"&gt;// ↑ At this point:&lt;/span&gt;
    &lt;span class="c"&gt;// 1. Creates socket&lt;/span&gt;
    &lt;span class="c"&gt;// 2. Sends request&lt;/span&gt;
    &lt;span class="c"&gt;// 3. Registers with poller&lt;/span&gt;
    &lt;span class="c"&gt;// 4. Yields to scheduler (gopark)&lt;/span&gt;
    &lt;span class="c"&gt;// 5. G1 marked as "waiting"&lt;/span&gt;
    &lt;span class="c"&gt;// 6. Scheduler switches to G2&lt;/span&gt;

    &lt;span class="c"&gt;// Later, when response arrives:&lt;/span&gt;
    &lt;span class="c"&gt;// 1. Network poller detects it&lt;/span&gt;
    &lt;span class="c"&gt;// 2. Scheduler marks G1 as "runnable"&lt;/span&gt;
    &lt;span class="c"&gt;// 3. Scheduler switches back to G1&lt;/span&gt;
    &lt;span class="c"&gt;// 4. G1 continues here with response&lt;/span&gt;

    &lt;span class="n"&gt;fmt&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Println&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StatusCode&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c"&gt;// G2's code&lt;/span&gt;
&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="n"&gt;handler2&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c"&gt;// G2 runs when G1 yields&lt;/span&gt;
    &lt;span class="c"&gt;// Could do anything: CPU work, another HTTP call, etc.&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c"&gt;// G3's code&lt;/span&gt;
&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="n"&gt;handler3&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c"&gt;// G3 runs when G2 yields or completes&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Blocking Task
&lt;/h4&gt;

&lt;p&gt;Let’s say the goroutine (G2) waiting to be scheduled needs to perform a disk I/O like a write operation, here’s how the scheduler will work:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The goroutine (G2) writes some heavy data to disk (say in GBs)&lt;/li&gt;
&lt;li&gt;The thread is blocked in the kernel&lt;/li&gt;
&lt;li&gt;Scheduler cannot preempt it (it’s in kernel)&lt;/li&gt;
&lt;li&gt;Go creates a new OS thread for other goroutines&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbbpta3c29o9mkgkymgo8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbbpta3c29o9mkgkymgo8.png" width="800" height="421"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Visual Timeline:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv3qhhwr9z93se2i2bof5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv3qhhwr9z93se2i2bof5.png" width="800" height="419"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code-level view:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="c"&gt;// When you call os.Open()&lt;/span&gt;
&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="n"&gt;Open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;File&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c"&gt;// 1. Enter syscall&lt;/span&gt;
    &lt;span class="n"&gt;runtime&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;entersyscall&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;defer&lt;/span&gt; &lt;span class="n"&gt;runtime&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;exitsyscall&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="c"&gt;// 2. Make actual syscall&lt;/span&gt;
    &lt;span class="n"&gt;fd&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;syscall&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;...&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c"&gt;// 3. If blocking detected, scheduler&lt;/span&gt;
    &lt;span class="c"&gt;// already created new thread&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;newFile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;fd&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  The three components: G, M, P
&lt;/h3&gt;

&lt;p&gt;Go’s scheduler uses three main components:&lt;/p&gt;

&lt;p&gt;G = Goroutine (your code)&lt;/p&gt;

&lt;p&gt;M = Machine (OS thread)&lt;/p&gt;

&lt;p&gt;P = Processor (execution context)&lt;/p&gt;

&lt;h4&gt;
  
  
  What is a Processor (P)?
&lt;/h4&gt;

&lt;p&gt;A Processor (P) is an execution context that:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Holds a local run queue of goroutines&lt;/li&gt;
&lt;li&gt;Binds to an OS thread (M) to execute goroutines&lt;/li&gt;
&lt;li&gt;Manages resources for running goroutines&lt;/li&gt;
&lt;li&gt;Acts as a bridge between goroutines and threads&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2nc6srs89cnd7d72mtjz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2nc6srs89cnd7d72mtjz.png" width="800" height="784"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Overview of Processor (P) in Go&lt;/em&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  How Processor (P) manages OS threads?
&lt;/h4&gt;

&lt;p&gt;Consider two Goroutines; G1 and G2, both waiting in queue of P1 (Processor). P1 is binded to OS thread 1 which is running on Core 1. Now:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;G1 wants to performs a network call (non-blocking)&lt;/li&gt;
&lt;li&gt;G2 needs to open a file from disk (blocking)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Here’s how Go scheduler will manage:&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Runs G1 from P1 queue on OS thread 1&lt;/li&gt;
&lt;li&gt;Goes to Network poller (epoll/kqueue), G1 waiting, OS thread 1 is free&lt;/li&gt;
&lt;li&gt;Scheduler runs G2 from P1 queue on OS thread 1&lt;/li&gt;
&lt;li&gt;G2 performs syscall() and is blocked in kernal&lt;/li&gt;
&lt;li&gt;Scheduler detects OS thread 1 is blocked&lt;/li&gt;
&lt;li&gt;Scheduler creates new OS thread, call runtime.newosproc()&lt;/li&gt;
&lt;li&gt;Scheduler detaches P1 from OS thread 1 and binds to OS thread 2&lt;/li&gt;
&lt;li&gt;Now Go Scheduler runs G3…Gn from P1 on OS thread 2&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftobb30glqk7cex5fk0u8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftobb30glqk7cex5fk0u8.png" width="800" height="641"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example on a logical 4-core machine:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;No of Processor = GOMAXPROCS = No of logical CPU cores.&lt;/p&gt;

&lt;p&gt;So, on 4 core machine, it will have P1…P4, where each P is binded to OS thread which is equal to GOMAXPROCS = 4 (CPU cores).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcddjptlkmn6uhv3vhsf2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcddjptlkmn6uhv3vhsf2.png" width="800" height="213"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;P = Processsor, M = OS thread, G = Goroutine&lt;/em&gt;&lt;/p&gt;

</description>
      <category>multithreading</category>
      <category>softwaredevelopment</category>
      <category>operatingsystems</category>
      <category>linux</category>
    </item>
    <item>
      <title>A Complete Guide to Distributed Tracing in Kotlin and Spring Boot with OpenTelemetry and Grafana…</title>
      <dc:creator>Shehroz Ali</dc:creator>
      <pubDate>Thu, 23 Oct 2025 22:49:15 +0000</pubDate>
      <link>https://dev.to/sherredev/a-complete-guide-to-distributed-tracing-in-kotlin-and-spring-boot-with-opentelemetry-and-grafana-4gh0</link>
      <guid>https://dev.to/sherredev/a-complete-guide-to-distributed-tracing-in-kotlin-and-spring-boot-with-opentelemetry-and-grafana-4gh0</guid>
      <description>&lt;h3&gt;
  
  
  A Complete Guide to Distributed Tracing in Kotlin and Spring Boot with OpenTelemetry and Grafana Loki
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F701aqs0amm7rkojv8vle.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F701aqs0amm7rkojv8vle.png" width="800" height="534"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In this blog, I’ll teach you how you can build end-to-end distributed tracing for your backend microservices in &lt;strong&gt;Kotlin, OpenTelemetry, Spring Boot and Grafana Loki&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;By following the guide, you’ll be able to correlate hundreds and thousands of telemetry data (traces &amp;amp; logs) emitted by your backend services and achieve end-to-end observability to reduce your MTTRs for production incidents and make your developers life easier.&lt;/p&gt;

&lt;p&gt;To give you a basic understanding of what we’ll be going to achieve, here’s a system-level diagram to help you understand all the components functioning together:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx9pqbz9whe5yg3p88qz2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx9pqbz9whe5yg3p88qz2.png" width="800" height="410"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;end-to-end observability&lt;/em&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Capturing HTTP traffic within Spring Boot:
&lt;/h4&gt;

&lt;p&gt;First, we need to enable HTTP network logging for our Spring Boot app, for that we’re going to use &lt;a href="https://github.com/zalando/logbook" rel="noopener noreferrer"&gt;Zalando’s open-source &lt;strong&gt;Logbook&lt;/strong&gt;&lt;/a&gt; which automatically logs HTTP requests and responses on your (preferred output writer). It works by intercepting HTTP traffic in your app and writing TRACE log using &lt;strong&gt;Slf4j&lt;/strong&gt; which is a fascade for Logging in Java. In this example, we’re going to use &lt;strong&gt;Logback&lt;/strong&gt; which is a logging library with auto-configuration support that comes already with spring-starter-web dependency.&lt;/p&gt;

&lt;p&gt;First let’s configure Logback to capture INFO, ERROR and TRACE logs and outputs on console:&lt;/p&gt;
&lt;h4&gt;
  
  
  Setting up Logback appenders:
&lt;/h4&gt;

&lt;p&gt;First we added, Console-Info  to capture logs at INFO level which will likely be our classes/service logs emitting some business/user insights/actions (which later devs will use to understand code behaviour during debugging).&lt;/p&gt;

&lt;p&gt;We also added another Console-Error  to capture ERROR logs emitted by our classes to indicate an exception/error message in logs.&lt;/p&gt;

&lt;p&gt;Lastly we added Console-Trace  to capture TRACE logs. Since Zalando’s Logbook library write TRACE logs for HTTP traffic, we’ve passed the Console-Trace appender to logger org.zalando to use Console-Trace appender to write TRACE logs on System output. Here’s the final logback-logging.xml file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight xml"&gt;&lt;code&gt;&lt;span class="cp"&gt;&amp;lt;?xml version="1.0" encoding="UTF-8"?&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;included&amp;gt;&lt;/span&gt;

    &lt;span class="nt"&gt;&amp;lt;appender&lt;/span&gt; &lt;span class="na"&gt;name=&lt;/span&gt;&lt;span class="s"&gt;"Console-Info"&lt;/span&gt; &lt;span class="na"&gt;class=&lt;/span&gt;&lt;span class="s"&gt;"ch.qos.logback.core.ConsoleAppender"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
        &lt;span class="c"&gt;&amp;lt;!-- Targeting System.out --&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;target&amp;gt;&lt;/span&gt;System.out&lt;span class="nt"&gt;&amp;lt;/target&amp;gt;&lt;/span&gt;
        &lt;span class="c"&gt;&amp;lt;!-- JSON Log Formatting --&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;filter&lt;/span&gt; &lt;span class="na"&gt;class=&lt;/span&gt;&lt;span class="s"&gt;"ch.qos.logback.classic.filter.LevelFilter"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;level&amp;gt;&lt;/span&gt;INFO&lt;span class="nt"&gt;&amp;lt;/level&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;onMatch&amp;gt;&lt;/span&gt;ACCEPT&lt;span class="nt"&gt;&amp;lt;/onMatch&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;onMismatch&amp;gt;&lt;/span&gt;DENY&lt;span class="nt"&gt;&amp;lt;/onMismatch&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;/filter&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;encoder&lt;/span&gt; &lt;span class="na"&gt;class=&lt;/span&gt;&lt;span class="s"&gt;"net.logstash.logback.encoder.LogstashEncoder"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;fieldNames&amp;gt;&lt;/span&gt;
                &lt;span class="nt"&gt;&amp;lt;message&amp;gt;&lt;/span&gt;[ignore]&lt;span class="nt"&gt;&amp;lt;/message&amp;gt;&lt;/span&gt;
                &lt;span class="nt"&gt;&amp;lt;levelValue&amp;gt;&lt;/span&gt;[ignore]&lt;span class="nt"&gt;&amp;lt;/levelValue&amp;gt;&lt;/span&gt;
                &lt;span class="nt"&gt;&amp;lt;version&amp;gt;&lt;/span&gt;[ignore]&lt;span class="nt"&gt;&amp;lt;/version&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;/fieldNames&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;provider&lt;/span&gt; &lt;span class="na"&gt;class=&lt;/span&gt;&lt;span class="s"&gt;"net.logstash.logback.composite.loggingevent.LoggingEventPatternJsonProvider"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
                &lt;span class="nt"&gt;&amp;lt;pattern&amp;gt;&lt;/span&gt;
                    {
                    "message": "%replace(%.-150message){'${MASKPATTERNS}', ' *****'}"
                    }
                &lt;span class="nt"&gt;&amp;lt;/pattern&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;/provider&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;customFields&amp;gt;&lt;/span&gt;
                {"service": "${logHost}"}
            &lt;span class="nt"&gt;&amp;lt;/customFields&amp;gt;&lt;/span&gt; &lt;span class="c"&gt;&amp;lt;!-- Optional: Add custom fields --&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;/encoder&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;/appender&amp;gt;&lt;/span&gt;

    &lt;span class="nt"&gt;&amp;lt;appender&lt;/span&gt; &lt;span class="na"&gt;name=&lt;/span&gt;&lt;span class="s"&gt;"Console-Error"&lt;/span&gt; &lt;span class="na"&gt;class=&lt;/span&gt;&lt;span class="s"&gt;"ch.qos.logback.core.ConsoleAppender"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
        &lt;span class="c"&gt;&amp;lt;!-- Targeting System.out --&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;target&amp;gt;&lt;/span&gt;System.err&lt;span class="nt"&gt;&amp;lt;/target&amp;gt;&lt;/span&gt;
        &lt;span class="c"&gt;&amp;lt;!-- JSON Log Formatting --&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;filter&lt;/span&gt; &lt;span class="na"&gt;class=&lt;/span&gt;&lt;span class="s"&gt;"ch.qos.logback.classic.filter.LevelFilter"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;level&amp;gt;&lt;/span&gt;ERROR&lt;span class="nt"&gt;&amp;lt;/level&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;onMatch&amp;gt;&lt;/span&gt;ACCEPT&lt;span class="nt"&gt;&amp;lt;/onMatch&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;onMismatch&amp;gt;&lt;/span&gt;DENY&lt;span class="nt"&gt;&amp;lt;/onMismatch&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;/filter&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;encoder&lt;/span&gt; &lt;span class="na"&gt;class=&lt;/span&gt;&lt;span class="s"&gt;"net.logstash.logback.encoder.LogstashEncoder"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;fieldNames&amp;gt;&lt;/span&gt;
                &lt;span class="nt"&gt;&amp;lt;message&amp;gt;&lt;/span&gt;[ignore]&lt;span class="nt"&gt;&amp;lt;/message&amp;gt;&lt;/span&gt;
                &lt;span class="nt"&gt;&amp;lt;levelValue&amp;gt;&lt;/span&gt;[ignore]&lt;span class="nt"&gt;&amp;lt;/levelValue&amp;gt;&lt;/span&gt;
                &lt;span class="nt"&gt;&amp;lt;version&amp;gt;&lt;/span&gt;[ignore]&lt;span class="nt"&gt;&amp;lt;/version&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;/fieldNames&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;provider&lt;/span&gt; &lt;span class="na"&gt;class=&lt;/span&gt;&lt;span class="s"&gt;"net.logstash.logback.composite.loggingevent.LoggingEventPatternJsonProvider"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
                &lt;span class="nt"&gt;&amp;lt;pattern&amp;gt;&lt;/span&gt;
                    {
                    "message": "%replace(%.-150message){'${MASKPATTERNS}', ' *****'}",
                    "stack_trace": "%replace(%ex{short}){'${MASKPATTERNS}', ' *****'}"
                    }
                &lt;span class="nt"&gt;&amp;lt;/pattern&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;/provider&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;customFields&amp;gt;&lt;/span&gt;
                {"service": "${logHost}"}
            &lt;span class="nt"&gt;&amp;lt;/customFields&amp;gt;&lt;/span&gt; &lt;span class="c"&gt;&amp;lt;!-- Optional: Add custom fields --&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;/encoder&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;/appender&amp;gt;&lt;/span&gt;

    &lt;span class="nt"&gt;&amp;lt;appender&lt;/span&gt; &lt;span class="na"&gt;name=&lt;/span&gt;&lt;span class="s"&gt;"Console-Trace"&lt;/span&gt; &lt;span class="na"&gt;class=&lt;/span&gt;&lt;span class="s"&gt;"ch.qos.logback.core.ConsoleAppender"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
        &lt;span class="c"&gt;&amp;lt;!-- Targeting System.out --&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;target&amp;gt;&lt;/span&gt;System.out&lt;span class="nt"&gt;&amp;lt;/target&amp;gt;&lt;/span&gt;
        &lt;span class="c"&gt;&amp;lt;!-- JSON Log Formatting --&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;filter&lt;/span&gt; &lt;span class="na"&gt;class=&lt;/span&gt;&lt;span class="s"&gt;"ch.qos.logback.classic.filter.LevelFilter"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;level&amp;gt;&lt;/span&gt;TRACE&lt;span class="nt"&gt;&amp;lt;/level&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;onMatch&amp;gt;&lt;/span&gt;ACCEPT&lt;span class="nt"&gt;&amp;lt;/onMatch&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;onMismatch&amp;gt;&lt;/span&gt;DENY&lt;span class="nt"&gt;&amp;lt;/onMismatch&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;/filter&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;encoder&lt;/span&gt; &lt;span class="na"&gt;class=&lt;/span&gt;&lt;span class="s"&gt;"net.logstash.logback.encoder.LogstashEncoder"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;fieldNames&amp;gt;&lt;/span&gt;
                &lt;span class="nt"&gt;&amp;lt;message&amp;gt;&lt;/span&gt;[ignore]&lt;span class="nt"&gt;&amp;lt;/message&amp;gt;&lt;/span&gt;
                &lt;span class="nt"&gt;&amp;lt;levelValue&amp;gt;&lt;/span&gt;[ignore]&lt;span class="nt"&gt;&amp;lt;/levelValue&amp;gt;&lt;/span&gt;
                &lt;span class="nt"&gt;&amp;lt;version&amp;gt;&lt;/span&gt;[ignore]&lt;span class="nt"&gt;&amp;lt;/version&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;/fieldNames&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;provider&lt;/span&gt; &lt;span class="na"&gt;class=&lt;/span&gt;&lt;span class="s"&gt;"net.logstash.logback.composite.loggingevent.LoggingEventPatternJsonProvider"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
                &lt;span class="nt"&gt;&amp;lt;pattern&amp;gt;&lt;/span&gt;
                    {
                    "message": "#asJson{%message}"
                    }
                &lt;span class="nt"&gt;&amp;lt;/pattern&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;/provider&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;customFields&amp;gt;&lt;/span&gt;
                {"service": "${logHost}", "pod": "${podName}"}
            &lt;span class="nt"&gt;&amp;lt;/customFields&amp;gt;&lt;/span&gt; &lt;span class="c"&gt;&amp;lt;!-- Optional: Add custom fields --&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;/encoder&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;/appender&amp;gt;&lt;/span&gt;

    &lt;span class="c"&gt;&amp;lt;!-- LOG everything at INFO level --&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;root&lt;/span&gt; &lt;span class="na"&gt;level=&lt;/span&gt;&lt;span class="s"&gt;"info"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;appender-ref&lt;/span&gt; &lt;span class="na"&gt;ref=&lt;/span&gt;&lt;span class="s"&gt;"Console-Info"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;/root&amp;gt;&lt;/span&gt;

    &lt;span class="nt"&gt;&amp;lt;logger&lt;/span&gt; &lt;span class="na"&gt;name=&lt;/span&gt;&lt;span class="s"&gt;"org.zalando.logbook"&lt;/span&gt; &lt;span class="na"&gt;level=&lt;/span&gt;&lt;span class="s"&gt;"INFO"&lt;/span&gt; &lt;span class="na"&gt;additivity=&lt;/span&gt;&lt;span class="s"&gt;"false"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;appender-ref&lt;/span&gt; &lt;span class="na"&gt;ref=&lt;/span&gt;&lt;span class="s"&gt;"Console-Trace"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;/logger&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In application.yml file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;logging&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;config&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;classpath:logback-logging.xml'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Structured JSON Logging for Grafana Loki:
&lt;/h4&gt;

&lt;p&gt;Here comes the important detail, since we’re going to structure logs on Grafana Loki based on labels and fields for filtering and querying, we’re going to emit logs as JSON output. For that we need to add this dependency for Logback &lt;a href="http://mvnrepository.com/artifact/net.logstash.logback/logstash-logback-encoder" rel="noopener noreferrer"&gt;(maven central)&lt;/a&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight xml"&gt;&lt;code&gt;        &lt;span class="nt"&gt;&amp;lt;dependency&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;groupId&amp;gt;&lt;/span&gt;net.logstash.logback&lt;span class="nt"&gt;&amp;lt;/groupId&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;artifactId&amp;gt;&lt;/span&gt;logstash-logback-encoder&lt;span class="nt"&gt;&amp;lt;/artifactId&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;version&amp;gt;&lt;/span&gt;6.6&lt;span class="nt"&gt;&amp;lt;/version&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;/dependency&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This will allow us to output logs as JSON on system output and later we can relabel fields on Grafana Agent (or Alloy) for Loki.&lt;/p&gt;

&lt;h4&gt;
  
  
  Writing HTTP logs as JSON:
&lt;/h4&gt;

&lt;p&gt;Now since our Logback supports writing logs in JSON format, we’ve added this line in our Console-Trace  to write Zalando’s HTTP log as JSON on output. The %message captures the entire TRACE level log and output on console with label message like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
 &lt;/span&gt;&lt;span class="err"&gt;message:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;&amp;lt;zalando's&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;TRACE&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;log&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;including&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;response&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;headers&amp;gt;&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Later we can extract out individual fields like &lt;strong&gt;status, type, body, path&lt;/strong&gt; and display on Grafana Loki as fields for filtering and querying logs.&lt;/p&gt;

&lt;h4&gt;
  
  
  Configuring Logbook in Kotlin:
&lt;/h4&gt;

&lt;p&gt;Now, since our Logback is all ready to emit logs (INFO, ERROR, TRACE) on System output, now it’s finally time to configure Logbook for capturing HTTP traffic inside our app.&lt;/p&gt;

&lt;p&gt;Below is the code for configuring Logbook in Kotlin for Spring Boot, add dependency as well:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight xml"&gt;&lt;code&gt;        &lt;span class="nt"&gt;&amp;lt;dependency&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;groupId&amp;gt;&lt;/span&gt;org.zalando&lt;span class="nt"&gt;&amp;lt;/groupId&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;artifactId&amp;gt;&lt;/span&gt;logbook-spring-boot-starter&lt;span class="nt"&gt;&amp;lt;/artifactId&amp;gt;&lt;/span&gt;
            &lt;span class="nt"&gt;&amp;lt;version&amp;gt;&lt;/span&gt;3.10.0&lt;span class="nt"&gt;&amp;lt;/version&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;/dependency&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;As it is a spring-starter dependency it already comes with lots of things pre-configured, let’s tweak it according to our needs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight kotlin"&gt;&lt;code&gt;&lt;span class="nd"&gt;@Configuration&lt;/span&gt;
&lt;span class="k"&gt;open&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;LogbookAutoConfiguration&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;jsonBodyFilter&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;JSONBodyFilter&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;sinkConfig&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;SinkConfiguration&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;

    &lt;span class="nd"&gt;@Bean&lt;/span&gt;
    &lt;span class="k"&gt;open&lt;/span&gt; &lt;span class="k"&gt;fun&lt;/span&gt; &lt;span class="nf"&gt;logbook&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;filterConfig&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;MDCFilter&lt;/span&gt;&lt;span class="p"&gt;?,&lt;/span&gt; &lt;span class="n"&gt;objectMapper&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;ObjectMapper&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nc"&gt;Logbook&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;Logbook&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;builder&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
            &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;correlationId&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;filterConfig&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sink&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sinkConfig&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;defaultSink&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;objectMapper&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
            &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;bodyFilter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;BodyFilter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;none&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
            &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;condition&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;exclude&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;requestTo&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/actuator/**"&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt;
            &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;build&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We’ve excluded /actuator to log because this will be our Kubernetes liveness/readiness probe endpoint for health check, so k8s will hit this endpoint every after X mins/secs to keep pod up and healthy, therefore we excluded to avoid de-bloating our Loki storage.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight kotlin"&gt;&lt;code&gt;&lt;span class="nd"&gt;@Component&lt;/span&gt;
&lt;span class="k"&gt;open&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;JSONBodyFilter&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;

    &lt;span class="k"&gt;open&lt;/span&gt; &lt;span class="k"&gt;fun&lt;/span&gt; &lt;span class="nf"&gt;runFilter&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt; &lt;span class="nc"&gt;BodyFilter&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;JsonBodyFilters&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replaceJsonStringProperty&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="nf"&gt;setOf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"token"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"refreshToken"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"scopes"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
            &lt;span class="s"&gt;" ****"&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here, we’ve masked these fields in our response/request bodies as it contains JWT tokens of users as a security practice, you might don’t need it depending upon your infra/team security policies.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight kotlin"&gt;&lt;code&gt;&lt;span class="nd"&gt;@Component&lt;/span&gt;
&lt;span class="k"&gt;open&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;SinkConfiguration&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;

    &lt;span class="k"&gt;open&lt;/span&gt; &lt;span class="k"&gt;fun&lt;/span&gt; &lt;span class="nf"&gt;defaultSink&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;objectMapper&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;ObjectMapper&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nc"&gt;Sink&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;formatter&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;HttpLogFormatter&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;JsonHttpLogFormatter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;objectMapper&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;writer&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;HttpLogWriter&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;DefaultHttpLogWriter&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;DefaultSink&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;formatter&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;writer&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here we’re using Zalando’s default HTTP log writer to write TRACE logs on output.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight kotlin"&gt;&lt;code&gt;&lt;span class="nd"&gt;@Component&lt;/span&gt;
&lt;span class="k"&gt;open&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;MDCFilter&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;CorrelationId&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;override&lt;/span&gt; &lt;span class="k"&gt;fun&lt;/span&gt; &lt;span class="nf"&gt;generate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;HttpRequest&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;correlationId&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;UUID&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;randomUUID&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;toString&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

        &lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;httpMethod&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;method&lt;/span&gt;
        &lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;httpPath&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;
        &lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;userId&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"userId"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;?.&lt;/span&gt;&lt;span class="nf"&gt;first&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;clientKey&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"X-Client-Key"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;?.&lt;/span&gt;&lt;span class="nf"&gt;first&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;ipAddress&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"x-forwarded-for"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;?.&lt;/span&gt;&lt;span class="nf"&gt;toString&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;userAgent&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"user-agent"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;?.&lt;/span&gt;&lt;span class="nf"&gt;toString&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

        &lt;span class="nc"&gt;MDC&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;put&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"httpMethod"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;httpMethod&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nc"&gt;MDC&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;put&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"httpPath"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;httpPath&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nc"&gt;MDC&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;put&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"userId"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nc"&gt;MDC&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;put&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"correlationId"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;correlationId&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nc"&gt;MDC&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;put&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"ClientKey"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;clientKey&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nc"&gt;MDC&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;put&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"ipAddress"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ipAddress&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nc"&gt;MDC&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;put&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"userAgent"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;userAgent&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;correlationId&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is an important component, here we’ve extended CorrelationId class by Zalando and override generate() method, so whenever internally this method invokes for correlationId generation which will be attached to each HTTP log, we’re also adding the same correlationId in our MDC along with some other fields as well. This means whenever a HTTP request comes inn, here’s what will happen:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Generate unique UUID as correlationId&lt;/li&gt;
&lt;li&gt;Attach correlationId in HTTP request/response TRACE log&lt;/li&gt;
&lt;li&gt;Also populate same correlationId in MDC&lt;/li&gt;
&lt;li&gt;Whenever we log (info, error) at class-level, same correlationId is attached because MDC is shared, allowing us to correlation HTTP logs with class-level logs&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;And, that’s it! We’re good to go to start writing HTTP logs on console/output propagating HTTP path, ip address, method and your own unique header fields correlating with your class-level logs.&lt;/p&gt;

&lt;h4&gt;
  
  
  Using Grafana Agent to ingest Kubernetes Pod logs:
&lt;/h4&gt;

&lt;p&gt;Since, we’re running our Spring Boot app in Kubernetes and writing logs on stdout (system output), k8s maintain a file on the node with all logs data at /var/log/pods/__//0.log we can configure Grafana Agent to ingest this log file and perform some parsing to transform raw JSON logs into structure logs and forward it to Loki tenant.&lt;/p&gt;

&lt;p&gt;This is a sample config for Grafana Alloy (similar to Grafana Agent as well):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Discover pods running on this node&lt;/span&gt;
&lt;span class="nx"&gt;discovery&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;kubernetes&lt;/span&gt; &lt;span class="s2"&gt;"pods"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;role&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"pod"&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// Discover the actual log file paths for pods&lt;/span&gt;
&lt;span class="nx"&gt;local&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;file_match&lt;/span&gt; &lt;span class="s2"&gt;"pod_logs"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;paths&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"/var/log/pods/*/*/*.log"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// Relabel to add useful metadata&lt;/span&gt;
&lt;span class="nx"&gt;discovery&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;relabel&lt;/span&gt; &lt;span class="s2"&gt;"pod_labels"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;targets&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;discovery&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;kubernetes&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;pods&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;targets&lt;/span&gt;

  &lt;span class="c1"&gt;// Extract namespace, pod, and container from the file path&lt;/span&gt;
  &lt;span class="nx"&gt;rule&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;source_labels&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;" __path__"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="nx"&gt;regex&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"/var/log/pods/([^/]+)/([^_]+)_([^/]+)/(.+)/(.+)&lt;/span&gt;&lt;span class="err"&gt;\\&lt;/span&gt;&lt;span class="s2"&gt;.log"&lt;/span&gt;
    &lt;span class="nx"&gt;target_label&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"namespace"&lt;/span&gt;
    &lt;span class="nx"&gt;replacement&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"$1"&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="nx"&gt;rule&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;source_labels&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;" __path__"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="nx"&gt;regex&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"/var/log/pods/([^/]+)/([^_]+)_([^/]+)/(.+)/(.+)&lt;/span&gt;&lt;span class="err"&gt;\\&lt;/span&gt;&lt;span class="s2"&gt;.log"&lt;/span&gt;
    &lt;span class="nx"&gt;target_label&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"pod"&lt;/span&gt;
    &lt;span class="nx"&gt;replacement&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"$2"&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="nx"&gt;rule&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;source_labels&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;" __path__"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="nx"&gt;regex&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"/var/log/pods/([^/]+)/([^_]+)_([^/]+)/(.+)/(.+)&lt;/span&gt;&lt;span class="err"&gt;\\&lt;/span&gt;&lt;span class="s2"&gt;.log"&lt;/span&gt;
    &lt;span class="nx"&gt;target_label&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"container"&lt;/span&gt;
    &lt;span class="nx"&gt;replacement&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"$4"&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// Read logs and forward to Loki&lt;/span&gt;
&lt;span class="nx"&gt;loki&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;source&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;file&lt;/span&gt; &lt;span class="s2"&gt;"pod_logs"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;targets&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;local&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;file_match&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;pod_logs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;targets&lt;/span&gt;
  &lt;span class="nx"&gt;forward_to&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;loki&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;write&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;default&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;receiver&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// Loki write target&lt;/span&gt;
&lt;span class="nx"&gt;loki&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;write&lt;/span&gt; &lt;span class="s2"&gt;"default"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;endpoint&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;url&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"https://&amp;lt;your-loki-endpoint&amp;gt;/loki/api/v1/push"&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="c1"&gt;// Optional: authentication&lt;/span&gt;
  &lt;span class="c1"&gt;// basic_auth {&lt;/span&gt;
  &lt;span class="c1"&gt;// username = "&amp;lt;user&amp;gt;"&lt;/span&gt;
  &lt;span class="c1"&gt;// password = "&amp;lt;password&amp;gt;"&lt;/span&gt;
  &lt;span class="c1"&gt;// }&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Grafana Alloy (or Agent) runs as daemon sets on your k8s nodes, so each daemon set read pod log file on the node, labels it to pod, namespace and container and push it to Loki tenant over HTTPs (make sure your Loki is accessible within cluster) or if outside cluster use &lt;strong&gt;VPC PrivateLink&lt;/strong&gt; to connect to Loki if running in different AWS network/account.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcwpmjvcckqivv4fuizlk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcwpmjvcckqivv4fuizlk.png" width="800" height="464"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Grafana Alloy to Loki for logs&lt;/em&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Multi Tenant Loki Dashboards:
&lt;/h4&gt;

&lt;p&gt;For better developer experience, we’re going to push class-level logs (info, error) to “service-logs” tenant, and for HTTP logs we’re going to push in another tenant “http-logs” so devs can filter/query based on fields.&lt;/p&gt;

&lt;p&gt;Each Loki tenant is a &lt;strong&gt;“logical” partition&lt;/strong&gt; on your storage backend (we’re using S3 bucket for our logs storage), so it writes each tenant logs in its own separate chunk on S3. Each tenant is served separately on UI/frontend so engineers can debug faster and helps reduce cognitive load when scanning through logs.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fokaygtgni30b8iowsd03.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fokaygtgni30b8iowsd03.png" width="799" height="445"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Multi tenant Loki dashboards&lt;/em&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Pushing HTTP logs to separate tenant on Loki:
&lt;/h4&gt;

&lt;p&gt;Now since we understand the underneath architecture of tenants in Loki, we’ll configure our Grafana Alloy (or Agent) config to filter out HTTP logs and push it to new tenant on Loki.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="c1"&gt;#########################&lt;/span&gt;
&lt;span class="c1"&gt;# TEAM A — Only HTTP logs (org.zalando)&lt;/span&gt;
&lt;span class="c1"&gt;#########################&lt;/span&gt;
&lt;span class="nx"&gt;loki&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;source&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;file&lt;/span&gt; &lt;span class="s2"&gt;"team_a_logs"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;targets&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;
    &lt;span class="nx"&gt;__path__&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"/var/log/myapp/*.log"&lt;/span&gt;
  &lt;span class="p"&gt;}]&lt;/span&gt;

  &lt;span class="nx"&gt;pipeline&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c1"&gt;# Step 1: Parse JSON logs&lt;/span&gt;
    &lt;span class="nx"&gt;stage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;json&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;expressions&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nx"&gt;level&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"level"&lt;/span&gt;
        &lt;span class="nx"&gt;logger_name&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"logger_name"&lt;/span&gt;
        &lt;span class="nx"&gt;msg&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"message"&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="c1"&gt;# Step 2: Keep only logs where logger_name == org.zalando&lt;/span&gt;
    &lt;span class="nx"&gt;stage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;keep&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;expressions&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nx"&gt;logger_name&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"org.zalando"&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="c1"&gt;# Step 3: Optional — add labels for Loki queries&lt;/span&gt;
    &lt;span class="nx"&gt;stage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;label&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;values&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nx"&gt;level&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"level"&lt;/span&gt;
        &lt;span class="nx"&gt;logger&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"logger_name"&lt;/span&gt;
        &lt;span class="nx"&gt;tenant&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"team-A"&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="c1"&gt;# Step 4: Send to Loki tenant A&lt;/span&gt;
    &lt;span class="nx"&gt;stage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;output&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;forward_to&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;loki&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;write&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;http_logs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;receiver&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nx"&gt;loki&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;write&lt;/span&gt; &lt;span class="s2"&gt;"http_logs"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;endpoint&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;url&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"https://loki.example.com/loki/api/v1/push"&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="nx"&gt;tenant_id&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"team-A"&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;#########################&lt;/span&gt;
&lt;span class="c1"&gt;# TEAM B — Only class-level logs (INFO, DEBUG, ERROR)&lt;/span&gt;
&lt;span class="c1"&gt;#########################&lt;/span&gt;
&lt;span class="nx"&gt;loki&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;source&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;file&lt;/span&gt; &lt;span class="s2"&gt;"team_b_logs"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;targets&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;
    &lt;span class="nx"&gt;__path__&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"/var/log/myapp/*.log"&lt;/span&gt;
  &lt;span class="p"&gt;}]&lt;/span&gt;

  &lt;span class="nx"&gt;pipeline&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c1"&gt;# Step 1: Parse JSON&lt;/span&gt;
    &lt;span class="nx"&gt;stage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;json&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;expressions&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nx"&gt;level&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"level"&lt;/span&gt;
        &lt;span class="nx"&gt;logger_name&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"logger_name"&lt;/span&gt;
        &lt;span class="nx"&gt;msg&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"message"&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="c1"&gt;# Step 2: Keep only certain log levels&lt;/span&gt;
    &lt;span class="nx"&gt;stage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;keep&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;expressions&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nx"&gt;level&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"~^(INFO|DEBUG|ERROR)$"&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="c1"&gt;# Step 3: Add useful labels&lt;/span&gt;
    &lt;span class="nx"&gt;stage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;label&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;values&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nx"&gt;level&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"level"&lt;/span&gt;
        &lt;span class="nx"&gt;logger&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"logger_name"&lt;/span&gt;
        &lt;span class="nx"&gt;tenant&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"team-B"&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="c1"&gt;# Step 4: Send to Loki tenant B&lt;/span&gt;
    &lt;span class="nx"&gt;stage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;output&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;forward_to&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;loki&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;write&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;service_logs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;receiver&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nx"&gt;loki&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;write&lt;/span&gt; &lt;span class="s2"&gt;"service_logs"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;endpoint&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;url&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"https://loki.example.com/loki/api/v1/push"&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="nx"&gt;tenant_id&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"team-B"&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Tenant service_logs:&lt;/strong&gt; Kept only logs with log level INFO, DEBUG, ERROR dropped TRACE level logs (i.e. HTTP logs).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tenant http_logs:&lt;/strong&gt; Kept only logs with logger_name equals to org.zalando i.e. HTTP logs, dropped other class-level logs (info, debug, error).&lt;/p&gt;

&lt;h4&gt;
  
  
  Distributed Tracing: OpenTelemetry
&lt;/h4&gt;

&lt;p&gt;Okay, now, we’re almost about to wrap up. The last thing we need to add is OpenTelemetry in our Spring Boot app so we can enable distributed tracing throughout our request lifecycle across microservices.&lt;/p&gt;

&lt;p&gt;The good thing is OpenTelemetry for JVM offers automatic instrumentation by running a JAR agent inside your docker container. So before running our app JAR in docker, we can download OTel agent JAR and pass as a JVM arg with our app, this will enable automatic instrumentation for tracing inside our app and will also propagate trace_id in MDC which can be correlated with logs as well.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="c"&gt;# === Stage 1: Build the application ===&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;maven:3.9.8-eclipse-temurin-21&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;AS&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;build&lt;/span&gt;

&lt;span class="c"&gt;# Set working directory&lt;/span&gt;
&lt;span class="k"&gt;WORKDIR&lt;/span&gt;&lt;span class="s"&gt; /app&lt;/span&gt;

&lt;span class="c"&gt;# Copy source and build&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; pom.xml .&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; src ./src&lt;/span&gt;
&lt;span class="k"&gt;RUN &lt;/span&gt;mvn clean package &lt;span class="nt"&gt;-DskipTests&lt;/span&gt;

&lt;span class="c"&gt;# === Stage 2: Runtime image ===&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="s"&gt; eclipse-temurin:21-jre&lt;/span&gt;

&lt;span class="c"&gt;# Set working directory&lt;/span&gt;
&lt;span class="k"&gt;WORKDIR&lt;/span&gt;&lt;span class="s"&gt; /app&lt;/span&gt;

&lt;span class="c"&gt;# Copy the Spring Boot fat JAR from the build stage&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; --from=build /app/target/*.jar app.jar&lt;/span&gt;

&lt;span class="c"&gt;# Download the latest OpenTelemetry Java agent&lt;/span&gt;
&lt;span class="c"&gt;# (Alternatively, you can include a specific version in your repo for consistency)&lt;/span&gt;
&lt;span class="k"&gt;ADD&lt;/span&gt;&lt;span class="s"&gt; https://github.com/open-telemetry/opentelemetry-java-instrumentation/releases/latest/download/opentelemetry-javaagent.jar /app/opentelemetry-javaagent.jar&lt;/span&gt;

&lt;span class="c"&gt;# Expose the application port&lt;/span&gt;
&lt;span class="k"&gt;EXPOSE&lt;/span&gt;&lt;span class="s"&gt; 8080&lt;/span&gt;

&lt;span class="c"&gt;# Environment variables for OpenTelemetry&lt;/span&gt;
&lt;span class="k"&gt;ENV&lt;/span&gt;&lt;span class="s"&gt; OTEL_SERVICE_NAME=my-spring-service \&lt;/span&gt;
    OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4317 \
    OTEL_RESOURCE_ATTRIBUTES=deployment.environment=prod,team=backend \
    OTEL_INSTRUMENTATION_LOGBACK_MDC_ENABLE=true

&lt;span class="c"&gt;# Start the app with the OTel Java Agent attached&lt;/span&gt;
&lt;span class="k"&gt;ENTRYPOINT&lt;/span&gt;&lt;span class="s"&gt; ["java", "-javaagent:/app/opentelemetry-javaagent.jar", "-jar", "/app/app.jar"]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Transporting Otel Telemetry over OTLP protocol to Grafana Tempo:
&lt;/h4&gt;

&lt;p&gt;The Otel agent JAR transports its telemetry data over OTLP protocol. Here we can either transport telemetry data to an Otel Collector acting as a central fascade for various tracing backends like (Tempo or Jaegar) or we can directly transport telemetry over OTLP protocol to Tempo as it supports OTLP for ingestion.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqb0ol5spzx11ku2aoohh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqb0ol5spzx11ku2aoohh.png" width="799" height="338"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Emitting Otel telemetry over OTLP to Tempo&lt;/em&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Testing:
&lt;/h4&gt;

&lt;p&gt;Done! That’s all, now finally we’re done with setting up our Grafana stack (Alloy, Tempo, Loki), we can start capturing and emitting HTTP layer and business-logic layer logs to Grafana Loki correlated with OpenTelemetry traces on Tempo based on trace_id allowing us to achieve end-to-end distributed tracing across entire system.&lt;/p&gt;

&lt;p&gt;Here how it looks like:&lt;/p&gt;

&lt;p&gt;HTTP log (request):&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8114pf6x04rmle8ut8hu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8114pf6x04rmle8ut8hu.png" width="800" height="675"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Grafana Loki (http_logs tenant)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;HTTP log (response):&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdo0dsvnlxz7ivq7pptca.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdo0dsvnlxz7ivq7pptca.png" width="800" height="393"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Grafana Loki (http_logs tenant)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Class-level log (business-logic layer):&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbefyhsxhnxe3liod119l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbefyhsxhnxe3liod119l.png" width="800" height="203"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Grafana Loki (service_logs tenant)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Full request trace on Grafana Tempo using trace_id&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyqe3lrskuih1wzukwhch.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyqe3lrskuih1wzukwhch.png" width="800" height="475"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Grafana Tempo&lt;/em&gt;&lt;/p&gt;

</description>
      <category>distributedsystems</category>
      <category>opentelemetry</category>
      <category>springboot</category>
      <category>observability</category>
    </item>
    <item>
      <title>From Monolith to Modularity: Modernizing Identity Platform at Scale</title>
      <dc:creator>Shehroz Ali</dc:creator>
      <pubDate>Wed, 11 Jun 2025 17:06:44 +0000</pubDate>
      <link>https://dev.to/sherredev/from-monolith-to-modularity-modernizing-identity-platform-at-scale-51a9</link>
      <guid>https://dev.to/sherredev/from-monolith-to-modularity-modernizing-identity-platform-at-scale-51a9</guid>
      <description>&lt;p&gt;We rebuilt the heart of our identity system — live, at scale, and without a single disruption. Here’s how we transformed a legacy monolith into a modular, high-trust authentication platform for millions.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5yz1xyd3x7su223az8fs.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5yz1xyd3x7su223az8fs.jpeg" width="800" height="534"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Introduction
&lt;/h3&gt;

&lt;p&gt;Bazaar’s Identity Platform is the authentication nucleus of our product ecosystem — orchestrating access for every user touchpoint, from our Grocery app and Rider interface to internal admin portals. It’s not just a login service; it governs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Sign-up, login, OTP verification&lt;/li&gt;
&lt;li&gt;Session lifecycle management&lt;/li&gt;
&lt;li&gt;Role-based access control&lt;/li&gt;
&lt;li&gt;OpenID Connect (OIDC) integration&lt;/li&gt;
&lt;li&gt;Multi-factor authentication&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At peak, it processes &lt;strong&gt;7000+ requests per minute&lt;/strong&gt; , handling millions of sessions across guest and registered users. Its stability defines the health of the entire platform. Any change to token behavior would ripple across every consumer product and backend — making correctness, reliability, and backward compatibility non-negotiable.&lt;/p&gt;

&lt;h3&gt;
  
  
  ⚠️ The Problem: Token Renewal at Scale Was Cracking
&lt;/h3&gt;

&lt;p&gt;Historically, our identity system issued JWT access and refresh tokens to both &lt;strong&gt;guests and registered users&lt;/strong&gt;. Clients relied on a shared API endpoint to refresh tokens every &lt;em&gt;n&lt;/em&gt; minutes.&lt;/p&gt;

&lt;p&gt;This approach began to show its cracks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🚨 &lt;strong&gt;Guest token renewals overwhelmed&lt;/strong&gt; the identity-service with unauthenticated traffic.&lt;/li&gt;
&lt;li&gt;🧩 &lt;strong&gt;Token lifecycle logic became entangled&lt;/strong&gt; with authentication and session logic.&lt;/li&gt;
&lt;li&gt;🧱 &lt;strong&gt;Domain boundaries blurred&lt;/strong&gt; , making the system increasingly hard to evolve.&lt;/li&gt;
&lt;li&gt;⚙️ Scaling stress — over &lt;strong&gt;1.9M+ active guest sessions&lt;/strong&gt; stored and rotated in DB.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This wasn’t just a performance bottleneck — it was an architectural red flag. The system needed to evolve or it would constrain our future growth.&lt;/p&gt;

&lt;h3&gt;
  
  
  Inspired by Global-Scale Identity Platforms
&lt;/h3&gt;

&lt;p&gt;We studied battle-tested identity systems to inform our design:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Amazon&lt;/strong&gt; : Session-token IDs for opaque anonymous flows&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Spotify&lt;/strong&gt; : Isolation between guest discovery and user-authenticated APIs&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Uber&lt;/strong&gt; : JWT enrichment at the edge via Envoy to decouple downstream auth&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These insights helped define our guiding principles.&lt;/p&gt;

&lt;h3&gt;
  
  
  Design Goals
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Reduce guest token renewal traffic on identity-service&lt;/li&gt;
&lt;li&gt;Introduce stateless guest sessions to remove DB dependency&lt;/li&gt;
&lt;li&gt;Preserve full backward compatibility for legacy clients&lt;/li&gt;
&lt;li&gt;Establish clear token domains: Guest, User, and Service&lt;/li&gt;
&lt;li&gt;Safely deliver change using TDD, BDD, and progressive rollout&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  From Monolith to Modular: Restructuring Identity as a Scalable Platform
&lt;/h3&gt;

&lt;p&gt;Before any token migration could succeed, we needed to modernize the very structure of our Identity Platform.&lt;/p&gt;

&lt;p&gt;What started as a monolithic mudball — entangling session logic, token management, login flows, and user models — was reimagined into a vertically sliced architecture, with clear boundaries, modularity, and domain ownership.&lt;/p&gt;

&lt;h4&gt;
  
  
  🎯 Why Vertical Slicing?
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;Our legacy system suffered from tight coupling between concerns — login flows touched session storage, token validation logic was scattered, and shared models leaked across contexts.&lt;/li&gt;
&lt;li&gt;Engineers had a hard time making localised changes without risking unrelated functionality.&lt;/li&gt;
&lt;li&gt;Adding support for new flows (like partner logins or stateless sessions) meant touching unrelated parts of the system.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  🧩 Our Modular Design Approach
&lt;/h4&gt;

&lt;p&gt;We restructured the platform into &lt;strong&gt;independent, domain-aligned vertical slices&lt;/strong&gt; , such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GuestService&lt;/li&gt;
&lt;li&gt;UserService&lt;/li&gt;
&lt;li&gt;SignupService&lt;/li&gt;
&lt;li&gt;OTPFlowHandler&lt;/li&gt;
&lt;li&gt;SessionLifecycleManager&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each module:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Owned its models, rules, business logic, and persistence.&lt;/li&gt;
&lt;li&gt;Was exposed via explicit interfaces (REST or internal contracts).&lt;/li&gt;
&lt;li&gt;Was testable, deployable, and evolvable in isolation.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This modularization accelerated development, simplified testing, and reduced the blast radius of changes — giving us the foundation to implement the new token lifecycle with confidence.&lt;/p&gt;

&lt;h4&gt;
  
  
  🔄 Future-Ready Foundation
&lt;/h4&gt;

&lt;p&gt;This shift wasn’t just about scaling what we had — it enabled:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Easier onboarding of engineers by navigating clear module boundaries.&lt;/li&gt;
&lt;li&gt;A pathway to microservices, as each vertical slice could become its own service.&lt;/li&gt;
&lt;li&gt;Faster incident resolution due to localised ownership and observability.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  🔄 Our Engineering Approach (Token Migration)
&lt;/h3&gt;

&lt;p&gt;This wasn’t a feature refactor — it was a &lt;strong&gt;core identity infrastructure rewrite&lt;/strong&gt; , under live, high-scale traffic. We chose to evolve the system incrementally and surgically.&lt;/p&gt;

&lt;h4&gt;
  
  
  🔁 1. Stateless Guest Session Tokens
&lt;/h4&gt;

&lt;p&gt;Instead of persisting or rotating refresh tokens for guests:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;We issued &lt;strong&gt;opaque session tokens&lt;/strong&gt; signed with a platform-wide secret&lt;/li&gt;
&lt;li&gt;Tokens had a &lt;strong&gt;long but bounded TTL&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Result: 70%+ drop in guest renewal traffic to identity-service&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  2. Enriched JWTs via Spring Gateway
&lt;/h4&gt;

&lt;p&gt;Based on the accessLevel claim (GUEST, USER, SERVICE), we:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Injected custom scopes and metadata into headers at the Gateway&lt;/li&gt;
&lt;li&gt;Allowed downstream services to &lt;strong&gt;stay stateless&lt;/strong&gt; and avoid JWT parsing&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  3. Middleware Unification
&lt;/h4&gt;

&lt;p&gt;All backend services received a &lt;strong&gt;uniform header structure&lt;/strong&gt; , regardless of caller type. This centralised auth logic while keeping the gateway lean.&lt;/p&gt;

&lt;h4&gt;
  
  
  4. Seamless Backward Compatibility
&lt;/h4&gt;

&lt;p&gt;Legacy clients still expecting guest refresh tokens (refreshToken = "GUEST") were gracefully intercepted and served valid tokens from the new guest_session API — no app updates required.&lt;/p&gt;

&lt;h3&gt;
  
  
  Engineering Practices That Enabled Safe Change
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;TDD &amp;amp; BDD&lt;/strong&gt; : Outside-in test design enabled refactoring behind robust test coverage&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Trunk-based development&lt;/strong&gt; : Incremental PRs with clear ownership&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ping-pong pairing&lt;/strong&gt; : Rapid peer iteration kept velocity high without compromising safety&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stateless-first mindset&lt;/strong&gt; : Reduced session state complexity at scale&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;accessLevel claims&lt;/strong&gt; : Cleanly routed and enforced logic boundaries across the stack&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Outcomes &amp;amp; Impact
&lt;/h3&gt;

&lt;p&gt;Despite refactoring one of the most sensitive systems in our infrastructure, we shipped without a glitch. Here’s what we achieved:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Zero downtime during rollout&lt;/li&gt;
&lt;li&gt;70% reduction in guest traffic to identity-service&lt;/li&gt;
&lt;li&gt;Full backward compatibility — no client-side changes&lt;/li&gt;
&lt;li&gt;Simplified operations — no need to manage guest session DB entries&lt;/li&gt;
&lt;li&gt;Clear domain boundaries for Guest, User, and Service tokens&lt;/li&gt;
&lt;li&gt;Enabled partner token flows with SERVICE scoped JWTs&lt;/li&gt;
&lt;li&gt;Robust test suite for long-term maintainability&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Final Thoughts
&lt;/h3&gt;

&lt;p&gt;Rebuilding identity wasn’t about rewriting a few endpoints.&lt;br&gt;&lt;br&gt;
 It was about evolving a monolith into a &lt;strong&gt;modular, testable, and scalable platform&lt;/strong&gt;.&lt;br&gt;&lt;br&gt;
 It was about building trust — at &lt;strong&gt;platform level&lt;/strong&gt; , at &lt;strong&gt;engineering level&lt;/strong&gt; , and at &lt;strong&gt;user level&lt;/strong&gt;.&lt;/p&gt;

</description>
      <category>springboot</category>
      <category>domaindrivendesign</category>
      <category>java</category>
      <category>cleanarchitecture</category>
    </item>
    <item>
      <title>How we solved SignUp/Login problem for Mint App using Uber’s USL — Unified SignUp Login Model</title>
      <dc:creator>Shehroz Ali</dc:creator>
      <pubDate>Sat, 09 Jul 2022 12:17:01 +0000</pubDate>
      <link>https://dev.to/sherredev/how-we-solved-signuplogin-problem-for-mint-app-using-ubers-usl-unified-signup-login-model-17m5</link>
      <guid>https://dev.to/sherredev/how-we-solved-signuplogin-problem-for-mint-app-using-ubers-usl-unified-signup-login-model-17m5</guid>
      <description>&lt;h3&gt;
  
  
  How Uber’s engineering model inspired me to increase customer acquisition for our app by 6% weekly
&lt;/h3&gt;

&lt;p&gt;So lately, we have ran into a questionnaire based problem, and here is the User story overview:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Ali discovered about Mint from Internet&lt;/p&gt;

&lt;p&gt;Ali likes to be the part of rockstar Mint Family as he likes Mint’s vision on bettering cities and urban climate infrastructure&lt;/p&gt;

&lt;p&gt;Ali headed towards for Sign Up for Mint App on &lt;em&gt;mymintrewards.com/signup&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Ali successfully subscribed to beta version of Mint App as an early adopter&lt;/p&gt;

&lt;p&gt;Weeks later, Mint team emailed Ali for onboarding on Mint’s App via &lt;em&gt;Google Play download link&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Ali successfully downloaded Mint App&lt;/p&gt;

&lt;p&gt;What will Ali see?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Now this is what which has made us to stop and operate for past 5 months and we ended up scratching our heads — i.e. how to keep wait-listing still live (persistently storing data somewhere which we can be use later to allow user to login) till we finish launching our pilot technology. This wasn’t an easy task as we need to figure out a way somehow that can keep both things up-and-gun running in parallel till our other teams finish setting up plans and actions on how to operate and signing up MOUs with initial Brands and Companies.&lt;/p&gt;

&lt;h4&gt;
  
  
  The Problem
&lt;/h4&gt;

&lt;p&gt;Problem is that, if a user has already signed up for Mint account via website than how we can unlock user’s Mint account for beta app using the same data? i.e. converting website form data into real-time user data. What will user see when he first time open the app?&lt;/p&gt;

&lt;h4&gt;
  
  
  Nested Problem
&lt;/h4&gt;

&lt;p&gt;Wait — here’s an another problem. We can’t directly ask for some sort of unique &lt;em&gt;auth key (say password, OAuth, etc.)&lt;/em&gt; which can also be used later to grant access to user to login into Mint app. The reason was we have some time launching the full-scale app till we fulfill our technological resources so it can be able handle strong continuous traffic and we also want to minimize latency so we can drive excellent in-app user experience. This was the main purpose of attaching subscriber form on the website so we can start onboarding users with moderate traffic giving us space to breath so we can make better informed decisions. Even we built form using no-code SaaS app to save our time plus resources and abstract out all the integration layers so we can really focus on the Business instead.&lt;/p&gt;

&lt;p&gt;In a nutshell, our goal is to convert that incoming form responses and turning into real-time relational database for allowing users to access their Mint account via our app.&lt;/p&gt;

&lt;p&gt;So how we solved the problem?&lt;/p&gt;

&lt;h4&gt;
  
  
  Unified SignUp Model (USL)
&lt;/h4&gt;

&lt;p&gt;We taken inspiration from Uber’s USL model derived by their excellent engineering team (a round of applause for them). We thoroughly studied their model and learnt how Uber solved their Login/SignUp problem for their food, delivery and riding apps seamlessly using USL pattern and successfully built an all-in-one smooth in-app experience whether you are a new user or just returning back. This was designed specially for Uber Eats, Uber Ride, Uber Delivery and other Uber products specially those requiring faster database communications for speedy transactions. We observed that somehow Uber had been experiencing the same problem within its multiple apps which we were facing at Mint and i.e. making Business Logic layer decides where to navigate user whether he/she is a completely new user or already signed up via &lt;em&gt;mymintrewards.com/signup&lt;/em&gt; or just returning back — this was quite similar to what Uber was also experiencing. We replicated the model and it worked.&lt;/p&gt;

&lt;p&gt;Let me give you a high-level overview of how we altered Uber’s USL according to our Business Logic and built a seamless onboarding experience for Mint App.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fin9g5lw9lf6fscqdgl1x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fin9g5lw9lf6fscqdgl1x.png" width="800" height="402"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;This is what I’ve come up with&lt;/em&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  The Final Solution
&lt;/h4&gt;

&lt;p&gt;You open the app — you will be asked to enter your email or phone number. Once you entered, our Business logic layer will decided where to navigate you and you will automatically be taken to where you belong. Boom, done. We made it simple for you. We dropped all those fancy fields, buttons and screens to minimize the end-burden on front-end client and improve the overall user experience enabling users quickly access their Mint account within 2-steps. Usability done right.&lt;/p&gt;

&lt;h4&gt;
  
  
  Nobody wants to wait — Be quick, Be seamless
&lt;/h4&gt;

&lt;p&gt;Today, end-user looks for value. He didn’t care about all of those underneath things going around. He didn’t care about your &lt;em&gt;technical debt&lt;/em&gt;. He has nothing to do with it. The only goal end-user wants to achieve from a &lt;em&gt;successful&lt;/em&gt; product is the solution to his problem aka value — real value. This should be his first word of mouth when first time opening an app — Woah!. I believe achieving this couldn’t be hard as technology today allows us if you rightly putted all the puzzle pieces together. The only thing you need to look for is converting Business model into actual solution — seamless, smooth and quick.&lt;/p&gt;

&lt;p&gt;Seeyou, again.&lt;/p&gt;

</description>
      <category>uber</category>
      <category>engineering</category>
    </item>
    <item>
      <title>Today is not about writing code — Today is all about ‘Velocity’</title>
      <dc:creator>Shehroz Ali</dc:creator>
      <pubDate>Thu, 23 Jun 2022 13:36:00 +0000</pubDate>
      <link>https://dev.to/sherredev/today-is-not-about-writing-code-today-is-all-about-velocity-51f6</link>
      <guid>https://dev.to/sherredev/today-is-not-about-writing-code-today-is-all-about-velocity-51f6</guid>
      <description>&lt;h3&gt;
  
  
  Today is not about writing code — Today is all about ‘Velocity’
&lt;/h3&gt;

&lt;p&gt;Yes, ‘Velocity’ — the same term that you used to be taught in your late high-school or mid-school physics classes. Maybe, some of you don’t like physics (just like me too) — well, but sometimes it doesn’t mean that it carries the same exact meaning to the actual context where we pointing. Right? So let’s talk about what ‘Velocity’ and ‘Software Engineering’ has to do all about and what their relationship tells us.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fao56f3ebzsfpewxvn7jd.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fao56f3ebzsfpewxvn7jd.jpeg" width="800" height="533"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Source: Pexels&lt;/em&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Referencing
&lt;/h4&gt;

&lt;p&gt;Amazon famously commits code to production every 11.6 seconds and they’re always learning really fast. Pause for a second here and imagine the velocity they moving. Seems more faster than last time I pulled my phone to book an Uber? Exactly. We don’t know how much lines do every merge contains (until we ourselves works with the team at Amazon) but surprisingly what we see is their velocity to ship to production faster, 11.6 seconds. How is that even possible? Let’s try to break and see what’s happening underneath the hood.&lt;/p&gt;

&lt;h4&gt;
  
  
  Business vs. Tech — Really?
&lt;/h4&gt;

&lt;p&gt;Right tooling. Yes. This is where I want you all to focus. If you really want to build better Continuous Delivery pipelines so you can test, build and release faster than you once again really need to think-off your delivery pipelines. DevOps enable teams to set things to automation and ships the new build to production environment through various CI and CD tools (such as Jenkins, Fastlane, ArgoCD, Ansible, Terraform, etc.). In 2022 today, building shouldn’t be considered as the only primary goal for a startup, it’s secondary, but what’s important today is your ‘Velocity’ from both business and technical perspective. Not every founder is a also a &lt;em&gt;‘technical’&lt;/em&gt; founder and not every founder cares a lot about those underlaying &lt;em&gt;‘technical debt’&lt;/em&gt; or your excellent-driven technology stack. Technology is just a tool we use today to build businesses for tomorrow. Technology helps us drive better customer acquisition by delivering values. It’s really the business first. Technology helps us to eliminate painful tasks by automating them and enabling teams to move faster and safer. It really help us achieve those defined business goals to increase profitability and scale faster.&lt;/p&gt;

&lt;h4&gt;
  
  
  Wonders did by the amazing Software Community
&lt;/h4&gt;

&lt;p&gt;So, listening to all of this, I would say that it’s the right time to re-pay a visit to the wonderful SaaS market and pick the right tools for the right job. Thanks to the wonderful beautiful community of developers and engineers that empowers other engineers and developers to put puzzle pieces together to get things done quickly in seconds or minutes.&lt;/p&gt;

&lt;h4&gt;
  
  
  For the trade-offs people
&lt;/h4&gt;

&lt;p&gt;Like, as always, I use to add the word &lt;em&gt;‘trade-offs’&lt;/em&gt; so you need to be-aware of it too. I mean, let’s be honest, its depends on us where gonna compromise and how much &lt;em&gt;‘technical debt’&lt;/em&gt; we gonna bear. I mean maybe that could charge an extra buck from your bank account but on the other side it is really helping teams to get things done quickly (such as SonarCloud) enabling teams to move forward and ship quality-code faster. If you really getting more value than you might expected and is flourishing your business model leading to bigger customer acquisition and much bigger rounds then I must say that you can mark it as a &lt;em&gt;green signal.&lt;/em&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  It’s all about shipping value today — not code
&lt;/h4&gt;

&lt;p&gt;Today is all about how much value we getting. As being a technical in nature, every new product or app that I tries left different impressions on me. Some might really excites me. Some might doesn’t and some might doesn’t sounds any meaning to me. Regardless of experiencing and building multiple apps and software, I must add here that your product should at the end delivers some value — value in terms of real value that makes people genuinely happy and stick to your product.&lt;/p&gt;

&lt;p&gt;That’s a wrap! Feel free to drop your thoughts below. I would be really happy listening to all your feedbacks.&lt;/p&gt;

&lt;p&gt;Until we met again, Seeyou. Have a wonderful day!&lt;/p&gt;

</description>
      <category>continuousdelivery</category>
      <category>softwareengineering</category>
      <category>velocity</category>
    </item>
  </channel>
</rss>
