<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Nahum Litvin</title>
    <description>The latest articles on DEV Community by Nahum Litvin (@nahumlitvin).</description>
    <link>https://dev.to/nahumlitvin</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1568834%2Fea8e22b6-6fd8-4427-a139-751bed8ff2a0.jpg</url>
      <title>DEV Community: Nahum Litvin</title>
      <link>https://dev.to/nahumlitvin</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/nahumlitvin"/>
    <language>en</language>
    <item>
      <title>Claude Code mod: themes engine, charts, graphs, colors. All you ever wanted.</title>
      <dc:creator>Nahum Litvin</dc:creator>
      <pubDate>Fri, 02 Oct 2026 14:43:06 +0000</pubDate>
      <link>https://dev.to/nahumlitvin/claude-code-mod-themes-engine-charts-graphs-colors-all-you-ever-wanted-2mff</link>
      <guid>https://dev.to/nahumlitvin/claude-code-mod-themes-engine-charts-graphs-colors-all-you-ever-wanted-2mff</guid>
      <description>&lt;p&gt;&lt;strong&gt;Claude Code mod: themes engine, charts, graphs, colors. All you ever wanted.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude plugin marketplace add NahumLitvin/prismantis
claude plugin &lt;span class="nb"&gt;install &lt;/span&gt;prismantis@prismantis
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://github.com/NahumLitvin/prismantis" rel="noopener noreferrer"&gt;https://github.com/NahumLitvin/prismantis&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmedia.daily.dev%2Fimage%2Fupload%2Fs--UNt3ZFs7--%2Ff_auto%2Fv1790951587%2Fugc%2Fcontent_43e8bc06-ed34-4b4e-9d72-80848301526e%3F_a%3DBAMAMicg0" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmedia.daily.dev%2Fimage%2Fupload%2Fs--UNt3ZFs7--%2Ff_auto%2Fv1790951587%2Fugc%2Fcontent_43e8bc06-ed34-4b4e-9d72-80848301526e%3F_a%3DBAMAMicg0" alt="prismantis" width="800" height="429"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
    </item>
    <item>
      <title>One lost signal, five days stuck, 45,000 frozen threads: fixing a gVisor hang upstream</title>
      <dc:creator>Nahum Litvin</dc:creator>
      <pubDate>Tue, 29 Sep 2026 00:00:00 +0000</pubDate>
      <link>https://dev.to/nahumlitvin/one-lost-signal-five-days-stuck-45000-frozen-threads-fixing-a-gvisor-hang-upstream-3di2</link>
      <guid>https://dev.to/nahumlitvin/one-lost-signal-five-days-stuck-45000-frozen-threads-fixing-a-gvisor-hang-upstream-3di2</guid>
      <description>&lt;p&gt;It looked like a deadlock in gVisor. It was a single missed wake-up, hiding behind a symptom teams have been reporting since 2021, and two strangers and a maintainer fixed it in about a week.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The hang in one picture: one signal, no retry, and a deadline check whose only action is a log line.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;On a Tuesday morning our alerting claimed that 559 pods were stuck Terminating in production. The real number was one. The rule counts kubelet failed-kill events rather than pods, and a sandbox that never dies gets retried forever, so a single wedged pod read as a fleet on fire. That pod turned out to be the interesting part.&lt;/p&gt;

&lt;p&gt;At Wix, our Velo grid runs users' backend JavaScript as untrusted code inside gVisor sandboxes on Amazon EKS. gVisor puts a userspace kernel between that code and the host kernel, so untrusted code never talks to the host kernel directly. Close to two thousand sandboxes per node group, short lifecycles, heavy churn. Deleting a pod should take seconds. This one had been Terminating since 07:11. kubelet's kill requests got DeadlineExceeded every 2 minutes and would have continued forever. We intervened by hand after about 90 minutes. A week earlier, eight pods on four nodes had sat like that for days.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why nobody's timeout saved us
&lt;/h2&gt;

&lt;p&gt;Before the forensics, the shape of the problem. Deleting a pod is a relay: each layer asks the next one, politely, to make something die. Every layer should hold two more things beyond the polite ask: a bound on how long it waits, and an escalation that works &lt;em&gt;without the cooperation of the thing it is waiting on&lt;/em&gt;. In our case the relay runs from kubelet (the Kubernetes node agent) to containerd (the container runtime) to the gVisor shim (the per-sandbox process containerd talks to) to the sandbox itself. Here is that stack, graded on all three:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Three arrows per hop: blue attempts, amber detects, violet recovers. Blue failed once, at the bottom. Amber either did not exist, fired into a log line, or fired and could only re-send blue. Violet existed nowhere, so one lost message at the bottom of the stack became the whole stack's hang.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Look at the violet column. Every escalation asked the very component being escalated against to cooperate. containerd's dead-shim cleanup only fires when the shim disconnects. Ours kept answering Connect and State, which never touched the lock, while Kill and Stats blocked behind it. The connection stayed up, so cleanup never fired. Alive enough to prevent its own cleanup.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a frozen sandbox looks like
&lt;/h2&gt;

&lt;p&gt;That userspace kernel is called the sentry. On systrap, the platform we run, the app's threads run in stripped-down stub processes and the sentry coordinates them. Our wedged pod had a sentry that was alive but answering nothing. runsc kill, gVisor's own command for stopping a sandbox, hung. runsc debug --stacks, the tool that is supposed to tell you why something hangs, also hung. The panic log was 0 bytes.&lt;/p&gt;

&lt;p&gt;The node told the rest of the story. Load average climbing about 3 per hour while actual CPU sat near 30%. Hundreds of sentry threads parked, waiting on shared memory. Stub processes accumulating. On the worst nodes, load kept climbing for days with the machine mostly idle. The only remediation that worked was SIGKILL to the whole sandbox process tree, by hand, over AWS SSM, a remote shell into the node. That became a runbook. Runbooks like that are a debt.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Stub processes pile up behind the frozen sentry: load average keeps climbing while the CPU does almost nothing. Threads stuck waiting, not working: the rising top panel against the flat bottom one is the bug.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;So we filed &lt;a href="https://github.com/google/gvisor/issues/14408" rel="noopener noreferrer"&gt;gvisor#14408&lt;/a&gt; with what we had: the outside view, two affected releases, and an honest admission that we could not get stacks because the debug tooling was wedged along with the sandbox. Our proposed fix, PR #14201, had been open since the week before.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then a stranger showed up with the other half
&lt;/h2&gt;

&lt;p&gt;The same day, another engineer filed &lt;a href="https://github.com/google/gvisor/issues/14405" rel="noopener noreferrer"&gt;gvisor#14405&lt;/a&gt;. Same bug at a different company, and they had the one thing we could not get: a full dump of the sentry's goroutines, Go's lightweight threads, from inside a frozen sentry. Their dump showed one of the sentry's internal worker threads sitting in the same wait loop for over five days, waiting for a stub thread that had missed a single wake-up signal.&lt;/p&gt;

&lt;p&gt;In their shim, about 45,000 waiting threads were queued behind one lock held by a kill that would not finish, and their containerd and kubelet memory grew until nodes ran out.&lt;/p&gt;

&lt;p&gt;The bug itself is painfully simple. The sentry sends the stub one interrupt signal. One. If that signal is lost, and nobody has yet pinned down why it sometimes is, the sentry keeps waiting for the answer with no retry and no escape. There is even a 30 second deadline in the code that notices the wait is too long. Here is the whole thing it did about it:&lt;/p&gt;

&lt;p&gt;if time.Now().After(deadline) { log.Warningf("Systrap task goroutine has been waiting on " + "ThreadContext.State futex too long. ...") // no retry, no escalation }&lt;/p&gt;

&lt;p&gt;Where did 45,000 waiting threads come from? cAdvisor kept polling every container's Stats. Each request added a goroutine that blocked behind the hung kill's lock. Its caller timed out, but a mutex wait cannot observe cancellation, so the goroutine stayed. Over five days, 45,000 accumulated: roughly one every 10 seconds, the scrape interval fossilized in a thread dump. The metrics system asking "how are you?" every 10 seconds had pushed the shim to about 600MB.&lt;/p&gt;

&lt;p&gt;A timeout that only logs a warning is not a timeout. It is a diary.&lt;/p&gt;

&lt;p&gt;And this exact wait sits on the teardown path. Killing a sandbox starts with a freeze-everything step that waits for every worker thread to park, so one stuck worker means the kill command can never answer, the shim call hangs, kubelet gets DeadlineExceeded, and each retry parks another goroutine behind the first one. The pod stays Terminating for days.&lt;/p&gt;

&lt;h2&gt;
  
  
  So, how did we fix it?
&lt;/h2&gt;

&lt;p&gt;The merged fix escalates in stages, gentlest first:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Tap again.&lt;/strong&gt; The interrupt is resent on each 5 second checkup wakeup instead of once. When the resend lands, a lost signal costs the workload a stall of a few seconds, then it carries on.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Photograph the scene.&lt;/strong&gt; If the stub is still unresponsive after the 30 second deadline, the sentry dumps every internal stack trace to the log. A frozen sandbox can not be debugged from outside unless it was started with --panic-signal; ours were not, theirs were, which is where their dump came from. So this failure now documents itself. The next team to hit anything like this gets for free what took two companies a week to assemble.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Remove only the broken part.&lt;/strong&gt; Then the sentry kills just the one stuck subprocess, through the same path it already uses when a stub dies naturally. The blocked task unwinds, teardown can finish, and the healthy subprocesses in the sandbox are untouched.&lt;/p&gt;

&lt;p&gt;My first version killed the entire sentry. gVisor maintainer Konstantin Bogomolov pushed for killing only the stuck subprocess, sparing healthy neighbors. He was right. The other reporter had raised the same concern earlier that day. Konstantin also caught a dropped continue, which led me to a race: a context that recovered during the wait could still get killed. Three review rounds in under a week took the PR from nobody having looked at it to merged on master.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The merged fix, staged from least to most disruptive. My first version jumped straight to killing the whole sandbox; review scoped it down.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that bothers me
&lt;/h2&gt;

&lt;p&gt;After the merge, we found reports since 2021 from a Cloudflare engineer, a GKE user, Knative, Talos Linux, Arista and others on the containerd tracker. Roughly eight teams over five years. Causes differed, and some cases got fixes. But the pattern kept returning: pods stuck Terminating, debugging tools hanging, someone killing processes by hand, and more often than not the thread ends without a stack trace from inside the sandbox.&lt;/p&gt;

&lt;p&gt;Two teams filing in the same week broke that pattern: our production evidence and proposed fix met the other team's internal stack dump.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Five years of the same symptom, with different root causes underneath. Most threads ended with the manual workaround, until two half-reports completed each other.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The relay diagram near the top of this post still isn't fully green. The merged fix repairs the layer where the hang started. We filed gvisor#14548 to bound shim kill waits, then SIGKILL the sandbox tree, and containerd#14081 to SIGKILL the shim after repeated kill RPC timeouts. In review of the shim PR, gvisor#14549, the same maintainer asked to narrow it to bounded waits on Kill, Stats and Status, because killing the sandbox from the shim would take out healthy containers too, the same over-reaction he flagged in the first fix. That revision is in progress, and escalation stays open. Either escalation, as originally filed, would have capped our incident at minutes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons learned
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt; &lt;strong&gt;A watchdog that only barks is not a watchdog.&lt;/strong&gt; If code detects a should-never-happen state, it must act: retry, self-report, fail. Grep your own codebase for deadline checks whose only body is a log line. This one surfaced because a sandbox sat stuck for five days behind it.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;File the issue even with half the evidence.&lt;/strong&gt; Our report had no stacks and said so plainly. Within a day a stranger's independent report supplied them, and their dump changed the fix. The half-report you are embarrassed to file is someone else's missing half.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Remediate the smallest thing that unblocks you.&lt;/strong&gt; Kill the subprocess, not the sandbox. The instinct under pressure is the big hammer. Review cut the kill from the whole sentry to one subprocess. The same restraint belongs in your incident runbooks.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Audit your escalation paths for circular dependencies.&lt;/strong&gt; Walk your stack with that diagram's two questions: bounded wait, and escalation that works without the target's cooperation. If the answer to the second bottoms out at "a human with SSH", write that down, because that is your actual design.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The fix shipped in gVisor release-20260831.0. Until our fleet is on it, the SSM kill runbook stays, but this one has an ending, and the next report of this hang will arrive with its stacks attached.&lt;/p&gt;

&lt;p&gt;This post was written by Nahum Litvin.&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>devops</category>
      <category>go</category>
      <category>opensource</category>
    </item>
    <item>
      <title>A node that reported Ready and started nothing: four parts, one containerd fix</title>
      <dc:creator>Nahum Litvin</dc:creator>
      <pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate>
      <link>https://dev.to/nahumlitvin/a-node-that-reported-ready-and-started-nothing-four-parts-one-containerd-fix-5gdk</link>
      <guid>https://dev.to/nahumlitvin/a-node-that-reported-ready-and-started-nothing-four-parts-one-containerd-fix-5gdk</guid>
      <description>&lt;h2&gt;
  
  
  What happened
&lt;/h2&gt;

&lt;p&gt;In July one of our Kubernetes nodes stopped starting pods for 4.5 hours. It reported Ready the entire time. The scheduler kept placing pods on it, every new container hung in creation, and nothing on the node's health path looked wrong.&lt;/p&gt;

&lt;p&gt;We run untrusted user code at Wix, one site per sandbox, so pods come and go all day. A node that silently stops starting them is worse than a node that dies: the scheduler keeps feeding it.&lt;/p&gt;

&lt;p&gt;The easy answer was a file descriptor limit. The snapshotter logs were full of "too many open files". Raise the limit, restart, move on. I almost shipped that.&lt;/p&gt;

&lt;p&gt;This is the whole story in one place: the outage, the two bugs, the fix that got reverted, and the one that stays.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it happens
&lt;/h2&gt;

&lt;p&gt;Two bugs lined up.&lt;/p&gt;

&lt;p&gt;The first was in containerd. During snapshot garbage collection, containerd calls the snapshotter's Remove while holding that snapshotter's write lock in containerd's metadata layer, with no deadline. It even drops the caller's cancellation. If the snapshotter never answers, the lock is held forever, and every snapshot operation on the node queues behind it, including the one that starts a new pod sandbox. Nothing on the readiness path takes that lock, which is why the node kept saying Ready.&lt;/p&gt;

&lt;p&gt;Readiness is a separate path. Kubelet asks the runtime for its status, the runtime answers from memory, and nobody touches the snapshot metadata. So the node was healthy by every check we had, and useless by the only one that mattered: can it start a container.&lt;/p&gt;

&lt;p&gt;The second was the trigger, in AWS's soci-snapshotter, the plugin that lazily loads container images. Every file verified on first access opened a reader from the span cache and never closed it. The sibling function about 45 lines up closes the exact same reader. One leaked file descriptor per file, on every node running lazy-loaded images. Under a burst of pod starts the snapshotter hit its limit of 65,535 and stopped answering. containerd was waiting on it, holding the lock.&lt;/p&gt;

&lt;p&gt;A proxy snapshotter like soci runs in its own process, and containerd talks to it over gRPC. That boundary is where the real problem lived, though it took me two more PRs to see it.&lt;/p&gt;

&lt;p&gt;The goroutine dumps told the story once we knew where to look. Hundreds of goroutines parked on that snapshotter lock, and one GC goroutine holding it, waiting for a gRPC header from the snapshotter that was never coming. The node was not broken. It was waiting, politely and forever.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we got wrong first
&lt;/h2&gt;

&lt;p&gt;I got the diagnosis wrong twice before I got it right. After the storm the file descriptor count looked low, 151 out of 65,535, so I concluded the limit theory was wrong. Then I flip-flopped back. The leaked descriptors had been reclaimed once Go's finalizers ran, so the steady state hid the peak.&lt;/p&gt;

&lt;p&gt;I did most of that investigation with Claude, reading goroutine dumps from the wedged node and two unfamiliar codebases side by side. It didn't write the fix. It made the investigation cheap enough to do properly, including catching my own wrong conclusion.&lt;/p&gt;

&lt;p&gt;The bigger mistake came later. My first containerd fix put a timeout on the one call that hung on my nodes, the GC's Remove. It was correct for my incident. It got merged after 40 days and 30 review rounds, and I wrote about it as a win.&lt;/p&gt;

&lt;p&gt;Those 30 rounds were the real engineering. The timeout's config key got renamed twice, because in a project with hundreds of keys a name is an API commitment. The default went from 30 minutes to 5 and back to 30. I pushed hard for 5. A maintainer pushed back: the old behavior was unbounded, so on upgrade the least surprising default is the big one, and anyone running remote snapshotters can tune it down. He was thinking about thousands of clusters upgrading. I was thinking about mine.&lt;/p&gt;

&lt;p&gt;My test got rewritten around Go's synctest, so it runs in milliseconds on a fake clock instead of sleeping and mutating global state. And when CI wedged in the merge queue, Michael Brown restarted the buckets, herded reviewers and queued it again until it landed.&lt;/p&gt;

&lt;p&gt;Then I opened a second PR to cover another call that could hang the same way. Working through it with the maintainers, we saw that I was patching symptoms one call at a time. The thing that hangs is an external plugin, and it can hang on any of its calls. With a release about to be cut, Michael Brown reverted my first fix 13 days after it merged, so it never shipped. I agreed with the call.&lt;/p&gt;

&lt;p&gt;The suggestion that changed the plan came in that second PR's thread: stop bounding components one by one, and put the bound where containerd talks to the plugin. Michael agreed. Reverting my first fix was the price of doing it that way, because two mechanisms bounding the same call would have been worse than one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;soci-snapshotter: close the reader in file Verify. One line, defer r.Close(), merged upstream in July.&lt;/li&gt;
&lt;li&gt;containerd, first attempt: bound the GC path's Remove with a configurable timeout. Merged, then reverted before release, because it fixed one call out of many.&lt;/li&gt;
&lt;li&gt;containerd, the fix that stays: a default_timeout setting on proxy plugins. When it is set, every call containerd makes to a proxy snapshotter gets that deadline if the incoming context has none. A caller that brings its own deadline keeps it. Leave it empty and nothing changes on upgrade.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The proxy snapshotter has 10 gRPC methods and none of them had a deadline unless the caller already set one. Putting the bound at the client covers all 10 at once, including the GC path my first fix targeted. It also covers paths I never hit in production, which is the point.&lt;/p&gt;

&lt;p&gt;Streaming calls needed a decision. Walk gets one deadline for the whole stream, not one per message, so time spent in callbacks counts against it. The content store proxy has the same gap but a different shape: a writer outlives the call that created it, so a per-call deadline would cut an ingest that is still being written. That one needs its own setting and its own PR.&lt;/p&gt;

&lt;p&gt;The review on the final PR was shorter, and it still taught me something. I had added an options argument to the exported constructor. Compatible, I thought. The maintainer pointed out it still breaks anyone who stores that function as a value, so the old constructor stays untouched and a new one takes the options.&lt;/p&gt;

&lt;p&gt;The SOCI side was the quick part. The PR went in on July 15 and AWS merged it five days later. The containerd side took three PRs and two months. The difference is not the size of the change. A one-line fix to an obvious leak has one right answer. A timeout in a runtime that thousands of clusters run has many plausible answers, and the review is where you find out which one the project can live with.&lt;/p&gt;

&lt;p&gt;It doesn't fix everything, and the PR says so. A GC pass over many slow snapshots can still add up while it holds the lock, because each call finishes just under the limit. That needs the lock itself reworked, and it is the next piece.&lt;/p&gt;

&lt;p&gt;Four posts, three PRs, one revert. The fix that finally lands is smaller than the one that got reverted, and it covers more. That is usually how it goes when the review is doing its job.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;A patch that is correct for your incident and a patch that is correct for everyone running the project are different things. The review is the distance between them.&lt;/li&gt;
&lt;li&gt;Review doesn't stop at merge. Mine got reverted, and the revert was the right engineering.&lt;/li&gt;
&lt;li&gt;Put the bound at the boundary. One deadline where containerd talks to the plugin covers every call; one deadline per call never ends.&lt;/li&gt;
&lt;li&gt;The steady state lies. Look at the peak, not the numbers after the storm.&lt;/li&gt;
&lt;li&gt;Maintainers do invisible work daily, for free, for strangers. Say so when they do it for you.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Credits and links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Michael Brown and the containerd maintainers, for 30 review rounds, a revert I agreed with, and a better fix&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/containerd/containerd/pull/13799" rel="noopener noreferrer"&gt;containerd/containerd#13799&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/containerd/containerd/pull/14187" rel="noopener noreferrer"&gt;containerd/containerd#14187&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/awslabs/soci-snapshotter/pull/2043" rel="noopener noreferrer"&gt;awslabs/soci-snapshotter#2043&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>kubernetes</category>
      <category>containers</category>
      <category>aws</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
