<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Adreet Gogoi</title>
    <description>The latest articles on DEV Community by Adreet Gogoi (@adreetgog).</description>
    <link>https://dev.to/adreetgog</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4119706%2F3df3689a-ed75-4790-aff1-f31c4fe9c73b.png</url>
      <title>DEV Community: Adreet Gogoi</title>
      <link>https://dev.to/adreetgog</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/adreetgog"/>
    <language>en</language>
    <item>
      <title>[Boost]</title>
      <dc:creator>Adreet Gogoi</dc:creator>
      <pubDate>Mon, 21 Sep 2026 14:10:44 +0000</pubDate>
      <link>https://dev.to/adreetgog/-549a</link>
      <guid>https://dev.to/adreetgog/-549a</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/remdore/how-much-traffic-can-a-6-server-handle-i-measured-it-9737-requests-a-second-24il" class="crayons-story__hidden-navigation-link"&gt;How much traffic can a $6 server handle? I measured it: 9,737 requests a second&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/remdore" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F374495%2F8514d87d-49b5-4b8f-865f-f24a5cc1c29c.png" alt="remdore profile" class="crayons-avatar__image"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/remdore" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Remdore
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Remdore
                
                
              
              &lt;div id="story-author-preview-content-4707740" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/remdore" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F374495%2F8514d87d-49b5-4b8f-865f-f24a5cc1c29c.png" class="crayons-avatar__image" alt=""&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Remdore&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/remdore/how-much-traffic-can-a-6-server-handle-i-measured-it-9737-requests-a-second-24il" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Sep 21&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/remdore/how-much-traffic-can-a-6-server-handle-i-measured-it-9737-requests-a-second-24il" id="article-link-4707740"&gt;
          How much traffic can a $6 server handle? I measured it: 9,737 requests a second
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/performance"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;performance&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/devops"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;devops&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/webdev"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;webdev&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/nginx"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;nginx&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/remdore/how-much-traffic-can-a-6-server-handle-i-measured-it-9737-requests-a-second-24il" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/exploding-head-daceb38d627e6ae9b730f36a1e390fca556a4289d5a41abb2c35068ad3e2c4b5.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/multi-unicorn-b44d6f8c23cdd00964192bedc38af3e82463978aa611b4365bd33a0f1f4f3e97.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;7&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/remdore/how-much-traffic-can-a-6-server-handle-i-measured-it-9737-requests-a-second-24il#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            6 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>Capability, Constraint, and the Real Work of Building AI</title>
      <dc:creator>Adreet Gogoi</dc:creator>
      <pubDate>Mon, 21 Sep 2026 08:14:20 +0000</pubDate>
      <link>https://dev.to/adreetgog/capability-constraint-and-the-real-work-of-building-ai-1d9m</link>
      <guid>https://dev.to/adreetgog/capability-constraint-and-the-real-work-of-building-ai-1d9m</guid>
      <description>&lt;p&gt;Science fiction gave us the templates. Helpful machines that expand human possibility. Systems that turn calm competence into quiet catastrophe. Early researchers carried both visions—some projected rapid intellectual progress toward better outcomes for everyone; others began asking harder questions about control, goals, and what happens when the systems we build become harder to steer. Those questions never went away. They simply became more concrete.&lt;/p&gt;

&lt;p&gt;In 2026 the conversation sharpened again when researchers close to the frontier started speaking more bluntly. Jacob Coxon, who had worked on pretraining at both OpenAI and Anthropic, resigned and said the leading labs were racing toward self-improving systems while “gambling with our lives.” He claimed many people inside the labs earnestly believe the technology could pose an existential threat within the decade. A colleague publicly put a personal estimate of greater than 10 percent on the chance of catastrophic outcomes in that window. These statements are not proof that any particular future is locked in. They are signals that the people closest to the work see real, non-zero chances of large-scale failure modes that current methods do not yet fully address.&lt;/p&gt;

&lt;p&gt;At the same time, the systems themselves are already constrained by the physical world. Energy demand from data centers continues to climb. Firm, low-carbon power is scarce. Water for cooling, land, materials, and grid interconnection timelines all act as binding limits. Efficiency improvements in hardware and algorithms help, yet overall demand has so far grown faster. Tokens and subscriptions remain the visible economic layer, but underneath them sit harder parameters: the quality and reliability of the electricity used, the carbon and water intensity of each unit of useful output, the utilization of existing hardware, and the true cost of scaling. Human physical activity as a way to “earn” compute remains a marginal idea—interesting as a cultural or health incentive, negligible as an energy source at the relevant scale.&lt;/p&gt;

&lt;p&gt;The picture that is emerging is neither pure utopia nor inevitable doom. It is a system under pressure, where capability keeps rising while multiple constraints tighten. One plausible shape looks something like this: frontier systems concentrate where firm, relatively clean power can be secured; most everyday AI runs on smaller, more efficient models closer to the user; training and inference become more carefully scheduled against grid conditions and resource intensity; pricing begins to reflect more of the real costs. Residual risk remains because alignment and control problems are not fully solved, and competitive dynamics still reward speed.&lt;/p&gt;

&lt;h3&gt;
  
  
  The real job
&lt;/h3&gt;

&lt;p&gt;In that environment the actual work of AI and systems design looks less like pure research and more like operating a high-stakes production system. The habits that matter most are familiar to anyone who has kept complex software alive under load:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Understanding the blast radius. What breaks if this model, this training run, or this deployment goes wrong—and how far does the damage travel?&lt;/li&gt;
&lt;li&gt;Making tradeoffs under pressure. Speed versus safety, capability versus energy cost, openness versus control. Every choice has costs that someone will eventually pay.&lt;/li&gt;
&lt;li&gt;Knowing what can fail. Not just the optimistic path, but the realistic failure modes: misaligned goals, unexpected agency, resource competition, cascading effects on infrastructure or society.&lt;/li&gt;
&lt;li&gt;Rolling back safely. Having the technical and organizational ability to slow, contain, or reverse course when evidence shows the system is behaving in ways that were not intended.&lt;/li&gt;
&lt;li&gt;Taking ownership when production breaks. The people closest to the systems carry a responsibility that cannot be fully outsourced to regulation, public statements, or future research.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are procedural disciplines. They do not guarantee any particular outcome. They raise the chance that capability can be expanded without quietly accumulating irreversible risk. They also force honesty about the limits of current methods. When insiders say the stakes feel civilizational, the productive response is not dismissal or panic. It is to treat the claim as a serious hypothesis that needs better evidence, better measurement, and better coordination—while continuing to ship useful systems under real constraints.&lt;/p&gt;

&lt;h3&gt;
  
  
  Parameters that actually move the needle
&lt;/h3&gt;

&lt;p&gt;Energy quality matters more than raw quantity. Dispatchable low-carbon sources—nuclear restarts, small modular reactors, geothermal, hybrids of renewables and storage—will shape where the most capable systems can be trained and run. Efficiency across layers (specialized hardware, model design, operational scheduling, utilization) multiplies the value of every watt. Water, land, and materials act as co-equal constraints. Economics that begin to price intensity into the cost of tokens and access will change incentives. Measurement systems that track energy or carbon per unit of useful work, rather than tokens alone, create feedback that design teams can actually use.&lt;/p&gt;

&lt;p&gt;None of these parameters eliminates risk. Together they define the feasible region in which capability can grow. The organizations that treat them as first-class design inputs—rather than externalities to be managed later—are more likely to scale without exhausting the physical and social license they depend on.&lt;/p&gt;

&lt;h3&gt;
  
  
  A working posture
&lt;/h3&gt;

&lt;p&gt;The future of AI will probably look stratified and constrained rather than unlimited. Some systems will be extraordinarily capable and tightly coupled to large energy and compute resources. Many others will be smaller, more specialized, and more widely distributed. Progress on alignment and control will continue, but so will the competitive pressure that makes deliberate slowdowns difficult. Insider warnings about large-scale failure modes will keep surfacing because the underlying technical and incentive problems have not been fully resolved.&lt;/p&gt;

&lt;p&gt;The useful stance is procedural rather than prophetic. Understand the blast radius of what is being built. Make the tradeoffs explicit. Catalog the ways things can fail. Design for safe rollback. Take ownership when the system behaves outside its intended envelope. Treat energy, water, efficiency, measurement, and control as co-equal design parameters alongside accuracy and speed.&lt;/p&gt;

&lt;p&gt;Science fiction prepared the imagination for both the helpful partner and the uncontrolled optimizer. The present is teaching that both possibilities remain open, and that the difference between them will be determined less by grand declarations than by the daily discipline of systems that can fail—and of the people who choose to own that failure mode before it becomes irreversible.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>systemdesign</category>
      <category>sustainability</category>
      <category>science</category>
    </item>
    <item>
      <title>[Boost]</title>
      <dc:creator>Adreet Gogoi</dc:creator>
      <pubDate>Mon, 21 Sep 2026 07:51:39 +0000</pubDate>
      <link>https://dev.to/adreetgog/-2n97</link>
      <guid>https://dev.to/adreetgog/-2n97</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/devteam/join-the-sanity-challenge-2500-in-prizes-for-five-winners-514m" class="crayons-story__hidden-navigation-link"&gt;Join the Sanity Challenge: $2,500 in prizes for FIVE winners!&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;
          &lt;a class="crayons-logo crayons-logo--l" href="/devteam"&gt;
            &lt;img alt="The DEV Team logo" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Forganization%2Fprofile_image%2F1%2Fd908a186-5651-4a5a-9f76-15200bc6801f.jpg" class="crayons-logo__image" width="800" height="800"&gt;
          &lt;/a&gt;

          &lt;a href="/heyitsjem" class="crayons-avatar  crayons-avatar--s absolute -right-2 -bottom-2 border-solid border-2 border-base-inverted  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1875101%2Ffed41ffa-3228-4a95-8381-2562d87d01ba.jpg" alt="heyitsjem profile" class="crayons-avatar__image" width="300" height="301"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/heyitsjem" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Jem
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Jem
                
                &lt;img alt="Community Curator" class="community-leader-icon" src="https://assets.dev.to/assets/community-leader-icon-b72c9e74eff54916e5c46c962f47ba40c9f611a71f8b157511f9613f69c0001b.svg" width="20" height="20"&gt;
              
              &lt;div id="story-author-preview-content-4616854" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/heyitsjem" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1875101%2Ffed41ffa-3228-4a95-8381-2562d87d01ba.jpg" class="crayons-avatar__image" alt="" width="300" height="301"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Jem&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

            &lt;span&gt;
              &lt;span class="crayons-story__tertiary fw-normal"&gt; for &lt;/span&gt;&lt;a href="/devteam" class="crayons-story__secondary fw-medium"&gt;The DEV Team&lt;/a&gt;
            &lt;/span&gt;
          &lt;/div&gt;
          &lt;a href="https://dev.to/devteam/join-the-sanity-challenge-2500-in-prizes-for-five-winners-514m" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Sep 18&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/devteam/join-the-sanity-challenge-2500-in-prizes-for-five-winners-514m" id="article-link-4616854"&gt;
          Join the Sanity Challenge: $2,500 in prizes for FIVE winners!
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/sanitychallenge"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;sanitychallenge&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/devchallenge"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;devchallenge&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/agents"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;agents&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/webdev"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;webdev&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/devteam/join-the-sanity-challenge-2500-in-prizes-for-five-winners-514m" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/multi-unicorn-b44d6f8c23cdd00964192bedc38af3e82463978aa611b4365bd33a0f1f4f3e97.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/fire-f60e7a582391810302117f987b22a8ef04a2fe0df7e3258a5f49332df1cec71e.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;150&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;span class="crayons-story__favorited"&gt;
              &lt;span class="favorited-marker"&gt;
                &lt;span&gt;
                  

                &lt;/span&gt;
                &lt;span class="hidden"&gt;
                  

                &lt;/span&gt;
              &lt;/span&gt;
            &lt;/span&gt;
            &lt;a href="https://dev.to/devteam/join-the-sanity-challenge-2500-in-prizes-for-five-winners-514m#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              25&lt;span class="hidden s:inline"&gt;&amp;nbsp;comments&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            7 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>The old SRE is dead.</title>
      <dc:creator>Adreet Gogoi</dc:creator>
      <pubDate>Thu, 17 Sep 2026 10:58:00 +0000</pubDate>
      <link>https://dev.to/adreetgog/the-old-sre-is-dead-nk8</link>
      <guid>https://dev.to/adreetgog/the-old-sre-is-dead-nk8</guid>
      <description>&lt;p&gt;If you’re in SRE, DevOps, Platform, or Cloud Engineering and still thinking in the old loop of&lt;br&gt;&lt;br&gt;
&lt;strong&gt;Monitor → Alert → Investigate → Fix&lt;/strong&gt;,&lt;br&gt;&lt;br&gt;
you’re already behind the systems you’re supposed to keep reliable.&lt;/p&gt;

&lt;p&gt;The traditional model assumed systems were mostly deterministic.&lt;br&gt;&lt;br&gt;
Something broke → you got an alert → you found the root cause → you fixed it.&lt;br&gt;&lt;br&gt;
That world is disappearing.&lt;/p&gt;

&lt;p&gt;Here’s the new operating model that is quietly taking over:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Observe → Understand → Predict → Decide → Automate → Verify&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This isn’t just a nicer flowchart. It’s a fundamental shift in how reliability is achieved.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. AI workloads broke the old assumptions
&lt;/h3&gt;

&lt;p&gt;LLMs, agents, and inference pipelines are not “experiments” anymore. They are production systems with non-deterministic behavior.&lt;/p&gt;

&lt;p&gt;Traditional monitoring still works for CPU, memory, and latency.&lt;br&gt;&lt;br&gt;
It completely fails when the failure mode is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The model started hallucinating more after a quiet dependency update
&lt;/li&gt;
&lt;li&gt;Token cost exploded because of a prompt change no one noticed
&lt;/li&gt;
&lt;li&gt;Output quality drifted even though every service is “green”
&lt;/li&gt;
&lt;li&gt;An agent started looping or making irrational decisions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You can no longer treat model accuracy, hallucination rate, output consistency, and cost-per-request as second-class metrics. They are now first-class reliability signals — the same way availability and error rate used to be.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Observability is no longer just visibility. It’s becoming the control plane
&lt;/h3&gt;

&lt;p&gt;Classic monitoring answers: “CPU is at 90%.”&lt;br&gt;&lt;br&gt;
Modern observability must answer:  &lt;/p&gt;

&lt;p&gt;“Why did this start happening 17 minutes ago?&lt;br&gt;&lt;br&gt;
Which service or model version is involved?&lt;br&gt;&lt;br&gt;
Did a deployment, data drift, or external API change cause it?&lt;br&gt;&lt;br&gt;
What is the business impact right now?”&lt;/p&gt;

&lt;p&gt;The future stack is not just Logs + Metrics + Traces.&lt;br&gt;&lt;br&gt;
It is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Logs + Metrics + Traces + Topology + AI signals + Business context → one intelligence layer&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When observability can reason across all of these, it stops being a passive dashboard and starts becoming the system that decides &lt;em&gt;when&lt;/em&gt; and &lt;em&gt;how&lt;/em&gt; to act.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. SLOs must expand or they become irrelevant
&lt;/h3&gt;

&lt;p&gt;A 99.95% availability SLO still matters.&lt;br&gt;&lt;br&gt;
But for AI systems it is no longer sufficient.&lt;/p&gt;

&lt;p&gt;You now need multi-dimensional SLOs that include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Accuracy / quality
&lt;/li&gt;
&lt;li&gt;Latency (including inference latency)
&lt;/li&gt;
&lt;li&gt;Reliability of the decision itself
&lt;/li&gt;
&lt;li&gt;Cost efficiency
&lt;/li&gt;
&lt;li&gt;Output consistency over time
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Because AI systems can change their behavior &lt;em&gt;after&lt;/em&gt; deployment without any code change, continuous verification becomes mandatory. You don’t just ship and monitor. You continuously test whether the system is still behaving the way it was supposed to.&lt;/p&gt;




&lt;p&gt;The old SRE mindset optimized for &lt;em&gt;keeping things up&lt;/em&gt;.&lt;br&gt;&lt;br&gt;
The new SRE mindset optimizes for &lt;em&gt;keeping things correct, predictable, and economically sane&lt;/em&gt; in a world where the systems themselves are learning and changing.&lt;/p&gt;

&lt;p&gt;This is not optional knowledge for people who want to stay relevant in reliability engineering.&lt;br&gt;&lt;br&gt;
This is the new baseline.&lt;/p&gt;

&lt;p&gt;If you’re still only watching infrastructure metrics, you’re watching the wrong dashboard.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>sre</category>
      <category>sla</category>
      <category>opensource</category>
    </item>
    <item>
      <title>From Physical Racks to Intelligent Modules: How the Meaning of Infrastructure Has Fundamentally Shifted</title>
      <dc:creator>Adreet Gogoi</dc:creator>
      <pubDate>Sun, 13 Sep 2026 08:54:41 +0000</pubDate>
      <link>https://dev.to/adreetgog/from-physical-racks-to-intelligent-modules-how-the-meaning-of-infrastructure-has-fundamentally-2ldd</link>
      <guid>https://dev.to/adreetgog/from-physical-racks-to-intelligent-modules-how-the-meaning-of-infrastructure-has-fundamentally-2ldd</guid>
      <description>&lt;p&gt;Infrastructure used to be something you could walk up to, touch, and hear humming in a raised-floor room. Capacity planning dominated the conversation: How many servers? How much power and cooling headroom? How many years until the next major refresh? Today, for many architects, infrastructure is less about forecasting static capacity and more about composing modular building blocks intelligently—and extracting far more value from every single one of them.&lt;/p&gt;

&lt;p&gt;That shift is not merely technological. It is a redefinition of what “infrastructure” even means.&lt;/p&gt;

&lt;h3&gt;
  
  
  Then: Infrastructure as Place, Asset, and Capacity Plan
&lt;/h3&gt;

&lt;p&gt;In the early days of enterprise computing, infrastructure was concrete. Mainframes occupied climate-controlled rooms. Later, dedicated data centers became the physical manifestation of IT capability. Power, cooling, cabling, racks, switches, and servers were the primary concerns. Network design meant physical topology and latency budgets measured in meters of copper or fiber. Capacity planning meant ordering hardware months (sometimes years) in advance and living with those decisions for a long time.&lt;/p&gt;

&lt;p&gt;Cost efficiency lived in procurement negotiations, power usage effectiveness (PUE), and the art of densifying racks without tripping thermal or structural limits. Reliability was engineered through redundancy at every layer. The architect’s job was largely to design a resilient &lt;em&gt;place&lt;/em&gt; that could host applications and to size it correctly so it neither ran out of capacity nor sat wastefully oversized.&lt;/p&gt;

&lt;p&gt;Ownership was clear. You bought or leased the hardware, controlled the physical security perimeter, and absorbed the constraints of that physical plant.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Intermediate Layer: Virtualization and the First Real Modularity
&lt;/h3&gt;

&lt;p&gt;Virtualization (and later containers and orchestration) began the great decoupling. The same physical server could host many logical machines. Storage became pools. Networks acquired overlays. Suddenly the unit of work moved upward from the individual server or switch to more flexible, shareable resources.&lt;/p&gt;

&lt;p&gt;This period introduced the first meaningful tension in definition: was infrastructure the hardware, or the virtual resources carved from it? Most organizations answered “both.” Hybrid operating models emerged. The physical layer remained critical for latency-sensitive, regulated, or high-throughput workloads, while the logical layer accelerated delivery and improved utilization. Capacity planning was still important, but it started to operate at multiple levels—physical headroom underneath, and more elastic logical capacity on top.&lt;/p&gt;

&lt;h3&gt;
  
  
  Now: Infrastructure as Composable Modules and Intelligent System Design
&lt;/h3&gt;

&lt;p&gt;Today the dominant mental model is infrastructure as a set of modular, programmable building blocks. Cloud platforms, software-defined everything, Infrastructure as Code, and platform engineering have turned what used to be long-lived assets into composable capabilities. The conversation has moved from “Do we have enough capacity?” toward “How do we assemble the right modules—with the right isolation, locality, performance, cost, and policy characteristics—for this workload?”&lt;/p&gt;

&lt;p&gt;Yet the market conversation still often defaults to compute—more cores, more GPUs, bigger clusters. That focus is incomplete. The more interesting shift is happening &lt;em&gt;inside&lt;/em&gt; the box and across the system: recent technologies and engineering practices are focused on extracting dramatically more useful work from the same GPU, the same server, or the same node. Better scheduling, smarter resource sharing, finer-grained isolation, improved memory and interconnect utilization, and careful co-design of software with hardware allow teams to balance and run more applications—or significantly heavier ones—on the same underlying compute.&lt;/p&gt;

&lt;p&gt;This is fundamentally a system-design problem, not just a capacity problem:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;From capacity forecasting to modular composition and density.&lt;/strong&gt; Classic capacity planning still matters at the physical substrate (especially for AI clusters, high-density compute, or regulated environments). But for most teams the daily work is assembling and governing modules while maximizing what each module can deliver. Understanding the blocks deeply—their performance envelopes, coupling points, operational behaviors, and true cost drivers—matters more than simply adding more of them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;From raw compute to intelligent utilization.&lt;/strong&gt; The same GPU or server can now support higher concurrent workloads, better multi-tenancy, or more demanding applications because the surrounding system (orchestration, networking, storage paths, scheduling, and application design) has become more sophisticated. Efficiency gains come from system-level thinking, not just faster silicon.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;From static topology to intent, constraints, and balance.&lt;/strong&gt; Modern designs declare desired outcomes (latency envelopes, data residency, blast-radius limits, cost ceilings, fairness across workloads) and rely on control planes and careful architecture to realize them. Balancing multiple heavy applications on shared infrastructure requires deliberate system design: isolation boundaries, priority and preemption policies, observability that surfaces contention, and architectures that avoid hidden coupling.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;From pure technical sizing to multi-dimensional intelligence.&lt;/strong&gt; Choosing and assembling modules now requires understanding not just raw capacity but utilization efficiency, carbon intensity, data-transfer economics, operational complexity, and the cognitive load on the teams that will run the system. Intelligence at the modular and system level means knowing which combinations create fragile contention and which create resilient, high-utilization platforms.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Physical infrastructure has not disappeared. Hyperscale data centers, specialized AI training clusters, and carefully engineered on-premises environments still demand deep hardware, power, cooling, and interconnect expertise. The difference is that these physical realities increasingly sit underneath a modular abstraction layer, and the highest leverage often comes from designing the system so that each physical or virtual block is used more completely and more intelligently.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What Has Not Changed&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Some fundamentals remain stubbornly physical. Electrons need paths. Heat must be removed. Light travels at finite speed. Regulatory boundaries still force locality decisions. Large-scale AI infrastructure continues to push power delivery, cooling, and fabric design in ways that feel very much like classic data-center problems—only denser and more expensive.&lt;/p&gt;

&lt;p&gt;The best practitioners keep one foot in each world. They can still design a resilient physical fabric when required, and they can also reason clearly about modular boundaries, interfaces, composition, and the system-level techniques that extract more useful work from every GPU and every server. They treat capacity not as the primary question but as one constraint among many that intelligent system design must satisfy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A Working Definition for Today&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Infrastructure is the set of modular capabilities—compute, storage, networking, identity, observability, and policy—required to run workloads reliably, securely, and economically. These capabilities are delivered through whatever combination of physical assets, virtual resources, and managed services best meets the constraints of the moment. The architect’s core responsibility is no longer primarily capacity planning or simply adding more compute. It is understanding the blocks deeply enough to compose them intelligently, designing the system so that the same compute power can productively support more (and heavier) applications, keeping failure domains contained, and evolving the platform without constant reinvention.&lt;/p&gt;

&lt;p&gt;It is no longer primarily a place or a fixed capacity plan. It is a set of well-understood modules and the intelligent system design that assembles and balances them under real-world constraints.&lt;/p&gt;

&lt;p&gt;The definition of infrastructure has expanded and become more abstract. The need for clear thinking about modularity, utilization, system design, and the physical realities that still underpin everything has not.&lt;/p&gt;

&lt;p&gt;What does “infrastructure” mean in your organization today—and how much of the real leverage now comes from smarter system design rather than simply more compute?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>infrastructure</category>
      <category>capacity</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>AI Makes Us Faster. Systems Thinking Makes Us Better.</title>
      <dc:creator>Adreet Gogoi</dc:creator>
      <pubDate>Fri, 11 Sep 2026 12:38:54 +0000</pubDate>
      <link>https://dev.to/adreetgog/ai-makes-us-faster-systems-thinking-makes-us-better-dl</link>
      <guid>https://dev.to/adreetgog/ai-makes-us-faster-systems-thinking-makes-us-better-dl</guid>
      <description>&lt;p&gt;AI is changing the way we work. But I think the bigger change is happening somewhere else:&lt;/p&gt;

&lt;p&gt;It is changing the way we need to think about technology.&lt;/p&gt;

&lt;p&gt;For years, many of us learned technology from the top down.&lt;/p&gt;

&lt;p&gt;Application → Database → Operating System → Virtual Machine → Infrastructure → Network.&lt;/p&gt;

&lt;p&gt;You write code, deploy an application, connect a database, and only look deeper when something breaks. AI is slowly forcing us to look in the opposite direction. Because an AI workload is not just an application.&lt;/p&gt;

&lt;p&gt;It can be:&lt;/p&gt;

&lt;p&gt;AI model → GPU → Server → Network → Storage → Power → Cooling → Data Center → Physical Infrastructure&lt;/p&gt;

&lt;p&gt;Suddenly, the layers underneath the application matter a lot more and this isn't only an infrastructure problem. It is becoming a technology problem for everyone.AI doesn't remove the system.&lt;/p&gt;

&lt;p&gt;Imagine an application suddenly becomes slow.&lt;/p&gt;

&lt;p&gt;The first thought might be:&lt;/p&gt;

&lt;p&gt;"Something is wrong with the application."&lt;/p&gt;

&lt;p&gt;But what if the database is waiting on storage? What if storage latency is caused by an overloaded host? What if the host is competing for resources? What if network congestion is involved?&lt;br&gt;
What if the workload has simply outgrown the infrastructure?&lt;/p&gt;

&lt;p&gt;AI can help investigate all of these possibilities incredibly quickly.&lt;/p&gt;

&lt;p&gt;It can read logs. It can correlate symptoms. It can generate commands.It can suggest hypotheses.It can compare architectures.&lt;br&gt;
It can even automate parts of the investigation.&lt;/p&gt;

&lt;p&gt;But there is still one important question:&lt;/p&gt;

&lt;p&gt;Does the person using AI understand what they are looking at?&lt;/p&gt;

&lt;p&gt;That is where systems thinking becomes extremely valuable.&lt;/p&gt;

&lt;p&gt;The new skill isn't just prompting. There is a lot of discussion around "prompt engineering."&lt;/p&gt;

&lt;p&gt;But in technical work, I think something deeper matters.&lt;/p&gt;

&lt;p&gt;System understanding.&lt;/p&gt;

&lt;p&gt;A good engineer can tell AI:&lt;/p&gt;

&lt;p&gt;what the architecture looks like&lt;br&gt;
what changed&lt;br&gt;
what is expected&lt;br&gt;
what is actually happening&lt;br&gt;
which metrics matter&lt;br&gt;
which logs are relevant&lt;br&gt;
what constraints exist&lt;br&gt;
what has already been tested&lt;br&gt;
what cannot be changed&lt;br&gt;
and what the risk of a wrong decision is&lt;/p&gt;

&lt;p&gt;That creates a very different interaction with AI.&lt;/p&gt;

&lt;p&gt;Instead of:&lt;/p&gt;

&lt;p&gt;"My server is slow. Fix it."&lt;/p&gt;

&lt;p&gt;You can say:&lt;/p&gt;

&lt;p&gt;"The application latency increased after deployment. CPU is normal, memory is stable, but database queries are showing increased I/O wait. Storage latency has also increased on the underlying VM host. Here are the relevant metrics and logs. What are the most likely failure paths, and what evidence would distinguish them?"&lt;/p&gt;

&lt;p&gt;Now AI becomes much more useful.&lt;/p&gt;

&lt;p&gt;Because context is intelligence.&lt;/p&gt;

&lt;p&gt;For systems, network and DevOps engineers, this doesn't mean everyone needs to become an expert in everything.&lt;/p&gt;

&lt;p&gt;A systems engineer doesn't need to become a database specialist.&lt;/p&gt;

&lt;p&gt;A network engineer doesn't need to become a kernel developer.&lt;/p&gt;

&lt;p&gt;A software engineer doesn't need to become a data-center engineer.&lt;/p&gt;

&lt;p&gt;But understanding how the pieces interact is becoming increasingly important.&lt;/p&gt;

&lt;p&gt;For example, engineers working around infrastructure should increasingly be comfortable thinking about:&lt;/p&gt;

&lt;p&gt;TCP/IP and network behaviour&lt;br&gt;
DNS, routing and load balancing&lt;br&gt;
Linux internals and resource limits&lt;br&gt;
CPU, memory and I/O behaviour&lt;br&gt;
virtualization and containers&lt;br&gt;
storage performance and failure modes&lt;br&gt;
databases and replication&lt;br&gt;
observability and telemetry&lt;br&gt;
distributed systems&lt;br&gt;
cloud architecture&lt;br&gt;
security boundaries&lt;br&gt;
automation and IaC&lt;br&gt;
CI/CD pipelines&lt;br&gt;
power, cooling and physical capacity&lt;br&gt;
and increasingly, GPUs and AI infrastructure&lt;/p&gt;

&lt;p&gt;You don't need mastery of every layer.&lt;/p&gt;

&lt;p&gt;You need to understand where your layer ends and another one begins.&lt;/p&gt;

&lt;p&gt;AI is also creating infrastructure&lt;/p&gt;

&lt;p&gt;There is an interesting contradiction happening.&lt;/p&gt;

&lt;p&gt;AI can automate parts of existing jobs.&lt;/p&gt;

&lt;p&gt;At the same time, AI is creating completely new technical requirements.&lt;/p&gt;

&lt;p&gt;More models mean more GPUs.&lt;/p&gt;

&lt;p&gt;More GPUs mean more servers.&lt;/p&gt;

&lt;p&gt;More servers mean more networking, storage, power and cooling.&lt;/p&gt;

&lt;p&gt;That creates demand for new infrastructure, new platforms, new tooling, new observability systems, new security models and new engineering roles.&lt;/p&gt;

&lt;p&gt;So the question isn't simply:&lt;/p&gt;

&lt;p&gt;"Will AI take jobs?"&lt;/p&gt;

&lt;p&gt;The harder question is:&lt;/p&gt;

&lt;p&gt;"Can new human demand and new economic activity grow fast enough to balance the work that automation removes?"&lt;/p&gt;

&lt;p&gt;That's not purely a technology question.&lt;/p&gt;

&lt;p&gt;It's an economic one.&lt;/p&gt;

&lt;p&gt;Productivity can increase dramatically, but people still need income and purchasing power to create demand for the products and services that new technology enables.&lt;/p&gt;

&lt;p&gt;Automation changes supply.&lt;br&gt;
Human purchasing power creates demand.&lt;/p&gt;

&lt;p&gt;The balance between the two matters.&lt;/p&gt;

&lt;p&gt;And this is where the conversation becomes much bigger than "AI versus humans."&lt;/p&gt;

&lt;p&gt;So where does the human fit?&lt;/p&gt;

&lt;p&gt;Maybe the future isn't about humans doing everything manually.&lt;/p&gt;

&lt;p&gt;And it probably isn't about AI doing everything either.&lt;/p&gt;

&lt;p&gt;It may be about human interpretation + machine acceleration.&lt;/p&gt;

&lt;p&gt;AI can generate ten possible explanations in seconds.&lt;/p&gt;

&lt;p&gt;The engineer has to determine which one actually makes sense.&lt;/p&gt;

&lt;p&gt;AI can generate a command.&lt;/p&gt;

&lt;p&gt;The engineer has to decide whether running it in production is safe.&lt;/p&gt;

&lt;p&gt;AI can propose an architecture.&lt;/p&gt;

&lt;p&gt;The engineer has to understand whether it fits the real system, budget, constraints and failure scenarios.&lt;/p&gt;

&lt;p&gt;AI can analyse thousands of lines of logs.&lt;/p&gt;

&lt;p&gt;The engineer still needs to know what question to ask.&lt;/p&gt;

&lt;p&gt;That's the gap.&lt;/p&gt;

&lt;p&gt;Not necessarily intelligence.&lt;/p&gt;

&lt;p&gt;Context. Responsibility. Judgment.&lt;/p&gt;

&lt;p&gt;"Jitna tez AI, utni tez galti bhi ho sakti hai."&lt;/p&gt;

&lt;p&gt;There is a simple idea here:&lt;/p&gt;

&lt;p&gt;If you don't understand the system, AI can simply help you make mistakes faster.&lt;/p&gt;

&lt;p&gt;That's why technical fundamentals aren't becoming obsolete.&lt;/p&gt;

&lt;p&gt;They may actually become more valuable.&lt;/p&gt;

&lt;p&gt;The engineer who understands systems can use AI as a force multiplier.&lt;/p&gt;

&lt;p&gt;The engineer who doesn't may simply generate more things they don't understand.&lt;/p&gt;

&lt;p&gt;The direction I see&lt;/p&gt;

&lt;p&gt;Perhaps we should stop thinking about technology as isolated layers.&lt;/p&gt;

&lt;p&gt;Instead, think about it as one connected system:&lt;/p&gt;

&lt;p&gt;Physical → Power &amp;amp; Cooling → Hardware → Network → Storage → OS → Virtualization → Platform → Database → Application → AI&lt;/p&gt;

&lt;p&gt;You don't have to become an expert at every level.&lt;/p&gt;

&lt;p&gt;But knowing that these levels exist—and understanding how failures travel between them—is becoming increasingly important.&lt;/p&gt;

&lt;p&gt;Because a slow application might not be an application problem.&lt;/p&gt;

&lt;p&gt;A failed AI workload might not be a model problem.&lt;/p&gt;

&lt;p&gt;A network issue might not be a network problem.&lt;/p&gt;

&lt;p&gt;And an infrastructure problem might eventually turn out to be a physical capacity problem.&lt;/p&gt;

&lt;p&gt;What you see is often just the surface.&lt;/p&gt;

&lt;p&gt;The deeper you understand the system, the better questions you can ask.&lt;/p&gt;

&lt;p&gt;And the better questions you ask, the more useful AI becomes.&lt;/p&gt;

&lt;p&gt;Maybe that's the real shift.&lt;/p&gt;

&lt;p&gt;AI makes us faster.&lt;br&gt;
Systems thinking makes us better.&lt;/p&gt;

&lt;p&gt;The future may not belong to people who know every technology.&lt;/p&gt;

&lt;p&gt;It may belong to people who understand how the technologies connect—and know when to let the machine help, and when to question it.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>systemdesign</category>
      <category>infrastructure</category>
      <category>data</category>
    </item>
    <item>
      <title>When a Tableau TSM HTTP 500 Wasn't Really an HTTP Problem</title>
      <dc:creator>Adreet Gogoi</dc:creator>
      <pubDate>Thu, 10 Sep 2026 17:33:20 +0000</pubDate>
      <link>https://dev.to/adreetgog/when-a-tableau-tsm-http-500-wasnt-really-an-http-problem-1kf9</link>
      <guid>https://dev.to/adreetgog/when-a-tableau-tsm-http-500-wasnt-really-an-http-problem-1kf9</guid>
      <description>&lt;p&gt;A real-world troubleshooting journey through Tableau Server, ZooKeeper, systemd, and a misleading No space left on device error.&lt;/p&gt;

&lt;p&gt;Sometimes infrastructure incidents are easy.&lt;/p&gt;

&lt;p&gt;A service is down.&lt;br&gt;
A port is closed.&lt;br&gt;
The logs tell you exactly what happened.&lt;/p&gt;

&lt;p&gt;And then there are incidents where everything looks almost healthy.&lt;/p&gt;

&lt;p&gt;That was the case with a Tableau Server environment I was troubleshooting.&lt;/p&gt;

&lt;p&gt;The symptom was straightforward:&lt;/p&gt;

&lt;p&gt;Tableau Services Manager (TSM) was returning HTTP 500 on port 8850.&lt;/p&gt;

&lt;p&gt;The interesting part was discovering that the HTTP 500 wasn't really the problem.&lt;/p&gt;

&lt;p&gt;The Architecture&lt;/p&gt;

&lt;p&gt;At a simplified level, the dependency chain looked like this:&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;              Administrator
                   │
                   ▼
          Tableau TSM :8850
                   │
                   ▼
          TabadminController
                   │
                   ▼
         Tableau coordination
                   │
                   ▼
              ZooKeeper
                :8866
                   │
                   ▼
         Persistent storage
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;So rather than immediately restarting Tableau, I started at the symptom and worked down the dependency chain.&lt;/p&gt;

&lt;p&gt;Step 1: Is TSM Actually Down?&lt;/p&gt;

&lt;p&gt;The TSM CLI was reporting:&lt;/p&gt;

&lt;p&gt;Creating ServerApi object - host:ubuntu, port:8850&lt;/p&gt;

&lt;p&gt;Checking server health at &lt;a href="https://ubuntu:8850/health" rel="noopener noreferrer"&gt;https://ubuntu:8850/health&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Client request:&lt;br&gt;
GET &lt;a href="https://ubuntu:8850/api/0.5/status" rel="noopener noreferrer"&gt;https://ubuntu:8850/api/0.5/status&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;TabadminController did not return a TsmResponse.&lt;/p&gt;

&lt;p&gt;500 - Server Error&lt;/p&gt;

&lt;p&gt;The first question was simple:&lt;/p&gt;

&lt;p&gt;Is anything actually listening on 8850?&lt;/p&gt;

&lt;p&gt;sudo ss -lntp | grep 8850&lt;/p&gt;

&lt;p&gt;It showed Tableau's tabadmincontrol process listening on the port.&lt;/p&gt;

&lt;p&gt;So this wasn't a simple:&lt;/p&gt;

&lt;p&gt;Port closed → service down&lt;/p&gt;

&lt;p&gt;Something was alive.&lt;/p&gt;

&lt;p&gt;Step 2: Does the HTTPS Service Respond?&lt;/p&gt;

&lt;p&gt;I tested the endpoint locally:&lt;/p&gt;

&lt;p&gt;curl -k -s &lt;a href="https://localhost:8850/health" rel="noopener noreferrer"&gt;https://localhost:8850/health&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It returned:&lt;/p&gt;

&lt;p&gt;{&lt;br&gt;
  "status": 403,&lt;br&gt;
  "error": "Forbidden"&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;A 403 might initially look like another failure.&lt;/p&gt;

&lt;p&gt;But it actually told us something important:&lt;/p&gt;

&lt;p&gt;The TabadminController was reachable and responding.&lt;/p&gt;

&lt;p&gt;So the investigation moved further down the stack.&lt;/p&gt;

&lt;p&gt;This is one of the most useful troubleshooting habits:&lt;/p&gt;

&lt;p&gt;Don't just look at the error code. Ask what the response proves.&lt;/p&gt;

&lt;p&gt;In this case, 403 proved that the HTTPS service was alive.&lt;/p&gt;

&lt;p&gt;Step 3: Follow the Dependency Chain&lt;/p&gt;

&lt;p&gt;Tableau Server depends on several internal services, including its coordination layer.&lt;/p&gt;

&lt;p&gt;The Tableau-managed ZooKeeper instance was:&lt;/p&gt;

&lt;p&gt;appzookeeper_0&lt;/p&gt;

&lt;p&gt;I checked its Tableau service status:&lt;/p&gt;

&lt;p&gt;sudo /var/opt/tableau/tableau_server/data/tabsvc/services/appzookeeper_0.20253.26.0206.0336/status.sh&lt;/p&gt;

&lt;p&gt;The important part of the output was:&lt;/p&gt;

&lt;p&gt;"details" : {&lt;br&gt;
  "message" : "Connection refused"&lt;br&gt;
},&lt;br&gt;
"processStatus" : "DOWN"&lt;/p&gt;

&lt;p&gt;That was our first strong indication that the problem was deeper than TSM.&lt;/p&gt;

&lt;p&gt;Next:&lt;/p&gt;

&lt;p&gt;sudo ss -lntp | grep 8866&lt;/p&gt;

&lt;p&gt;Nothing was listening.&lt;/p&gt;

&lt;p&gt;So we now had:&lt;/p&gt;

&lt;p&gt;TSM :8850&lt;br&gt;
     │&lt;br&gt;
     ▼&lt;br&gt;
TabadminController&lt;br&gt;
     │&lt;br&gt;
     ▼&lt;br&gt;
ZooKeeper&lt;br&gt;
     │&lt;br&gt;
     ✕&lt;br&gt;
   :8866&lt;/p&gt;

&lt;p&gt;The HTTP 500 was beginning to look like a downstream dependency failure.&lt;/p&gt;

&lt;p&gt;Step 4: The Logs Finally Told the Story&lt;/p&gt;

&lt;p&gt;This was the turning point.&lt;/p&gt;

&lt;p&gt;Instead of continuing to restart services, I went to the ZooKeeper logs.&lt;/p&gt;

&lt;p&gt;There it was:&lt;/p&gt;

&lt;p&gt;Severe unrecoverable error, from thread : SyncThread:0&lt;/p&gt;

&lt;p&gt;java.io.IOException: No space left on device&lt;/p&gt;

&lt;p&gt;The stack trace pointed into ZooKeeper's transaction-log handling:&lt;/p&gt;

&lt;p&gt;org.apache.zookeeper.server.persistence.FileTxnLog.append(...)&lt;/p&gt;

&lt;p&gt;Immediately afterwards:&lt;/p&gt;

&lt;p&gt;SyncThread:0 ... Thread exits, error code 1&lt;/p&gt;

&lt;p&gt;Then:&lt;/p&gt;

&lt;p&gt;SingleNodeZookeeperWrapper - Attempting to shutdown zookeeper instance.&lt;/p&gt;

&lt;p&gt;And finally:&lt;/p&gt;

&lt;p&gt;NettyServerCnxnFactory - shutdown called ... :8866&lt;/p&gt;

&lt;p&gt;Now the failure chain was clear:&lt;/p&gt;

&lt;p&gt;ZooKeeper tries to write transaction log&lt;br&gt;
                │&lt;br&gt;
                ▼&lt;br&gt;
       Filesystem returns ENOSPC&lt;br&gt;
                │&lt;br&gt;
                ▼&lt;br&gt;
       ZooKeeper critical thread exits&lt;br&gt;
                │&lt;br&gt;
                ▼&lt;br&gt;
          ZooKeeper shuts down&lt;br&gt;
                │&lt;br&gt;
                ▼&lt;br&gt;
          Port 8866 disappears&lt;br&gt;
                │&lt;br&gt;
                ▼&lt;br&gt;
 Tableau coordination becomes unavailable&lt;br&gt;
                │&lt;br&gt;
                ▼&lt;br&gt;
       TSM API request fails&lt;br&gt;
                │&lt;br&gt;
                ▼&lt;br&gt;
             HTTP 500&lt;/p&gt;

&lt;p&gt;That was the real story.&lt;/p&gt;

&lt;p&gt;The Strange Part: The Disk Wasn't Full&lt;/p&gt;

&lt;p&gt;Naturally, the next command was:&lt;/p&gt;

&lt;p&gt;df -h&lt;/p&gt;

&lt;p&gt;And the result was surprising:&lt;/p&gt;

&lt;p&gt;/dev/vda1    991G    181G    811G    19%    /&lt;/p&gt;

&lt;p&gt;There was 811 GB free.&lt;/p&gt;

&lt;p&gt;So what exactly did No space left on device mean?&lt;/p&gt;

&lt;p&gt;This is where the investigation needs to be precise.&lt;/p&gt;

&lt;p&gt;ENOSPC does not necessarily mean the entire filesystem has reached 100% capacity.&lt;/p&gt;

&lt;p&gt;Possible explanations include:&lt;/p&gt;

&lt;p&gt;inode exhaustion&lt;br&gt;
filesystem quotas&lt;br&gt;
a full temporary filesystem&lt;br&gt;
application-specific storage limits&lt;br&gt;
deleted-but-open files&lt;br&gt;
filesystem or storage errors&lt;/p&gt;

&lt;p&gt;So I checked additional evidence rather than simply declaring:&lt;/p&gt;

&lt;p&gt;"The disk was full."&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;df -i&lt;/p&gt;

&lt;p&gt;checks inode usage.&lt;/p&gt;

&lt;p&gt;And:&lt;/p&gt;

&lt;p&gt;sudo lsof +L1&lt;/p&gt;

&lt;p&gt;can reveal deleted files that processes still have open.&lt;/p&gt;

&lt;p&gt;In this environment, Tableau processes were holding deleted files open, including a large deleted Hyper temporary file.&lt;/p&gt;

&lt;p&gt;However, that alone was not enough to prove that those files directly caused ZooKeeper's ENOSPC event.&lt;/p&gt;

&lt;p&gt;That distinction matters.&lt;/p&gt;

&lt;p&gt;The evidence proves:&lt;/p&gt;

&lt;p&gt;ZooKeeper's transaction-log write received ENOSPC.&lt;/p&gt;

&lt;p&gt;It does not prove exactly which underlying storage condition produced that response.&lt;/p&gt;

&lt;p&gt;Good troubleshooting means knowing the difference between evidence and hypothesis.&lt;/p&gt;

&lt;p&gt;Another Trap: systemd Said ZooKeeper Was Running&lt;/p&gt;

&lt;p&gt;There was another confusing clue.&lt;/p&gt;

&lt;p&gt;Checking the Tableau user-level systemd service:&lt;/p&gt;

&lt;p&gt;sudo -u tableau \&lt;br&gt;
XDG_RUNTIME_DIR=/run/user/999 \&lt;br&gt;
systemctl --user status appzookeeper_0&lt;/p&gt;

&lt;p&gt;reported:&lt;/p&gt;

&lt;p&gt;Active: active (running)&lt;br&gt;
Main PID: 881&lt;/p&gt;

&lt;p&gt;But:&lt;/p&gt;

&lt;p&gt;sudo ss -lntp | grep 8866&lt;/p&gt;

&lt;p&gt;showed nothing.&lt;/p&gt;

&lt;p&gt;And Tableau's own status script said:&lt;/p&gt;

&lt;p&gt;processStatus : DOWN&lt;/p&gt;

&lt;p&gt;So which one was correct?&lt;/p&gt;

&lt;p&gt;The answer is: they were measuring different things.&lt;/p&gt;

&lt;p&gt;systemd knew that the process existed.&lt;/p&gt;

&lt;p&gt;Tableau knew that the application wasn't functioning.&lt;/p&gt;

&lt;p&gt;The port check knew that nothing was accepting connections.&lt;/p&gt;

&lt;p&gt;The logs explained why.&lt;/p&gt;

&lt;p&gt;This is why a single:&lt;/p&gt;

&lt;p&gt;systemctl status&lt;/p&gt;

&lt;p&gt;should never be treated as the complete definition of application health.&lt;/p&gt;

&lt;p&gt;I like to think of it as four independent signals:&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;      Process
         +
       Port
         +
  Application health
         +
        Logs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;When those disagree, investigate the disagreement.&lt;/p&gt;

&lt;p&gt;One More Lesson: start Is Not Always restart&lt;/p&gt;

&lt;p&gt;Because systemd believed the service was already running, simply doing:&lt;/p&gt;

&lt;p&gt;systemctl --user start appzookeeper_0&lt;/p&gt;

&lt;p&gt;didn't necessarily give us a fresh application process.&lt;/p&gt;

&lt;p&gt;The service manager already believed:&lt;/p&gt;

&lt;p&gt;Active: active (running)&lt;/p&gt;

&lt;p&gt;So the appropriate operational action was a restart, executed in the Tableau user's systemd context:&lt;/p&gt;

&lt;p&gt;sudo -u tableau \&lt;br&gt;
XDG_RUNTIME_DIR=/run/user/999 \&lt;br&gt;
systemctl --user restart appzookeeper_0&lt;/p&gt;

&lt;p&gt;The important part here isn't the command itself.&lt;/p&gt;

&lt;p&gt;It's understanding who owns the service.&lt;/p&gt;

&lt;p&gt;Using:&lt;/p&gt;

&lt;p&gt;sudo systemctl ...&lt;/p&gt;

&lt;p&gt;and:&lt;/p&gt;

&lt;p&gt;sudo -u tableau systemctl --user ...&lt;/p&gt;

&lt;p&gt;are not equivalent.&lt;/p&gt;

&lt;p&gt;A user-level systemd service belongs to its user session, so operating it in the wrong context can produce errors such as:&lt;/p&gt;

&lt;p&gt;Failed to connect to bus&lt;br&gt;
What I Took Away From the Incident&lt;/p&gt;

&lt;p&gt;The actual fix was only one part of the exercise.&lt;/p&gt;

&lt;p&gt;The bigger lesson was the troubleshooting methodology.&lt;/p&gt;

&lt;p&gt;When TSM reported:&lt;/p&gt;

&lt;p&gt;HTTP 500&lt;/p&gt;

&lt;p&gt;I could have immediately restarted Tableau.&lt;/p&gt;

&lt;p&gt;Instead, I asked:&lt;/p&gt;

&lt;p&gt;Is 8850 listening?&lt;br&gt;
        ↓&lt;br&gt;
Yes.&lt;/p&gt;

&lt;p&gt;Does the HTTPS service respond?&lt;br&gt;
        ↓&lt;br&gt;
Yes.&lt;/p&gt;

&lt;p&gt;Is ZooKeeper healthy?&lt;br&gt;
        ↓&lt;br&gt;
No.&lt;/p&gt;

&lt;p&gt;Is 8866 listening?&lt;br&gt;
        ↓&lt;br&gt;
No.&lt;/p&gt;

&lt;p&gt;Why did ZooKeeper stop?&lt;br&gt;
        ↓&lt;br&gt;
Read the logs.&lt;/p&gt;

&lt;p&gt;What happened?&lt;br&gt;
        ↓&lt;br&gt;
ENOSPC while writing the transaction log.&lt;/p&gt;

&lt;p&gt;Every step reduced the search space.&lt;/p&gt;

&lt;p&gt;What I Would Monitor in Production&lt;/p&gt;

&lt;p&gt;After an incident like this, monitoring should go beyond CPU, memory and overall disk usage.&lt;/p&gt;

&lt;p&gt;I'd monitor:&lt;/p&gt;

&lt;p&gt;Filesystem capacity&lt;br&gt;
df -h&lt;br&gt;
Inodes&lt;br&gt;
df -i&lt;br&gt;
Deleted-but-open files&lt;br&gt;
sudo lsof +L1&lt;br&gt;
Critical Tableau ports&lt;br&gt;
8850 → TSM&lt;br&gt;
8866 → ZooKeeper&lt;br&gt;
Application health&lt;/p&gt;

&lt;p&gt;A process being alive isn't enough.&lt;/p&gt;

&lt;p&gt;A port being open isn't enough.&lt;/p&gt;

&lt;p&gt;A service showing active (running) isn't enough.&lt;/p&gt;

&lt;p&gt;The application must actually be functional.&lt;/p&gt;

&lt;p&gt;The Bigger Lesson: Troubleshooting Is an Art&lt;/p&gt;

&lt;p&gt;This incident started with:&lt;/p&gt;

&lt;p&gt;"Tableau TSM is returning HTTP 500."&lt;/p&gt;

&lt;p&gt;But the real failure was several layers below it.&lt;/p&gt;

&lt;p&gt;That's what makes infrastructure troubleshooting interesting.&lt;/p&gt;

&lt;p&gt;The goal isn't to memorize every possible command.&lt;/p&gt;

&lt;p&gt;It's to know what question to ask next.&lt;/p&gt;

&lt;p&gt;Don't assume an HTTP error is an HTTP problem.&lt;/p&gt;

&lt;p&gt;Don't assume a running process is a healthy application.&lt;/p&gt;

&lt;p&gt;Don't assume df -h showing free space means ENOSPC is impossible.&lt;/p&gt;

&lt;p&gt;And don't confuse a possible explanation with a proven root cause.&lt;/p&gt;

&lt;p&gt;Instead:&lt;/p&gt;

&lt;p&gt;Follow the evidence.&lt;/p&gt;

&lt;p&gt;Check the port.&lt;/p&gt;

&lt;p&gt;Check the process.&lt;/p&gt;

&lt;p&gt;Check the dependency.&lt;/p&gt;

&lt;p&gt;Read the logs.&lt;/p&gt;

&lt;p&gt;Challenge your assumptions.&lt;/p&gt;

&lt;p&gt;Then fix the layer that's actually broken.&lt;/p&gt;

&lt;p&gt;Because sometimes the most misleading part of an incident is the very first error you see.&lt;/p&gt;

&lt;p&gt;In this case:&lt;/p&gt;

&lt;p&gt;HTTP 500&lt;/p&gt;

&lt;p&gt;was merely the messenger.&lt;/p&gt;

&lt;p&gt;The real story was:&lt;/p&gt;

&lt;p&gt;ZooKeeper&lt;br&gt;
   ↓&lt;br&gt;
Transaction log write&lt;br&gt;
   ↓&lt;br&gt;
ENOSPC&lt;br&gt;
   ↓&lt;br&gt;
ZooKeeper shutdown&lt;br&gt;
   ↓&lt;br&gt;
Coordination failure&lt;br&gt;
   ↓&lt;br&gt;
TSM HTTP 500&lt;/p&gt;

&lt;p&gt;And that, to me, is the essence of everyday systems troubleshooting:&lt;/p&gt;

&lt;p&gt;The art isn't knowing every answer.&lt;br&gt;
The art is knowing where to look next.&lt;/p&gt;

</description>
      <category>tableau</category>
      <category>ai</category>
      <category>sysad</category>
      <category>sre</category>
    </item>
  </channel>
</rss>
