<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sahan</title>
    <description>The latest articles on DEV Community by Sahan (@sahan).</description>
    <link>https://dev.to/sahan</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F222556%2F4502fb4c-7962-40f9-a1d2-f699e8b5a164.jpeg</url>
      <title>DEV Community: Sahan</title>
      <link>https://dev.to/sahan</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sahan"/>
    <language>en</language>
    <item>
      <title>I Grew an Engineering Blog from 0 to 463,000 Pageviews - Here's What Worked</title>
      <dc:creator>Sahan</dc:creator>
      <pubDate>Tue, 04 Aug 2026 01:43:00 +0000</pubDate>
      <link>https://dev.to/sahan/i-grew-an-engineering-blog-from-0-to-463000-pageviews-heres-what-worked-gmj</link>
      <guid>https://dev.to/sahan/i-grew-an-engineering-blog-from-0-to-463000-pageviews-heres-what-worked-gmj</guid>
      <description>&lt;p&gt;I published the first posts on &lt;a href="https://www.sahansera.dev/" rel="noopener noreferrer"&gt;sahansera.dev&lt;/a&gt; in December 2019.&lt;/p&gt;

&lt;p&gt;There was no launch campaign, existing audience, or reliable stream of visitors waiting for them. Like most new personal sites, the blog started at zero. I wrote about problems I had encountered, shared the posts where I could, and hoped somebody searching for the same answers would eventually find them.&lt;/p&gt;

&lt;p&gt;By 1 August 2026, the blog had accumulated &lt;strong&gt;463,362 recorded pageviews&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Along the way, it has reached readers in &lt;strong&gt;at least 184 countries and territories&lt;/strong&gt;. Here is where readers came from last month:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxkto2zn16pzpqq2o7deh.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxkto2zn16pzpqq2o7deh.webp" alt="World map showing readers reaching the blog from countries and territories around the world" width="800" height="458"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Earlier geography data is not included, so the lifetime reach may be broader.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That number is personally meaningful, but it is not the most interesting part of the story. The useful part is what happened underneath it: a small collection of practical engineering articles generated most of the traffic, some posts kept helping people for years, and many things I assumed would matter barely moved the numbers at all.&lt;/p&gt;

&lt;p&gt;This is what I learned.&lt;/p&gt;

&lt;h2&gt;
  
  
  The short version
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Specific, practical articles generated most of the traffic.&lt;/li&gt;
&lt;li&gt;Evergreen posts kept growing for years after publication.&lt;/li&gt;
&lt;li&gt;The blog continued reaching readers during a long publishing break.&lt;/li&gt;
&lt;li&gt;Traffic revealed demand, but it did not tell me whether readers returned.&lt;/li&gt;
&lt;li&gt;Comparing posts fairly requires consistent time windows, not raw lifetime totals.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  From zero to 463,362 pageviews
&lt;/h2&gt;

&lt;p&gt;I calculated the lifetime total using two consecutive analytics exports with no overlapping dates.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3lfg8de3ptvdsjwssxv7.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3lfg8de3ptvdsjwssxv7.webp" alt="Timeline showing the blog launching with zero visitors in December 2019, reaching 307,476 pageviews in its first three and a half years, then adding 155,886 views over the next three years for a total of 463,362 recorded pageviews" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A few posts did most of the work
&lt;/h2&gt;

&lt;p&gt;The distribution of traffic surprised me more than the total. Just five articles account for roughly 46% of all the pageviews the blog has recorded since launch.&lt;/p&gt;

&lt;p&gt;The leading articles are remarkably consistent in what they offer:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Article&lt;/th&gt;
&lt;th&gt;Lifetime pageviews&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.sahansera.dev/in-memory-caching-aspcore-dotnet/" rel="noopener noreferrer"&gt;Simple In-Memory Caching in .NET with &lt;code&gt;IMemoryCache&lt;/code&gt;&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;86,975&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.sahansera.dev/understanding-websockets-with-aspnetcore-5/" rel="noopener noreferrer"&gt;Understanding WebSockets with ASP.NET&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;43,272&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.sahansera.dev/distributed-caching-aspnet-core-redis/" rel="noopener noreferrer"&gt;Distributed Caching in ASP.NET Core with Redis&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;31,911&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.sahansera.dev/dotnet-core-ioc-container/" rel="noopener noreferrer"&gt;Having Fun with Microsoft IoC Container for .NET Core&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;27,812&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.sahansera.dev/dotnet-core-generic-host/" rel="noopener noreferrer"&gt;Understanding the .NET Generic Host Model&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;23,040&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Other durable performers cover securing the Hangfire dashboard, Kubernetes commands and arguments, and gRPC across Go and .NET.&lt;/p&gt;

&lt;p&gt;These are not broad opinion pieces. Each one helps a developer understand a specific concept or complete a specific task.&lt;/p&gt;

&lt;p&gt;The lesson is not simply that these technologies are popular. It is that &lt;strong&gt;clear intent compounds&lt;/strong&gt;. A post answering a concrete question can remain useful every day for years, even when I am not actively promoting it.&lt;/p&gt;


&lt;div class="crayons-card c-embed"&gt;

  &lt;br&gt;
&lt;strong&gt;💡 The pattern:&lt;/strong&gt; Useful, specific posts can compound for years.&lt;br&gt;

&lt;/div&gt;


&lt;h2&gt;
  
  
  Evergreen technical writing compounds slowly
&lt;/h2&gt;

&lt;p&gt;My highest-traffic article is an introduction to in-memory caching that I published in January 2020. Years later, it is still the largest entry point to the site.&lt;/p&gt;

&lt;p&gt;That changed how I think about the return on writing.&lt;/p&gt;

&lt;p&gt;A social post has a short distribution window. It may reach many people immediately and then disappear. A useful technical article behaves differently. It may receive very little attention on its first day, but it can be discovered repeatedly through search, links, code repositories, and recommendations.&lt;/p&gt;

&lt;p&gt;The early results can feel underwhelming because the compounding is almost invisible. The article needs to be indexed. It needs to answer the query well enough for people to stay. Other pages need to link to it. Search engines need time to understand whether it is useful.&lt;/p&gt;

&lt;p&gt;None of my successful posts felt like a breakthrough when I pressed publish. Their value accumulated quietly.&lt;/p&gt;

&lt;h2&gt;
  
  
  The blog kept working while life took priority
&lt;/h2&gt;

&lt;p&gt;The growth was not driven by a perfectly consistent publishing schedule.&lt;/p&gt;

&lt;p&gt;In 2023, I became a first-time dad. My family also moved into a new home, and I changed jobs. I wrote about that season in &lt;a href="https://www.sahansera.dev/my-plans-for-sahanseradev-2024/" rel="noopener noreferrer"&gt;my plans for sahansera.dev in 2024&lt;/a&gt;, acknowledging that blogging had taken a back seat while I focused on being the best dad I could be.&lt;/p&gt;

&lt;p&gt;The publication dates tell the story plainly. I published two posts in 2023, none in 2024, and returned with seven posts in 2025. I had hoped to resume a regular schedule sooner, but life had a different rhythm.&lt;/p&gt;

&lt;p&gt;I am glad I did not treat the pause as a reason to abandon the blog. While I was not publishing, the existing articles continued answering questions, appearing in search results, and bringing new readers to the site. The work I had already done kept compounding when I did not have the time or energy to add more.&lt;/p&gt;

&lt;p&gt;That makes the lifetime total more meaningful to me. It did not come from operating a content machine or forcing myself to publish through every season of life. It came from building a useful body of work, letting it breathe, and returning when I had something worthwhile to share.&lt;/p&gt;


&lt;div class="crayons-card c-embed"&gt;

  &lt;br&gt;
&lt;strong&gt;💡 The pause:&lt;/strong&gt; The blog kept growing even when publishing had to wait.&lt;br&gt;

&lt;/div&gt;


&lt;p&gt;When I started writing again in 2025, I explored streaming APIs, HTTP internals, Python environments, and home-lab Kubernetes. The break had not erased the audience. It gave me a chance to return with different experiences and better questions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practical specificity beats broad ambition
&lt;/h2&gt;

&lt;p&gt;The best-performing titles make a small promise:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Configure in-memory caching.&lt;/li&gt;
&lt;li&gt;Build a gRPC server or client.&lt;/li&gt;
&lt;li&gt;Run Kafka locally for testing.&lt;/li&gt;
&lt;li&gt;Secure a Hangfire dashboard.&lt;/li&gt;
&lt;li&gt;Understand how WebSockets work.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The posts are narrow enough for the reader to know why they should click, but substantial enough to teach the surrounding concepts.&lt;/p&gt;

&lt;p&gt;This balance matters. A title such as "Everything You Need to Know About Distributed Systems" sounds ambitious but does not reveal which problem it solves. "Building a gRPC Server in Go" is less grand and much more useful to the person who needs exactly that.&lt;/p&gt;

&lt;p&gt;My better articles tend to combine three layers:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A direct answer to a practical problem.&lt;/li&gt;
&lt;li&gt;An explanation of what is happening underneath.&lt;/li&gt;
&lt;li&gt;A working implementation readers can adapt.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That combination has become the clearest description of what I want this blog to be.&lt;/p&gt;

&lt;h2&gt;
  
  
  One successful article should become a cluster
&lt;/h2&gt;

&lt;p&gt;For a long time, I treated each post as an isolated piece of work. The analytics show why that leaves value on the table.&lt;/p&gt;

&lt;p&gt;The audience for an in-memory caching tutorial is likely to care about distributed caching, Redis, invalidation, cache stampedes, testing, and production failure modes. Someone building a gRPC server may next need a client, authentication, retries, deadlines, streaming, observability, or Kubernetes deployment guidance.&lt;/p&gt;


&lt;div class="crayons-card c-embed"&gt;

  &lt;br&gt;
&lt;strong&gt;💡 The opportunity:&lt;/strong&gt; A successful article is evidence that a useful topic cluster exists.&lt;br&gt;

&lt;/div&gt;


&lt;p&gt;The gRPC series already demonstrates this. Its introduction, Go server and client, .NET server and client, and deployment posts reinforce one another. The individual posts can satisfy focused searches while the series gives interested readers a natural route through the broader subject.&lt;/p&gt;

&lt;p&gt;I want to apply the same model to three areas:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;.NET caching and application reliability&lt;/li&gt;
&lt;li&gt;Kafka and event-driven system failure modes&lt;/li&gt;
&lt;li&gt;Kubernetes operations and production troubleshooting&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This does not mean publishing minor variations of the same article. Each post still needs a distinct problem and search intent. The connection between them should help a reader progress from a basic implementation to the difficult production questions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Traffic is not the same as an audience
&lt;/h2&gt;

&lt;p&gt;The historical reports showed that most visitors left after reading a single page. That sounds alarming until the context is considered.&lt;/p&gt;

&lt;p&gt;Many visitors arrive from search, find a code sample or explanation, solve their immediate problem, and leave. For a reference-style engineering article, that can be a successful visit rather than a rejection.&lt;/p&gt;

&lt;p&gt;At the same time, the data exposes a real weakness: I made it easy to consume one answer but did not always make the next useful step obvious.&lt;/p&gt;

&lt;p&gt;Chronological previous-and-next links are not enough. A reader on a caching article probably does not want the post I happened to publish immediately afterward. They want the most relevant continuation of the problem they are already solving.&lt;/p&gt;

&lt;p&gt;The improvements I am making are straightforward:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Add contextual links to related articles within the explanation.&lt;/li&gt;
&lt;li&gt;Show a clear next step at the end of high-traffic posts.&lt;/li&gt;
&lt;li&gt;Organise related material into visible series and topic hubs.&lt;/li&gt;
&lt;li&gt;Keep the email and RSS subscription options easy to find.&lt;/li&gt;
&lt;li&gt;Link runnable examples to maintained repositories.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The goal is not to trap somebody on the site. It is to make the site more useful when they want to go deeper.&lt;/p&gt;


&lt;div class="crayons-card c-embed"&gt;

  &lt;br&gt;
&lt;strong&gt;💡 The real goal:&lt;/strong&gt; Traffic becomes an audience only when readers have a reason and a way to return.&lt;br&gt;

&lt;/div&gt;


&lt;h2&gt;
  
  
  I measured traffic but not outcomes
&lt;/h2&gt;

&lt;p&gt;Another uncomfortable lesson is that I collected a lot of traffic data without defining what success should mean beyond pageviews.&lt;/p&gt;

&lt;p&gt;My current analytics setup contains no configured conversion events. I can see that people read an article, but I cannot reliably answer questions such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Did the article lead someone to another useful post?&lt;/li&gt;
&lt;li&gt;Did they subscribe by email or RSS?&lt;/li&gt;
&lt;li&gt;Did they visit the example repository?&lt;/li&gt;
&lt;li&gt;Did they copy a code sample?&lt;/li&gt;
&lt;li&gt;Which landing pages create returning readers?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These questions matter more now than the raw total.&lt;/p&gt;

&lt;p&gt;Pageviews helped me understand which subjects have demand. The next stage is to measure whether the blog creates a relationship with the reader. I plan to treat a confirmed email subscription as the primary conversion, then track supporting actions such as RSS clicks, repository visits, code copying, deep scrolling, and movement between related articles.&lt;/p&gt;

&lt;p&gt;Not every personal blog needs a conversion funnel. But if I want to improve something, I need to be explicit about what "better" means.&lt;/p&gt;

&lt;h2&gt;
  
  
  Analytics data needs maintenance too
&lt;/h2&gt;

&lt;p&gt;Just over 35% of the views in the recent export have the page title &lt;code&gt;(not set)&lt;/code&gt;. The file also contains fragmented title variants and at least one suspicious spam-like title.&lt;/p&gt;

&lt;p&gt;That is a useful reminder that analytics is not automatically a source of truth just because it contains precise-looking numbers.&lt;/p&gt;

&lt;p&gt;Collection can break. Titles can change. A migration can alter definitions. Filters can create misleading exports. Bots and spam can contaminate reports. Privacy settings can change how users and sessions are identified.&lt;/p&gt;

&lt;p&gt;I will use page paths as the canonical dimension for content reporting, investigate why page titles are missing, filter known noise, and verify that a single page view is recorded for each navigation. I also want Google Search Console beside Analytics so I can see queries, impressions, rankings, and click-through rates - not only the visits that already happened.&lt;/p&gt;

&lt;p&gt;Measurement is part of maintaining the site, not something completed by pasting in a tracking ID once.&lt;/p&gt;

&lt;h2&gt;
  
  
  New posts need a fair evaluation window
&lt;/h2&gt;

&lt;p&gt;Several of my newer articles have far fewer lifetime views than posts published in 2020 or 2021. That does not necessarily mean they failed.&lt;/p&gt;

&lt;p&gt;An article published last month should not be compared directly with one that has accumulated search traffic for six years. Lifetime totals reward age.&lt;/p&gt;

&lt;p&gt;For new work, I am moving toward a smaller set of time-normalised measures:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Search impressions and clicks during the first 28 and 90 days&lt;/li&gt;
&lt;li&gt;Views per 30 days since publication&lt;/li&gt;
&lt;li&gt;Engagement and code-copy actions&lt;/li&gt;
&lt;li&gt;Movement to another related article&lt;/li&gt;
&lt;li&gt;Email or RSS subscription actions&lt;/li&gt;
&lt;li&gt;Whether traffic continues growing after the initial promotion window&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This should make it easier to distinguish a promising article that needs time from one whose topic, title, or distribution genuinely missed the mark.&lt;/p&gt;

&lt;h2&gt;
  
  
  What 463,000 pageviews means to me
&lt;/h2&gt;

&lt;p&gt;Four hundred and sixty-three thousand is not a huge number on the scale of the internet. It is huge compared with the zero visitors I had in December 2019.&lt;/p&gt;

&lt;p&gt;More importantly, it represents individual moments when somebody had a problem and something I wrote may have helped them move forward.&lt;/p&gt;

&lt;p&gt;The experience has changed my view of successful technical writing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You do not need a large initial audience.&lt;/li&gt;
&lt;li&gt;Useful, specific posts can compound for years.&lt;/li&gt;
&lt;li&gt;A handful of durable articles may matter more than a constant publishing schedule.&lt;/li&gt;
&lt;li&gt;Updating and connecting existing work can be more valuable than always starting from zero.&lt;/li&gt;
&lt;li&gt;Honest measurement is more useful than the largest possible headline.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you are starting an engineering blog with no visitors, that is normal. Write down the problem you just solved. Explain enough of the underlying system that the solution remains useful. Include the details you wish had been available when you were searching.&lt;/p&gt;

&lt;p&gt;Then publish it and give it time.&lt;/p&gt;

&lt;p&gt;That is more or less how this blog went from zero to 463,000 pageviews - one specific problem at a time.&lt;/p&gt;

&lt;p&gt;If you write technical articles, what has surprised you most about the posts that keep finding readers? I would love to compare notes in the comments.&lt;/p&gt;

</description>
      <category>blogging</category>
      <category>writing</category>
      <category>webdev</category>
      <category>career</category>
    </item>
    <item>
      <title>How We Replaced a Critical Data Path Without a Flag Day</title>
      <dc:creator>Sahan</dc:creator>
      <pubDate>Mon, 03 Aug 2026 12:55:00 +0000</pubDate>
      <link>https://dev.to/sahan/how-we-replaced-a-critical-data-path-without-a-flag-day-1n2</link>
      <guid>https://dev.to/sahan/how-we-replaced-a-critical-data-path-without-a-flag-day-1n2</guid>
      <description>&lt;p&gt;Replacing an API call is easy. Replacing the source of truth behind an automated decision is not.&lt;/p&gt;

&lt;p&gt;I was reminded of this while migrating a critical workflow from a legacy event feed to a canonical state API. Both systems appeared to answer the same question: &lt;em&gt;should the workflow act on this record now?&lt;/em&gt; But they had different schemas, different update timings, and, more importantly, slightly different models of the same lifecycle.&lt;/p&gt;

&lt;p&gt;This was not a path where we could deploy the new code on Friday and watch the error rate. A false negative could leave work undone. A false positive could trigger an irreversible action from stale data. The HTTP request succeeding told us almost nothing about whether the new path was making the right decision.&lt;/p&gt;

&lt;p&gt;So we did not treat it as a normal code replacement. We treated it as a controlled transfer of authority.&lt;/p&gt;

&lt;p&gt;The migration used a query-only mode, shadow comparisons, discrepancy alerts, and a deliberately boring cutover. None of those techniques are particularly novel. What mattered was how we combined them, what we chose to compare, and how we decided the old path was finally safe to delete.&lt;/p&gt;

&lt;p&gt;This post walks through that process.&lt;/p&gt;

&lt;h2&gt;
  
  
  The dangerous assumption: same data, newer API
&lt;/h2&gt;

&lt;p&gt;The legacy path consumed a source built specifically around one class of lifecycle event. The replacement exposed a broader canonical record containing current state, relevant dates, and other attributes used by several workflows.&lt;/p&gt;

&lt;p&gt;On paper, the migration looked like this:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyzmhvgt48a3vl1b035vc.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyzmhvgt48a3vl1b035vc.webp" alt="The legacy event feed and canonical state API both drive the same decision logic and downstream action" width="799" height="373"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That diagram hides the risky part. The two sources did not merely encode the same fact using different field names.&lt;/p&gt;

&lt;p&gt;A purpose-built event feed tends to answer an event-shaped question: &lt;em&gt;which transitions were recorded?&lt;/em&gt; A canonical API tends to answer a state-shaped question: &lt;em&gt;what is true about this entity now?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Those questions overlap, but they are not identical.&lt;/p&gt;

&lt;p&gt;An entity can have an old end-state event and later return to an active state. A record can contain several dates, each valid in its own business context. One source may update immediately while another catches up later. Missing data can mean “not applicable,” “not received yet,” or “something is broken.”&lt;/p&gt;

&lt;p&gt;If we had translated fields one-for-one and switched traffic, the code would have looked correct while preserving none of those semantics.&lt;/p&gt;

&lt;p&gt;The first useful decision was therefore to stop calling this an API migration. We were migrating a &lt;strong&gt;business decision&lt;/strong&gt; from one model of the world to another.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start by writing down the invariants
&lt;/h2&gt;

&lt;p&gt;Before building the new path, we wrote down what the workflow must continue to guarantee.&lt;/p&gt;

&lt;p&gt;The important invariants were roughly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;An active entity must never be processed because of a stale historical event.&lt;/li&gt;
&lt;li&gt;A legitimate lifecycle transition must not be missed because one optional field is absent.&lt;/li&gt;
&lt;li&gt;The effective date must have the same business meaning before and after migration.&lt;/li&gt;
&lt;li&gt;Ambiguous or contradictory data must fail safely and become visible.&lt;/li&gt;
&lt;li&gt;Reprocessing the same record must not produce duplicate downstream actions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This step sounds obvious, but it changed the review entirely. Instead of asking whether the new client correctly parsed a payload, we could ask whether it preserved the rules the system existed to enforce.&lt;/p&gt;

&lt;p&gt;It also exposed a subtle problem with “the old system is the source of truth.” If the legacy path had known defects, perfect agreement would reproduce them. The old output was a baseline, not an oracle.&lt;/p&gt;

&lt;p&gt;That meant every mismatch needed investigation, but it did not mean the new path was automatically wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Phase one: make the new path incapable of causing damage
&lt;/h2&gt;

&lt;p&gt;The first version of the new integration was deliberately incomplete. It could query the canonical state API, normalize the response, and calculate what action it &lt;em&gt;would&lt;/em&gt; take. It could not perform that action.&lt;/p&gt;

&lt;p&gt;I think of this as &lt;strong&gt;query-only mode&lt;/strong&gt; :&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;decision&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;evaluate_canonical_record&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;record&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;query_only&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;decision&lt;/span&gt;

&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The real implementation had more safeguards than this, but the boundary was just as explicit. Read and decide on one side; mutate on the other.&lt;/p&gt;

&lt;p&gt;Query-only mode gave us a few useful properties:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;We could run the code with realistic production data.&lt;/li&gt;
&lt;li&gt;We could inspect decisions without creating tickets, sending notifications, or changing access.&lt;/li&gt;
&lt;li&gt;We could debug authentication, pagination, missing fields, and schema assumptions separately from cutover risk.&lt;/li&gt;
&lt;li&gt;We had a reusable operational tool for investigating individual records.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This last point was unexpectedly valuable. Migration controls are often treated as temporary scaffolding, but a safe read-only execution mode is also a good diagnostic interface. It lets an engineer ask, &lt;em&gt;“What would the system do with this input right now?”&lt;/em&gt; without having to reproduce the entire workflow locally.&lt;/p&gt;

&lt;p&gt;There was one rule we kept firm: query-only could not mean “mostly read-only.” If a code path still emitted an event or called a downstream service before checking the flag, the control was cosmetic. The no-side-effect guarantee had to sit at the boundary where side effects began.&lt;/p&gt;

&lt;h2&gt;
  
  
  Phase two: one writer, two decision-makers
&lt;/h2&gt;

&lt;p&gt;Once the new path could evaluate real records safely, we ran it alongside the legacy implementation.&lt;/p&gt;

&lt;p&gt;Only the legacy path was allowed to perform actions. The new path observed the same logical input and produced a candidate decision. We then compared the two.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy4yv3xon2pthqhxjgg7f.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy4yv3xon2pthqhxjgg7f.webp" alt="Production input runs through both the legacy and canonical evaluation paths, while only the legacy path can perform the automated action" width="799" height="453"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is usually called a shadow migration or dark launch. The new code sees production-shaped traffic but does not own production effects.&lt;/p&gt;

&lt;p&gt;The obvious implementation compares two Boolean values:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;legacy_should_process&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;new_should_process&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is useful, but not enough. When the values differ, a Boolean tells you that the migration is unsafe and nothing about why.&lt;/p&gt;

&lt;p&gt;We made the decision explain itself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;Decision&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;should_process&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;effective_date&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;reason&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;currently_active&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;source_record_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the shadow result could compare:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Whether each path would act.&lt;/li&gt;
&lt;li&gt;Which effective date it would use.&lt;/li&gt;
&lt;li&gt;Why it reached that conclusion.&lt;/li&gt;
&lt;li&gt;Which source record contributed to the decision.&lt;/li&gt;
&lt;li&gt;Whether either path considered the input incomplete or ambiguous.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is the part of the migration I would reuse everywhere: &lt;strong&gt;compare normalized decisions, not raw responses&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Raw payload comparison is noisy. One API may use &lt;code&gt;null&lt;/code&gt; where another omits a field. Dates may use different time zones. Identifiers may refer to different resources. A hundred harmless representation differences can hide the one semantic difference that matters.&lt;/p&gt;

&lt;p&gt;The normalized decision is the contract your users actually experience.&lt;/p&gt;

&lt;h2&gt;
  
  
  A mismatch is a finding, not just an error
&lt;/h2&gt;

&lt;p&gt;As soon as the comparison ran against real data, discrepancies appeared. That was the point.&lt;/p&gt;

&lt;p&gt;It is tempting to turn every mismatch into a page. I would avoid that. Early shadow traffic can be noisy, and training people to ignore an alert stream is a poor way to launch a critical system.&lt;/p&gt;

&lt;p&gt;Instead, we recorded every mismatch with enough context to investigate it and grouped them into a small taxonomy:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Expected timing differences&lt;/strong&gt; — one source had updated before the other.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Representation differences&lt;/strong&gt; — the sources agreed, but normalization was incomplete.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Semantic differences&lt;/strong&gt; — both records were valid, but the business interpretation differed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data-quality problems&lt;/strong&gt; — records were missing, stale, or internally contradictory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Implementation defects&lt;/strong&gt; — our new code had selected the wrong field or applied the rule incorrectly.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This classification matters because each category leads to a different response.&lt;/p&gt;

&lt;p&gt;A timing difference may need a grace period or a later recheck. A representation problem belongs in the adapter. A semantic difference needs a product or domain decision. Bad source data needs a defensive rule and an escalation path. An implementation bug needs code and a regression test.&lt;/p&gt;

&lt;p&gt;Without a taxonomy, teams tend to “fix the diff” until the graphs become green. That can accidentally teach the new system to mimic legacy behaviour without understanding it.&lt;/p&gt;

&lt;p&gt;The goal was not zero differences at any cost. It was zero &lt;strong&gt;unexplained&lt;/strong&gt; differences.&lt;/p&gt;

&lt;h2&gt;
  
  
  The edge cases were the migration
&lt;/h2&gt;

&lt;p&gt;The happy path agreed quickly. That did not make the migration nearly complete.&lt;/p&gt;

&lt;p&gt;Two edge cases forced us to refine the new model.&lt;/p&gt;

&lt;p&gt;The first involved a historical end-state event for an entity whose current state had since changed. If we looked only for the existence of that event, we could produce a false positive. The canonical record gave us another signal: the entity’s current status. We added that verification before allowing the workflow to proceed.&lt;/p&gt;

&lt;p&gt;The second involved choosing the effective date. The new source exposed more than one plausible date, and the most conveniently named field was not necessarily the date the downstream process expected. We had to trace the business meaning through the old path and deliberately select the corresponding value.&lt;/p&gt;

&lt;p&gt;Neither bug was difficult to fix once understood. The hard part was creating a migration that allowed us to see them before they became actions.&lt;/p&gt;

&lt;p&gt;That is why I do not judge shadow migrations by the amount of traffic replayed. A million ordinary records can give more confidence than they deserve. One reactivation, one delayed update, or one contradictory date can tell you much more about whether the new model is correct.&lt;/p&gt;

&lt;p&gt;Coverage should be measured across &lt;strong&gt;business scenarios&lt;/strong&gt; , not only request counts.&lt;/p&gt;

&lt;p&gt;For this kind of workflow I want an explicit scenario set:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Normal lifecycle transition.&lt;/li&gt;
&lt;li&gt;Future-dated transition.&lt;/li&gt;
&lt;li&gt;Reactivation or status reversal.&lt;/li&gt;
&lt;li&gt;Missing optional attributes.&lt;/li&gt;
&lt;li&gt;Conflicting dates.&lt;/li&gt;
&lt;li&gt;Duplicate input.&lt;/li&gt;
&lt;li&gt;Source timeout or partial response.&lt;/li&gt;
&lt;li&gt;A record that changes while being processed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Some scenarios will occur naturally during the shadow period. Rare but dangerous ones should be exercised with fixtures or controlled replay rather than waiting for production to provide them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deciding when to cut over
&lt;/h2&gt;

&lt;p&gt;“The dashboards look fine” is not a cutover criterion.&lt;/p&gt;

&lt;p&gt;Before transferring authority to the new path, we wanted evidence in several dimensions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Correctness:&lt;/strong&gt; no unexplained decision or effective-date mismatches.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scenario coverage:&lt;/strong&gt; important lifecycle transitions had been observed or tested.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stability:&lt;/strong&gt; the comparison stayed clean across a meaningful observation window, not just one quiet day.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Failure behaviour:&lt;/strong&gt; timeouts, missing records, and contradictory data failed safely.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Observability:&lt;/strong&gt; we could tell whether the new path queried, decided, skipped, failed, or acted.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rollback:&lt;/strong&gt; restoring the old authority was understood and quick.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There is no universal percentage or number of days that makes a migration safe. The right window depends on the frequency and consequence of the events you are trying to observe.&lt;/p&gt;

&lt;p&gt;If an important edge case happens once a month, a clean afternoon tells you nothing about it. If the cost of a false positive is high, the threshold should reflect that asymmetry.&lt;/p&gt;

&lt;p&gt;The cutover itself was intentionally uneventful. We changed which path was authoritative while retaining the ability to compare and roll back. We did not bundle unrelated cleanup into the same release. Boring is a feature when transferring control of a critical workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do not leave the old path “just in case”
&lt;/h2&gt;

&lt;p&gt;After the new path had operated successfully, we removed the legacy implementation and the shadow comparison.&lt;/p&gt;

&lt;p&gt;This can feel premature. Keeping the old path around appears to preserve a fallback. In reality, a dormant fallback decays quickly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Its credentials and dependencies still need maintenance.&lt;/li&gt;
&lt;li&gt;Engineers must continue reasoning about two implementations.&lt;/li&gt;
&lt;li&gt;Future changes may update one path but not the other.&lt;/li&gt;
&lt;li&gt;Someone can accidentally reactivate code that has not been exercised in months.&lt;/li&gt;
&lt;li&gt;The temporary feature flag becomes permanent architecture.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A rollback path is valuable during migration. A second production system with no clear retirement date is technical debt.&lt;/p&gt;

&lt;p&gt;We treated deletion as a planned migration phase rather than optional cleanup:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7dmdlqv1r2migvla2nj8.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7dmdlqv1r2migvla2nj8.webp" alt="The migration moves through build, observe, reconcile, cut over, soak, and delete phases, with rollback available through the soak period" width="800" height="307"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The order matters. Deleting before the soak period removes your quickest recovery option. Never deleting leaves you paying for the migration forever.&lt;/p&gt;

&lt;p&gt;The comparison infrastructure should also be removed or deliberately repurposed. Shadow code often doubles reads, emits high-cardinality logs, and contains branching that the main workflow no longer needs. Once its question has been answered, it should not quietly become part of the permanent request path.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would carry into the next migration
&lt;/h2&gt;

&lt;p&gt;The practical pattern is straightforward:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Define the invariants before translating fields.&lt;/li&gt;
&lt;li&gt;Separate reads and decisions from side effects.&lt;/li&gt;
&lt;li&gt;Give the new path a genuine query-only mode.&lt;/li&gt;
&lt;li&gt;Keep exactly one writer while both paths evaluate.&lt;/li&gt;
&lt;li&gt;Compare normalized decisions and their reasons.&lt;/li&gt;
&lt;li&gt;Investigate and classify every meaningful mismatch.&lt;/li&gt;
&lt;li&gt;Measure coverage across business scenarios, not only traffic volume.&lt;/li&gt;
&lt;li&gt;Set evidence-based cutover and rollback criteria.&lt;/li&gt;
&lt;li&gt;Keep the cutover small.&lt;/li&gt;
&lt;li&gt;Delete the legacy path after a defined soak period.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The broader lesson is that migrations like this are not primarily data plumbing exercises. They are exercises in transferring trust.&lt;/p&gt;

&lt;p&gt;The old path has years of accumulated behaviour, including behaviour nobody documented because the code made it seem obvious. The new source may be cleaner and more canonical, but that does not make your interpretation of it correct. Confidence comes from forcing both systems to make their decisions in the open, then explaining every place they disagree.&lt;/p&gt;

&lt;p&gt;That takes longer than changing an endpoint. It is still much cheaper than discovering after cutover that a green deployment was making the wrong decision perfectly.&lt;/p&gt;

&lt;p&gt;Thanks for reading ✌️&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>distributedsystems</category>
      <category>sre</category>
      <category>backend</category>
    </item>
    <item>
      <title>When "no healthy upstream" isn't about the upstream you think</title>
      <dc:creator>Sahan</dc:creator>
      <pubDate>Thu, 23 Jul 2026 00:00:00 +0000</pubDate>
      <link>https://dev.to/sahan/when-no-healthy-upstream-isnt-about-the-upstream-you-think-1lf0</link>
      <guid>https://dev.to/sahan/when-no-healthy-upstream-isnt-about-the-upstream-you-think-1lf0</guid>
      <description>&lt;p&gt;&lt;code&gt;no healthy upstream&lt;/code&gt; is the kind of error that makes you expect wreckage.&lt;/p&gt;

&lt;p&gt;Then you open the dashboards and find… almost nothing. CPU is low. No pods have crashed. The last deployment was hours ago. By the time you refresh the page, the service has recovered by itself.&lt;/p&gt;

&lt;p&gt;That was the scene a few weeks ago when I started chasing an intermittent failure in a search backend. The eventual fix was only a few lines. The interesting part was getting there.&lt;/p&gt;

&lt;p&gt;We already had a confident root-cause analysis (RCA), complete with a tidy explanation and a one-line remedy. It was also pointing at the wrong subsystem. This post is about the gap between a plausible story and the evidence, and about a common failure mode in which a mostly healthy fleet slowly removes itself from service.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shape of the failure
&lt;/h2&gt;

&lt;p&gt;The setup was ordinary: a search backend behind a load balancer, with a fixed pool of worker processes on each instance.&lt;/p&gt;

&lt;p&gt;Every so often, with no obvious schedule, a burst of requests failed. The browser showed a bare &lt;code&gt;no healthy upstream&lt;/code&gt;, and a minute or two later everything worked again.&lt;/p&gt;

&lt;p&gt;One clue appeared every time. Backend p99 latency climbed to &lt;em&gt;almost exactly&lt;/em&gt; the load balancer timeout, stayed flat, and then dropped back to normal:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqmoff5wwgwnhhacu80tg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqmoff5wwgwnhhacu80tg.png" alt="Backend p99 latency rising sharply to the load balancer timeout, remaining flat during the failure window, and then returning to baseline" width="800" height="373"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That shape matters. Organic slowdowns tend to be uneven. This was a cliff, a flat top, and a recovery. Requests were not gradually becoming slower; they were running into a deadline and being cut off.&lt;/p&gt;

&lt;p&gt;The theory I inherited was &lt;strong&gt;CPU throttling&lt;/strong&gt;. A heavy periodic job supposedly pegged the pod’s CPU, the scheduler throttled it, request handling starved, and the load balancer eventually evicted the instance.&lt;/p&gt;

&lt;p&gt;It was coherent. Better still, it came with a one-line fix: raise the CPU limit. That is an attractive combination during an incident. But a root-cause theory makes predictions, and these predictions did not survive contact with the evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Treating the RCA as a hypothesis, not a conclusion
&lt;/h2&gt;

&lt;p&gt;Instead of treating the existing RCA as a conclusion, I treated it as a hypothesis: if CPU throttling caused the incidents, what else should I be able to observe?&lt;/p&gt;

&lt;p&gt;Three checks came back wrong.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;There was no CPU-bound work in the serving path.&lt;/strong&gt; The compute-heavy batch indexers ran on &lt;em&gt;separate&lt;/em&gt; machines and reached the datastore over the network. They never ran inside the pods serving requests. A process outside the pod’s cgroup cannot cause that pod to be CPU-throttled. The graphs agreed: container throttling counters stayed flat during every event.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The timing did not fit a scheduled trigger.&lt;/strong&gt; The batch jobs ran on a coarse schedule and produced one sustained utilisation bump. The incidents arrived at arbitrary minutes and happened far more often than the jobs ran. If a timer were responsible, the failures should have followed the timer. They did not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The failures ignored instance and version boundaries.&lt;/strong&gt; The same symptom appeared across different pods and deployment SHAs. A regression usually follows a version. A shared external event usually hits the fleet together. These failures did neither. The pattern looked more like each instance was doing something &lt;em&gt;to itself&lt;/em&gt;, triggered by its own traffic.&lt;/p&gt;

&lt;p&gt;To keep the CPU theory alive, I would have had to explain away the topology, the timing, and the distribution of failures. At that point the theory was creating more questions than it answered, so I dropped it and went back to the logs.&lt;/p&gt;

&lt;h2&gt;
  
  
  One line in the logs, and why it changes the model
&lt;/h2&gt;

&lt;p&gt;The event window contained one recurring error:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ConnectionTimeout: Connection timed out
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The stack ended in the datastore client, blocked on a socket read that never returned.&lt;/p&gt;

&lt;p&gt;That one line flipped the model.&lt;/p&gt;

&lt;p&gt;A CPU-throttled worker is ready to run but cannot get enough scheduler time. A worker blocked on I/O is off-CPU, waiting on the network while still occupying its worker slot. From the outside, both look like “latency went up, then requests timed out.” Underneath, they are opposites.&lt;/p&gt;

&lt;p&gt;Adding CPU to an I/O stall does not unblock the socket. At best, it gives you more workers to park behind the same slow dependency. This is why edge symptoms are a dangerous thing to tune against: resource saturation and dependency blocking can produce the same fever while needing completely different treatment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why one slow dependency saturates the whole instance
&lt;/h2&gt;

&lt;p&gt;The mechanism came down to two ordinary client settings that were dangerous in combination:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;No timeout on individual datastore calls.&lt;/strong&gt; One call could occupy a worker long after the user had given up.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retries on connection timeouts.&lt;/strong&gt; After waiting too long once, the worker would wait again, with backoff in between.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Nothing exotic. That is partly what makes this failure mode easy to miss.&lt;/p&gt;

&lt;p&gt;Little’s Law shows why it is fatal to a fixed worker pool. Concurrency is &lt;code&gt;L = λW&lt;/code&gt;: the arrival rate (&lt;code&gt;λ&lt;/code&gt;) multiplied by the average time each request spends in the system (&lt;code&gt;W&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;Suppose the service receives 200 requests per second and normally responds in 40 ms:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;L = 200 req/s × 0.04 s = 8 concurrent requests
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Eight concurrent requests are easy for the pool to absorb. But if the datastore slows down and requests wait for tens of seconds, &lt;code&gt;W&lt;/code&gt; increases by three orders of magnitude. Retries stretch it further. The required concurrency quickly exceeds the number of workers available.&lt;/p&gt;

&lt;p&gt;The pool then looks something like this:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fri82z29xylte2gmuyvmr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fri82z29xylte2gmuyvmr.png" alt="Comparison of a healthy worker pool with spare capacity and a saturated pool where every worker waits on the datastore while requests and health checks queue" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This creates &lt;strong&gt;head-of-line blocking&lt;/strong&gt;. Workers parked on datastore calls cannot serve the fast requests behind them, so latency rises for &lt;em&gt;everything&lt;/em&gt;, not only for requests that reached the slow dependency.&lt;/p&gt;

&lt;p&gt;Health checks are caught in the same queue. They time out, the load balancer removes the instance from rotation, and the remaining instances receive more traffic. Once enough instances fail their health checks, the load balancer has nowhere to send the next request. That is when the user sees &lt;code&gt;no healthy upstream&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;There was one more twist. The client’s total retry time could exceed the load balancer’s deadline:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0i8udb7skg7w41jwejt3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0i8udb7skg7w41jwejt3.png" alt="Timeline showing client retries continuing after the load balancer deadline and holding a worker after the user has gone" width="800" height="347"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The load balancer returned an error while the worker continued retrying a result nobody could receive. Every millisecond after the outer deadline was wasted work, and the wasted work held a scarce worker slot.&lt;/p&gt;

&lt;p&gt;This is a &lt;strong&gt;deadline-propagation&lt;/strong&gt; failure. Inner operations must finish inside the outer request deadline. Better still, pass the outer deadline through the call chain so every layer knows when its result has become useless.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it recovered by itself
&lt;/h2&gt;

&lt;p&gt;The recovery was initially reassuring. In hindsight, it was the worrying part.&lt;/p&gt;

&lt;p&gt;The datastore blip triggered retries. Those retries added load to the datastore while it was already struggling, which caused more timeouts and therefore more retries:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F869hast8hhod3x3f82pc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F869hast8hhod3x3f82pc.png" alt="Feedback loop where a slow datastore causes timeouts, retries, and additional datastore load that reinforces the slowdown" width="800" height="360"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That is the start of a &lt;strong&gt;metastable failure&lt;/strong&gt; : a brief trigger knocks the system out of its healthy state, then a feedback loop keeps it unhealthy after the original trigger has passed.&lt;/p&gt;

&lt;p&gt;We got lucky. The datastore blips were short enough that traffic fell below the tipping point before the retry loop became self-sustaining. A slightly longer blip could have kept the loop alive until we restarted the fleet or shed enough traffic to escape it.&lt;/p&gt;

&lt;p&gt;Self-recovery was not proof of resilience. It was a warning shot.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix, and why we kept it small
&lt;/h2&gt;

&lt;p&gt;The instinct during an availability incident is to add headroom: raise CPU limits, increase the worker pool, add replicas. That can help with genuine capacity problems. Here, it would only give the retry loop more workers to occupy.&lt;/p&gt;

&lt;p&gt;The useful fix was to put a hard bound on the cost of one request.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First, add an explicit timeout to every downstream call.&lt;/strong&gt; Choose it from the dependency’s healthy latency distribution rather than picking a pleasing round number. It should sit comfortably above healthy p99.9, but well below the load balancer’s deadline. A call that can wait forever is a worker you can lose forever.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Second, bound retries as a fraction of normal traffic, not as an unconditional count on every request.&lt;/strong&gt; A rule such as “retry three times” allows every client to multiply traffic precisely when the dependency is least able to handle it. A token-bucket retry budget keeps the added load bounded; the Google SRE guidance uses 10% as an example. Add jitter too, or clients can wake up and retry in synchronised waves.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Third, be especially reluctant to retry connection timeouts.&lt;/strong&gt; During saturation, a timeout often means the dependency is already over its limit. Another immediate attempt is unlikely to help. Retries make the most sense for independent, transient failures and only for idempotent operations. Search requests were idempotent, at least, so duplicate side effects were not another problem waiting for us.&lt;/p&gt;

&lt;p&gt;The decision for each failed call becomes straightforward:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6845gh2q6hg7gwo7o6no.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6845gh2q6hg7gwo7o6no.png" alt="Decision flow where a dependency call returns on success, retries only while deadline and budget remain, and otherwise fails fast to free the worker" width="800" height="507"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Circuit breakers, load shedding, and bulkheads could all strengthen this design:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A &lt;strong&gt;circuit breaker&lt;/strong&gt; stops callers repeatedly rediscovering the same outage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Load shedding&lt;/strong&gt; rejects excess work early enough to keep the service responsive.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;bulkhead&lt;/strong&gt; gives the dependency its own bounded concurrency pool, preventing it from occupying every worker.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All three add code, tuning, and operational state. We did not have evidence that we needed them yet. Once a small timeout and a bounded retry policy removed the amplifier, adding more machinery would have solved a hypothetical problem rather than the incident in front of us.&lt;/p&gt;

&lt;h2&gt;
  
  
  “Won’t failing fast just move the errors to the caller?”
&lt;/h2&gt;

&lt;p&gt;This was the immediate pushback: if the backend gives up sooner, won’t users simply see more errors?&lt;/p&gt;

&lt;p&gt;Only if we compare failing fast with a world in which every request eventually succeeds. That was not the world we had.&lt;/p&gt;

&lt;p&gt;With the worker pool full of hung calls, &lt;em&gt;every&lt;/em&gt; request eventually failed. The entire search feature became a &lt;code&gt;no healthy upstream&lt;/code&gt; page, including requests that never needed the slow dependency in the first place.&lt;/p&gt;

&lt;p&gt;Failing fast keeps the instances responsive and in rotation. It turns one correlated, fleet-wide outage into a smaller number of independent failures: the requests that actually hit the bad path. Those are failures a caller can absorb with a cached result, an empty state, or a retry button.&lt;/p&gt;

&lt;p&gt;A nonessential component should not be able to take down the whole page. Give it its own deadline and a graceful fallback, and a slow dependency degrades one part of the experience instead of replacing the entire document with an error page.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verifying the bound, not asserting it
&lt;/h2&gt;

&lt;p&gt;I did not want to ship a fix that worked only on a whiteboard, so the validation was intentionally mechanical.&lt;/p&gt;

&lt;p&gt;With a &lt;strong&gt;healthy dependency&lt;/strong&gt; , normal traffic should remain normal: the same latency distribution and no new errors. If the timeout clips healthy p99.9 requests, it is too aggressive and will create the very failures it is meant to prevent.&lt;/p&gt;

&lt;p&gt;With an &lt;strong&gt;unreachable dependency&lt;/strong&gt; , requests should fail near the configured bound and before the outer load balancer deadline. More importantly, worker occupancy and in-flight request counts should remain flat instead of climbing.&lt;/p&gt;

&lt;p&gt;That flat line under induced failure is the real acceptance test. It shows that the amplifier is gone.&lt;/p&gt;

&lt;p&gt;The saturating case needs its own load test. The happy path will never prove that a service behaves well when every downstream call is stuck.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I actually took from it
&lt;/h2&gt;

&lt;p&gt;The technical lesson is easy to audit for: an unbounded downstream call plus eager retries is a latent outage with a feedback loop attached. It can sit quietly for months, green on every dashboard, until a dependency has a bad thirty seconds. Then it turns that blip into a fleet-wide event.&lt;/p&gt;

&lt;p&gt;It is worth searching your own services for downstream calls without deadlines and retries without a shared budget. Those two settings deserve to be reviewed together, because together they can change the shape of a failure.&lt;/p&gt;

&lt;p&gt;The lesson that stayed with me, though, was about diagnosis.&lt;/p&gt;

&lt;p&gt;The CPU theory was clean. It was mechanistic. It came with a satisfying one-line fix. And it survived three pieces of contradictory evidence because we had started treating it as an answer instead of a claim.&lt;/p&gt;

&lt;p&gt;A useful root cause makes predictions you can check. When the topology, timing, and distribution of failures all disagree with the story, elegance stops counting. In this case, one dull line in a log file told us more than the tidy explanation we had already grown attached to.&lt;/p&gt;

&lt;p&gt;That was the expensive part of the incident: not the eventual three-line fix, but learning to let the evidence ruin a good story.&lt;/p&gt;

&lt;p&gt;Thanks for reading ✌️&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;a href="https://sigops.org/s/conferences/hotos/2021/papers/hotos21-s11-bronson.pdf" rel="noopener noreferrer"&gt;Metastable Failures in Distributed Systems - Bronson et al., HotOS ‘21&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://aws.amazon.com/builders-library/timeouts-retries-and-backoff-with-jitter/" rel="noopener noreferrer"&gt;Timeouts, retries, and backoff with jitter - Amazon Builders’ Library&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://sre.google/sre-book/addressing-cascading-failures/" rel="noopener noreferrer"&gt;Addressing Cascading Failures - Google SRE Book&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://sre.google/sre-book/handling-overload/" rel="noopener noreferrer"&gt;Handling Overload - Google SRE Book&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://research.google/pubs/pub40801/" rel="noopener noreferrer"&gt;The Tail at Scale - Dean &amp;amp; Barroso, CACM 2013&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.envoyproxy.io/docs/envoy/latest/faq/load_balancing/disable_circuit_breaking" rel="noopener noreferrer"&gt;no healthy upstream - Envoy load balancing FAQ&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://martinfowler.com/bliki/CircuitBreaker.html" rel="noopener noreferrer"&gt;Circuit Breaker - Martin Fowler&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>devops</category>
      <category>architecture</category>
      <category>backend</category>
      <category>performance</category>
    </item>
    <item>
      <title>The Acknowledgment Gap - How Event-Driven Systems Lose Messages Without Errors</title>
      <dc:creator>Sahan</dc:creator>
      <pubDate>Thu, 16 Jul 2026 01:05:00 +0000</pubDate>
      <link>https://dev.to/sahan/the-acknowledgment-gap-how-event-driven-systems-lose-messages-without-errors-5hmj</link>
      <guid>https://dev.to/sahan/the-acknowledgment-gap-how-event-driven-systems-lose-messages-without-errors-5hmj</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for &lt;a href="https://dev.to/bugsmash"&gt;DEV's Summer Bug Smash: Smash Stories&lt;/a&gt; powered by &lt;a href="https://sentry.io/" rel="noopener noreferrer"&gt;Sentry&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;If you have built event-driven systems for any length of time, you have probably internalized the mantra of &lt;em&gt;at-least-once delivery&lt;/em&gt;: keep retrying until the work is done, and design everything downstream to be idempotent. It’s good advice. But there’s a subtle failure mode that hides right underneath it - one where your system faithfully reports success, commits its progress, and quietly drops work on the floor. No exception, no alert, no dead letter. Just a message that was supposed to do something, and didn’t.&lt;/p&gt;

&lt;p&gt;I ran into this recently while debugging why a small percentage of events were mysteriously never being processed. Everything &lt;em&gt;looked&lt;/em&gt; healthy. The producer got a &lt;code&gt;2xx&lt;/code&gt;. The consumer committed its offset. The dashboards were green. And yet the work never happened. This post is about that gap - the space between “the system accepted my request” and “the system actually did the work” - and why it’s one of the more dangerous places for a distributed system to lose data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Motivation
&lt;/h2&gt;

&lt;p&gt;Most write-ups about message-processing reliability focus on the well-known hazards: duplicate delivery, out-of-order messages, poison pills, consumer lag. Those are real, and there’s plenty written about them. What I found much less discussed is the case where &lt;strong&gt;the acknowledgment itself is a lie&lt;/strong&gt; - where a component reports success for an operation that has only been &lt;em&gt;accepted&lt;/em&gt;, not &lt;em&gt;completed&lt;/em&gt;, and a second component treats that acknowledgment as permission to throw the original message away.&lt;/p&gt;

&lt;p&gt;This is a design smell that shows up across all sorts of stacks: a queue consumer that calls an async API, a workflow engine that enqueues a job, a service that hands off to a background worker. Any time you have a &lt;strong&gt;handoff across an asynchronous boundary&lt;/strong&gt; , you have the potential for this gap. So I want to walk through the anatomy of the bug from first principles, then talk about how to close it properly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Background: two kinds of “yes”
&lt;/h2&gt;

&lt;p&gt;Before we get to the bug, let’s be precise about acknowledgments, because the whole problem lives in some sloppy vocabulary.&lt;/p&gt;

&lt;p&gt;When a system replies to your request, it can mean one of two very different things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;“I have accepted your request.”&lt;/strong&gt; - I’ve durably recorded your intent, and I promise to &lt;em&gt;try&lt;/em&gt; to do the work. Think &lt;a href="https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/202" rel="noopener noreferrer"&gt;HTTP 202 Accepted&lt;/a&gt;. The work hasn’t happened yet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;“I have completed your request.”&lt;/strong&gt; - The work is done, and its effects are durable. Think &lt;code&gt;200 OK&lt;/code&gt; with a result body.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are worlds apart, and conflating them is the root of a lot of pain. The trouble is that many APIs return the &lt;em&gt;same status code&lt;/em&gt; for “accepted” whether or not the work will eventually succeed. A &lt;code&gt;202&lt;/code&gt; (or a &lt;code&gt;204 No Content&lt;/code&gt;, which is even more ambiguous) tells you the request was received. It tells you &lt;em&gt;nothing&lt;/em&gt; about whether the work will run.&lt;/p&gt;

&lt;p&gt;Now layer on the consumer side. A huge number of event-driven systems are built on brokers that use &lt;strong&gt;offset-based consumer groups&lt;/strong&gt; - &lt;a href="https://en.wikipedia.org/wiki/Apache_Kafka" rel="noopener noreferrer"&gt;Apache Kafka&lt;/a&gt; being the canonical example. If you want a primer, I wrote an &lt;a href="https://sahansera.dev/introduction-to-apache-kafka/" rel="noopener noreferrer"&gt;introduction to Apache Kafka&lt;/a&gt; a while back. The mental model is simple:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Messages in a partition have monotonically increasing offsets.&lt;/li&gt;
&lt;li&gt;Your consumer reads a message, does some work, and then &lt;strong&gt;commits&lt;/strong&gt; (or “marks”) the offset to say &lt;em&gt;“I’m done with everything up to here.”&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;If the consumer crashes before committing, the broker redelivers from the last committed offset. That’s what gives you at-least-once semantics.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;💡 The offset commit is a &lt;em&gt;promise about the past&lt;/em&gt;. When you commit offset N, you are telling the broker “every message up to and including N has been fully handled, and you never need to give them to me again.” If that statement isn’t actually true, you have manufactured data loss with your own hands.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hold onto those two ideas - the ambiguous “yes” and the offset-as-promise - because the bug is what happens when they collide.&lt;/p&gt;

&lt;h2&gt;
  
  
  Anatomy of the pipeline
&lt;/h2&gt;

&lt;p&gt;Let me describe a deliberately generic pipeline. Strip away the specific technologies and almost every async system looks like this:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyf0tq0ee976zmexiqrxh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyf0tq0ee976zmexiqrxh.png" alt="Event processing pipeline showing the source, broker, consumer, asynchronous job API, worker, and offset commit path" width="800" height="373"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The consumer’s job is to read an event and translate it into an &lt;em&gt;action&lt;/em&gt; by calling some downstream control plane - an async job API, a workflow trigger, a task queue. The control plane accepts the request and, at some later point, a worker actually executes it.&lt;/p&gt;

&lt;p&gt;The consumer’s loop, in pseudocode, looked roughly like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;msg&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;broker&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Read&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;dispatchJob&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c"&gt;// calls the async control plane&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;retryWithBackoff&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;dispatchJob&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="c"&gt;// Always commit so we never get stuck reprocessing a bad message.&lt;/span&gt;
    &lt;span class="n"&gt;broker&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Commit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At a glance this looks reasonable, even defensive. There’s a retry with backoff. There’s a comment explaining that we always commit to avoid getting wedged on a poison message. Someone clearly thought about failure here.&lt;/p&gt;

&lt;p&gt;And that is exactly what makes the bug so insidious.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bug: the acknowledgment gap
&lt;/h2&gt;

&lt;p&gt;Here’s the sequence that loses data.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The consumer reads a message and calls &lt;code&gt;dispatchJob&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;The control plane returns &lt;strong&gt;&lt;code&gt;204 No Content&lt;/code&gt;&lt;/strong&gt; - &lt;em&gt;“request accepted, a job has been created.”&lt;/em&gt; From the consumer’s point of view, this is success. &lt;code&gt;err&lt;/code&gt; is &lt;code&gt;nil&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Milliseconds later, and &lt;strong&gt;entirely outside the consumer’s view&lt;/strong&gt; , the control plane &lt;em&gt;rejects&lt;/em&gt; the job before it executes. Maybe a concurrency quota was exceeded. Maybe an admission controller said no. Maybe the queue was full. The job transitions straight to a terminal “rejected” state without a single line of business logic ever running.&lt;/li&gt;
&lt;li&gt;Back in the consumer, &lt;code&gt;dispatchJob&lt;/code&gt; returned &lt;code&gt;nil&lt;/code&gt;, so the retry loop never fires - there was nothing to retry, as far as it knows.&lt;/li&gt;
&lt;li&gt;The consumer commits the offset.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The offset moves forward. The broker will never redeliver that message. The job never ran. And nobody was told.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8raghqfmzp3r9lx03ywe.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8raghqfmzp3r9lx03ywe.png" alt="Sequence diagram showing a job being accepted, rejected before execution, and then lost when the consumer commits its offset" width="800" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is the acknowledgment gap. The failure happened in the &lt;strong&gt;window between “accepted” and “executed”&lt;/strong&gt; , and our success signal was wired to the wrong end of that window. We treated &lt;em&gt;“a job was created”&lt;/em&gt; as if it meant &lt;em&gt;“a job will run,”&lt;/em&gt; and those are not the same statement.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;💡 The most dangerous bugs aren’t the ones that throw. They’re the ones that return &lt;code&gt;nil&lt;/code&gt;. An exception is a gift - it’s the system telling you something is wrong. Silent, structurally-invisible loss gives you nothing to catch.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;What made it worse is that the “always commit” decision - added defensively to avoid an infinite reprocessing loop - turned a &lt;em&gt;recoverable&lt;/em&gt; failure into an &lt;em&gt;unrecoverable&lt;/em&gt; one. The one safety mechanism that could have saved us (letting the broker redeliver) was disabled precisely when we needed it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the usual instincts don’t save you
&lt;/h2&gt;

&lt;p&gt;When engineers first see this, they reach for familiar fixes. Most of them don’t actually close the gap:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;“Just check the status code.”&lt;/strong&gt; We did. It was &lt;code&gt;204&lt;/code&gt;. The status code describes the &lt;em&gt;acceptance&lt;/em&gt;, not the &lt;em&gt;outcome&lt;/em&gt;. The information we needed didn’t exist yet at the moment we got the response.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;“Add a retry.”&lt;/strong&gt; There was one. It only triggers on a failed &lt;em&gt;dispatch&lt;/em&gt;, not a failed &lt;em&gt;execution&lt;/em&gt;. You can’t retry something you don’t know failed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;“Make it idempotent.”&lt;/strong&gt; Idempotency is necessary but not sufficient here. Idempotency protects you from doing the work &lt;em&gt;twice&lt;/em&gt;; it does nothing to protect you from doing it &lt;em&gt;zero&lt;/em&gt; times.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;“Use exactly-once semantics.”&lt;/strong&gt; Setting aside the long debate about whether &lt;a href="https://en.wikipedia.org/wiki/Two_Generals%27_Problem" rel="noopener noreferrer"&gt;exactly-once is even a coherent goal&lt;/a&gt; across independent systems - the transactional guarantees of your broker do not extend into a third-party control plane you’re calling over HTTP. The moment you cross that boundary, you’re back to coordinating two independent systems with no shared transaction.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The real issue is architectural: &lt;strong&gt;we committed our durable progress based on a signal that didn’t actually confirm the work was durable.&lt;/strong&gt; No amount of tuning the individual pieces fixes that. You have to move the acknowledgment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fixing it: verify before you ack
&lt;/h2&gt;

&lt;p&gt;The core principle is a single sentence:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;💡 &lt;strong&gt;Never acknowledge a message until you have confirmed the work it represents has actually started (or completed) - not merely been accepted.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Everything else is mechanics. Let’s walk through them, because the mechanics are where the interesting distributed-systems problems hide.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Close the loop: confirm execution, don’t assume it
&lt;/h3&gt;

&lt;p&gt;Instead of trusting the &lt;code&gt;204&lt;/code&gt;, the consumer now &lt;em&gt;verifies&lt;/em&gt; that the dispatched job reached a real running (or terminal-success) state before committing. In practice that means polling the control plane’s read API after dispatch:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="n"&gt;started&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;dispatchAndVerify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;started&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;broker&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Commit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c"&gt;// safe: the work is genuinely underway&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c"&gt;// do NOT commit - let redelivery give us another shot&lt;/span&gt;
    &lt;span class="n"&gt;alert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;dispatchAndVerify&lt;/code&gt; dispatches, then polls: &lt;em&gt;did a job actually enter a non-rejected state?&lt;/em&gt; If it sees the tell-tale “rejected before executing” terminal state, or it can’t find the job at all within a bounded window, it treats that as a failure - which is the thing our original code could never see.&lt;/p&gt;

&lt;p&gt;This is really just applying &lt;strong&gt;read-after-write&lt;/strong&gt; thinking to a control plane. Don’t trust the write acknowledgment; go read the state back and confirm reality matches your intent.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. The correlation problem
&lt;/h3&gt;

&lt;p&gt;Here’s a genuinely tricky sub-problem that this exposes, and it’s a great example of why distributed systems are hard: &lt;strong&gt;the dispatch API often doesn’t tell you the ID of the thing it just created.&lt;/strong&gt; You fire a request, you get back &lt;code&gt;204 No Content&lt;/code&gt; - literally no content - and now you need to find “the job I just created” among all the jobs.&lt;/p&gt;

&lt;p&gt;You’re left correlating on secondary signals:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A &lt;strong&gt;creation timestamp window&lt;/strong&gt; (“a job created after time T”), which is racy under concurrency - two near-simultaneous dispatches can be ambiguous.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;business key&lt;/strong&gt; embedded into the job’s metadata at creation time, if the API lets you set something like a name or a label.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The robust fix is to make the work &lt;strong&gt;self-identifying&lt;/strong&gt; : stamp a correlation key you already own (the entity ID, a request UUID) into the job at dispatch time, so that when you read the state back you can match on it &lt;em&gt;exactly&lt;/em&gt; rather than guessing by time. If your control plane supports naming or tagging the work, use it. This is the async equivalent of propagating a trace ID, and it pays for itself the first time you have to debug a race.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;💡 Any time you hand work across an async boundary, ask: &lt;em&gt;“When this comes back, how will I know it’s mine?”&lt;/em&gt; If the answer is “by timestamp,” you have a race waiting to happen.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  3. Bounded redelivery, or: don’t trade loss for a hot loop
&lt;/h3&gt;

&lt;p&gt;The moment you say “don’t commit on failure so the broker redelivers,” someone will rightly point out the opposite failure mode: what if the work &lt;em&gt;keeps&lt;/em&gt; failing? Now you’ve built an infinite reprocessing loop, and you’re hammering a control plane that’s already unhappy. This is the eternal tension:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Commit too eagerly&lt;/strong&gt; → you lose messages (the original bug).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Never commit on failure&lt;/strong&gt; → you can wedge the consumer forever.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The answer is &lt;strong&gt;bounded retries with escalation&lt;/strong&gt;. Track how many times a given message has been through the wringer - keyed by its stable identity (partition + offset, or a business key) - and:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Retry with &lt;a href="https://en.wikipedia.org/wiki/Exponential_backoff" rel="noopener noreferrer"&gt;exponential backoff&lt;/a&gt; while attempts remain, so a transient quota exhaustion gets a chance to clear.&lt;/li&gt;
&lt;li&gt;Once you’ve exhausted the budget, &lt;strong&gt;stop, escalate loudly, and then commit&lt;/strong&gt; so a single doomed message can’t block the whole partition behind it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That final commit is not “giving up silently” - it’s a deliberate, &lt;em&gt;observable&lt;/em&gt; decision to route the message to a human (or a &lt;a href="https://en.wikipedia.org/wiki/Dead_letter_queue" rel="noopener noreferrer"&gt;dead-letter queue&lt;/a&gt;) instead of blocking the stream. The difference between this and the original bug is everything: the original dropped work with &lt;em&gt;zero&lt;/em&gt; signal; this drops it only after N visible, alarmed attempts.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Separate transient failures from terminal ones
&lt;/h3&gt;

&lt;p&gt;A subtle but important refinement: when you poll to verify, &lt;strong&gt;a failure to read the state is not the same as the work being rejected.&lt;/strong&gt; If your verification call itself hits a network blip or a &lt;code&gt;503&lt;/code&gt;, and you treat that as “the job failed,” you’ll re-dispatch and potentially create &lt;em&gt;duplicate&lt;/em&gt; work - trading a lost-message bug for a double-processing bug.&lt;/p&gt;

&lt;p&gt;So the verification loop needs to distinguish:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;“I confirmed the job was rejected”&lt;/strong&gt; → terminal, re-dispatch is warranted.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;“I couldn’t reach the control plane to check”&lt;/strong&gt; → transient, just retry the &lt;em&gt;read&lt;/em&gt;, don’t re-dispatch.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Only definitive answers should drive irreversible decisions. Everything else is a retryable read. This is the same discipline as not making state transitions on ambiguous signals - you wait until you actually &lt;em&gt;know&lt;/em&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Make the invisible visible
&lt;/h3&gt;

&lt;p&gt;The reason this bug survived in production is that it was &lt;strong&gt;structurally unobservable&lt;/strong&gt;. So the last piece is observability, and it’s not optional:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Emit a metric/event whenever a dispatched unit of work fails to start.&lt;/li&gt;
&lt;li&gt;Alert when the retry budget is exhausted and a message is dropped.&lt;/li&gt;
&lt;li&gt;Log the correlation key, the attempt count, and a link to the rejected work.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If your system is going to make a hard decision like “I’m dropping this after five tries,” that decision must be &lt;em&gt;the loudest thing in the room&lt;/em&gt;, not a silent commit. A good rule of thumb: &lt;strong&gt;every place your code can decide to discard work should be capable of paging a human.&lt;/strong&gt; You may choose not to page - but the capability being there forces you to consciously design the failure path instead of falling into one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stepping back: the general principle
&lt;/h2&gt;

&lt;p&gt;If you zoom out from the specific mechanics, this whole class of bug reduces to a few reusable principles that are worth carrying into any distributed system you build:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Distinguish “accepted” from “completed” - always.&lt;/strong&gt; Treat them as different events with different names, different metrics, and different handling. Never let one masquerade as the other.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anchor your durable acknowledgment to the durable outcome.&lt;/strong&gt; Commit your offset (or delete your message, or mark your row done) based on confirmation of the &lt;em&gt;effect you care about&lt;/em&gt;, not on a transport-level receipt.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A &lt;code&gt;nil&lt;/code&gt; error is a claim, not a fact.&lt;/strong&gt; Verify claims that cross trust boundaries, especially async ones. Read-after-write is cheap insurance.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bound every retry, and escalate at the boundary.&lt;/strong&gt; Unbounded retries and silent drops are two sides of the same coin; the cure for both is a visible, finite budget with a loud exit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Only act irreversibly on unambiguous signals.&lt;/strong&gt; Transient “I don’t know” should never trigger a decision that assumes “no.”&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;None of these are novel on their own. What’s interesting is how a single missing distinction - &lt;em&gt;accepted vs. executed&lt;/em&gt; - cascades into silent data loss when it meets an offset commit. It’s a good reminder that in distributed systems, the bugs rarely live inside a component. They live in the &lt;strong&gt;seams between components&lt;/strong&gt; , where two reasonable local decisions add up to one unreasonable global one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tradeoffs
&lt;/h2&gt;

&lt;p&gt;Nothing here is free, and I’d be doing you a disservice to pretend otherwise.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What you gain&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;No more silent loss - the failure mode that’s hardest to detect and most damaging to trust.&lt;/li&gt;
&lt;li&gt;A verifiable, observable processing pipeline where “done” actually means done.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What it costs&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Extra latency and load.&lt;/strong&gt; Verifying execution means additional reads against the control plane per message. Poll intervals and budgets need tuning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;More moving parts.&lt;/strong&gt; Correlation keys, attempt tracking, and escalation paths are code you now own and test.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You must embrace at-least-once for real.&lt;/strong&gt; Verify-and-redeliver &lt;em&gt;will&lt;/em&gt; occasionally produce duplicates (e.g., if a job actually started but your confirmation read failed). Idempotency downstream stops being optional - but that was always true; this just makes it honest.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For low-value, high-volume telemetry you might happily accept silent loss and skip all of this. For work where &lt;em&gt;every single message must result in an action&lt;/em&gt;, the cost is obviously worth it. As always, the right answer depends on what the data is worth.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The most memorable bugs are the ones that teach you to distrust a word you’d been using carelessly. For me, that word was “success.” A &lt;code&gt;2xx&lt;/code&gt; is not success. A committed offset is not success. Success is &lt;em&gt;the effect you actually wanted, confirmed to be durable.&lt;/em&gt; Everything else is just a system being polite.&lt;/p&gt;

&lt;p&gt;If you take one thing away: go look at your event-driven pipelines and ask where you commit progress based on an &lt;em&gt;acknowledgment&lt;/em&gt; rather than a &lt;em&gt;confirmation&lt;/em&gt;. If those two things are wired together, you probably have an acknowledgment gap hiding in there too - quietly green on every dashboard, right up until someone asks where their data went.&lt;/p&gt;

&lt;p&gt;Thanks for reading ✌️&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;a href="https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/202" rel="noopener noreferrer"&gt;HTTP 202 Accepted - MDN&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://en.wikipedia.org/wiki/Apache_Kafka" rel="noopener noreferrer"&gt;Apache Kafka - Wikipedia&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://en.wikipedia.org/wiki/Exponential_backoff" rel="noopener noreferrer"&gt;Exponential backoff - Wikipedia&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://en.wikipedia.org/wiki/Dead_letter_queue" rel="noopener noreferrer"&gt;Dead letter queue - Wikipedia&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://en.wikipedia.org/wiki/Two_Generals%27_Problem" rel="noopener noreferrer"&gt;The Two Generals’ Problem - Wikipedia&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>devchallenge</category>
      <category>bugsmash</category>
      <category>architecture</category>
      <category>distributedsystems</category>
    </item>
    <item>
      <title>Introducing gh-weekly-updates - Automate Your Weekly GitHub Impact Summaries</title>
      <dc:creator>Sahan</dc:creator>
      <pubDate>Sat, 21 Mar 2026 22:22:00 +0000</pubDate>
      <link>https://dev.to/sahan/introducing-gh-weekly-updates-automate-your-weekly-github-impact-summaries-1f1c</link>
      <guid>https://dev.to/sahan/introducing-gh-weekly-updates-automate-your-weekly-github-impact-summaries-1f1c</guid>
      <description>&lt;p&gt;If you are anything like me, you’ve probably spent a Friday afternoon trying to remember everything you did that week. Maybe it’s for a standup, a 1:1 with your manager, or just to keep track of your own progress. You end up clicking through PRs, issues, and Slack threads, trying to piece together a coherent story. It’s tedious, and honestly, it’s time you could spend doing actual work.&lt;/p&gt;

&lt;p&gt;That’s why I built &lt;a href="https://github.com/sahansera/gh-weekly-updates" rel="noopener noreferrer"&gt;&lt;strong&gt;gh-weekly-updates&lt;/strong&gt;&lt;/a&gt; - a CLI tool that automatically collects your GitHub activity and generates a structured weekly summary using AI.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fvc015gr9a7oohb9o2m8r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fvc015gr9a7oohb9o2m8r.png" alt="pypi" width="798" height="179"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;💡 &lt;strong&gt;GitHub repo&lt;/strong&gt; : &lt;a href="https://github.com/sahansera/gh-weekly-updates" rel="noopener noreferrer"&gt;github.com/sahansera/gh-weekly-updates&lt;/a&gt;. It’s open source and available on PyPI!&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;As engineers, we’re constantly shipping code, reviewing PRs, filing issues, and jumping into discussions. But when it comes time to reflect on the week, all that context is scattered across repos. I wanted something that could:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Pull all my GitHub activity into one place&lt;/li&gt;
&lt;li&gt;Summarise it in a way that highlights what actually matters&lt;/li&gt;
&lt;li&gt;Run on a schedule so I don’t have to think about it&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I couldn’t find anything that did exactly this, so I built it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What It Does
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;gh-weekly-updates&lt;/code&gt; connects to the GitHub API, collects your activity for a given period, and sends it to an AI model (via &lt;a href="https://github.com/marketplace/models" rel="noopener noreferrer"&gt;GitHub Models&lt;/a&gt;) to produce a structured Markdown summary.&lt;/p&gt;

&lt;p&gt;Here’s what it picks up:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Pull requests&lt;/strong&gt; you authored (with merge status, additions/deletions, changed files)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pull requests&lt;/strong&gt; you reviewed&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Issues&lt;/strong&gt; you created&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Issue comments&lt;/strong&gt; you left&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Discussions&lt;/strong&gt; you started or participated in&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The output is grouped by project or theme and structured into sections like Wins, Challenges, and What’s Next. You can also customise the prompt to match whatever format your team uses.&lt;/p&gt;

&lt;h2&gt;
  
  
  Getting Started
&lt;/h2&gt;

&lt;p&gt;It’s a Python CLI tool, so you can install it with pip:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;gh-weekly-updates
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then just run it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# If you're already logged in with the GitHub CLI&lt;/span&gt;
gh auth login
gh-weekly-updates
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That’s it. It will auto-discover repos you contributed to in the past week and generate a summary.&lt;/p&gt;

&lt;p&gt;You can also point it at specific repos and date ranges:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gh-weekly-updates &lt;span class="nt"&gt;--since&lt;/span&gt; 2026-02-09 &lt;span class="nt"&gt;--until&lt;/span&gt; 2026-02-16 &lt;span class="nt"&gt;--repos&lt;/span&gt; my-org/my-repo
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Configuration
&lt;/h2&gt;

&lt;p&gt;For more control, you can create a &lt;code&gt;config.yaml&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;org&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;my-org&lt;/span&gt;

&lt;span class="na"&gt;repos&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;my-org/api-service&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;my-org/web-app&lt;/span&gt;

&lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;openai/gpt-4.1&lt;/span&gt;

&lt;span class="c1"&gt;# Automatically push the summary to a repo&lt;/span&gt;
&lt;span class="na"&gt;push_repo&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;my-user/my-weekly-updates&lt;/span&gt;

&lt;span class="c1"&gt;# Customise the AI prompt&lt;/span&gt;
&lt;span class="na"&gt;prompt_file&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;my-prompt.txt&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The config supports everything from repo lists to custom prompts. You can even swap out the AI model if you have a preference.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running on a Schedule with GitHub Actions
&lt;/h2&gt;

&lt;p&gt;This is where it gets really useful. You can set up a GitHub Actions workflow to run it every Monday morning and push the summary to a repo automatically:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Weekly Summary&lt;/span&gt;

&lt;span class="na"&gt;on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;schedule&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;cron&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;0&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;9&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;1'&lt;/span&gt; &lt;span class="c1"&gt;# Every Monday at 9am UTC&lt;/span&gt;
  &lt;span class="na"&gt;workflow_dispatch&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;

&lt;span class="na"&gt;jobs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;summarise&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-latest&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/checkout@v4&lt;/span&gt;

      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/setup-python@v5&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;python-version&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;3.12'&lt;/span&gt;

      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;pip install gh-weekly-updates&lt;/span&gt;

      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Generate summary&lt;/span&gt;
        &lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;GITHUB_TOKEN&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.GH_PAT }}&lt;/span&gt; &lt;span class="c1"&gt;# must be named GITHUB_TOKEN&lt;/span&gt;
        &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;gh-weekly-updates --config config.yaml&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now every Monday, you get a fresh summary committed to your repo. No manual effort required.&lt;/p&gt;

&lt;h2&gt;
  
  
  Custom Prompts
&lt;/h2&gt;

&lt;p&gt;The default prompt produces a summary with Wins, Challenges, and What’s Next sections. But you can tailor it to your needs. For example, if your team does impact-style updates, you might want sections like Strategic Influence or Next Steps.&lt;/p&gt;

&lt;p&gt;Just create a text file with your prompt and reference it in your config:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;prompt_file&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;my-custom-prompt.txt&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The prompt receives all your raw activity data as context, so you can shape the output however you like.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Open Source?
&lt;/h2&gt;

&lt;p&gt;I initially built this for myself to automate my own weekly updates at work. But I figured other engineers probably have the same problem, so I cleaned it up and open-sourced it. The tool is intentionally simple - it does one thing and tries to do it well.&lt;/p&gt;

&lt;p&gt;If you find it useful, give it a ⭐ on &lt;a href="https://github.com/sahansera/gh-weekly-updates" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;. And if you have ideas for improvements, PRs and issues are always welcome!&lt;/p&gt;

&lt;h2&gt;
  
  
  What’s Next
&lt;/h2&gt;

&lt;p&gt;A few things I’m thinking about for future releases:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;More activity sources&lt;/strong&gt; : Picking up commit messages, release notes, and code review comments&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multiple output formats&lt;/strong&gt; : Slack messages, email digests, Notion pages&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Team summaries&lt;/strong&gt; : Aggregate activity across a whole team, not just one person&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If any of these sound interesting to you, feel free to open an issue or start a discussion on the repo.&lt;/p&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GitHub&lt;/strong&gt; : &lt;a href="https://github.com/sahansera/gh-weekly-updates" rel="noopener noreferrer"&gt;github.com/sahansera/gh-weekly-updates&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;PyPI&lt;/strong&gt; : &lt;a href="https://pypi.org/project/gh-weekly-updates/" rel="noopener noreferrer"&gt;pypi.org/project/gh-weekly-updates&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GitHub Models&lt;/strong&gt; : &lt;a href="https://github.com/marketplace/models" rel="noopener noreferrer"&gt;github.com/marketplace/models&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Thanks for reading! If you have any questions, feel free to reach out on &lt;a href="https://twitter.com/_SahanSera" rel="noopener noreferrer"&gt;Twitter&lt;/a&gt; or drop a comment below. 🤗&lt;/p&gt;

</description>
      <category>github</category>
      <category>python</category>
      <category>productivity</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Deploying GitHub Self-Hosted Runners on Your Home Kubernetes Cluster with ARC</title>
      <dc:creator>Sahan</dc:creator>
      <pubDate>Tue, 05 Aug 2025 12:28:00 +0000</pubDate>
      <link>https://dev.to/sahan/deploying-github-self-hosted-runners-on-your-home-kubernetes-cluster-with-arc-3gaj</link>
      <guid>https://dev.to/sahan/deploying-github-self-hosted-runners-on-your-home-kubernetes-cluster-with-arc-3gaj</guid>
      <description>&lt;p&gt;If you followed my last post, &lt;a href="https://sahansera.dev/building-home-lab-kubernetes-cluster-old-hardware-k3s/" rel="noopener noreferrer"&gt;Building a Home Lab Kubernetes Cluster with Old Hardware and k3s&lt;/a&gt;, you now have a proper x86 Kubernetes cluster humming away on your old laptops. So, what’s next? Time to put that cluster to work—let’s run GitHub Actions jobs on your own hardware!&lt;/p&gt;

&lt;p&gt;Why? Because if you have got decent hardware - self-hosting your runners might be faster, gives you full control (no GitHub minutes limit!), and lets you run bigger jobs (CI, builds, ML, you name it) on your home infra. And with &lt;a href="https://github.com/actions/actions-runner-controller" rel="noopener noreferrer"&gt;Actions Runner Controller - ARC&lt;/a&gt;, managing runners at scale on Kubernetes is surprisingly easy.&lt;/p&gt;

&lt;p&gt;Here’s how to set it all up.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is ARC and Why Should You Care?
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/actions/actions-runner-controller" rel="noopener noreferrer"&gt;ARC (Actions Runner Controller)&lt;/a&gt; is an open-source Kubernetes operator from GitHub. It spins up and manages GitHub Actions runners as Kubernetes pods—no more manually registering runners, no more pets, just cattle. Runners auto-scale up and down as jobs arrive. It’s perfect for CI/CD, especially on clusters you own.&lt;/p&gt;

&lt;p&gt;Here’s a high-level view of how it works under the hood&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Feie9qn3jfwmtedxba37s.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Feie9qn3jfwmtedxba37s.png" alt="ARC Architecture" width="800" height="514"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Working Kubernetes cluster (see &lt;a href="https://sahansera.dev/building-home-lab-kubernetes-cluster-old-hardware-k3s/" rel="noopener noreferrer"&gt;previous post&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;kubectl&lt;/code&gt; and &lt;a href="https://helm.sh/" rel="noopener noreferrer"&gt;&lt;code&gt;helm&lt;/code&gt;&lt;/a&gt; installed on your machine&lt;/li&gt;
&lt;li&gt;A GitHub Personal Access Token (PAT) with &lt;code&gt;repo&lt;/code&gt; and &lt;code&gt;admin:org&lt;/code&gt; scopes&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  1️⃣ Pre-Setup: Quick Checks
&lt;/h2&gt;

&lt;p&gt;Make sure you have what you need:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;which helm
kubectl version &lt;span class="nt"&gt;--client&lt;/span&gt;
helm list &lt;span class="nt"&gt;-A&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If those commands work, you’re good to go.&lt;/p&gt;

&lt;h2&gt;
  
  
  2️⃣ Install ARC Controller
&lt;/h2&gt;

&lt;p&gt;Let’s install the ARC controller into your control plane namespace:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;NAMESPACE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"actions-runner-controller"&lt;/span&gt;
helm &lt;span class="nb"&gt;install &lt;/span&gt;arc &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--namespace&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;NAMESPACE&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--create-namespace&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  oci://ghcr.io/actions/actions-runner-controller-charts/gha-runner-scale-set-controller
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This deploys the controller which manages your runners.&lt;/p&gt;

&lt;h2&gt;
  
  
  3️⃣ Deploy a Runner Scale Set
&lt;/h2&gt;

&lt;p&gt;Time to create the runners that will actually do the work. Replace the example GitHub URL and PAT with your details:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;INSTALLATION_NAME&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"arc-runner-set"&lt;/span&gt;
&lt;span class="nv"&gt;NAMESPACE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"arc-runners"&lt;/span&gt;
&lt;span class="nv"&gt;GITHUB_CONFIG_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"https://github.com/youruser/yourrepo"&lt;/span&gt;
&lt;span class="nv"&gt;GITHUB_PAT&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"ghp_123456..."&lt;/span&gt;

helm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;INSTALLATION_NAME&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--namespace&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;NAMESPACE&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--create-namespace&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--set&lt;/span&gt; &lt;span class="nv"&gt;githubConfigUrl&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;GITHUB_CONFIG_URL&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--set&lt;/span&gt; githubConfigSecret.github_token&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;GITHUB_PAT&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  oci://ghcr.io/actions/actions-runner-controller-charts/gha-runner-scale-set
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;GITHUB_CONFIG_URL&lt;/code&gt;: The repo or org you want to run jobs for.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;GITHUB_PAT&lt;/code&gt;: Your Personal Access Token.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  4️⃣ Check That It’s Working
&lt;/h2&gt;

&lt;p&gt;Verify the controller and runner pods are up:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Controller&lt;/span&gt;
kubectl get pods &lt;span class="nt"&gt;-n&lt;/span&gt; actions-runner-controller

&lt;span class="c"&gt;# Runners&lt;/span&gt;
kubectl get pods &lt;span class="nt"&gt;-n&lt;/span&gt; arc-runners

&lt;span class="c"&gt;# Runner set status&lt;/span&gt;
kubectl get AutoscalingRunnerSet &lt;span class="nt"&gt;-A&lt;/span&gt;
kubectl describe AutoscalingRunnerSet arc-runner-set &lt;span class="nt"&gt;-n&lt;/span&gt; arc-runners
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You should see your runners show up as pods. If you trigger a workflow in your GitHub repo, you’ll see a pod spin up, do the job, and then shut down—magic.&lt;/p&gt;

&lt;h2&gt;
  
  
  5️⃣ Testing It Out
&lt;/h2&gt;

&lt;p&gt;Here’s the fun part. Create a simple GitHub Actions workflow in your repo to test the runners:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Test ARC Runners&lt;/span&gt;
&lt;span class="na"&gt;on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;push&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;

&lt;span class="na"&gt;jobs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;build&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;arc-runners&lt;/span&gt; &lt;span class="c1"&gt;# This tells GitHub to use your self-hosted runners. Use the NAMESPACE name you defined in step 3.&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Checkout code&lt;/span&gt;
        &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/checkout@v2&lt;/span&gt;

      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Run a script&lt;/span&gt;
        &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;echo "Hello from ARC Runner!"&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;List files&lt;/span&gt;
        &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ls -la&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here’s ci.yaml from my githubstats repo if you need a working example: &lt;a href="https://github.com/sahansera/githubstats/blob/main/.github/workflows/ci.yml" rel="noopener noreferrer"&gt;ci.yaml&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Make a commit to trigger the workflow. You should see the runner pod spin up, execute the job, and then terminate.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fvaaybcapar2anb3sr9a6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fvaaybcapar2anb3sr9a6.png" alt="arc self hosted github runners k8s 1" width="456" height="392"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;center&gt;
Here's the workflow running on my ARC self-hosted runner
&lt;/center&gt;

&lt;h2&gt;
  
  
  6️⃣ Monitoring (Bonus: Grafana)
&lt;/h2&gt;

&lt;p&gt;Want to geek out and monitor your runners? If you’ve set up Prometheus/Grafana (see my upcoming post if not!), you can:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Check pod CPU/memory usage&lt;/li&gt;
&lt;li&gt;Track how many runners are running&lt;/li&gt;
&lt;li&gt;See logs for each pod&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Handy queries for Grafana dashboards:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# Number of ARC runner pods
kube_pod_status_phase{namespace="arc-runners", phase="Running"}

# Pod CPU usage
rate(container_cpu_usage_seconds_total{namespace="arc-runners"}[5m])
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here's mine:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fjmeo8nezem3f4fvqiv15.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fjmeo8nezem3f4fvqiv15.png" alt="arc self hosted github runners k8s 2" width="800" height="193"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;center&gt;
Here's what the workflow looks like in action
&lt;/center&gt;

&lt;h2&gt;
  
  
  6️⃣ Useful Commands for Day-to-Day Ops
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Watch runner scaling in real-time&lt;/span&gt;
kubectl get AutoscalingRunnerSet &lt;span class="nt"&gt;-n&lt;/span&gt; arc-runners &lt;span class="nt"&gt;-w&lt;/span&gt;

&lt;span class="c"&gt;# See events and troubleshoot&lt;/span&gt;
kubectl get events &lt;span class="nt"&gt;-n&lt;/span&gt; arc-runners &lt;span class="nt"&gt;--sort-by&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;.metadata.creationTimestamp

&lt;span class="c"&gt;# Pod logs (for a specific runner)&lt;/span&gt;
kubectl logs &lt;span class="nt"&gt;-n&lt;/span&gt; arc-runners &amp;lt;pod-name&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Troubleshooting Tips
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Runner not connecting?&lt;/strong&gt; Double-check your PAT and network access.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pods stuck or crash-looping?&lt;/strong&gt; Check logs for clues and make sure your cluster has enough resources.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Can’t see runners in GitHub?&lt;/strong&gt; Make sure the config URL matches your repo/org and the PAT has correct scopes.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  That’s It! You’re Running GitHub Actions on Your Own Cluster
&lt;/h2&gt;

&lt;p&gt;You now have GitHub Actions jobs running &lt;em&gt;at home&lt;/em&gt; on your cluster, scaling up and down automatically. No more slow or limited runners. Your home lab just levelled up—CI/CD, builds, ML, you name it.&lt;/p&gt;

&lt;p&gt;Stay tuned for my next post where I’ll show you how to get beautiful observability dashboards and set up alerting for your home cluster.&lt;br&gt;&lt;br&gt;
Happy automating!&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Questions? Want to show off your setup? Ping me on &lt;a href="https://twitter.com/_SahanSera" rel="noopener noreferrer"&gt;X (Twitter)&lt;/a&gt; or drop a comment below!&lt;/em&gt;&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>github</category>
      <category>devops</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Building a Home Lab Kubernetes Cluster with Old Hardware and k3s</title>
      <dc:creator>Sahan</dc:creator>
      <pubDate>Fri, 01 Aug 2025 10:30:00 +0000</pubDate>
      <link>https://dev.to/sahan/building-a-home-lab-kubernetes-cluster-with-old-hardware-and-k3s-f76</link>
      <guid>https://dev.to/sahan/building-a-home-lab-kubernetes-cluster-with-old-hardware-and-k3s-f76</guid>
      <description>&lt;p&gt;If you've read my previous &lt;a href="https://sahansera.dev/building-your-own-private-kubernetes-cluster-on-a-raspberry-pi-4-with-k3s/" rel="noopener noreferrer"&gt;post&lt;/a&gt; about building a Raspberry Pi k3s cluster, you know I'm a huge fan of home labs. There's something uniquely satisfying about getting distributed systems running on a bunch of hardware you already own. This time, though, I wanted something a bit more powerful-a cluster that could handle not just learning and tinkering, but also heavier dev, CI/CD, and even ML workloads.&lt;/p&gt;

&lt;p&gt;And as it turns out, there's a ton you can do with a handful of old laptops, a simple switch, and Ubuntu Server. If you're thinking about upgrading your home cluster or want to avoid vendor lock-in, this one's for you.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Move From Raspberry Pi to x86?
&lt;/h2&gt;

&lt;p&gt;Don't get me wrong-RPi clusters are awesome for learning, hacking, and even some light home automation. But if you've ever tried running real dev pipelines, CI/CD, or any data-heavy ML stuff, you'll hit those limits &lt;em&gt;fast&lt;/em&gt;. Plus, WiFi can get a bit flaky when you're trying to keep nodes connected under load.&lt;/p&gt;

&lt;p&gt;I had a few spare laptops sitting around, and it made perfect sense to give them a second life and push them to their limits. Bonus: with Ethernet and a decent switch, you get much more reliable connectivity.&lt;/p&gt;

&lt;h2&gt;
  
  
  Network Setup: Keep It Simple (and Reliable)
&lt;/h2&gt;

&lt;p&gt;For this build, I kept things straightforward: all nodes are wired to a simple network switch. This gives much better performance and reliability compared to WiFi, but honestly, you could still pull this off over wireless if needed.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F8j0sn4wz9bupla03rrkg.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F8j0sn4wz9bupla03rrkg.jpg" alt="Network Switch" width="800" height="343"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;center&gt;
   My network switch - how it all started 🛜
&lt;/center&gt;

&lt;p&gt;&lt;strong&gt;A quick tip:&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Reserve static IPs for each node using your router's DHCP reservations.
&lt;/li&gt;
&lt;li&gt;This makes node management, SSH, and Kubernetes networking so much easier.
&lt;/li&gt;
&lt;li&gt;Check your DHCP table and make sure every machine has a unique, predictable IP address.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Getting Ubuntu Server Set Up on Each Node
&lt;/h2&gt;

&lt;p&gt;Here's a high level diagram of what my setup looks like:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fd5y197omtv8e22ipk8m8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fd5y197omtv8e22ipk8m8.png" alt="Home Lab Setup Diagram" width="800" height="338"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Let's talk about prepping each machine. My old laptops had a new lease on life running Ubuntu Server. I recommend Ubuntu Server LTS for its simplicity and compatibility. If you're using VMs, just make sure to set the NIC to bridged mode so each VM acts as a full member of your home LAN.&lt;/p&gt;

&lt;p&gt;Here's my go-to checklist for each node:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Update the system and install SSH:&lt;/strong&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   &lt;span class="nb"&gt;sudo &lt;/span&gt;apt update &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;sudo &lt;/span&gt;apt upgrade &lt;span class="nt"&gt;-y&lt;/span&gt;
   &lt;span class="nb"&gt;sudo &lt;/span&gt;apt &lt;span class="nb"&gt;install &lt;/span&gt;openssh-server &lt;span class="nt"&gt;-y&lt;/span&gt;
   &lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl &lt;span class="nb"&gt;enable &lt;/span&gt;ssh
   &lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl start ssh
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Set a memorable hostname:&lt;/strong&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   &lt;span class="nb"&gt;sudo &lt;/span&gt;hostnamectl set-hostname master   &lt;span class="c"&gt;# or worker-1, etc.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;(Optional) Update /etc/hosts for local resolution:&lt;/strong&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   &lt;span class="nb"&gt;sudo &lt;/span&gt;nano /etc/hosts
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Change any lines like &lt;code&gt;127.0.1.1 old-hostname&lt;/code&gt; to your new hostname.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Find your node's IP address:&lt;/strong&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   ip a
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use this IP for your DHCP reservation so it always stays the same.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Test SSH from another machine:&lt;/strong&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   ssh &amp;lt;username&amp;gt;@&amp;lt;node-ip&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Installing k3s: Lightweight, Powerful, Familiar
&lt;/h2&gt;

&lt;p&gt;This process feels pretty magical the first time you see it all come together. I stuck with &lt;a href="https://k3s.io/" rel="noopener noreferrer"&gt;k3s&lt;/a&gt;-lightweight and perfect for home or edge clusters.&lt;/p&gt;

&lt;h3&gt;
  
  
  On the Control Plane Node ("master")
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Install k3s:&lt;/strong&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   curl &lt;span class="nt"&gt;-sfL&lt;/span&gt; https://get.k3s.io | sh -
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Check that k3s is running:&lt;/strong&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   &lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl status k3s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Get your cluster join token:&lt;/strong&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   &lt;span class="nb"&gt;sudo cat&lt;/span&gt; /var/lib/rancher/k3s/server/node-token
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Copy this somewhere safe-you'll need it for your worker nodes.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Get the node's LAN IP:&lt;/strong&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   ip a
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Let's say your master node IP is &lt;code&gt;192.168.0.200&lt;/code&gt;.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Check your cluster status:&lt;/strong&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   kubectl get nodes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;If you ever want to reset/reinstall k3s, just run:&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo&lt;/span&gt; /usr/local/bin/k3s-uninstall.sh
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  On Each Worker Node ("worker-1", "worker-2", etc.)
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Join the cluster:&lt;/strong&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   curl &lt;span class="nt"&gt;-sfL&lt;/span&gt; https://get.k3s.io | &lt;span class="nv"&gt;K3S_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;https://192.168.0.200:6443 &lt;span class="nv"&gt;K3S_TOKEN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&amp;lt;your-token-here&amp;gt; &lt;span class="nv"&gt;INSTALL_K3S_EXEC&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"--node-ip=&amp;lt;worker-ip&amp;gt;"&lt;/span&gt; sh -
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Replace &lt;code&gt;192.168.0.200&lt;/code&gt; with your control plane's IP.&lt;br&gt;&lt;br&gt;
   Replace &lt;code&gt;&amp;lt;your-token-here&amp;gt;&lt;/code&gt; with the token from your master node.&lt;br&gt;&lt;br&gt;
   Replace &lt;code&gt;&amp;lt;worker-ip&amp;gt;&lt;/code&gt; with this worker's IP, e.g., &lt;code&gt;192.168.0.201&lt;/code&gt;.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;That's it!&lt;/strong&gt; The node will auto-register with your cluster. No manual kubeconfig needed.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Check the cluster again from the master:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   kubectl get nodes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You should see both &lt;code&gt;master&lt;/code&gt; and your worker(s) as &lt;code&gt;Ready&lt;/code&gt;!&lt;/p&gt;




&lt;h2&gt;
  
  
  Troubleshooting and Gotchas
&lt;/h2&gt;

&lt;p&gt;As with any home lab project, there are always a few snags-here's what I ran into and how to fix it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Node not joining?&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Double-check your token (no spaces, copy the whole thing).&lt;/li&gt;
&lt;li&gt;Make sure you can &lt;code&gt;ping&lt;/code&gt; the control plane from your worker node.&lt;/li&gt;
&lt;li&gt;Firewalls can get in the way. Ensure port 6443 is open between nodes.&lt;/li&gt;
&lt;li&gt;If you get hostname conflicts, just change the hostname and restart the k3s agent (&lt;code&gt;sudo systemctl restart k3s-agent&lt;/code&gt;).&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Cluster state looks weird after hostname change?&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You might see both the old and new hostnames in &lt;code&gt;kubectl get nodes&lt;/code&gt;. Just delete the old one:
&lt;/li&gt;
&lt;/ul&gt;

&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl delete node &amp;lt;old-node-name&amp;gt;
&lt;/code&gt;&lt;/pre&gt;


&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;IP address keeps changing?&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Make a DHCP reservation for each node's MAC address in your router.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;WiFi unreliable?&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Ethernet is always the best bet for clusters, especially for heavy workloads.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What's Next?
&lt;/h2&gt;

&lt;p&gt;With the basics in place, you're ready to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Run and test real-world apps and dev environments&lt;/li&gt;
&lt;li&gt;Build your own CI/CD pipelines&lt;/li&gt;
&lt;li&gt;Experiment with ML workloads on dedicated nodes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This way, you get all the power of a real Kubernetes lab, but full control and no monthly surprises from a cloud provider.&lt;/p&gt;




&lt;h2&gt;
  
  
  Lessons Learned
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;RPi clusters are great for edge computing and learning,&lt;/strong&gt; but once you need real muscle, x86 hardware makes a &lt;em&gt;huge&lt;/em&gt; difference.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;WiFi is convenient, but Ethernet is still king&lt;/strong&gt; for stable clusters-especially if you're doing anything performance-sensitive.&lt;/li&gt;
&lt;li&gt;Old laptops/desktops are a goldmine for home lab builds. Don't let them collect dust!&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Final Thoughts
&lt;/h2&gt;

&lt;p&gt;If you have old machines and a bit of curiosity, you can build a cluster that's surprisingly capable-no Raspberry Pis required.&lt;br&gt;&lt;br&gt;
And if you've already done it with RPis, this is the perfect upgrade path.&lt;br&gt;&lt;br&gt;
Keep an eye out for my next post where I'll deep-dive into adding observability and tooling!&lt;/p&gt;

&lt;p&gt;Happy clustering! 🫡&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>tutorial</category>
      <category>distributedsystems</category>
      <category>linux</category>
    </item>
    <item>
      <title>How Does the Python Virtual Environment Work?</title>
      <dc:creator>Sahan</dc:creator>
      <pubDate>Mon, 28 Jul 2025 00:17:00 +0000</pubDate>
      <link>https://dev.to/sahan/how-does-the-python-virtual-environment-work-2e1l</link>
      <guid>https://dev.to/sahan/how-does-the-python-virtual-environment-work-2e1l</guid>
      <description>&lt;p&gt;When you start working with Python, one of the first recommendations you’ll hear is to use a “virtual environment.” But what exactly is a Python virtual environment, and how does it work under the hood?&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem: Dependency Hell
&lt;/h2&gt;

&lt;p&gt;Python projects often rely on third-party libraries. If you install packages globally, different projects can end up fighting over package versions. This is called “dependency hell.” For example, Project A might require &lt;code&gt;requests==2.25&lt;/code&gt;, while Project B needs &lt;code&gt;requests==2.31&lt;/code&gt;. Installing both globally can cause conflicts and break your projects.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Solution: Virtual Environments
&lt;/h2&gt;

&lt;p&gt;A virtual environment is an isolated workspace for your Python project. It lets you install packages locally, so each project can have its own dependencies, regardless of what’s installed elsewhere on your system.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Does It Work?
&lt;/h2&gt;

&lt;p&gt;When you create a virtual environment (using &lt;code&gt;python -m venv myenv&lt;/code&gt; or &lt;code&gt;virtualenv myenv&lt;/code&gt;), Python does the following:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Creates a Dedicated Directory Structure&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A bin/ or Scripts/ directory with a Python executable and activation scripts&lt;/li&gt;
&lt;li&gt;A lib/ directory with a copy of the Python standard library&lt;/li&gt;
&lt;li&gt;A pyvenv.cfg config file for metadata
&lt;/li&gt;
&lt;/ul&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;myenv/
├── bin/                &lt;span class="c"&gt;# (Note: Scripts\ on Windows)&lt;/span&gt;
│   ├── activate        &lt;span class="c"&gt;# Shell script to activate the environment (Unix)&lt;/span&gt;
│   ├── activate.bat    &lt;span class="c"&gt;# Batch script (Windows CMD)&lt;/span&gt;
│   ├── Activate.ps1    &lt;span class="c"&gt;# PowerShell script (Windows PowerShell)&lt;/span&gt;
│   ├── pip             &lt;span class="c"&gt;# Environment-specific pip&lt;/span&gt;
│   └── python          &lt;span class="c"&gt;# Environment-specific Python interpreter&lt;/span&gt;
├── lib/
│   └── pythonX.Y/
│       └── site-packages/  &lt;span class="c"&gt;# Installed packages go here&lt;/span&gt;
├── pyvenv.cfg          &lt;span class="c"&gt;# Configuration file with environment metadata&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Configures a Standalone Python Interpreter:&lt;/strong&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The environment includes its own Python executable (or a symlink to it), ensuring that all commands run from within the environment use the correct interpreter.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;On most systems, this is a &lt;strong&gt;symlink or copy&lt;/strong&gt; of the base Python binary&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;This interpreter respects only the packages installed within the environment&lt;br&gt;
&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;which python
/path/to/myenv/bin/python
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Sets Up Local Package Management:&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Each environment gets its own site-packages directory:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;When you run pip install, packages go here instead of the global location&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;This isolation prevents version conflicts and makes dependency management predictable&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Creates Activation Scripts:&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Activation scripts help you &lt;em&gt;enter&lt;/em&gt; the environment by:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Modifying your &lt;code&gt;$PATH&lt;/code&gt; so that python and pip point to the virtual environment&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Optionally updating your shell prompt (e.g., showing (myenv))&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Ensuring commands are scoped to the environment&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These scripts are OS-specific:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Unix/macOS: &lt;code&gt;source myenv/bin/activate&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Windows CMD: &lt;code&gt;myenv\Scripts\activate.bat&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;PowerShell: &lt;code&gt;myenv\Scripts\Activate.ps1&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Includes a Configuration File:&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The &lt;code&gt;pyvenv.cfg&lt;/code&gt; file records metadata about the environment, including the Python version and the location of the base interpreter.&lt;/p&gt;

&lt;p&gt;This file stores:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The Python version used&lt;/li&gt;
&lt;li&gt;The path to the base interpreter&lt;/li&gt;
&lt;li&gt;Whether system site packages are accessible (default: no)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This metadata is used when running the environment to preserve consistent behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  Command Resolution
&lt;/h2&gt;

&lt;p&gt;So what happens when you are in a venv as opposed to running python command globally? The diagram below illustrates how Python and pip commands are resolved with and without a virtual environment. When no virtual environment is active, your system’s PATH directs these commands to the globally installed Python interpreter and packages.&lt;/p&gt;

&lt;p&gt;However, once a virtual environment is activated, the PATH is modified to point to the environment’s own executables. This ensures that all Python commands and package installations stay isolated within the virtual environment, avoiding conflicts with system-wide installations.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fyheyolmuaa7jq0zfioq6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fyheyolmuaa7jq0zfioq6.png" alt="Command Resolution" width="800" height="635"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Is This Powerful?
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Isolation:&lt;/strong&gt; Each project gets its own dependencies and versions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reproducibility:&lt;/strong&gt; You can lock dependencies with a &lt;code&gt;requirements.txt&lt;/code&gt; or &lt;code&gt;pyproject.toml&lt;/code&gt; file, making it easy for others (or yourself in the future) to recreate the environment.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No Admin Rights Needed:&lt;/strong&gt; You don’t need system-wide permissions to install packages.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Advanced Use Cases
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Multiple Python Versions:&lt;/strong&gt; Use virtual environments to test your code against different Python versions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Custom Activation Scripts:&lt;/strong&gt; Modify the activation script to set environment variables specific to your project.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Integration with CI/CD:&lt;/strong&gt; Virtual environments are essential for setting up isolated builds in CI/CD pipelines.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Under the Hood: What’s Really Happening?
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The virtual environment is just a directory with a specific structure.&lt;/li&gt;
&lt;li&gt;No containers or VMs are involved-just clever manipulation of paths and environment variables.&lt;/li&gt;
&lt;li&gt;Deleting the virtual environment directory removes all installed packages for that project.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Debugging Tips
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;If activation doesn’t work, check your shell configuration.&lt;/li&gt;
&lt;li&gt;Use &lt;code&gt;python -m site&lt;/code&gt; to inspect the site-packages directory.&lt;/li&gt;
&lt;li&gt;Verify the &lt;code&gt;pyvenv.cfg&lt;/code&gt; file for any misconfigurations.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Alternative Tools:&lt;/strong&gt; &lt;code&gt;venv&lt;/code&gt; is standard for Python 3.3+, but tools like &lt;code&gt;virtualenv&lt;/code&gt;, &lt;code&gt;conda&lt;/code&gt;, or &lt;code&gt;pipenv&lt;/code&gt; exist for advanced use cases or older Python versions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Python virtual environments are a foundational tool for modern Python development. They solve the problem of dependency conflicts, make projects more portable, and keep your system clean. Whether you’re building a quick script or a large application, understanding how virtual environments work will save you countless headaches down the road.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.python.org/3/library/venv.html" rel="noopener noreferrer"&gt;Official Python venv Documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://peps.python.org/pep-0405/" rel="noopener noreferrer"&gt;PEP 405: Python Virtual Environments&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>python</category>
      <category>programming</category>
      <category>fundamentals</category>
    </item>
    <item>
      <title>Deep Dive - How Chunked Transfer Encoding Works</title>
      <dc:creator>Sahan</dc:creator>
      <pubDate>Fri, 04 Apr 2025 05:51:00 +0000</pubDate>
      <link>https://dev.to/sahan/deep-dive-how-chunked-transfer-encoding-works-4o9n</link>
      <guid>https://dev.to/sahan/deep-dive-how-chunked-transfer-encoding-works-4o9n</guid>
      <description>&lt;p&gt;Chunked transfer encoding is a key HTTP/1.1 feature that allows servers to stream data incrementally without knowing the total size of the response upfront. It’s particularly useful in streaming APIs, live updates, and large or dynamically-generated responses.&lt;/p&gt;

&lt;p&gt;In this post, we’ll practically explore how chunked transfer encoding works using the backend we developed in my previous blog post on &lt;a href="https://sahansera.dev/streaming-apis-python-nextjs-part2/" rel="noopener noreferrer"&gt;Streaming APIs with FastAPI and Next.js — Part 2&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  🤔 What is Chunked Transfer Encoding?
&lt;/h2&gt;

&lt;p&gt;Chunked transfer encoding modifies HTTP responses into a series of &lt;strong&gt;chunks&lt;/strong&gt; , each prefixed with its size in bytes. It allows servers to start sending response data immediately, without having to calculate the full content length beforehand.&lt;/p&gt;

&lt;p&gt;When &lt;code&gt;Transfer-Encoding: chunked&lt;/code&gt; is present, the client receives data &lt;strong&gt;incrementally&lt;/strong&gt; and knows the response has ended when a &lt;strong&gt;zero-length&lt;/strong&gt; chunk appears.&lt;/p&gt;

&lt;h2&gt;
  
  
  💻 Hands-on Example
&lt;/h2&gt;

&lt;p&gt;Let’s use the &lt;a href="https://github.com/sahansera/streaming-apis/tree/main/backend" rel="noopener noreferrer"&gt;FastAPI backend&lt;/a&gt; we built.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;make start-backend
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fuzpb8h76qu1vujzibmqr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fuzpb8h76qu1vujzibmqr.png" alt="understanding chunked transfer encoding 1" width="800" height="288"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Let’s hit the &lt;code&gt;/stream&lt;/code&gt; endpoint with &lt;code&gt;curl&lt;/code&gt; to see how chunked transfer encoding works in practice.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="nt"&gt;--raw&lt;/span&gt; http://localhost:8000/stream
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F34i8vtmna38udcolr4e2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F34i8vtmna38udcolr4e2.png" alt="understanding chunked transfer encoding-2" width="800" height="571"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;-i&lt;/code&gt;: Include headers in output.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;--raw&lt;/code&gt;: Disable curl’s automatic decoding, revealing raw chunked encoding.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Expected output:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="nt"&gt;--raw&lt;/span&gt; localhost:8000/stream
HTTP/1.1 200 OK
&lt;span class="nb"&gt;date&lt;/span&gt;: Mon, 31 Mar 2025 09:51:47 GMT
server: uvicorn
content-type: text/plain&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nv"&gt;charset&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;utf-8
Transfer-Encoding: chunked

1f
Waiting &lt;span class="k"&gt;for &lt;/span&gt;new log entries...

1f
Waiting &lt;span class="k"&gt;for &lt;/span&gt;new log entries...

1f
Waiting &lt;span class="k"&gt;for &lt;/span&gt;new log entries...

1f
Waiting &lt;span class="k"&gt;for &lt;/span&gt;new log entries...

1f
Waiting &lt;span class="k"&gt;for &lt;/span&gt;new log entries...

30
Simulated log entry at Mon Mar 31 20:21:53 2025

0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here’s a diagram to help visualize the chunked transfer encoding:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fzuokatgjbcg4qrp3p8ep.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fzuokatgjbcg4qrp3p8ep.png" alt="understanding chunked transfer encoding 3" width="800" height="599"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here’s what’s happening:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Each chunk starts with its length in hexadecimal (&lt;code&gt;1f&lt;/code&gt; = 31 bytes).&lt;/li&gt;
&lt;li&gt;The data follows the length, and the next chunk starts after a newline.&lt;/li&gt;
&lt;li&gt;The chunk with 30 represents a simulated log entry (&lt;code&gt;30&lt;/code&gt; = 48 bytes).&lt;/li&gt;
&lt;li&gt;The response ends with a zero-length chunk (&lt;code&gt;0&lt;/code&gt;).&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Note:&lt;/em&gt; This aligns with the techniques demonstrated in my previous blog series &lt;a href="https://sahansera.dev/streaming-apis-python-nextjs-part1/" rel="noopener noreferrer"&gt;Streaming APIs with FastAPI and Next.js (Part 1)&lt;/a&gt; and &lt;a href="https://sahansera.dev/streaming-apis-python-nextjs-part2/" rel="noopener noreferrer"&gt;Part 2&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  🛠️ &lt;strong&gt;Step-by-Step Breakdown of Chunking&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;We’ll be using the &lt;a href="https://github.com/sahansera/streaming-apis/blob/main/backend/api/index.py" rel="noopener noreferrer"&gt;index.py&lt;/a&gt;. Here’s exactly what’s happening under the hood:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Generator (&lt;code&gt;yield&lt;/code&gt;)&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Every time &lt;code&gt;yield&lt;/code&gt; is executed, the Starlette framework (used internally by FastAPI) receives a new piece of data to stream to the client.&lt;/li&gt;
&lt;li&gt;Each yielded data segment corresponds &lt;strong&gt;directly to one HTTP chunk&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example, the &lt;a href="https://github.com/sahansera/streaming-apis/blob/main/backend/api/index.py#L34" rel="noopener noreferrer"&gt;yielded line&lt;/a&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Waiting for new log entries...&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;is packaged into one HTTP chunk.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Starlette’s StreamingResponse Handling&lt;/strong&gt; :&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The &lt;code&gt;StreamingResponse&lt;/code&gt; from Starlette wraps the async generator.&lt;/li&gt;
&lt;li&gt;Starlette doesn’t wait until the generator finishes (which might be infinite). Instead, it immediately pushes each yielded chunk to the underlying ASGI server, typically Uvicorn.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;3. Uvicorn’s Chunk Formatting&lt;/strong&gt; :&lt;/p&gt;

&lt;p&gt;Uvicorn (the ASGI server you’re using) receives the yielded chunk from Starlette and &lt;strong&gt;formats it according to the HTTP/1.1 chunked transfer encoding specification&lt;/strong&gt; :&lt;/p&gt;

&lt;p&gt;Each chunk is transmitted as follows:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&amp;lt;chunk-size &lt;span class="k"&gt;in &lt;/span&gt;hexadecimal&amp;gt;&lt;span class="se"&gt;\r\n&lt;/span&gt;
&amp;lt;chunk-data&amp;gt;&lt;span class="se"&gt;\r\n&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here’s how one of your actual data chunks might look:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;1f&lt;span class="se"&gt;\r\n&lt;/span&gt;
Waiting &lt;span class="k"&gt;for &lt;/span&gt;new log entries...&lt;span class="se"&gt;\n\r\n&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;1f&lt;/code&gt; = 31 bytes, the exact length of &lt;code&gt;"Waiting for new log entries...\n"&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;4. Continuous Chunk Transmission&lt;/strong&gt; :&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Uvicorn immediately sends each formatted chunk down the TCP connection.&lt;/li&gt;
&lt;li&gt;Your client (like &lt;code&gt;curl&lt;/code&gt;) receives each chunk as soon as it’s sent, which allows incremental processing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;5. Ending the Stream&lt;/strong&gt; :&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If your generator ever completes (or if the server shuts down the connection), Uvicorn sends a special &lt;strong&gt;zero-length chunk&lt;/strong&gt; (&lt;code&gt;0\r\n\r\n&lt;/code&gt;) to indicate that transmission has ended.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Example final chunk:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;0&lt;span class="se"&gt;\r\n&lt;/span&gt;
&lt;span class="se"&gt;\r\n&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  🙋 What about HTTP/2 and HTTP/3?
&lt;/h2&gt;

&lt;p&gt;The short answer: HTTP/2+ does not use chunked encoding at all. In fact, the HTTP/2 specification explicitly forbids the use of the &lt;code&gt;Transfer-Encoding: chunked&lt;/code&gt; header; if a client incorrectly tries to send it, it’s considered a protocol error.&lt;/p&gt;

&lt;p&gt;Instead, HTTP/2 uses a more efficient binary framing layer that allows multiplexing multiple streams over a single connection. This means that chunked transfer encoding is not necessary in HTTP/2 and HTTP/3, as the protocol itself handles streaming more efficiently.&lt;/p&gt;




&lt;h2&gt;
  
  
  Key Takeaways
&lt;/h2&gt;

&lt;p&gt;Through these practical examples, you’ve seen firsthand how chunked transfer encoding enables incremental streaming of data:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Responses are sent as a series of chunks, each with a defined size.&lt;/li&gt;
&lt;li&gt;The end of data transmission is indicated by a zero-length chunk.&lt;/li&gt;
&lt;li&gt;Tools like &lt;code&gt;curl&lt;/code&gt;, Python frameworks like FastAPI, and browser developer tools help visualize and debug chunked encoding.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Understanding this helps you build better streaming APIs and debug complex HTTP interactions effectively.&lt;/p&gt;

&lt;p&gt;Happy Streaming! 🚀&lt;/p&gt;




&lt;h2&gt;
  
  
  Further Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://datatracker.ietf.org/doc/html/rfc9112#section-7.1" rel="noopener noreferrer"&gt;RFC 9112 – HTTP/1.1 Specification&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Transfer-Encoding" rel="noopener noreferrer"&gt;MDN Web Docs: Transfer-Encoding&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://sahansera.dev/streaming-apis-python-nextjs-part1/" rel="noopener noreferrer"&gt;Streaming APIs with FastAPI and Next.js (Part 1)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://sahansera.dev/streaming-apis-python-nextjs-part2/" rel="noopener noreferrer"&gt;Streaming APIs with FastAPI and Next.js (Part 2)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>http</category>
      <category>python</category>
      <category>fundamentals</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Upgrading sahansera.dev to Gatsby 5</title>
      <dc:creator>Sahan</dc:creator>
      <pubDate>Wed, 02 Apr 2025 08:10:00 +0000</pubDate>
      <link>https://dev.to/sahan/upgrading-sahanseradev-to-gatsby-5-3p99</link>
      <guid>https://dev.to/sahan/upgrading-sahanseradev-to-gatsby-5-3p99</guid>
      <description>&lt;p&gt;The time is now 1:45 AM and I’m finally done upgrading my blog to Gatsby V5. It’s been one heck of a ride, but it was worth it.&lt;/p&gt;

&lt;p&gt;This blog post is a reflection of my experience upgrading my blog to Gatsby V5. I’ll cover the challenges I faced, the solutions I found, and the lessons learned along the way.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Upgrade Journey
&lt;/h2&gt;

&lt;p&gt;After a two year long hiatus from blogging, I decided to dust off my old blog and give it a much-needed facelift. When I tried to run &lt;code&gt;gatsby develop&lt;/code&gt;, I was greeted with a slew of warnings and errors. It was clear that my blog was long overdue for an upgrade.&lt;/p&gt;

&lt;p&gt;I took a crack at it a few months ago, but I quickly got overwhelmed by the number of breaking changes and decided to put it on hold. This time it’s different because I was determined to fully utilize LLMs for my advantage.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F1kivd29wcna5ip0iw0e5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F1kivd29wcna5ip0iw0e5.png" alt="upgrading gatsby 5 1" width="800" height="78"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Sneak peak of my last 3 commits&lt;/em&gt; 🤣&lt;/p&gt;

&lt;p&gt;Being a weekend, which is filled with family time and chores, I decided to timebox it just for a few hours. I thought, “How hard can it be? Just a few package updates and some code tweaks, right?” Little did I know that this would turn into a full-blown adventure.&lt;/p&gt;

&lt;h3&gt;
  
  
  Dependency Overhauls
&lt;/h3&gt;

&lt;p&gt;One of the first tasks was to ensure all my plugins were updated to their latest non-breaking minor versions. I used &lt;a href="https://www.npmjs.com/package/npm-check-updates" rel="noopener noreferrer"&gt;&lt;code&gt;npm-check-updates&lt;/code&gt;&lt;/a&gt; to help with this.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ncu &lt;span class="nt"&gt;--interactive&lt;/span&gt; &lt;span class="nt"&gt;--format&lt;/span&gt; group
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fy6m9vy8d1v93lk9qvp7t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fy6m9vy8d1v93lk9qvp7t.png" alt="Example of ncu CLI tool in action" width="800" height="764"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Example of ncu CLI tool in action&lt;/em&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;💡 Tip: Always stay on a separate branch, use commits for each change you make, while doing updates.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;There were bunch of warnings, but nothing major. Next, I upgraded React and ReactDOM to v18. This was a bit tricky because I had to ensure that all my dependencies were compatible with React 18. I also had to update my Babel configuration to support the new JSX transform.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install &lt;/span&gt;react@^18.0.0 react-dom@^18.0.0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After that, I bumped up the Gatsby version to &lt;code&gt;^5.0.0&lt;/code&gt; and ran &lt;code&gt;npm install&lt;/code&gt;. This is where the fun really began 😅&lt;/p&gt;

&lt;h3&gt;
  
  
  Plugin Upgrades
&lt;/h3&gt;

&lt;p&gt;Many of my plugins were still on versions designed for Gatsby V4. I had to update or replace:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;gatsby-plugin-catch-links, gatsby-plugin-feed, gatsby-plugin-google-gtag, gatsby-plugin-layout, gatsby-plugin-manifest, gatsby-plugin-offline, gatsby-plugin-react-helmet, gatsby-plugin-sharp, and gatsby-transformer-sharp etc..&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Each of these required an upgrade to versions like &lt;code&gt;^5.14.0&lt;/code&gt; or &lt;code&gt;^6.14.0&lt;/code&gt; to resolve the peer dependency conflicts with Gatsby V5.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Image Handling:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
I migrated from the deprecated &lt;code&gt;gatsby-image&lt;/code&gt; to the modern &lt;a href="https://www.gatsbyjs.com/docs/reference/release-notes/image-migration/" rel="noopener noreferrer"&gt;&lt;code&gt;gatsby-plugin-image&lt;/code&gt;&lt;/a&gt;. This required updating my GraphQL queries from using &lt;code&gt;fluid&lt;/code&gt; and &lt;code&gt;fixed&lt;/code&gt; fragments to using the &lt;code&gt;gatsbyImageData&lt;/code&gt; field. I updated imports, replaced &lt;code&gt;&amp;lt;Img&amp;gt;&lt;/code&gt; with &lt;code&gt;&amp;lt;GatsbyImage&amp;gt;&lt;/code&gt;, and used &lt;code&gt;getImage()&lt;/code&gt; to extract image data.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Much of this is covered in the official migration &lt;a href="https://www.gatsbyjs.com/docs/reference/release-notes/migrating-from-v4-to-v5/" rel="noopener noreferrer"&gt;guide&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Then I ran into the infamous TypeComposer error.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;Cannot create as TypeComposer the following value:
  GraphQLScalarType&lt;span class="o"&gt;({&lt;/span&gt; name: &lt;span class="s2"&gt;"Date"&lt;/span&gt;, description: &lt;span class="s2"&gt;"A date string, such as 2007-12-03, compliant with the
 ISO 8601 standard for representation of dates and times using the Gregorian calendar."&lt;/span&gt;,
specifiedByURL: undefined, serialize: &lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="k"&gt;function &lt;/span&gt;String], parseValue: &lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="k"&gt;function &lt;/span&gt;String], parseLiteral:
&lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="k"&gt;function &lt;/span&gt;parseLiteral], extensions: &lt;span class="o"&gt;{&lt;/span&gt; &lt;span class="o"&gt;}&lt;/span&gt;, astNode: undefined, extensionASTNodes: &lt;span class="o"&gt;[]&lt;/span&gt; &lt;span class="o"&gt;})&lt;/span&gt;&lt;span class="nb"&gt;.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;💡 The tl;dr here is there were multiple versions of &lt;code&gt;graphql&lt;/code&gt; being used by other dependencies. I had to ensure that all my dependencies were using the same version of &lt;code&gt;graphql&lt;/code&gt;. I did this by adding a &lt;code&gt;resolutions&lt;/code&gt; field in my &lt;code&gt;package.json&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;
&lt;/blockquote&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"resolutions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"graphql"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"^16.6.0"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then I ran &lt;code&gt;rm -rf node_modules package-lock.json &amp;amp;&amp;amp; npm install&lt;/code&gt; again to ensure all dependencies were using the same version of &lt;code&gt;graphql&lt;/code&gt;. TypeComposer error was gone, but I still had a few warnings.&lt;/p&gt;

&lt;p&gt;Next, I transformed my old GraphQL queries to the new format using the Gatsby codemod.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx gatsby-codemods@latest sort-and-aggr-graphql &lt;span class="nb"&gt;.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This worked like a charm and fixed most of the query related issues. Everything else, I just asked ChatGPT to help me with. I had to tweak a few queries here and there, but nothing major.&lt;/p&gt;

&lt;h3&gt;
  
  
  Styling Challenges
&lt;/h3&gt;

&lt;p&gt;I also relied heavily on &lt;a href="https://github.com/vercel/styled-jsx" rel="noopener noreferrer"&gt;styled-jsx&lt;/a&gt; for component-scoped CSS. However, Gatsby’s official plugin, &lt;code&gt;gatsby-plugin-styled-jsx&lt;/code&gt;, only supports styled-jsx v3—and I needed styled-jsx v5 for React 18 compatibility. After some consideration, I decided to remove the plugin entirely and instead configured Babel to handle styled-jsx directly.&lt;/p&gt;

&lt;p&gt;This is where ChatGPT had most of problems. It went through a diamond dependency resolution problem and got stuck in a loop. I had to manually intervene and guide it through the process.&lt;/p&gt;

&lt;h4&gt;
  
  
  Enabling Nested CSS with styled-jsx
&lt;/h4&gt;

&lt;p&gt;Even after removing the old plugin, I ran into issues when my &lt;code&gt;&amp;lt;style jsx&amp;gt;&lt;/code&gt; blocks used nested CSS rules. By default, styled-jsx doesn’t support nesting. I resolved this by integrating &lt;a href="https://github.com/vercel/styled-jsx-plugin-postcss" rel="noopener noreferrer"&gt;styled-jsx-plugin-postcss&lt;/a&gt; along with the PostCSS Nested plugin:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Installed the Packages:&lt;/strong&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Created a Configuration File:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
I added a &lt;code&gt;styled-jsx.config.js&lt;/code&gt; at the project root:&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Updated Babel Config:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
In my &lt;code&gt;.babelrc&lt;/code&gt;, I ensured the styled-jsx plugin was configured to use the PostCSS plugin:&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This allowed my nested CSS in components like the header, footer, and blog items to compile correctly. I would say this the part where LLM helped me the most! 💪&lt;/p&gt;

&lt;h3&gt;
  
  
  PostCSS Configuration Adjustments
&lt;/h3&gt;

&lt;p&gt;During the upgrade, I also encountered warnings about duplicate autoprefixer instances. My initial &lt;code&gt;postcss.config.js&lt;/code&gt; was using &lt;code&gt;postcss-cssnext&lt;/code&gt;, which is now deprecated. After some experiments and research, I updated the config to use &lt;code&gt;postcss-preset-env&lt;/code&gt;—a modern alternative that handles vendor prefixes efficiently.&lt;/p&gt;




&lt;h2&gt;
  
  
  Lessons Learned
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Upgrade in Small Steps:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Tackling dependency conflicts one by one and verifying functionality helps isolate problems.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Read Plugin Changelogs:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Understanding what’s changed in plugin APIs (like the migration from &lt;code&gt;fluid&lt;/code&gt;/&lt;code&gt;fixed&lt;/code&gt; to &lt;code&gt;gatsbyImageData&lt;/code&gt;) is crucial.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Be Ready to Reconfigure:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Sometimes it’s necessary to remove old plugins (like &lt;code&gt;gatsby-plugin-styled-jsx&lt;/code&gt;) and configure Babel or PostCSS directly to maintain compatibility with the latest React and Gatsby versions.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Test Thoroughly:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Both in development mode and via production builds (&lt;code&gt;gatsby build&lt;/code&gt; and &lt;code&gt;gatsby serve&lt;/code&gt;), to ensure SSR and dynamic behaviors work as expected.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Final Thoughts
&lt;/h2&gt;

&lt;p&gt;Upgrading to Gatsby V5 has not only improved performance and introduced new features, but it also forced me to re-examine and modernize the entire toolchain—from image handling and styling to PostCSS configurations.&lt;/p&gt;

&lt;p&gt;If you’re planning a similar upgrade, take your time, tackle one dependency at a time, and don’t hesitate to experiment with configurations until everything clicks.&lt;/p&gt;

&lt;p&gt;Prepare to do some A/B testing by having a develop branch and a production branch. This way, you can compare the performance and functionality of your old setup with the new one.&lt;/p&gt;

&lt;p&gt;Happy coding, and enjoy your faster, modernized Gatsby blog! 💜&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.gatsbyjs.com/docs/reference/release-notes/migrating-from-v4-to-v5/" rel="noopener noreferrer"&gt;Gatsby V5 Migration Guide&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;Feel free to leave a comment if you have any questions or if you’d like to share your upgrade experiences!&lt;/p&gt;

</description>
      <category>javascript</category>
      <category>gatsby</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Streaming APIs with FastAPI and Next.js — Part 2</title>
      <dc:creator>Sahan</dc:creator>
      <pubDate>Mon, 31 Mar 2025 08:10:00 +0000</pubDate>
      <link>https://dev.to/sahan/streaming-apis-with-fastapi-and-nextjs-part-2-2jof</link>
      <guid>https://dev.to/sahan/streaming-apis-with-fastapi-and-nextjs-part-2-2jof</guid>
      <description>&lt;p&gt;In &lt;a href="https://sahansera.dev/streaming-apis-python-nextjs-part1" rel="noopener noreferrer"&gt;Part 1&lt;/a&gt;, we explored how to stream data into a React component using modern browser APIs. Now it’s time to build the other half: the &lt;strong&gt;FastAPI backend&lt;/strong&gt; that makes it all work.&lt;/p&gt;

&lt;p&gt;In this post, we’ll walk through setting up a streaming endpoint using FastAPI and discuss how chunked transfer encoding works on the server side.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;💡 &lt;strong&gt;Code Repository&lt;/strong&gt; : The complete code is available on &lt;a href="https://github.com/sahansera/streaming-apis" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;. You can clone it and run it locally to follow along.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  🛠️ Building the Streaming Backend with FastAPI
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fwibei3s0difnbdlijti3.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fwibei3s0difnbdlijti3.jpg" alt="How the data will flow from Backend to the Frontend" width="800" height="494"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;FastAPI makes it really easy to return a streaming response using its &lt;a href="https://fastapi.tiangolo.com/advanced/custom-response/#streamingresponse?" rel="noopener noreferrer"&gt;&lt;code&gt;StreamingResponse&lt;/code&gt;&lt;/a&gt; class from &lt;code&gt;starlette.responses&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Let’s build a &lt;code&gt;/stream&lt;/code&gt; endpoint in &lt;a href="https://github.com/sahansera/streaming-apis/blob/main/backend/api/index.py" rel="noopener noreferrer"&gt;&lt;code&gt;index.py&lt;/code&gt;&lt;/a&gt; that simulates real-time data like server logs or chat messages.&lt;/p&gt;

&lt;h3&gt;
  
  
  🚀 Simulating a Real-Time Log Stream
&lt;/h3&gt;

&lt;p&gt;Here’s a minimal FastAPI app with a streaming endpoint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# backend/index.py
&lt;/span&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;typing&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Any&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Generator&lt;/span&gt;

&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;fastapi&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;FastAPI&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;fastapi.responses&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;StreamingResponse&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;fastapi.middleware.cors&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;CORSMiddleware&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;uvicorn&lt;/span&gt; &lt;span class="c1"&gt;# Import uvicorn for running the server
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;threading&lt;/span&gt;

&lt;span class="n"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;FastAPI&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="c1"&gt;# Add CORS middleware to allow cross-origin requests
# ...
&lt;/span&gt;
&lt;span class="n"&gt;LOG_FILE_PATH&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;logs/server.log&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;log_stream&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;log_file_path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;Generator&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;log_file_path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;log_file&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="c1"&gt;# Move to the end of the file
&lt;/span&gt;            &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;log_file&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;seek&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;SEEK_END&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;line&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;log_file&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;readline&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
                &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt;
                &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Waiting for new log entries...&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="c1"&gt;# Heartbeat message
&lt;/span&gt;                    &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;# Wait for new lines to be written
&lt;/span&gt;    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;FileNotFoundError&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Log file not found.&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Error reading log file: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;simulate_log_generation&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Simulate log entries being written to the log file.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;LOG_FILE_PATH&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;log_file&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;log_file&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Simulated log entry at &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ctime&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;# Write a new log entry every 5 seconds
&lt;/span&gt;
&lt;span class="nd"&gt;@app.on_event&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;startup&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;start_log_simulation&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Start the log simulation in a background thread.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;threading&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;target&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;simulate_log_generation&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;daemon&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;start&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="nd"&gt;@app.get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/stream&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;StreamingResponse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;log_stream&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;LOG_FILE_PATH&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;media_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text/plain&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  🧠 How It Works
&lt;/h2&gt;

&lt;p&gt;Let’s break it down:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. &lt;strong&gt;Synchronous Generator&lt;/strong&gt;
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;log_stream&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;log_file_path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;Generator&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;log_file_path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;log_file&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;log_file&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;seek&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;SEEK_END&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;line&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;log_file&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;readline&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
                &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt;
                &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Waiting for new log entries...&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                    &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;FileNotFoundError&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Log file not found.&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Error reading log file: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We define a synchronous generator function that yields data one chunk at a time. Each &lt;code&gt;yield&lt;/code&gt; becomes a chunk sent to the client. The &lt;code&gt;time.sleep(1)&lt;/code&gt; simulates delay between events (like logs being written).&lt;/p&gt;

&lt;h3&gt;
  
  
  2. &lt;strong&gt;StreamingResponse&lt;/strong&gt;
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;StreamingResponse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;log_stream&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;LOG_FILE_PATH&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;media_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text/plain&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;StreamingResponse&lt;/code&gt; tells FastAPI to send the data as it becomes available, rather than waiting for the entire response to be generated. The &lt;code&gt;media_type&lt;/code&gt; is optional but helps inform the browser how to handle the content.&lt;/p&gt;

&lt;h2&gt;
  
  
  ⚙️ Running the Server
&lt;/h2&gt;

&lt;p&gt;To run the server:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;make start-backend
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then hit:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;http://localhost:8000/stream
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fwww.sahansera.dev%2Fstatic%2Fe8560d80474db7fc252b8f2e7cd3d05f%2F5a190%2Fstreaming-apis-python-nextjs-part2-2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fwww.sahansera.dev%2Fstatic%2Fe8560d80474db7fc252b8f2e7cd3d05f%2F5a190%2Fstreaming-apis-python-nextjs-part2-2.png" title="streaming-apis-python-nextjs-part2-2.png" alt="streaming-apis-python-nextjs-part2-2.png" width="800" height="411"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You should see logs appear one line at a time in the terminal or browser, depending on how you call the endpoint.&lt;/p&gt;




&lt;h2&gt;
  
  
  💡 Tips for Production
&lt;/h2&gt;

&lt;h3&gt;
  
  
  ✅ &lt;strong&gt;Keep the Stream Alive&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;In a real app, your data stream might be longer-running. Our example already implements this with a heartbeat:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt;
&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Waiting for new log entries...&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="c1"&gt;# Heartbeat message
&lt;/span&gt;    &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;# Wait for new lines to be written
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This ensures the client knows the connection is still alive even when there’s no new data.&lt;/p&gt;

&lt;h3&gt;
  
  
  🧹 &lt;strong&gt;Handle Disconnects Gracefully&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;If the client disconnects, your generator should stop yielding. Starlette handles this internally, but you can wrap the stream in a &lt;code&gt;try/except&lt;/code&gt; block to catch cancellations if needed.&lt;/p&gt;

&lt;h3&gt;
  
  
  🔐 &lt;strong&gt;Secure Your Stream&lt;/strong&gt;
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Add authentication if you’re streaming sensitive data.&lt;/li&gt;
&lt;li&gt;Rate-limit the endpoint to avoid abuse.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  🧪 Testing with Curl
&lt;/h2&gt;

&lt;p&gt;You can test the streaming endpoint with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://localhost:8000/stream
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You should see each message appear line by line, every second.&lt;/p&gt;




&lt;h2&gt;
  
  
  🔗 Hooking it up with the Frontend
&lt;/h2&gt;

&lt;p&gt;Now that our backend is streaming correctly, the frontend from Part 1 will handle it smoothly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;http://localhost:8000/stream&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;reader&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getReader&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="c1"&gt;// ...&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;As each &lt;code&gt;yield&lt;/code&gt; in the backend emits data, your React UI will update in near real-time. ✨&lt;/p&gt;

&lt;h2&gt;
  
  
  🧠 Recap
&lt;/h2&gt;

&lt;p&gt;In this post, we built a real-time streaming API using FastAPI with just a few lines of code. Here’s what we did:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Built a synchronous generator to simulate streaming logs&lt;/li&gt;
&lt;li&gt;Used &lt;code&gt;StreamingResponse&lt;/code&gt; to stream text over HTTP&lt;/li&gt;
&lt;li&gt;Connected it to our frontend from Part 1&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Together, the backend and frontend make a simple but powerful full-stack streaming system.&lt;/p&gt;

&lt;h2&gt;
  
  
  🧱 Next Steps
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;🔌 Add dynamic data (e.g. logs from a file or DB)&lt;/li&gt;
&lt;li&gt;📦 Stream structured data (like JSON Lines)&lt;/li&gt;
&lt;li&gt;📈 Use this setup for real-time dashboards or log viewers&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  🛠️ Useful Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://fastapi.tiangolo.com/advanced/custom-response/#streamingresponse" rel="noopener noreferrer"&gt;FastAPI: StreamingResponse&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.python.org/3/glossary.html#term-generator" rel="noopener noreferrer"&gt;Python Generators&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.uvicorn.org/" rel="noopener noreferrer"&gt;Uvicorn Docs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/sahansera/streaming-apis" rel="noopener noreferrer"&gt;GitHub Repo&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;If you enjoyed this series, feel free to star the repo ⭐ or share it with a friend. Got ideas or feedback? Hit me up on &lt;a href="https://twitter.com/_sahansera" rel="noopener noreferrer"&gt;Twitter&lt;/a&gt; or drop an issue in the GitHub repo.&lt;/p&gt;

&lt;p&gt;Happy streaming! 🚀&lt;/p&gt;

</description>
      <category>python</category>
      <category>backend</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Streaming APIs with FastAPI and Next.js — Part 1</title>
      <dc:creator>Sahan</dc:creator>
      <pubDate>Sun, 30 Mar 2025 09:00:00 +0000</pubDate>
      <link>https://dev.to/sahan/streaming-apis-with-fastapi-and-nextjs-part-1-3ndj</link>
      <guid>https://dev.to/sahan/streaming-apis-with-fastapi-and-nextjs-part-1-3ndj</guid>
      <description>&lt;p&gt;Streaming data in the browser is one of those things that feels magical the first time you see it: data appears live — no need to wait for the full response to load. In this two-part series, we’ll walk through building a small full-stack app that uses &lt;strong&gt;FastAPI&lt;/strong&gt; to stream data and a &lt;strong&gt;Next.js&lt;/strong&gt; frontend to consume and render it in real-time.&lt;/p&gt;

&lt;p&gt;There are many ways to stream data in the browser, from WebSockets to Server-Sent Events (SSE). But in this post, we’ll focus on a lesser-known method: &lt;a href="https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Headers/Transfer-Encoding" rel="noopener noreferrer"&gt;chunked transfer encoding&lt;/a&gt;. This technique allows the server to send data in small, manageable chunks, which the browser can process as they arrive.&lt;/p&gt;

&lt;p&gt;This post focuses on the &lt;strong&gt;frontend&lt;/strong&gt; bit. In Part 2, we’ll dive into building the &lt;strong&gt;FastAPI backend&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  🔧 The Setup
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fox66rxuf1tmjli33aipy.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fox66rxuf1tmjli33aipy.jpg" alt="How the data will flow from Backend to the Frontend" width="800" height="494"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Let’s say you have a streaming API running locally at:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;http://localhost:8000/stream
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This endpoint sends back a stream of text data — think server logs, chat messages, or real-time updates. The goal is to connect to this endpoint from a React component and display data &lt;strong&gt;as it arrives&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Here’s the React component (&lt;code&gt;stream.tsx&lt;/code&gt;) that handles this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;useEffect&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;useState&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;react&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;StreamPage&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;dataChunks&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;setDataChunks&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;useState&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;([]);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;isLoading&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;setIsLoading&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;useState&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;setError&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;useState&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="nf"&gt;useEffect&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;fetchStream&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;http://localhost:8000/stream&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

        &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
          &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`HTTP error! Status: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
          &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Response body is null&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;reader&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getReader&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
        &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;decoder&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;TextDecoder&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;utf-8&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

        &lt;span class="k"&gt;while &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
          &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;done&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;reader&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
          &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;done&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;break&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

          &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;chunk&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;decoder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
          &lt;span class="nf"&gt;setDataChunks&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;prev&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[...&lt;/span&gt;&lt;span class="nx"&gt;prev&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="nf"&gt;setIsLoading&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nf"&gt;setError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt; &lt;span class="k"&gt;instanceof&lt;/span&gt; &lt;span class="nb"&gt;Error&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
        &lt;span class="nf"&gt;setIsLoading&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="nf"&gt;fetchStream&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;[]);&lt;/span&gt;

  &lt;span class="k"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;div&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;
      &lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;h1&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="nx"&gt;Streaming&lt;/span&gt; &lt;span class="nx"&gt;Data&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="sr"&gt;/h1&lt;/span&gt;&lt;span class="err"&gt;&amp;gt;
&lt;/span&gt;      &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="nb"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="sr"&gt;/p&amp;gt;&lt;/span&gt;&lt;span class="err"&gt;}
&lt;/span&gt;      &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;isLoading&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;dataChunks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="nx"&gt;Loading&lt;/span&gt; &lt;span class="nx"&gt;stream&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;...&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="sr"&gt;/p&lt;/span&gt;&lt;span class="err"&gt;&amp;gt;
&lt;/span&gt;      &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;pre&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;dataChunks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="sr"&gt;/pre&lt;/span&gt;&lt;span class="err"&gt;&amp;gt;
&lt;/span&gt;      &lt;span class="p"&gt;)}&lt;/span&gt;
    &lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="sr"&gt;/div&lt;/span&gt;&lt;span class="err"&gt;&amp;gt;
&lt;/span&gt;  &lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  🧠 What’s Really Happening?
&lt;/h2&gt;

&lt;p&gt;Let’s break this down and understand the key concepts behind streaming in the browser:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. &lt;strong&gt;Fetching the Stream&lt;/strong&gt;
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;http://localhost:8000/stream&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Unlike regular &lt;code&gt;fetch()&lt;/code&gt; calls that wait for the entire response before handing it over, this gives you access to the &lt;strong&gt;ReadableStream&lt;/strong&gt; — letting you process data chunk-by-chunk.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. &lt;strong&gt;Getting a Reader&lt;/strong&gt;
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;reader&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getReader&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This gives us a &lt;code&gt;ReadableStreamDefaultReader&lt;/code&gt;, which lets us manually pull chunks of data from the response. This is part of the &lt;a href="https://developer.mozilla.org/en-US/docs/Web/API/Streams_API" rel="noopener noreferrer"&gt;Streams API&lt;/a&gt;, now supported in all major browsers.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. &lt;strong&gt;Reading Chunks&lt;/strong&gt;
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;done&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;reader&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;value&lt;/code&gt;: a &lt;code&gt;Uint8Array&lt;/code&gt; representing a chunk of binary data.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;done&lt;/code&gt;: &lt;code&gt;true&lt;/code&gt; when the stream is finished.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We loop until &lt;code&gt;done&lt;/code&gt; becomes &lt;code&gt;true&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. &lt;strong&gt;Decoding the Text&lt;/strong&gt;
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;decoder&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;TextDecoder&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;utf-8&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;chunk&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;decoder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Streaming responses may split characters across chunks, especially for multi-byte encodings like UTF-8. &lt;code&gt;TextDecoder&lt;/code&gt; handles this for us, making sure our text is correctly reconstructed.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. &lt;strong&gt;Updating the UI&lt;/strong&gt;
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nf"&gt;setDataChunks&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;prev&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[...&lt;/span&gt;&lt;span class="nx"&gt;prev&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We append each new chunk to our state array. React re-renders the component with every update, giving us that real-time feel.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. &lt;strong&gt;One-Time Effect&lt;/strong&gt;
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nf"&gt;useEffect&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;[]);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The fetch logic runs once when the component mounts. Perfect for one-time side effects like opening a connection.&lt;/p&gt;




&lt;h2&gt;
  
  
  ✅ Recap
&lt;/h2&gt;

&lt;p&gt;With just a few lines of code, we’ve created a streaming experience in the browser using modern Web APIs. The key things were:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;fetch()&lt;/code&gt; with a streaming response&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;ReadableStream&lt;/code&gt; + &lt;code&gt;TextDecoder&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Updating React state to progressively display data&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In Part 2, we’ll build the &lt;strong&gt;FastAPI backend&lt;/strong&gt; to power this stream — including how to set up a streaming route and control flush timing for chunked responses.&lt;/p&gt;




&lt;h2&gt;
  
  
  💡 Gotchas to Watch Out For
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;CORS&lt;/strong&gt; : Make sure your FastAPI server has CORS enabled if frontend and backend are on different ports/domains.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Buffering&lt;/strong&gt; : Some servers (or even browsers like Safari) buffer streaming responses — so test in Chrome/Edge for best results.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cleanup&lt;/strong&gt; : If you’re adding WebSocket or long-running fetches, remember to cancel or clean them up in &lt;code&gt;useEffect&lt;/code&gt; cleanup.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  📚 References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://developer.mozilla.org/en-US/docs/Web/API/Streams_API?utm_source=sahansera.dev" rel="noopener noreferrer"&gt;Streams API&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developer.mozilla.org/en-US/docs/Web/API/TextDecoder?utm_source=sahansera.dev" rel="noopener noreferrer"&gt;TextDecoder&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developer.mozilla.org/en-US/docs/Web/API/ReadableStream?utm_source=sahansera.dev" rel="noopener noreferrer"&gt;ReadableStream&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://fastapi.tiangolo.com/advanced/custom-response/#streamingresponse?utm_source=sahansera.dev" rel="noopener noreferrer"&gt;Streaming Response&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>nextjs</category>
      <category>tutorial</category>
      <category>python</category>
    </item>
  </channel>
</rss>
