<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: George Michalakis</title>
    <description>The latest articles on DEV Community by George Michalakis (@thegm26).</description>
    <link>https://dev.to/thegm26</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3823163%2Fad37f758-bf59-43eb-8bb7-537b099906ff.jpeg</url>
      <title>DEV Community: George Michalakis</title>
      <link>https://dev.to/thegm26</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/thegm26"/>
    <language>en</language>
    <item>
      <title>Switch Codex accounts seamlessly with AgentHop</title>
      <dc:creator>George Michalakis</dc:creator>
      <pubDate>Sat, 19 Sep 2026 15:28:48 +0000</pubDate>
      <link>https://dev.to/thegm26/switch-codex-accounts-seamlessly-with-agenthop-11b7</link>
      <guid>https://dev.to/thegm26/switch-codex-accounts-seamlessly-with-agenthop-11b7</guid>
      <description>&lt;h2&gt;
  
  
  Story time
&lt;/h2&gt;

&lt;p&gt;I use separate, legitimate Codex profiles for distinct contexts: personal work, a client workspace, and experiments that should not inherit the same local setup. Keeping those contexts separate is useful; repeatedly navigating between them was not.&lt;/p&gt;

&lt;p&gt;Each time I needed a different profile, I would log out, change the active setup, reload Codex, and authenticate again. That routine was repetitive, slow, and disruptive.&lt;/p&gt;

&lt;p&gt;Worse, it pulled me out of the task just to answer ordinary operational questions: &lt;em&gt;which profile is selected? Is it signed in? What status did Codex last report?&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  First, let's make the switching less painful
&lt;/h2&gt;

&lt;p&gt;The first (obvious?) answer was a small CLI script. It gave each context an isolated local profile and made the normal sign-out/re-auth cycle unnecessary for routine switching. That solved the most visible friction: I could select the identity and configuration used by the CLI without redoing the entire setup.&lt;/p&gt;

&lt;p&gt;And this definitely made the switch faster. However, it did not make the decision easier.&lt;/p&gt;

&lt;p&gt;When you are deep in work, I still needed to know:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Which profile is signed in?&lt;/li&gt;
&lt;li&gt;Which one Codex reports as ready, close, or blocked, and when will it reset?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I could ask the CLI for status, but checking several isolated profiles by hand was still manual and easy to get wrong.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frmhik932rnvnjtvfmw8n.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frmhik932rnvnjtvfmw8n.gif" alt="Terminal showing AgentHop's CLI profile-status output for several local Codex profiles, including their active state and reset information." width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;AgentHop's CLI makes each local profile's current status visible at a glance.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  DaaS (Dashboard as a Service)
&lt;/h2&gt;

&lt;p&gt;So the script became a dashboard (wow right). AgentHop reads the local, supported status for each profile and turns it into a small operational view:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;ready&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;close&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;blocked&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;plus readable reset times. Profiles that are currently usable appear first; a blocked profile remains muted until the provider reports it as usable again.&lt;/p&gt;

&lt;p&gt;That was the first payoff. Instead of bouncing between profile directories and terminal output, I could refresh once and see the current local picture.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F35awhddelsnnvupd5apr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F35awhddelsnnvupd5apr.png" alt="AgentHop dashboard showing account status cards, remaining 5-hour and weekly usage, reset times, and the active profile." width="799" height="454"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The dashboard keeps the fuller account picture available when a quick status check is not enough.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;But a dashboard still asks you to open a dashboard. For a check that happens several times in a day, that is one window too many.&lt;/p&gt;

&lt;h2&gt;
  
  
  The control surface moved to the tray
&lt;/h2&gt;

&lt;p&gt;Now AgentHop starts quietly in the Linux system tray. Clicking its icon gives me the short list I actually need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the active profile;&lt;/li&gt;
&lt;li&gt;a suggested available profile when one is reported; then&lt;/li&gt;
&lt;li&gt;profiles grouped by current status.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A blocked profile stays muted; an enabled entry can be clicked to switch directly.&lt;/p&gt;

&lt;p&gt;The full dashboard is still there for onboarding, inspecting status, or preparing a new command, but it is no longer a mandatory stop between work and the profile I need.&lt;/p&gt;

&lt;p&gt;And this distinction matters. The tray is for the frequent, low-friction question—&lt;em&gt;which profile do I need right now?&lt;/em&gt; The dashboard is for the less frequent, higher-context tasks—&lt;em&gt;add an identity, inspect the details, or start a new session.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcn52575y52y0wf031b10.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcn52575y52y0wf031b10.gif" alt="AgentHop workflow moving from CLI profile status to the Linux tray menu and then the dashboard, where an enabled profile can be selected." width="720" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;From CLI status to a tray action and the full dashboard without losing the operational context.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Then I found the problem that mattered more
&lt;/h2&gt;

&lt;p&gt;The first time I switched profiles and tried to continue a real task, I expected to run &lt;code&gt;resume&lt;/code&gt; and pick up the conversation. Instead, the resume list looked empty.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The account switch had worked. My working memory had not followed me.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That discovery changed the project. Pretty much the local profile directory was doing two jobs at once: it held account-specific identity, but it also held the transcripts, indexes, snapshots, and other state that made a session discoverable. Pointing &lt;code&gt;CODEX_HOME&lt;/code&gt; at a new profile did not just choose a different account; it could make the CLI look at a different history.&lt;/p&gt;

&lt;p&gt;As a result I found myself asking a much more interesting question: how do you keep identity isolated without turning every account switch into a separate universe?&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/Thegm26/AgentHop" rel="noopener noreferrer"&gt;AgentHop&lt;/a&gt; is the MVP that came out of that question. How?&lt;/p&gt;

&lt;p&gt;It is a local-first dashboard for managing AI coding CLI profiles while keeping resumable work close at hand. It:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;separates profile identity from continuity;&lt;/li&gt;
&lt;li&gt;shows the status Codex makes available; and&lt;/li&gt;
&lt;li&gt;prepares the next Codex command without treating a profile change as a new project.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Independent project:&lt;/strong&gt; AgentHop is unofficial, local software. It is not affiliated with, endorsed by, or supported by OpenAI. It does not create accounts, combine subscriptions, bypass limits, or expose credentials to the browser. It does not decide whether any account setup or use complies with provider policy.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Use only accounts you are authorized to use and follow the provider's terms. For OpenAI services, see the current &lt;a href="https://openai.com/policies/terms-of-use/" rel="noopener noreferrer"&gt;Terms of Use&lt;/a&gt; and &lt;a href="https://help.openai.com/en/articles/20001068" rel="noopener noreferrer"&gt;account-switching help&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The architecture: identity is NOT continuity
&lt;/h2&gt;

&lt;p&gt;Most quick account switchers point a CLI at a different home directory. That works for isolated credentials, but a home directory often contains much more: configuration, transcripts, archived sessions, SQLite indexes, locks, and snapshots. Change the whole directory and you may change all of those at once.&lt;/p&gt;

&lt;p&gt;The central design decision is deliberately simple:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Identity is isolated.&lt;/strong&gt; Each account has its own local profile home.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Continuity is shared.&lt;/strong&gt; Resumable session state lives in a canonical shared location defined by the provider adapter, including the discoverability metadata and indexes that let the CLI find it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The provider owns provider details.&lt;/strong&gt; The dashboard should not need to know how a particular CLI stores credentials, models usage, or indexes threads.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4gyizpjwwn8odx0v7xp8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4gyizpjwwn8odx0v7xp8.png" alt="Diagram showing the selected account providing isolated credentials and shared continuity state to the Codex CLI." width="800" height="432"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is the difference between “switch who is making the next request” and “switch into a separate universe with no history.”&lt;/p&gt;

&lt;h2&gt;
  
  
  What AgentHop does today
&lt;/h2&gt;

&lt;p&gt;AgentHop currently ships with an OpenAI Codex adapter and a local React + FastAPI control surface. It can:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;discover the default and named local profiles;&lt;/li&gt;
&lt;li&gt;show profile state as &lt;strong&gt;ready&lt;/strong&gt;, &lt;strong&gt;close&lt;/strong&gt;, or &lt;strong&gt;blocked&lt;/strong&gt;, rather than hiding two independent status windows behind one percentage;&lt;/li&gt;
&lt;li&gt;show human-readable 5-hour and weekly reset times when the provider supplies them;&lt;/li&gt;
&lt;li&gt;mute profiles the provider reports as blocked and group profiles by their current status; and&lt;/li&gt;
&lt;li&gt;create a new isolated profile, guide the user through the normal terminal sign-in, select a profile, reconcile shared continuity when required, and prepare a new or resume command.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The ordering is operational, not a policy judgment. Profiles reported as usable appear first. When a profile has more than one blocked window, AgentHop presents the later reset as its expected unblock time because that is the latest status change it can show.&lt;/p&gt;

&lt;p&gt;The practical sequence is refresh, choose the profile appropriate for the work, reconcile shared state when required, then copy the next command.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl36t35kz18vmhnx046dx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl36t35kz18vmhnx046dx.png" alt="Diagram showing the AgentHop flow: refresh profile status, choose the appropriate profile, reconcile continuity when needed, then copy a new or resume command." width="800" height="175"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The browser sees sanitized profile names, status, reset information, and command results. It never needs the contents of an authentication file.&lt;/p&gt;

&lt;h2&gt;
  
  
  A small architecture that leaves room to grow
&lt;/h2&gt;

&lt;p&gt;AgentHop is not “provider-neutral” because it has a generic dropdown. Different CLIs have different credential stores, session models, supported status interfaces, and safe launch semantics. The shared core stays small so adapters can own those facts.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Responsibility&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;React dashboard&lt;/td&gt;
&lt;td&gt;Local control surface; renders sanitized state and starts bounded actions.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FastAPI backend&lt;/td&gt;
&lt;td&gt;Validates input, owns the local API boundary, selects adapters, and returns safe results.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Provider adapter&lt;/td&gt;
&lt;td&gt;Discovers profiles, reads supported status, prepares commands, and owns continuity reconciliation behavior.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For Codex, AgentHop builds on documented &lt;code&gt;CODEX_HOME&lt;/code&gt;, authentication, and app-server interfaces. The browser does not host an interactive CLI. It gives you a quoted command to review and run in your own terminal.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it locally
&lt;/h2&gt;

&lt;p&gt;You need Python 3.11+, Node.js 18.19+ (Node 20+ recommended), npm, and the Codex CLI on your &lt;code&gt;PATH&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/Thegm26/AgentHop.git
&lt;span class="nb"&gt;cd &lt;/span&gt;AgentHop

python3 &lt;span class="nt"&gt;-m&lt;/span&gt; venv .venv
.venv/bin/pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-e&lt;/span&gt; &lt;span class="s1"&gt;'.[dev]'&lt;/span&gt;
npm &lt;span class="nt"&gt;--prefix&lt;/span&gt; frontend ci
./scripts/desktop.sh
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This repository launcher starts AgentHop in the local Linux system tray and builds the dashboard from the installed frontend dependencies. Click the icon to refresh or switch an enabled profile; choose &lt;strong&gt;Open dashboard…&lt;/strong&gt; for the full view. The launcher keeps the backend on loopback. The README also covers the optional desktop-menu installer and the alternate &lt;code&gt;.venv/bin/agenthop desktop&lt;/code&gt; launcher when frontend assets are already built.&lt;/p&gt;

&lt;p&gt;For frontend development only, use another terminal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd &lt;/span&gt;frontend
npm run dev
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Click &lt;strong&gt;Add account&lt;/strong&gt;, choose a local profile name, copy the displayed sign-in command, and complete the normal provider login in your browser. Then return to AgentHop and choose &lt;strong&gt;Refresh usage&lt;/strong&gt;. The app shows commands for review rather than running a shell for you. For desktop installation, APIs, and security details, see the &lt;a href="https://github.com/Thegm26/AgentHop#readme" rel="noopener noreferrer"&gt;README&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this is—and what it is not
&lt;/h2&gt;

&lt;p&gt;This is a local, single-user MVP: it does not create accounts, combine subscriptions, bypass provider limits, or send credentials to the dashboard. &lt;/p&gt;

&lt;p&gt;It is not a remote or multi-user service, and Codex is its first working adapter. See the README for the complete security and deployment notes.&lt;/p&gt;

&lt;p&gt;Repo: &lt;a href="https://github.com/Thegm26/AgentHop" rel="noopener noreferrer"&gt;https://github.com/Thegm26/AgentHop&lt;/a&gt;&lt;/p&gt;

</description>
      <category>react</category>
      <category>python</category>
      <category>productivity</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Open Tab: Let Your Spare Change Cover Someone Else's Checkout</title>
      <dc:creator>George Michalakis</dc:creator>
      <pubDate>Mon, 07 Sep 2026 05:21:34 +0000</pubDate>
      <link>https://dev.to/thegm26/open-tab-let-your-spare-change-cover-someone-elses-checkout-40h5</link>
      <guid>https://dev.to/thegm26/open-tab-let-your-spare-change-cover-someone-elses-checkout-40h5</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for &lt;a href="https://dev.to/challenges/weekend-2026-09-03"&gt;Weekend Challenge: Generosity Edition&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;Open Tab lets one customer's spare change help with someone else's purchase.&lt;/p&gt;

&lt;p&gt;A customer reaches checkout and chooses to add a small contribution. The money enters a shared pool. Another customer can use part of that pool during an ordinary checkout, with no application and no explanation required.&lt;/p&gt;

&lt;p&gt;I built it around a familiar moment: someone has enough for most of a purchase and could use a little help with the rest. Asking a stranger or an employee can feel exposing. Open Tab keeps that moment inside the same payment screen everyone uses.&lt;/p&gt;

&lt;p&gt;The weekend prototype connects two fictional local businesses, Café Sol and Bread &amp;amp; Butter Bakery. Each business has its own employee route and product catalogue. Both feed the same Open Tab pool and activity ledger.&lt;/p&gt;

&lt;p&gt;The contribution follows a predictable rule. A purchase moves toward the next 50-cent mark. When the total already ends in &lt;code&gt;.00&lt;/code&gt; or &lt;code&gt;.50&lt;/code&gt;, Open Tab adds €0.20 instead. The customer always sees the amount before paying.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why small amounts can work
&lt;/h2&gt;

&lt;p&gt;Research supports testing a small round-up prompt ahead of a flat donation request. A study in the &lt;em&gt;Journal of Consumer Psychology&lt;/em&gt; found that people responded more favourably to round-up requests even when both formats asked for the same amount. The researchers linked the effect to lower perceived pain of giving. An Open Tab pilot would still need to measure its own contribution rate. &lt;a href="https://doi.org/10.1002/jcpy.1064" rel="noopener noreferrer"&gt;Kelting et al., &lt;em&gt;Would You Like to Round Up and Donate the Difference?&lt;/em&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Everyday food purchases create repeated opportunities for that choice. The UK's nationally representative National Diet and Nutrition Survey found that 72% of participants had bought food or drink away from home during the previous seven days. Most did so once or twice that week. In its shorter food record, 17% reported an occasion involving a café, coffee shop, sandwich bar, or deli. These figures show recurring relevant transactions rather than expected Open Tab usage. &lt;a href="https://www.gov.uk/government/statistics/national-diet-and-nutrition-survey-2019-to-2023/national-diet-and-nutrition-survey-2019-to-2023-report" rel="noopener noreferrer"&gt;UK National Diet and Nutrition Survey, 2019 to 2023&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Digital checkout provides an existing place for Open Tab across the euro area. The ECB's SPACE 2024 study drew on 50,000 consumers. It reports continued growth in digital payments, with cards remaining the most popular digital method. Electronic payment acceptance increased across every euro-area country. &lt;a href="https://www.ecb.europa.eu/stats/ecb_surveys/space/html/index.en.html" rel="noopener noreferrer"&gt;European Central Bank, SPACE 2024&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The demo menu is an illustrative local price book. It ranges from a €1.80 butter croissant to a €10.50 pasta bowl, while most drinks cost between €2.50 and €4.50. Pret's public UK delivery menu listed an all-butter croissant at £2.70 and a flatbread at £6.15 on September 7, 2026. This comparison supports the general purchase scale rather than a claim about average European prices. &lt;a href="https://www.pret.co.uk/en-GB/pret-delivers/menu" rel="noopener noreferrer"&gt;Pret A Manger delivery menu&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;At that scale, cents accumulate quickly. If 100 eligible checkouts contributed an average of €0.25, the pool would receive €25 in one day. With 200 eligible checkouts at the same average, it would receive €50. These examples are arithmetic scenarios. Real results depend on checkout volume and customer participation, with local pricing and programme rules also shaping the total.&lt;/p&gt;

&lt;p&gt;For a participating business, Open Tab could create community affinity without requiring a points programme. Regular customers can see that their small contribution has a local use, giving them a reason to feel connected to the venue. That loyalty effect remains a product hypothesis. A pilot would need to measure merchant willingness alongside customer return behaviour.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/ab9WEPEpiS8" width="710" height="399"&gt;
  &lt;/iframe&gt;
)&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffnqwokqm9m94ylkkpo25.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffnqwokqm9m94ylkkpo25.png" alt="Interaction Image" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Live employee views:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://open-tab-chi.vercel.app/admin/cafe" rel="noopener noreferrer"&gt;Café Sol&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://open-tab-chi.vercel.app/admin/bakery" rel="noopener noreferrer"&gt;Bread &amp;amp; Butter Bakery&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Customer terminal: &lt;a href="https://open-tab-chi.vercel.app/checkout" rel="noopener noreferrer"&gt;Open Customer Checkout&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Keep an employee view open and place Customer Checkout beside it.&lt;/p&gt;
&lt;h3&gt;
  
  
  Contribute to Open Tab
&lt;/h3&gt;

&lt;p&gt;Choose products on the employee screen and create a checkout. The customer terminal receives the order. Select the round-up option and watch the shared pool increase on the employee dashboard.&lt;/p&gt;
&lt;h3&gt;
  
  
  Use the shared pool
&lt;/h3&gt;

&lt;p&gt;Create a fresh checkout and select &lt;strong&gt;Use Open Tab&lt;/strong&gt; on the customer terminal. Available choices adapt to the purchase and current pool balance. Apply one choice, pay the remainder, then return to the employee screen to see the receivable and updated activity.&lt;/p&gt;

&lt;p&gt;Payments are simulated for this challenge. The persisted accounting lifecycle models how contributions and assisted purchases move through the system.&lt;/p&gt;
&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;Source: &lt;a href="https://github.com/Thegm26/open-tab" rel="noopener noreferrer"&gt;github.com/Thegm26/open-tab&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The key accounting transition is authorization. It locks the shared scenario and order before checking available funds. An idempotency record makes a safe retry return its original result.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;into&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="k"&gt;public&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;demo_scenarios&lt;/span&gt;
&lt;span class="k"&gt;where&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;scenario_id&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="k"&gt;update&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;into&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="k"&gt;public&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;
&lt;span class="k"&gt;where&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;order_id&lt;/span&gt; &lt;span class="k"&gt;and&lt;/span&gt; &lt;span class="n"&gt;scenario_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="k"&gt;update&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;out&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;public&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ot_idempotency_replay&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'authorize'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;p_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;p_hash&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="n"&gt;if&lt;/span&gt; &lt;span class="k"&gt;out&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt; &lt;span class="k"&gt;then&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;out&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;end&lt;/span&gt; &lt;span class="n"&gt;if&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;settled_pool_cents&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;coalesce&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="k"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount_cents&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;filter&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;where&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'authorized'&lt;/span&gt; &lt;span class="k"&gt;and&lt;/span&gt; &lt;span class="n"&gt;expires_at&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;v_now&lt;/span&gt;
  &lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;into&lt;/span&gt; &lt;span class="n"&gt;v_available&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="k"&gt;public&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;fund_reservations&lt;/span&gt;
&lt;span class="k"&gt;where&lt;/span&gt; &lt;span class="n"&gt;scenario_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="n"&gt;if&lt;/span&gt; &lt;span class="n"&gt;p_amount&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;v_available&lt;/span&gt; &lt;span class="k"&gt;then&lt;/span&gt;
  &lt;span class="n"&gt;raise&lt;/span&gt; &lt;span class="n"&gt;exception&lt;/span&gt; &lt;span class="s1"&gt;'INSUFFICIENT_POOL'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;end&lt;/span&gt; &lt;span class="n"&gt;if&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;Open Tab uses Next.js with TypeScript. Employee routes provide the point-of-sale controls for each business. Customer Checkout follows the newest order and updates as its payment state changes.&lt;/p&gt;

&lt;p&gt;Supabase provides persistent Postgres storage. Every monetary value uses integer cents, which avoids floating-point accounting errors. Transactional RPC functions own the important balance changes.&lt;/p&gt;

&lt;p&gt;A round-up creates a pool credit only after simulated payment succeeds. When someone requests help, Open Tab creates a short-lived claim and reserves the chosen amount. The purchase total sets the ceiling. Pool availability and server policy can reduce it further.&lt;/p&gt;

&lt;p&gt;An assisted purchase creates a merchant receivable for the amount covered by Open Tab. Employees can settle outstanding receivables from the dashboard. Refunds preserve the split between customer money and pool money. Eligible funds return to the pool, while a refund after settlement creates recovery debt.&lt;/p&gt;

&lt;p&gt;The app uses idempotency keys with request hashes for payment-sensitive operations. A retry with the same input returns the recorded result. Reusing that key with changed input fails safely.&lt;/p&gt;

&lt;p&gt;The interface is responsive across phone and desktop layouts. Activity pagination keeps a longer ledger readable, while separate customer and employee views make the live checkout loop easy to follow.&lt;/p&gt;

&lt;p&gt;The automated suite currently passes 40 tests across 7 files. It covers the reservation lifecycle, refund accounting, settlement behavior, and repeated requests. The production build also passes TypeScript validation.&lt;/p&gt;

&lt;p&gt;This remains a focused hackathon simulation. It uses one shared terminal lane for the newest order. A production POS integration would assign terminal lanes to specific merchant devices and connect them to a payment processor.&lt;/p&gt;

&lt;p&gt;The part I cared about most was preserving dignity at checkout. Customers see a quiet, ordinary choice. Underneath that choice, the ledger protects funds that another person may rely on.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
    </item>
    <item>
      <title>K8s: Node Maintenance &amp; Eviction</title>
      <dc:creator>George Michalakis</dc:creator>
      <pubDate>Sun, 06 Sep 2026 17:17:06 +0000</pubDate>
      <link>https://dev.to/thegm26/k8s-node-maintenance-eviction-1mkm</link>
      <guid>https://dev.to/thegm26/k8s-node-maintenance-eviction-1mkm</guid>
      <description>&lt;p&gt;If you read my previous posts on &lt;a href="https://dev.to/thegm26/k8s-topology-spread-1hcn"&gt;topology spread&lt;/a&gt; and &lt;a href="https://dev.to/thegm26/k8s-the-affinity-club-4c2"&gt;the affinity club&lt;/a&gt;, we mostly talked about Pod placement: which nodes can accept a new Pod, and how topology spread constraints keep replicas balanced.&lt;/p&gt;

&lt;p&gt;But clusters consist of nodes, and nodes need maintenance too. A worker node may need an operating-system upgrade, a kernel update, more capacity, or replacement hardware.&lt;/p&gt;

&lt;p&gt;“Okay, then delete it lol,” right? &lt;/p&gt;

&lt;p&gt;Sure, if you are okay with the people calling those services starting to scream &amp;gt;:)&lt;/p&gt;

&lt;p&gt;Before taking a node out of service, we need to move its workload safely.&lt;/p&gt;

&lt;p&gt;That raises three related questions:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;How do we stop new Pods arriving on the node?&lt;/li&gt;
&lt;li&gt;How do the Pods already running there leave?&lt;/li&gt;
&lt;li&gt;How many of those Pods is it safe to interrupt at once?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The first two are handled by &lt;code&gt;cordon&lt;/code&gt;, &lt;code&gt;drain&lt;/code&gt;, and eviction. &lt;br&gt;
For the third, read the PodDisruptionBudget post ;)&lt;/p&gt;
&lt;h2&gt;
  
  
  Cordon: stop new placements
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl cordon worker-1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Cordoning marks a node as unschedulable. The scheduler will not place new Pods there, but the Pods already on the node keep running.&lt;/p&gt;

&lt;p&gt;Think of it as closing a hotel to new guests: the guests already in their rooms are still there.&lt;/p&gt;

&lt;p&gt;To make a node schedulable again later:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl uncordon worker-1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Drain: prepare a node to go out of service
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl drain worker-1 &lt;span class="nt"&gt;--ignore-daemonsets&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;kubectl drain&lt;/code&gt; first cordons the node, then tries to evict its eligible Pods. Controllers such as Deployments and StatefulSets then create replacement Pods, which the scheduler can place on other suitable nodes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft0svfi2n3x9v5prazeuw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft0svfi2n3x9v5prazeuw.png" alt="Draining a node" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Why &lt;code&gt;--ignore-daemonsets&lt;/code&gt;?
&lt;/h3&gt;

&lt;p&gt;In our case, &lt;code&gt;kubectl drain&lt;/code&gt; stops before evicting the web Pods because it encounters DaemonSet Pods: &lt;code&gt;kindnet&lt;/code&gt; and &lt;code&gt;kube-proxy&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;DaemonSet Pods are intended to run on every applicable node, so drain refuses to remove them by default. &lt;/p&gt;

&lt;p&gt;&lt;code&gt;--ignore-daemonsets&lt;/code&gt; tells drain to leave those Pods in place and continue evicting eligible workload Pods.&lt;/p&gt;

&lt;p&gt;Those DaemonSet Pods remain until the node itself goes offline or is removed. That is exactly what we want during the drain:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;networking / node-level services stay available while the "regular" workload leaves.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But wait... did you see what the command says?&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Evicting pod...&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Why evicting instead of deleting?&lt;/p&gt;

&lt;h2&gt;
  
  
  Eviction: removing pods in a graceful / policy-aware manner
&lt;/h2&gt;

&lt;p&gt;When possible, &lt;code&gt;kubectl drain&lt;/code&gt; uses Kubernetes’ &lt;a href="https://kubernetes.io/docs/concepts/scheduling-eviction/api-eviction/" rel="noopener noreferrer"&gt;Eviction API&lt;/a&gt; rather than directly deleting a Pod.&lt;/p&gt;

&lt;p&gt;An eviction asks Kubernetes to terminate a Pod gracefully and subject to cluster policy. &lt;/p&gt;

&lt;p&gt;The Pod receives its configured &lt;code&gt;terminationGracePeriodSeconds&lt;/code&gt;, giving the application time to shut down. Whether it stops receiving traffic and completes cleanup safely depends on the workload’s readiness handling, lifecycle hooks, and termination behavior.&lt;/p&gt;

&lt;p&gt;Most importantly for the PodDisruptionBudget post, the API checks whether a &lt;a href="https://kubernetes.io/docs/concepts/workloads/pods/disruptions/" rel="noopener noreferrer"&gt;PodDisruptionBudget&lt;/a&gt; permits the disruption.&lt;/p&gt;

&lt;p&gt;In a maintenance scenario, the tool should use eviction so availability safeguards are respected.&lt;/p&gt;

&lt;h2&gt;
  
  
  Planned versus unplanned disruption
&lt;/h2&gt;

&lt;p&gt;Node maintenance is a &lt;strong&gt;voluntary disruption&lt;/strong&gt; since an administrator intentionally asks Pods to leave a healthy node. &lt;/p&gt;

&lt;p&gt;However:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A node crash&lt;/li&gt;
&lt;li&gt;power failure&lt;/li&gt;
&lt;li&gt;network partitions &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;are all &lt;strong&gt;involuntary disruptions&lt;/strong&gt;... &lt;/p&gt;

&lt;p&gt;At this case Kubernetes cannot ask permission before a failed node becomes unavailable.&lt;/p&gt;

&lt;p&gt;PodDisruptionBudgets protect only against voluntary disruptions. They cannot prevent a machine from failing, but they help us avoid making an existing incident worse while performing planned work.&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>maintenance</category>
      <category>distributedsystems</category>
      <category>containers</category>
    </item>
    <item>
      <title>K8s: Topology Spread</title>
      <dc:creator>George Michalakis</dc:creator>
      <pubDate>Sun, 30 Aug 2026 16:27:11 +0000</pubDate>
      <link>https://dev.to/thegm26/k8s-topology-spread-1hcn</link>
      <guid>https://dev.to/thegm26/k8s-topology-spread-1hcn</guid>
      <description>&lt;p&gt;Following the previous article, where we were introduced to the &lt;a href="https://dev.to/thegm26/k8s-the-affinity-club-4c2"&gt;affinity club&lt;/a&gt;, we concluded that if:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;We apply pod anti-affinity to our critical service.&lt;/li&gt;
&lt;li&gt;We have only two worker nodes.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;We are “safe” if one node fails, but we cannot scale beyond two replicas during a traffic spike or another period of distress.&lt;/p&gt;

&lt;p&gt;This is where &lt;a href="https://kubernetes.io/docs/concepts/scheduling-eviction/topology-spread-constraints/" rel="noopener noreferrer"&gt;topology spread constraints&lt;/a&gt; come into play.&lt;/p&gt;

&lt;p&gt;Before continuing, it helps to remember what the Kubernetes scheduler does: whenever a new Pod needs a home, it filters out unsuitable nodes and then chooses among the remaining candidates.&lt;/p&gt;

&lt;p&gt;Topology spread constraints participate in that decision; they do not move Pods that are already running.&lt;/p&gt;

&lt;p&gt;In our &lt;code&gt;web&lt;/code&gt; Deployment, we add the following constraint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;topologySpreadConstraints&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;maxSkew&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;1&lt;/span&gt;
    &lt;span class="na"&gt;topologyKey&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;kubernetes.io/hostname&lt;/span&gt;
    &lt;span class="na"&gt;whenUnsatisfiable&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;DoNotSchedule&lt;/span&gt;
    &lt;span class="na"&gt;nodeAffinityPolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Honor&lt;/span&gt;
    &lt;span class="na"&gt;nodeTaintsPolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Honor&lt;/span&gt;
    &lt;span class="na"&gt;labelSelector&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;matchLabels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;app&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;web&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Let’s dissect it one field at a time.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;code&gt;maxSkew: 1&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;maxSkew&lt;/code&gt; defines the largest permitted difference in the number of matching Pods between topology domains. In our case, each worker node is a topology domain.&lt;/p&gt;

&lt;p&gt;For two worker nodes, the possible distributions look like this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;code&gt;worker&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;&lt;code&gt;worker2&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;Allowed?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;❌&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;❌&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If the difference is greater than &lt;code&gt;1&lt;/code&gt;, the distribution is not allowed. A new Pod that would violate this rule remains &lt;code&gt;Pending&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;code&gt;topologyKey: kubernetes.io/hostname&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;The topology key tells the scheduler how to divide the cluster into topology domains.&lt;/p&gt;

&lt;p&gt;Because &lt;code&gt;kubernetes.io/hostname&lt;/code&gt; normally has a unique value on each node, the scheduler spreads matching Pods across individual nodes. In a larger cluster, we could instead spread across zones by using a label such as &lt;code&gt;topology.kubernetes.io/zone&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Node inclusion policies
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;nodeAffinityPolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Honor&lt;/span&gt;
&lt;span class="na"&gt;nodeTaintsPolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Honor&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These fields control whether node affinity and &lt;a href="https://kubernetes.io/docs/concepts/scheduling-eviction/taint-and-toleration/" rel="noopener noreferrer"&gt;taints&lt;/a&gt; are respected when the scheduler decides which topology domains participate in the skew calculation.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;nodeAffinityPolicy: Honor&lt;/code&gt; has no practical effect in this experiment because the &lt;code&gt;web&lt;/code&gt; Pods do not define a &lt;code&gt;nodeSelector&lt;/code&gt; or node affinity. &lt;code&gt;nodeTaintsPolicy: Honor&lt;/code&gt;, however, matters because our nodes have taints.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;nodeTaintsPolicy&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;It controls whether node taints are considered when the scheduler calculates the topology spread.&lt;/p&gt;

&lt;p&gt;In our Kind cluster, the control-plane node has a &lt;code&gt;NoSchedule&lt;/code&gt; taint that the &lt;code&gt;web&lt;/code&gt; Pods do not tolerate.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu2itkdd4qeq9sg7ml1ke.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu2itkdd4qeq9sg7ml1ke.png" alt="Control-plane NoSchedule taint" width="799" height="139"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;With &lt;code&gt;nodeTaintsPolicy: Honor&lt;/code&gt;, the scheduler respects that taint and excludes the control-plane node from the spread calculation. It therefore calculates the distribution using only &lt;code&gt;worker&lt;/code&gt; and &lt;code&gt;worker2&lt;/code&gt;.&lt;/p&gt;

&lt;h4&gt;
  
  
  Without Honor
&lt;/h4&gt;

&lt;p&gt;Without &lt;code&gt;Honor&lt;/code&gt;, the default policy is &lt;code&gt;Ignore&lt;/code&gt;. The control-plane node could then be counted as an empty topology domain even though the &lt;code&gt;web&lt;/code&gt; Pods cannot actually run there.&lt;/p&gt;

&lt;p&gt;A distribution such as &lt;code&gt;1 / 1 / 0&lt;/code&gt; could make the next Pod violate &lt;code&gt;maxSkew: 1&lt;/code&gt; and remain &lt;code&gt;Pending&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Honor&lt;/code&gt; does not add a taint or evict existing Pods. It only controls which nodes participate in the topology-spread calculation.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;code&gt;whenUnsatisfiable: DoNotSchedule&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;If scheduling a Pod would violate &lt;code&gt;maxSkew&lt;/code&gt;, the scheduler leaves it &lt;code&gt;Pending&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;In other words, &lt;code&gt;DoNotSchedule&lt;/code&gt; prioritizes satisfying the spread constraint over scheduling the Pod immediately.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;code&gt;labelSelector&lt;/code&gt;
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;labelSelector&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;matchLabels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;app&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;web&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The label selector determines which Pods are counted when calculating the distribution.&lt;/p&gt;

&lt;p&gt;We are adding the constraint to a Deployment, so shouldn’t Kubernetes automatically count all Pods from that Deployment? No. The scheduler counts Pods whose labels match this selector. The &lt;code&gt;web&lt;/code&gt; Pod template must therefore carry the matching &lt;code&gt;app: web&lt;/code&gt; label.&lt;/p&gt;

&lt;p&gt;I know it feels a little weird. You can inspect the result visually with &lt;a href="https://github.com/kubernetes-sigs/headlamp" rel="noopener noreferrer"&gt;Headlamp&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Let’s see what we did:&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FThegm26%2Fdrainlab-k8%2Fmain%2Fassets%2Ftopology-spread-recording-1-light.gif%3Fv%3D2" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FThegm26%2Fdrainlab-k8%2Fmain%2Fassets%2Ftopology-spread-recording-1-light.gif%3Fv%3D2" alt="Four replicas spread across two workers" width="720" height="407"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;With four replicas, the scheduler places two Pods on each worker.&lt;/p&gt;

&lt;p&gt;Now suppose &lt;code&gt;worker2&lt;/code&gt; needs maintenance. &lt;code&gt;kubectl cordon worker2&lt;/code&gt; prevents new Pods from being scheduled there; &lt;code&gt;kubectl drain worker2&lt;/code&gt; would cordon it and then evict eligible Pods. In this demo, we recreate the &lt;code&gt;web&lt;/code&gt; Pods after cordoning to force the scheduler to make the placement decision again.&lt;/p&gt;

&lt;p&gt;In our cluster, the cordoned node receives the &lt;code&gt;node.kubernetes.io/unschedulable:NoSchedule&lt;/code&gt; taint. With &lt;code&gt;nodeTaintsPolicy: Honor&lt;/code&gt;, the scheduler excludes that node from the spread calculation.&lt;/p&gt;

&lt;p&gt;Because we have not set &lt;code&gt;minDomains&lt;/code&gt;, &lt;code&gt;worker&lt;/code&gt; becomes the only eligible topology domain. If it has enough capacity, all replacement Pods can be scheduled there. &lt;code&gt;maxSkew&lt;/code&gt; is therefore still satisfied: there is only one eligible domain to compare.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FThegm26%2Fdrainlab-k8%2Fmain%2Fassets%2Ftopology-spread-recording-2-light.gif%3Fv%3D2" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FThegm26%2Fdrainlab-k8%2Fmain%2Fassets%2Ftopology-spread-recording-2-light.gif%3Fv%3D2" alt="Replacement Pods on the eligible worker" width="599" height="339"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  But in the second case, did we hide the issue under the carpet?
&lt;/h2&gt;

&lt;p&gt;Yes and no.&lt;/p&gt;

&lt;p&gt;With &lt;code&gt;maxSkew: 1&lt;/code&gt;, we wanted to say: &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If one worker disappears, do not keep stacking replicas on the other one.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But that is not what the manifest currently says. It only limits skew across the topology domains that participate in the calculation.&lt;/p&gt;

&lt;p&gt;To require at least two eligible domains, we add &lt;code&gt;minDomains: 2&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;topologySpreadConstraints&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;maxSkew&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;1&lt;/span&gt;
    &lt;span class="na"&gt;minDomains&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;2&lt;/span&gt; &lt;span class="c1"&gt;# Require at least two eligible domains (workers)&lt;/span&gt;
    &lt;span class="na"&gt;topologyKey&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;kubernetes.io/hostname&lt;/span&gt;
    &lt;span class="na"&gt;whenUnsatisfiable&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;DoNotSchedule&lt;/span&gt;
    &lt;span class="na"&gt;nodeAffinityPolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Honor&lt;/span&gt;
    &lt;span class="na"&gt;nodeTaintsPolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Honor&lt;/span&gt;
    &lt;span class="na"&gt;labelSelector&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;matchLabels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;app&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;web&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When fewer than two eligible domains remain, Kubernetes treats the global minimum Pod count as zero when calculating skew.&lt;/p&gt;

&lt;p&gt;Below, we scale the Deployment to five replicas and compare two scenarios. &lt;/p&gt;

&lt;p&gt;First, &lt;code&gt;worker2&lt;/code&gt; is cordoned. With only one eligible domain and &lt;code&gt;minDomains: 2&lt;/code&gt;, Kubernetes schedules one Pod on &lt;code&gt;worker&lt;/code&gt;; the other four remain &lt;code&gt;Pending&lt;/code&gt; because placing another Pod would violate &lt;code&gt;maxSkew: 1&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;After I uncordon &lt;code&gt;worker2&lt;/code&gt;, both domains become eligible again and all five Pods run with a balanced &lt;code&gt;3/2&lt;/code&gt; distribution.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FThegm26%2Fdrainlab-k8%2Fmain%2Fassets%2Ftopology-spread-min-domains-demo.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FThegm26%2Fdrainlab-k8%2Fmain%2Fassets%2Ftopology-spread-min-domains-demo.gif" alt="Topology spread with minDomains" width="720" height="402"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is the difference between &lt;code&gt;maxSkew&lt;/code&gt; and &lt;code&gt;minDomains&lt;/code&gt;:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;maxSkew&lt;/code&gt; controls balance across eligible domains, while &lt;code&gt;minDomains&lt;/code&gt; specifies how many eligible domains must exist before Kubernetes permits further placement. It preserves the intended topology, not application availability by itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  What comes next?
&lt;/h2&gt;

&lt;p&gt;We now know how to control where new Pods may be scheduled, but maintenance introduces another question: how many running Pods may Kubernetes voluntarily evict at the same time?&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/thegm26/k8s-node-maintenance-eviction-1mkm"&gt;Eviction? Maintenance? Why do we care?&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Read about PodDisruptionBudgets (PDBs), and let's find out together :D&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>openshift</category>
      <category>distributedsystems</category>
    </item>
    <item>
      <title>K8s: The Affinity Club</title>
      <dc:creator>George Michalakis</dc:creator>
      <pubDate>Sun, 16 Aug 2026 10:26:02 +0000</pubDate>
      <link>https://dev.to/thegm26/k8s-the-affinity-club-4c2</link>
      <guid>https://dev.to/thegm26/k8s-the-affinity-club-4c2</guid>
      <description>&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;If you have fiddled with Kubernetes long enough to worry about which node your Pods are actually running on, you can probably relate.&lt;/p&gt;

&lt;p&gt;First: what does &lt;em&gt;affinity&lt;/em&gt; even mean?&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Affinity:&lt;/strong&gt; a strong feeling that you understand or like someone or something; a close relationship between people or things with similar qualities.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;p&gt;Okay. Fair enough.&lt;/p&gt;

&lt;p&gt;In Kubernetes, affinity is a family of scheduling rules:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Node affinity&lt;/strong&gt; defines hard requirements and soft preferences for which nodes can run a Pod.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pod affinity&lt;/strong&gt; lets us place a Pod near other Pods.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pod anti-affinity&lt;/strong&gt; lets us keep Pods away from other Pods.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So, loosely, if I were a Pod, I would say:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“I have an affinity for being near or away from these Pods.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But why care?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;At the end of the day, I can always scale to more replicas. I’m safe… right?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fap7g99un2wouabo5li9e.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fap7g99un2wouabo5li9e.gif" alt="Asking are you sure about that?" width="250" height="250"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Let’s see it in action
&lt;/h3&gt;

&lt;p&gt;Let’s take the worst-case scenario.&lt;/p&gt;

&lt;p&gt;Imagine a cluster with two worker nodes. We have a critical service, but all of its Pods happen to be scheduled on &lt;code&gt;worker-1&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Traffic grows, so we scale the Deployment. But every new replica still lands on &lt;code&gt;worker-1&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpeyzeipzp7yq3wnsvq6r.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpeyzeipzp7yq3wnsvq6r.gif" alt="Lab 01 placement visual" width="800" height="393"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;After scaling, two things become obvious:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;worker-1&lt;/code&gt; is becoming stressed.&lt;/li&gt;
&lt;li&gt;If &lt;code&gt;worker-1&lt;/code&gt; fails, every replica fails with it—and the application is down.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The scheduler is not random, but without a placement rule, &lt;strong&gt;it has no obligation&lt;/strong&gt;* to spread our replicas across workers.&lt;/p&gt;

&lt;p&gt;In this scenario, &lt;code&gt;worker-2&lt;/code&gt; is right there... &lt;strong&gt;empty&lt;/strong&gt;..&lt;/p&gt;

&lt;p&gt;Ok. Can we force Kubernetes to keep replicas of the same service on different workers?&lt;/p&gt;




&lt;p&gt;Yes. That is exactly what pod anti-affinity is for.&lt;/p&gt;

&lt;p&gt;Below, we apply pod anti-affinity to our critical service’s Deployment.&lt;/p&gt;

&lt;h3&gt;
  
  
  Before: no placement rule
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;apps/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Deployment&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;web&lt;/span&gt;
  &lt;span class="na"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;drainlab&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;replicas&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;2&lt;/span&gt;
  &lt;span class="na"&gt;selector&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;matchLabels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;app&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;web&lt;/span&gt;
  &lt;span class="na"&gt;template&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;app&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;web&lt;/span&gt;
    &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
          &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx:1.27-alpine&lt;/span&gt;
          &lt;span class="na"&gt;resources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="na"&gt;requests&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="na"&gt;cpu&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;250m&lt;/span&gt;
              &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;128Mi&lt;/span&gt;
          &lt;span class="na"&gt;ports&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;containerPort&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;80&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  After: hard pod anti-affinity
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;apps/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Deployment&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;web&lt;/span&gt;
  &lt;span class="na"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;drainlab&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;replicas&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;2&lt;/span&gt;
  &lt;span class="na"&gt;selector&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;matchLabels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;app&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;web&lt;/span&gt;
  &lt;span class="na"&gt;template&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;app&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;web&lt;/span&gt;
    &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;affinity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="c1"&gt;# &amp;lt;-Magic starts to happen here&lt;/span&gt;
        &lt;span class="na"&gt;podAntiAffinity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; 
          &lt;span class="na"&gt;requiredDuringSchedulingIgnoredDuringExecution&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;labelSelector&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                &lt;span class="na"&gt;matchLabels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                  &lt;span class="na"&gt;app&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;web&lt;/span&gt;
              &lt;span class="na"&gt;topologyKey&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;kubernetes.io/hostname&lt;/span&gt; &lt;span class="c1"&gt;# &amp;lt;-finishes here&lt;/span&gt;
      &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
          &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx:1.27-alpine&lt;/span&gt;
          &lt;span class="na"&gt;resources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="na"&gt;requests&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="na"&gt;cpu&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;250m&lt;/span&gt;
              &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;128Mi&lt;/span&gt;
          &lt;span class="na"&gt;ports&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;containerPort&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;80&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This tells the scheduler:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;requiredDuringScheduling “A new &lt;code&gt;app: web&lt;/code&gt; Pod cannot be scheduled onto a node that already runs another matching &lt;code&gt;app: web&lt;/code&gt; Pod.”&lt;/p&gt;

&lt;p&gt;If every eligible worker already runs one, the new Pod stays &lt;code&gt;Pending&lt;/code&gt; rather than placing two replicas on the same node.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;More specifically:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;requiredDuringScheduling&lt;/code&gt;: if no valid worker exists, the new Pod stays &lt;code&gt;Pending&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;IgnoredDuringExecution&lt;/code&gt;: Kubernetes does not evict an already-running Pod merely because later changes violate the anti-affinity condition.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;topologyKey: kubernetes.io/hostname&lt;/code&gt;: each worker node is treated as a separate failure domain.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwzqz3wjpix4gr2wac7o2.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwzqz3wjpix4gr2wac7o2.gif" alt="01lab-02" width="719" height="353"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;But wait… if I have two eligible worker nodes and need four replicas, do the other two remain &lt;code&gt;Pending&lt;/code&gt;?&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwabmxep9e7ahawrn4et3.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwabmxep9e7ahawrn4et3.gif" alt="Shocked Cat" width="220" height="209"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Before anti-affinity, all four Pods could run on &lt;code&gt;worker-1&lt;/code&gt;, even though one node failure would take down the entire service.&lt;/p&gt;

&lt;p&gt;Now, Kubernetes protects us from that false sense of safety but hard pod anti-affinity &lt;strong&gt;limits&lt;/strong&gt; us to one matching Pod per worker.&lt;/p&gt;

&lt;p&gt;So we gained failure isolation, but we gave up that &lt;strong&gt;easy&lt;/strong&gt; scaling?&lt;/p&gt;

&lt;p&gt;No, This is where topology spread constraints come in.&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;*Kubernetes may already prefer to spread Pods using soft scheduling preferences. Exact placement depends on available resources and scheduler configuration. This example demonstrates why relying on a preference is different from declaring a hard availability rule.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;Labs / Visualizations for this article can be found &lt;a href="https://github.com/Thegm26/drainlab-k8" rel="noopener noreferrer"&gt;here&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>devops</category>
      <category>sre</category>
    </item>
    <item>
      <title>Save yourself from Architectural Amnesia: ADRs</title>
      <dc:creator>George Michalakis</dc:creator>
      <pubDate>Sat, 14 Mar 2026 02:03:07 +0000</pubDate>
      <link>https://dev.to/thegm26/save-yourself-from-architectural-amnesia-adrs-506p</link>
      <guid>https://dev.to/thegm26/save-yourself-from-architectural-amnesia-adrs-506p</guid>
      <description>&lt;h2&gt;
  
  
  When you see it...
&lt;/h2&gt;

&lt;p&gt;From smaller teams to larger ones, from the senior SWE managing the entire backlog to EOs, POs, and SMs juggling overlapping responsibilities, I have often found myself in meetings or PR discussions nitpicking scope details and suddenly realizing one of two things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;I disagree with the broader implementation picture.&lt;/li&gt;
&lt;li&gt;I do not even remember whether I agreed with it in the first place.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That second one is worse.&lt;/p&gt;

&lt;p&gt;For context, I am working in a SAFe Scrum setup, so these roles and handoffs are a very real part of day-to-day delivery.&lt;/p&gt;

&lt;h2&gt;
  
  
  Half-measures (Best-case scenario)
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;"Uhh, there is a Confluence page about this decision. We had a meeting for that."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;Clicks the link.&lt;/em&gt;&lt;br&gt;&lt;br&gt;
&lt;em&gt;Scans a page with 10 comments&lt;/em&gt; (8 resolved, 2 still open).&lt;br&gt;&lt;br&gt;
OK, this makes a bit more sense now.&lt;/p&gt;

&lt;p&gt;Still, I do not remember what we said. I do not remember how long it took, and it probably took longer than it should have.&lt;/p&gt;

&lt;p&gt;Even in this best-case scenario, I am still left with two problems:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;I do not remember the back-and-forth, and some of it was probably important.&lt;/li&gt;
&lt;li&gt;If I now have a different opinion, I really do not want to reopen the topic. That usually means more meetings, more pages, and more comments. Atlassian may be doing a great job with version control, but I do not want to be stuck restoring version 23 of a Confluence page that nobody will care about two weeks later.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;
  
  
  Give credit where it is due
&lt;/h2&gt;

&lt;p&gt;That does not mean Confluence is useless.&lt;/p&gt;

&lt;p&gt;Architects, PMs, and business stakeholders need version control, comments, and discussions too. You cannot expect a product owner:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;To log in to GitHub every time they want to sync with the architect on whether the team is migrating the DB.&lt;/li&gt;
&lt;li&gt;To align on the confidence level of PI objectives with different POs inside a pull request.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Confluence has its place.&lt;/p&gt;

&lt;p&gt;But at the same time, nobody wants to spend all of this just to discuss something like a new or existing PR label:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;10 minutes to raise it during the daily stand-up&lt;/li&gt;
&lt;li&gt;5 minutes to schedule a call&lt;/li&gt;
&lt;li&gt;20 minutes to prepare a page explaining the current state and the proposal&lt;/li&gt;
&lt;li&gt;35 minutes for the meeting itself&lt;/li&gt;
&lt;li&gt;20 minutes for the notes afterward&lt;/li&gt;
&lt;li&gt;20 minutes updating the page with the latest comments and decisions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And yes, those numbers are completely realistic.&lt;/p&gt;
&lt;h2&gt;
  
  
  Architecture Decision Records
&lt;/h2&gt;

&lt;p&gt;This is where ADRs help.&lt;/p&gt;

&lt;p&gt;An ADR is a lightweight record of an important technical decision: why it was needed, what options were considered, what was chosen, and what tradeoffs came with that choice.&lt;/p&gt;

&lt;p&gt;At its simplest, it is just a markdown file living in your repo, somewhere like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;docs/
  adrs/
    0001-use-topic-a.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is basically it.&lt;/p&gt;

&lt;p&gt;A small document, committed with the code, with sections such as:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Context: what problem are we trying to solve?&lt;/li&gt;
&lt;li&gt;Decision: what did we choose?&lt;/li&gt;
&lt;li&gt;Alternatives considered: what were the other realistic options?&lt;/li&gt;
&lt;li&gt;Consequences: what do we gain, and what do we give up?&lt;/li&gt;
&lt;li&gt;Status: proposed, accepted, superseded, deprecated&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Nothing fancy. Just enough to preserve the reasoning behind a technical decision close to the codebase where engineers will actually look for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Basic rules
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Keep ADRs for technical topics that are likely to be searched while someone is working in the repo. If you used a specific pattern, tool, or extension, I would much rather &lt;code&gt;Ctrl+F&lt;/code&gt; the repo and find the reasoning in the docs than dig through a GitHub page or old meeting notes.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Make the person opening the PR the driver of the discussion. They should gather feedback, collect comments, help the team converge on a decision, and eventually merge or close the PR with the ADR alongside the change.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Do not optimize for unanimous agreement. Optimize for a clear decision with explicit tradeoffs and enough context that the next person can understand why it happened.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That is the real value: not perfect documentation, but recorded reasoning close to the codebase.&lt;/p&gt;

&lt;p&gt;If you want a good collection of ADR templates, look &lt;a href="https://github.com/joelparkerhenderson/architecture-decision-record" rel="noopener noreferrer"&gt;here&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;PS1: Use this as a starting point for introducing ADRs to your team.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;PS2: An agent helped with syntax refactoring. The expressions, the main writing, the flow and the pain is mine :)&lt;/em&gt;&lt;/p&gt;

</description>
      <category>adr</category>
      <category>documentation</category>
      <category>architecture</category>
    </item>
  </channel>
</rss>
