<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Renato Silva</title>
    <description>The latest articles on DEV Community by Renato Silva (@renato_silva_71eef0fc385f).</description>
    <link>https://dev.to/renato_silva_71eef0fc385f</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1698909%2F98e99ca4-6bff-40dc-9314-cf98388fbbf3.jpg</url>
      <title>DEV Community: Renato Silva</title>
      <link>https://dev.to/renato_silva_71eef0fc385f</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/renato_silva_71eef0fc385f"/>
    <language>en</language>
    <item>
      <title>Riverpod vs Bloc: Fixing the FutureBuilder Anti-Pattern</title>
      <dc:creator>Renato Silva</dc:creator>
      <pubDate>Thu, 01 Oct 2026 15:55:35 +0000</pubDate>
      <link>https://dev.to/renato_silva_71eef0fc385f/riverpod-vs-bloc-fixing-the-futurebuilder-anti-pattern-30d4</link>
      <guid>https://dev.to/renato_silva_71eef0fc385f/riverpod-vs-bloc-fixing-the-futurebuilder-anti-pattern-30d4</guid>
      <description>&lt;h2&gt;
  
  
  🔧 The Problem
&lt;/h2&gt;

&lt;p&gt;If you've worked on a Flutter codebase for more than a few sprints, you've seen this pattern:&lt;/p&gt;

&lt;p&gt;dart&lt;br&gt;
class UserProfileScreen extends StatelessWidget {&lt;br&gt;
  final String userId;&lt;br&gt;
  const UserProfileScreen({required this.userId, super.key});&lt;/p&gt;

&lt;p&gt;Future _fetchUser() =&amp;gt; UserRepository().getUser(userId);&lt;/p&gt;

&lt;p&gt;&lt;a class="mentioned-user" href="https://dev.to/override"&gt;@override&lt;/a&gt;&lt;br&gt;
  Widget build(BuildContext context) {&lt;br&gt;
    return FutureBuilder(&lt;br&gt;
      future: _fetchUser(),&lt;br&gt;
      builder: (context, snapshot) {&lt;br&gt;
        if (snapshot.connectionState == ConnectionState.waiting) {&lt;br&gt;
          return const CircularProgressIndicator();&lt;br&gt;
        }&lt;br&gt;
        if (snapshot.hasError) {&lt;br&gt;
          return Text('Error: ${snapshot.error}');&lt;br&gt;
        }&lt;br&gt;
        return UserCard(user: snapshot.data!);&lt;br&gt;
      },&lt;br&gt;
    );&lt;br&gt;
  }&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;It looks harmless. It even looks idiomatic — it's in every Flutter tutorial. But &lt;code&gt;_fetchUser()&lt;/code&gt; is called &lt;strong&gt;inside &lt;code&gt;build()&lt;/code&gt;&lt;/strong&gt;, which means it reruns on every rebuild that touches this widget. Scroll the parent &lt;code&gt;ListView&lt;/code&gt;, trigger a &lt;code&gt;setState&lt;/code&gt; two levels up, rotate the device — and congratulations, you just refetched the user, flashed a loading spinner over good data, and maybe hit your API rate limit.&lt;/p&gt;

&lt;p&gt;The usual "fix" is to hoist the future into a &lt;code&gt;late final&lt;/code&gt; field or &lt;code&gt;initState&lt;/code&gt;, which helps — until the widget gets rebuilt by its parent with a new key, or the &lt;code&gt;userId&lt;/code&gt; changes and nothing refreshes because the future was already memoized. Now you've got stale data instead of redundant fetches. Either way, the async lifecycle is coupled to the widget lifecycle, and widget lifecycles are not a reliable place to put business logic.&lt;/p&gt;

&lt;h2&gt;
  
  
  🕵️ Why This Keeps Biting Teams
&lt;/h2&gt;

&lt;p&gt;The core issue isn't &lt;code&gt;FutureBuilder&lt;/code&gt; itself — it's a well-built widget. The issue is &lt;strong&gt;where the future is created&lt;/strong&gt;. Widgets rebuild for reasons that have nothing to do with your data: theme changes, &lt;code&gt;MediaQuery&lt;/code&gt; updates, a sibling's &lt;code&gt;setState&lt;/code&gt;, hot reload. None of those should restart a network call, but if the future lives in &lt;code&gt;build()&lt;/code&gt;, they will.&lt;/p&gt;

&lt;p&gt;The real fix is architectural: async operations belong in a layer that outlives individual widget rebuilds and has an explicit, observable lifecycle. That's exactly what Riverpod and Bloc are for — they just get there differently.&lt;/p&gt;

&lt;h2&gt;
  
  
  🧩 Riverpod: AsyncValue as a First-Class Citizen
&lt;/h2&gt;

&lt;p&gt;Riverpod's &lt;code&gt;FutureProvider&lt;/code&gt; (or &lt;code&gt;AsyncNotifier&lt;/code&gt; for anything more than a one-shot fetch) moves the async call out of the widget tree entirely. The provider is keyed by its arguments, cached, and only re-executes when its dependencies actually change.&lt;/p&gt;

&lt;p&gt;dart&lt;br&gt;
@riverpod&lt;br&gt;
Future user(UserRef ref, String userId) {&lt;br&gt;
  return ref.watch(userRepositoryProvider).getUser(userId);&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;dart&lt;br&gt;
class UserProfileScreen extends ConsumerWidget {&lt;br&gt;
  final String userId;&lt;br&gt;
  const UserProfileScreen({required this.userId, super.key});&lt;/p&gt;

&lt;p&gt;&lt;a class="mentioned-user" href="https://dev.to/override"&gt;@override&lt;/a&gt;&lt;br&gt;
  Widget build(BuildContext context, WidgetRef ref) {&lt;br&gt;
    final userAsync = ref.watch(userProvider(userId));&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;return userAsync.when(
  data: (user) =&amp;gt; UserCard(user: user),
  loading: () =&amp;gt; const CircularProgressIndicator(),
  error: (err, stack) =&amp;gt; Text('Error: $err'),
);
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;}&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;The win here is subtle but important: &lt;code&gt;ref.watch(userProvider(userId))&lt;/code&gt; doesn't re-run the future on rebuild. The provider is cached against the &lt;code&gt;userId&lt;/code&gt; argument. If a parent rebuilds and this widget rebuilds with it, Riverpod just hands back the existing &lt;code&gt;AsyncValue&lt;/code&gt; — no new HTTP call. If &lt;code&gt;userId&lt;/code&gt; changes, Riverpod recognizes it as a different provider instance and fetches fresh data automatically. You get request deduplication and cache invalidation almost for free, and &lt;code&gt;AsyncValue.when&lt;/code&gt; forces you to handle loading/error/data explicitly, which eliminates the "forgot to check &lt;code&gt;snapshot.hasError&lt;/code&gt;" class of bugs.&lt;/p&gt;

&lt;p&gt;The trade-off: you're leaning on code generation (&lt;code&gt;@riverpod&lt;/code&gt;) and a mental model of providers-as-dependency-graph, which has a real learning curve if your team hasn't used DI-flavored patterns before. Debugging "why did this provider rebuild" sometimes means reading &lt;code&gt;ref.watch&lt;/code&gt; chains across multiple files.&lt;/p&gt;

&lt;h2&gt;
  
  
  🧱 Bloc: Explicit States for Explicit Transitions
&lt;/h2&gt;

&lt;p&gt;Bloc takes a more ceremonial but arguably more auditable approach: you model every async phase as a discrete state, and a &lt;code&gt;Cubit&lt;/code&gt; or &lt;code&gt;Bloc&lt;/code&gt; owns the transition logic.&lt;/p&gt;

&lt;p&gt;dart&lt;br&gt;
sealed class UserState {}&lt;br&gt;
class UserInitial extends UserState {}&lt;br&gt;
class UserLoading extends UserState {}&lt;br&gt;
class UserLoaded extends UserState {&lt;br&gt;
  final User user;&lt;br&gt;
  UserLoaded(this.user);&lt;br&gt;
}&lt;br&gt;
class UserError extends UserState {&lt;br&gt;
  final String message;&lt;br&gt;
  UserError(this.message);&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;class UserCubit extends Cubit {&lt;br&gt;
  final UserRepository repository;&lt;br&gt;
  UserCubit(this.repository) : super(UserInitial());&lt;/p&gt;

&lt;p&gt;Future loadUser(String userId) async {&lt;br&gt;
    emit(UserLoading());&lt;br&gt;
    try {&lt;br&gt;
      final user = await repository.getUser(userId);&lt;br&gt;
      emit(UserLoaded(user));&lt;br&gt;
    } catch (e) {&lt;br&gt;
      emit(UserError(e.toString()));&lt;br&gt;
    }&lt;br&gt;
  }&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;dart&lt;br&gt;
class UserProfileScreen extends StatelessWidget {&lt;br&gt;
  final String userId;&lt;br&gt;
  const UserProfileScreen({required this.userId, super.key});&lt;/p&gt;

&lt;p&gt;&lt;a class="mentioned-user" href="https://dev.to/override"&gt;@override&lt;/a&gt;&lt;br&gt;
  Widget build(BuildContext context) {&lt;br&gt;
    return BlocProvider(&lt;br&gt;
      create: (_) =&amp;gt; UserCubit(context.read())..loadUser(userId),&lt;br&gt;
      child: BlocBuilder(&lt;br&gt;
        builder: (context, state) =&amp;gt; switch (state) {&lt;br&gt;
          UserLoading() || UserInitial() =&amp;gt; const CircularProgressIndicator(),&lt;br&gt;
          UserLoaded(:final user) =&amp;gt; UserCard(user: user),&lt;br&gt;
          UserError(:final message) =&amp;gt; Text('Error: $message'),&lt;br&gt;
        },&lt;br&gt;
      ),&lt;br&gt;
    );&lt;br&gt;
  }&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;Because &lt;code&gt;loadUser&lt;/code&gt; is called once in &lt;code&gt;create&lt;/code&gt;, not on every &lt;code&gt;build&lt;/code&gt;, the fetch is decoupled from widget rebuilds the same way Riverpod decouples it — just via explicit lifecycle methods instead of dependency-graph caching. &lt;code&gt;BlocProvider&lt;/code&gt; guarantees the &lt;code&gt;Cubit&lt;/code&gt; instance (and its in-flight future) survives rebuilds of the subtree, as long as you don't recreate the provider higher up the tree.&lt;/p&gt;

&lt;p&gt;The trade-off is verbosity: four state classes for one fetch, plus boilerplate if you need to refetch when &lt;code&gt;userId&lt;/code&gt; changes (you'd typically do that with a &lt;code&gt;didUpdateWidget&lt;/code&gt; check or a &lt;code&gt;BlocListener&lt;/code&gt; reacting to a parent-level event). In exchange, you get a state machine that's trivially testable with &lt;code&gt;bloc_test&lt;/code&gt; and a transition log you can replay for debugging — genuinely valuable in regulated or high-incident-rate codebases where "what sequence of states did we actually go through" matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  ⚖️ The Actual Trade-off
&lt;/h2&gt;

&lt;p&gt;Both solve the same root problem — stop constructing futures inside &lt;code&gt;build()&lt;/code&gt; — but they optimize for different failure modes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Riverpod&lt;/strong&gt; optimizes for &lt;em&gt;correctness with less code&lt;/em&gt;: caching, deduplication, and parameterized providers are handled by the framework. You pay for it with a steeper initial learning curve and less explicit control over state transitions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bloc&lt;/strong&gt; optimizes for &lt;em&gt;auditability and testability&lt;/em&gt;: every state change is a named, inspectable event. You pay for it with more files and manual wiring for things like argument-based refetching.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If your team is small, iterates fast, and wants async logic colocated with minimal ceremony, Riverpod's &lt;code&gt;AsyncNotifier&lt;/code&gt; usually wins. If you're on a larger team, need strict separation between "what happened" and "how the UI reacts," or already have Bloc conventions baked into code review checklists, stick with it — just make sure you're creating the &lt;code&gt;Bloc&lt;/code&gt;/&lt;code&gt;Cubit&lt;/code&gt; outside of &lt;code&gt;build()&lt;/code&gt;, same rule as Riverpod providers.&lt;/p&gt;

&lt;p&gt;What neither pattern tolerates is the original sin: constructing async work inside the widget's &lt;code&gt;build()&lt;/code&gt; method. That's the actual anti-pattern, not &lt;code&gt;FutureBuilder&lt;/code&gt; itself — &lt;code&gt;FutureBuilder&lt;/code&gt; is just the symptom you see first.&lt;/p&gt;

&lt;h2&gt;
  
  
  🗣️ Over to You
&lt;/h2&gt;

&lt;p&gt;Which pattern has actually survived contact with your production incidents — Riverpod's cached providers or Bloc's explicit state machine? Or have you found a third way (raw &lt;code&gt;ChangeNotifier&lt;/code&gt;, signals, &lt;code&gt;flutter_hooks&lt;/code&gt;) that handles this better than either?&lt;/p&gt;

</description>
      <category>flutter</category>
      <category>riverpod</category>
      <category>bloc</category>
      <category>dart</category>
    </item>
    <item>
      <title>Your AI Code Reviewer Needs a Test Suite Too</title>
      <dc:creator>Renato Silva</dc:creator>
      <pubDate>Thu, 24 Sep 2026 13:58:17 +0000</pubDate>
      <link>https://dev.to/renato_silva_71eef0fc385f/your-ai-code-reviewer-needs-a-test-suite-too-4g41</link>
      <guid>https://dev.to/renato_silva_71eef0fc385f/your-ai-code-reviewer-needs-a-test-suite-too-4g41</guid>
      <description>&lt;h2&gt;
  
  
  🤖 The Problem
&lt;/h2&gt;

&lt;p&gt;You plugged an LLM into your PR pipeline. It leaves comments about SQL injection, N+1 queries, missing null checks. Everyone's thrilled for about two sprints, and then someone notices it stopped catching the exact same bug pattern it used to flag. A prompt got tweaked. A model version got bumped upstream. Someone added "be more concise" to the system prompt and it started skipping the security section entirely.&lt;/p&gt;

&lt;p&gt;Nobody noticed because nobody was checking. The AI reviewer is running in production, making judgment calls on every PR in your org, and it has zero test coverage. That's the part that should bother you more than it probably does.&lt;/p&gt;

&lt;h2&gt;
  
  
  🕳️ Why This Gap Exists
&lt;/h2&gt;

&lt;p&gt;We test the code the AI reviews. We don't test the reviewer itself, because it feels fuzzy — "it's an LLM, how do you even assert against prose?" But that's a cop-out. You don't need to assert on exact wording. You need to assert on &lt;em&gt;behavior&lt;/em&gt;: given a diff with a known bug, does the reviewer flag it, in the right file, with the right severity, without three false positives burying the real issue?&lt;/p&gt;

&lt;p&gt;That's a testable claim. It just requires building fixtures the same way you'd build fixtures for any other system with non-deterministic-ish output — like testing a search ranking algorithm or a fraud-detection model.&lt;/p&gt;

&lt;h2&gt;
  
  
  🧪 Building a Golden-Diff Suite
&lt;/h2&gt;

&lt;p&gt;The core idea: collect real PRs (or synthetic ones) with known, labeled defects, and treat them like golden files. Run each one through your reviewer, then assert the output contains — or doesn't contain — specific findings.&lt;/p&gt;

&lt;p&gt;Here's a minimal structure using Python and pytest, since most review pipelines are just a script wrapping an LLM call:&lt;/p&gt;

&lt;p&gt;python&lt;/p&gt;

&lt;h1&gt;
  
  
  reviewer/client.py
&lt;/h1&gt;

&lt;p&gt;import os&lt;br&gt;
from openai import OpenAI&lt;/p&gt;

&lt;p&gt;client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])&lt;/p&gt;

&lt;p&gt;SYSTEM_PROMPT = open("prompts/review_system.md").read()&lt;/p&gt;

&lt;p&gt;def review_diff(diff_text: str) -&amp;gt; str:&lt;br&gt;
    response = client.chat.completions.create(&lt;br&gt;
        model="gpt-4.1",&lt;br&gt;
        temperature=0,&lt;br&gt;
        messages=[&lt;br&gt;
            {"role": "system", "content": SYSTEM_PROMPT},&lt;br&gt;
            {"role": "user", "content": f"Review this diff:\n\n{diff_text}"},&lt;br&gt;
        ],&lt;br&gt;
    )&lt;br&gt;
    return response.choices[0].message.content&lt;/p&gt;

&lt;p&gt;Note &lt;code&gt;temperature=0&lt;/code&gt;. You're not eliminating nondeterminism, but you're minimizing it — this matters a lot when you're about to assert against the output.&lt;/p&gt;

&lt;h2&gt;
  
  
  📁 Structuring the Fixtures
&lt;/h2&gt;

&lt;p&gt;Each fixture is a known-bad diff plus a manifest describing what &lt;em&gt;should&lt;/em&gt; get flagged:&lt;/p&gt;

&lt;p&gt;fixtures/&lt;br&gt;
  sql_injection_raw_query/&lt;br&gt;
    diff.patch&lt;br&gt;
    expected.yaml&lt;br&gt;
  missing_await_async_call/&lt;br&gt;
    diff.patch&lt;br&gt;
    expected.yaml&lt;br&gt;
  hardcoded_secret_in_config/&lt;br&gt;
    diff.patch&lt;br&gt;
    expected.yaml&lt;/p&gt;

&lt;p&gt;yaml&lt;/p&gt;

&lt;h1&gt;
  
  
  fixtures/sql_injection_raw_query/expected.yaml
&lt;/h1&gt;

&lt;p&gt;must_flag:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;category: security
file: app/db/queries.py
line_range: [42, 45]
keywords: ["injection", "parameteriz", "sanitiz"]
must_not_flag:&lt;/li&gt;
&lt;li&gt;category: style
keywords: ["variable naming"]
max_total_comments: 4&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The &lt;code&gt;must_not_flag&lt;/code&gt; block matters as much as &lt;code&gt;must_flag&lt;/code&gt;. A reviewer that comments on everything technically "catches" every bug, but it's useless noise. You want to pin down both precision and recall.&lt;/p&gt;

&lt;h2&gt;
  
  
  ✅ Writing the Assertions
&lt;/h2&gt;

&lt;p&gt;Since the output is prose, not structured data, you have two options: force structured output (JSON mode, function calling) or parse loosely with keyword/semantic matching. I'd push hard for structured output — it makes the whole test suite dramatically less brittle.&lt;/p&gt;

&lt;p&gt;python&lt;/p&gt;

&lt;h1&gt;
  
  
  reviewer/schema.py
&lt;/h1&gt;

&lt;p&gt;from pydantic import BaseModel&lt;/p&gt;

&lt;p&gt;class Finding(BaseModel):&lt;br&gt;
    category: str&lt;br&gt;
    file: str&lt;br&gt;
    line_start: int&lt;br&gt;
    line_end: int&lt;br&gt;
    message: str&lt;br&gt;
    severity: str&lt;/p&gt;

&lt;p&gt;class ReviewResult(BaseModel):&lt;br&gt;
    findings: list[Finding]&lt;/p&gt;

&lt;p&gt;python&lt;/p&gt;

&lt;h1&gt;
  
  
  tests/test_golden_diffs.py
&lt;/h1&gt;

&lt;p&gt;import yaml&lt;br&gt;
import pytest&lt;br&gt;
from pathlib import Path&lt;br&gt;
from reviewer.client import review_diff&lt;br&gt;
from reviewer.schema import ReviewResult&lt;/p&gt;

&lt;p&gt;FIXTURE_DIR = Path("fixtures")&lt;/p&gt;

&lt;p&gt;def load_fixtures():&lt;br&gt;
    for folder in FIXTURE_DIR.iterdir():&lt;br&gt;
        diff = (folder / "diff.patch").read_text()&lt;br&gt;
        expected = yaml.safe_load((folder / "expected.yaml").read_text())&lt;br&gt;
        yield pytest.param(diff, expected, id=folder.name)&lt;/p&gt;

&lt;p&gt;@pytest.mark.parametrize("diff,expected", load_fixtures())&lt;br&gt;
def test_reviewer_catches_known_bug(diff, expected):&lt;br&gt;
    raw = review_diff(diff)&lt;br&gt;
    result = ReviewResult.model_validate_json(raw)&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;for must in expected.get("must_flag", []):
    matches = [
        f for f in result.findings
        if f.category == must["category"]
        and f.file == must["file"]
        and any(kw in f.message.lower() for kw in must["keywords"])
    ]
    assert matches, f"Reviewer missed expected finding: {must}"

for forbidden in expected.get("must_not_flag", []):
    matches = [
        f for f in result.findings
        if any(kw in f.message.lower() for kw in forbidden["keywords"])
    ]
    assert not matches, f"Reviewer raised noise it shouldn't have: {forbidden}"

max_comments = expected.get("max_total_comments")
if max_comments is not None:
    assert len(result.findings) &amp;lt;= max_comments
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;This is a regression suite in the truest sense: every time someone edits the system prompt, swaps the model, or adjusts temperature, this runs and tells you exactly what broke. "Prompt change reduced recall on SQL injection cases from 100% to 60%" is a real, actionable CI failure — not a vibe.&lt;/p&gt;

&lt;h2&gt;
  
  
  🔄 Running It in CI
&lt;/h2&gt;

&lt;p&gt;The expensive part is the LLM calls, so don't run this on every commit to every branch. Run it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;On any PR that touches &lt;code&gt;prompts/&lt;/code&gt;, &lt;code&gt;reviewer/&lt;/code&gt;, or model config&lt;/li&gt;
&lt;li&gt;Nightly, to catch silent drift from provider-side model updates&lt;/li&gt;
&lt;li&gt;Before promoting a new prompt version to production, as a hard gate&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;yaml&lt;/p&gt;

&lt;h1&gt;
  
  
  .github/workflows/reviewer-regression.yml
&lt;/h1&gt;

&lt;p&gt;name: reviewer-regression&lt;br&gt;
on:&lt;br&gt;
  pull_request:&lt;br&gt;
    paths:&lt;br&gt;
      - "prompts/&lt;strong&gt;"&lt;br&gt;
      - "reviewer/&lt;/strong&gt;"&lt;br&gt;
  schedule:&lt;br&gt;
    - cron: "0 6 * * *"&lt;br&gt;
jobs:&lt;br&gt;
  golden-diffs:&lt;br&gt;
    runs-on: ubuntu-latest&lt;br&gt;
    steps:&lt;br&gt;
      - uses: actions/checkout@v4&lt;br&gt;
      - uses: actions/setup-python@v5&lt;br&gt;
        with:&lt;br&gt;
          python-version: "3.12"&lt;br&gt;
      - run: pip install -r requirements.txt&lt;br&gt;
      - run: pytest tests/test_golden_diffs.py -v&lt;br&gt;
        env:&lt;br&gt;
          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}&lt;/p&gt;

&lt;h2&gt;
  
  
  ⚖️ Trade-offs Worth Naming
&lt;/h2&gt;

&lt;p&gt;This isn't free, and it's not perfectly deterministic even with &lt;code&gt;temperature=0&lt;/code&gt; — different model versions can still shift slightly. A few honest trade-offs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Fixture rot&lt;/strong&gt;: real-world bug patterns evolve. Budget time to add new fixtures whenever a bug slips past the reviewer in production — that miss becomes your next test case.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost&lt;/strong&gt;: a suite of 50 fixtures run nightly against GPT-4-class models adds up. Consider a cheaper model for the regression suite if it's representative enough, and reserve the expensive model for prod.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;False confidence&lt;/strong&gt;: passing the golden-diff suite doesn't mean the reviewer is good, only that it hasn't regressed on the specific patterns you've thought to encode. Treat it as a floor, not a ceiling.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Structured output constraints&lt;/strong&gt;: forcing JSON schemas can sometimes make models slightly less thorough in free-form reasoning. Worth A/B testing structured vs. prose output against your fixture set before committing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these are reasons to skip this. They're reasons to scope it like any other test suite — start with the five bug patterns that have bitten you hardest in production, not fifty hypothetical ones.&lt;/p&gt;

&lt;h2&gt;
  
  
  🚀 Wrap Up
&lt;/h2&gt;

&lt;p&gt;If your AI reviewer has opinions about your code quality, it deserves the same scrutiny you'd apply to any other piece of logic sitting between a developer and a merge button. A golden-diff suite is cheap to start — a handful of real bugs pulled from your git history and a YAML file describing what "catching it" looks like.&lt;/p&gt;

&lt;p&gt;What's the bug pattern your AI reviewer has already let through that you haven't turned into a test case yet? That's usually fixture #1.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>codereview</category>
      <category>python</category>
    </item>
    <item>
      <title>NPU, DPU, QPU: Which One Actually Belongs in Your Stack</title>
      <dc:creator>Renato Silva</dc:creator>
      <pubDate>Mon, 21 Sep 2026 15:35:37 +0000</pubDate>
      <link>https://dev.to/renato_silva_71eef0fc385f/npu-dpu-qpu-which-one-actually-belongs-in-your-stack-3a07</link>
      <guid>https://dev.to/renato_silva_71eef0fc385f/npu-dpu-qpu-which-one-actually-belongs-in-your-stack-3a07</guid>
      <description>&lt;p&gt;Every hardware vendor keeps mailing you a new three-letter acronym like it's a subpoena. NPU. DPU. QPU. Somewhere a marketing team is very happy. Somewhere else, you're trying to figure out if any of this actually changes what you deploy on Tuesday.&lt;/p&gt;

&lt;p&gt;Short answer: two of these are already earning their keep in production today, and one is still mostly a research toy wearing a lab coat. Let's separate the hype from the workload.&lt;/p&gt;

&lt;h2&gt;
  
  
  🧠 NPU — Yes, If You're Doing On-Device Inference
&lt;/h2&gt;

&lt;p&gt;An NPU (Neural Processing Unit) is a chip optimized for the matrix-multiply-and-accumulate operations that dominate neural network inference. The pitch is real: NPUs deliver way better performance-per-watt than a CPU, and often better perf-per-watt than a GPU too, specifically for low-precision (int8/int4) inference workloads.&lt;/p&gt;

&lt;p&gt;Where it actually matters:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Mobile and edge apps&lt;/strong&gt; — on-device transcription, camera segmentation, keyword spotting. Apple's Neural Engine, Qualcomm's Hexagon NPU, and Intel's AI Boost are all built for exactly this.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Battery-constrained inference&lt;/strong&gt; — anything running continuously (wake-word detection, always-on vision) where GPU power draw would murder your battery life.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Privacy-sensitive workloads&lt;/strong&gt; — keeping inference local instead of round-tripping to a cloud GPU.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Where it doesn't matter:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Training.&lt;/strong&gt; NPUs are inference-first. If you're fine-tuning, you're still on GPU or TPU.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Server-side batch inference&lt;/strong&gt; where you can amortize GPU cost across many requests. A GPU with good batching will often out-throughput an NPU designed for single-stream, low-latency edge inference.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anything using fp32/fp64.&lt;/strong&gt; NPUs are built around quantized, low-precision math. If your model isn't quantization-friendly, you'll fight the toolchain more than you save on power.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Here's what actually shipping to an NPU looks like today, using ONNX Runtime with a hardware-specific execution provider:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
import onnxruntime as ort&lt;/p&gt;

&lt;h1&gt;
  
  
  Falls back gracefully, but explicitly requests NPU execution first
&lt;/h1&gt;

&lt;p&gt;providers = [&lt;br&gt;
    ("QNNExecutionProvider", {"backend_path": "QnnHtp.dll"}),  # Qualcomm NPU&lt;br&gt;
    "CPUExecutionProvider",&lt;br&gt;
]&lt;/p&gt;

&lt;p&gt;session = ort.InferenceSession("keyword_spotter_int8.onnx", providers=providers)&lt;/p&gt;

&lt;p&gt;output = session.run(&lt;br&gt;
    None,&lt;br&gt;
    {"audio_frame": frame.astype("int8")},&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;The honest trade-off: you're trading model flexibility and easy debugging for power efficiency and latency. If your model is a quantized int8 CNN or small transformer running continuously on a phone or embedded board, the NPU is the correct answer. If you're prototyping on fp32 weights and want fast iteration, stay on CPU/GPU until the model is locked.&lt;/p&gt;

&lt;h2&gt;
  
  
  🌐 DPU — Yes, If Your Network/Storage Path Is the Bottleneck
&lt;/h2&gt;

&lt;p&gt;DPUs (Data Processing Units — NVIDIA BlueField, AMD Pensando, Intel IPU) offload the stuff your CPU is bad at anyway: packet processing, TLS termination, storage virtualization, RDMA, and vSwitch logic. The core insight is that a modern 100/200Gbps NIC generates more interrupts and per-packet overhead than a general-purpose CPU core can chew through without falling over.&lt;/p&gt;

&lt;p&gt;Where it actually matters:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Multi-tenant infrastructure&lt;/strong&gt; — cloud providers offloading virtual switching so tenant VMs don't steal CPU cycles from the host for networking.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Storage-heavy clusters&lt;/strong&gt; — NVMe-oF, erasure coding, and compression offloaded so your storage nodes' CPUs are free for actual application logic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;High-throughput security&lt;/strong&gt; — inline TLS/IPsec at line rate without eating into compute you're paying for.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Where it doesn't matter:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A typical web service doing &amp;lt; 10Gbps.&lt;/strong&gt; Your kernel network stack and CPU are fine. A DPU here is solving a problem you don't have.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Small clusters without a dedicated infra team.&lt;/strong&gt; DPUs come with real operational overhead — you're now managing another programmable device with its own firmware, OS (often a stripped Linux running on ARM cores), and failure modes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A representative DPU workload is offloading flow classification with something like DPDK or P4, rather than handling it in kernel space:&lt;/p&gt;

&lt;p&gt;c&lt;br&gt;
// Simplified P4 match-action rule offloaded to DPU ASIC/FPGA pipeline&lt;br&gt;
table classify_flow {&lt;br&gt;
    key = {&lt;br&gt;
        hdr.ipv4.src_addr: exact;&lt;br&gt;
        hdr.ipv4.dst_addr: exact;&lt;br&gt;
        hdr.tcp.dst_port:  exact;&lt;br&gt;
    }&lt;br&gt;
    actions = {&lt;br&gt;
        forward_to_vm;&lt;br&gt;
        drop_flow;&lt;br&gt;
        mirror_to_ids;&lt;br&gt;
    }&lt;br&gt;
    size = 65536;&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;apply {&lt;br&gt;
    classify_flow.apply();&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;The trade-off here is capex and complexity versus CPU headroom. If your bottleneck is genuinely network/storage I/O stealing cycles from application logic, a DPU is a straightforward win. If you're not saturating a 25Gbps link, you're buying a Ferrari to sit in traffic.&lt;/p&gt;

&lt;h2&gt;
  
  
  ⚛️ QPU — Not Yet, and Probably Not Soon
&lt;/h2&gt;

&lt;p&gt;Here's where I'll be the buzzkill. QPUs are real, IBM, IonQ, and Rigetti will happily give you cloud access, and quantum algorithms like Shor's and Grover's are mathematically legitimate. But "belongs in your stack today" implies a production workload with a favorable cost/benefit versus classical hardware. That doesn't exist yet for essentially any commercial application.&lt;/p&gt;

&lt;p&gt;Current NISQ-era (Noisy Intermediate-Scale Quantum) hardware has:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Qubit counts in the hundreds, not the millions needed for meaningful error-corrected computation.&lt;/li&gt;
&lt;li&gt;Decoherence times measured in microseconds, meaning circuit depth is severely limited.&lt;/li&gt;
&lt;li&gt;No demonstrated quantum advantage on a problem anyone is actually paid to solve in production — optimization, chemistry simulation, and cryptanalysis demos are all still research-scale.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What you &lt;em&gt;can&lt;/em&gt; legitimately do today is experimentation and skill-building, which has real value if quantum ever matures on your timeline:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
from qiskit import QuantumCircuit&lt;br&gt;
from qiskit_aer import AerSimulator&lt;/p&gt;

&lt;p&gt;qc = QuantumCircuit(2, 2)&lt;br&gt;
qc.h(0)&lt;br&gt;
qc.cx(0, 1)&lt;br&gt;
qc.measure([0, 1], [0, 1])&lt;/p&gt;

&lt;p&gt;sim = AerSimulator()&lt;br&gt;
result = sim.run(qc, shots=1000).result()&lt;br&gt;
print(result.get_counts())&lt;/p&gt;

&lt;p&gt;Notice this ran on a simulator, not real quantum hardware — and for almost every "quantum experiment" blog post you'll see this year, that's the honest state of things. If you're doing quantum chemistry research or working at a place with a dedicated quantum team, real QPU access via cloud APIs (Qiskit Runtime, Braket) makes sense as R&amp;amp;D. For a normal product stack, a QPU line item is a research budget decision, not an engineering one.&lt;/p&gt;

&lt;h2&gt;
  
  
  🔧 The Actual Decision Framework
&lt;/h2&gt;

&lt;p&gt;Strip away the vendor decks and it's a simple filter:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Chip&lt;/th&gt;
&lt;th&gt;Real workload today&lt;/th&gt;
&lt;th&gt;Ask yourself&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;NPU&lt;/td&gt;
&lt;td&gt;Quantized inference at the edge&lt;/td&gt;
&lt;td&gt;Is this model quantized, latency-sensitive, and running on battery?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DPU&lt;/td&gt;
&lt;td&gt;Line-rate network/storage offload&lt;/td&gt;
&lt;td&gt;Is my CPU actually saturated by I/O, not application logic?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;QPU&lt;/td&gt;
&lt;td&gt;Research and algorithm exploration&lt;/td&gt;
&lt;td&gt;Am I doing this for a paper/PoC, or do I actually have a production problem it solves?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Most teams reading this don't need any of the three — a well-tuned GPU or even CPU inference path, a beefy NIC with kernel bypass, and zero quantum anything will outperform premature specialization. The chips earn their keep only when the workload characteristics (precision, throughput, or problem class) actually demand it.&lt;/p&gt;

&lt;p&gt;What's the acronym you've seen misapplied the hardest — teams reaching for specialized silicon before checking if the boring hardware was actually the bottleneck? Drop your war story below.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>hardware</category>
      <category>quantum</category>
      <category>networking</category>
    </item>
    <item>
      <title>SSH Tunnel Manager in Rust: CLI vs Native GUI Trade-offs</title>
      <dc:creator>Renato Silva</dc:creator>
      <pubDate>Thu, 17 Sep 2026 14:02:56 +0000</pubDate>
      <link>https://dev.to/renato_silva_71eef0fc385f/ssh-tunnel-manager-in-rust-cli-vs-native-gui-trade-offs-47e5</link>
      <guid>https://dev.to/renato_silva_71eef0fc385f/ssh-tunnel-manager-in-rust-cli-vs-native-gui-trade-offs-47e5</guid>
      <description>&lt;h2&gt;
  
  
  🔧 The Problem
&lt;/h2&gt;

&lt;p&gt;A few weeks ago a Swift-based macOS SSH tunnel manager started making the rounds here — a menu bar app that lets you spin up local/remote port forwards without touching a terminal. It's a genuinely nice pattern: SSH tunnels are one of those tools everyone reaches for constantly (jump boxes, database access, staging environments) but nobody wants to remember the flag syntax for.&lt;/p&gt;

&lt;p&gt;bash&lt;br&gt;
ssh -N -L 5432:db.internal:5432 -i ~/.ssh/jump_key jumpbox.example.com&lt;/p&gt;

&lt;p&gt;That command is fine until you have twelve of them across three environments, and you forget which one you killed last Tuesday.&lt;/p&gt;

&lt;p&gt;I wanted the same convenience but without being locked to macOS, so I built it twice in Rust: once as a CLI with a TOML-driven tunnel registry, and once as a native GUI using &lt;code&gt;tauri&lt;/code&gt;. This post is about what that comparison actually cost me — not in "which is better" terms, but in concrete trade-offs around distribution, process management, and platform integration.&lt;/p&gt;

&lt;h2&gt;
  
  
  🧱 The Core: Managing SSH as a Child Process
&lt;/h2&gt;

&lt;p&gt;Both versions share the same backend logic. Rust doesn't have a native SSH client library that's production-ready enough for arbitrary key/agent auth quirks, so both versions shell out to the system &lt;code&gt;ssh&lt;/code&gt; binary and manage it as a subprocess — same approach the Swift app uses under the hood, incidentally.&lt;/p&gt;

&lt;p&gt;rust&lt;br&gt;
use std::process::{Command, Child, Stdio};&lt;br&gt;
use std::collections::HashMap;&lt;/p&gt;

&lt;p&gt;pub struct Tunnel {&lt;br&gt;
    pub name: String,&lt;br&gt;
    pub local_port: u16,&lt;br&gt;
    pub remote_host: String,&lt;br&gt;
    pub remote_port: u16,&lt;br&gt;
    pub ssh_host: String,&lt;br&gt;
    process: Option,&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;impl Tunnel {&lt;br&gt;
    pub fn start(&amp;amp;mut self) -&amp;gt; std::io::Result&amp;lt;()&amp;gt; {&lt;br&gt;
        let forward = format!("{}:{}:{}", self.local_port, self.remote_host, self.remote_port);&lt;br&gt;
        let child = Command::new("ssh")&lt;br&gt;
            .args(["-N", "-L", &amp;amp;forward, &amp;amp;self.ssh_host])&lt;br&gt;
            .stdout(Stdio::null())&lt;br&gt;
            .stderr(Stdio::piped())&lt;br&gt;
            .spawn()?;&lt;br&gt;
        self.process = Some(child);&lt;br&gt;
        Ok(())&lt;br&gt;
    }&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;pub fn stop(&amp;amp;mut self) -&amp;gt; std::io::Result&amp;lt;()&amp;gt; {
    if let Some(mut child) = self.process.take() {
        child.kill()?;
        child.wait()?;
    }
    Ok(())
}

pub fn is_alive(&amp;amp;mut self) -&amp;gt; bool {
    match &amp;amp;mut self.process {
        Some(child) =&amp;gt; matches!(child.try_wait(), Ok(None)),
        None =&amp;gt; false,
    }
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;}&lt;/p&gt;

&lt;p&gt;pub type TunnelRegistry = HashMap;&lt;/p&gt;

&lt;p&gt;This part was identical effort in both builds. The divergence starts the moment you decide how a human is supposed to interact with it.&lt;/p&gt;

&lt;h2&gt;
  
  
  🖥️ The CLI Path
&lt;/h2&gt;

&lt;p&gt;The CLI version uses &lt;code&gt;clap&lt;/code&gt; for argument parsing and a TOML config file for tunnel definitions:&lt;/p&gt;

&lt;p&gt;toml&lt;br&gt;
[tunnels.staging_db]&lt;br&gt;
ssh_host = "jumpbox.example.com"&lt;br&gt;
local_port = 5432&lt;br&gt;
remote_host = "db.internal"&lt;br&gt;
remote_port = 5432&lt;/p&gt;

&lt;p&gt;[tunnels.staging_redis]&lt;br&gt;
ssh_host = "jumpbox.example.com"&lt;br&gt;
local_port = 6379&lt;br&gt;
remote_host = "cache.internal"&lt;br&gt;
remote_port = 6379&lt;/p&gt;

&lt;p&gt;bash&lt;br&gt;
tunlctl up staging_db&lt;br&gt;
tunlctl status&lt;br&gt;
tunlctl down staging_db&lt;br&gt;
tunlctl up --all&lt;/p&gt;

&lt;p&gt;What this bought me:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Distribution is trivial.&lt;/strong&gt; &lt;code&gt;cargo install tunlctl&lt;/code&gt;, or a single static binary via &lt;code&gt;cross&lt;/code&gt; for Linux/macOS/Windows. No code signing, no notarization, no App Store review queue.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scriptability for free.&lt;/strong&gt; People immediately asked for &lt;code&gt;tunlctl up staging --json&lt;/code&gt; to pipe into other tooling. A GUI can't casually do that.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Persistence via existing tools.&lt;/strong&gt; Backgrounding and keeping tunnels alive after the CLI exits meant either double-forking or leaning on &lt;code&gt;systemd&lt;/code&gt;/&lt;code&gt;launchd&lt;/code&gt; unit files I generate on demand. That's more plumbing than a GUI app gets automatically just by staying resident.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What it cost me:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;No visual "this tunnel died" signal. You either poll &lt;code&gt;tunlctl status&lt;/code&gt; or you don't find out until your app starts throwing connection refused errors.&lt;/li&gt;
&lt;li&gt;No system tray indicator, no click-to-toggle. For a tool used dozens of times a day, keystrokes and cognitive load add up in a way a menu bar icon avoids entirely.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  🪟 The GUI Path
&lt;/h2&gt;

&lt;p&gt;The Tauri version wraps the same &lt;code&gt;Tunnel&lt;/code&gt; struct behind commands, and Rust plus a lightweight HTML/CSS frontend for the tray menu and window:&lt;/p&gt;

&lt;p&gt;rust&lt;/p&gt;

&lt;h1&gt;
  
  
  [tauri::command]
&lt;/h1&gt;

&lt;p&gt;fn toggle_tunnel(name: String, state: tauri::State) -&amp;gt; Result {&lt;br&gt;
    let mut registry = state.registry.lock().map_err(|e| e.to_string())?;&lt;br&gt;
    let tunnel = registry.get_mut(&amp;amp;name).ok_or("tunnel not found")?;&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;if tunnel.is_alive() {
    tunnel.stop().map_err(|e| e.to_string())?;
    Ok(false)
} else {
    tunnel.start().map_err(|e| e.to_string())?;
    Ok(true)
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;}&lt;/p&gt;

&lt;p&gt;What this bought me:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Ambient status.&lt;/strong&gt; A tray icon that turns green/red per tunnel, updated by a background poller, is genuinely better UX than typing a status command. This is the single biggest reason the Swift app resonated with people.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Process lifetime is free.&lt;/strong&gt; The app itself is the long-running process; tunnels live and die with it, no daemon management needed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Native feel where it matters&lt;/strong&gt;, e.g. macOS keychain integration for SSH passphrases instead of relying on &lt;code&gt;ssh-agent&lt;/code&gt; being configured correctly.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What it cost me:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Distribution got real overhead.&lt;/strong&gt; Unsigned builds trigger Gatekeeper warnings on macOS and SmartScreen on Windows. Signing a macOS app means an Apple Developer account, notarization via &lt;code&gt;xcrun notarytool&lt;/code&gt;, and a CI step that didn't exist in the CLI world at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bundle size and build complexity jumped.&lt;/strong&gt; Tauri is far lighter than Electron, but you're still shipping a webview-based frontend, dealing with &lt;code&gt;tauri.conf.json&lt;/code&gt;, and debugging IPC serialization between Rust and JS for things that were plain function calls in the CLI.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cross-platform tray behavior is not actually uniform.&lt;/strong&gt; Linux tray icon support depends on the desktop environment having a functioning &lt;code&gt;StatusNotifierItem&lt;/code&gt; implementation. It's a smaller thing than you'd hope for a "cross-platform" story.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  ⚖️ The Actual Trade-off
&lt;/h2&gt;

&lt;p&gt;If I had to compress this into one sentence: &lt;strong&gt;the CLI wins on time-to-ship and the GUI wins on time-to-use.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For a tool I'm building for myself and a small team that already lives in a terminal, &lt;code&gt;tunlctl&lt;/code&gt; shipped in an afternoon and required zero ongoing signing infrastructure. For something meant to be handed to less terminal-comfortable teammates, or anyone who wants "click icon, tunnel appears," the GUI's ambient status indicator justified every bit of the notarization pain.&lt;/p&gt;

&lt;p&gt;My actual answer ended up being both, sharing the same core crate — the CLI is the source of truth and the automation-friendly interface, and the GUI is a thin, optional shell around it for people who want the tray icon. That's not a cop-out; it's the same reason &lt;code&gt;docker&lt;/code&gt; has both a CLI and Docker Desktop.&lt;/p&gt;

&lt;h2&gt;
  
  
  🙋 Over to You
&lt;/h2&gt;

&lt;p&gt;If you've built or maintained an internal dev tool that started as a CLI and grew a GUI (or vice versa) — where did the complexity actually show up for you? Was it distribution, state management, or something else entirely?&lt;/p&gt;

&lt;p&gt;If you want to poke at the code, the tunnel-management core described here is intentionally backend-agnostic — happy to expand on the daemon/&lt;code&gt;launchd&lt;/code&gt; integration in a follow-up if there's interest.&lt;/p&gt;

</description>
      <category>rust</category>
      <category>cli</category>
      <category>ssh</category>
      <category>macos</category>
    </item>
    <item>
      <title>I Profiled 1M Goroutines So You Don't Have to</title>
      <dc:creator>Renato Silva</dc:creator>
      <pubDate>Mon, 14 Sep 2026 15:35:12 +0000</pubDate>
      <link>https://dev.to/renato_silva_71eef0fc385f/i-profiled-1m-goroutines-so-you-dont-have-to-40ba</link>
      <guid>https://dev.to/renato_silva_71eef0fc385f/i-profiled-1m-goroutines-so-you-dont-have-to-40ba</guid>
      <description>&lt;h2&gt;
  
  
  🔧 The Problem
&lt;/h2&gt;

&lt;p&gt;Every Go dev has heard "goroutines are cheap." True, relative to OS threads. But cheap isn't free, and at scale the bill comes due in ways that don't show up until you're staring at a memory graph wondering why your service that spawns a goroutine per request just OOM'd at 800k concurrent connections.&lt;/p&gt;

&lt;p&gt;I wanted actual numbers, not vibes. So I spun up 1,000,000 goroutines in a few different shapes, profiled them with &lt;code&gt;go tool pprof&lt;/code&gt; and &lt;code&gt;runtime.MemStats&lt;/code&gt;, and tried to separate two costs that get lumped together as "goroutine overhead":&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Stack memory&lt;/strong&gt; — the growable per-goroutine stack&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scheduler bookkeeping&lt;/strong&gt; — the &lt;code&gt;g&lt;/code&gt; struct, &lt;code&gt;sudog&lt;/code&gt;s, run queue entries, and GC scanning overhead&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;These are different cost centers with different scaling behavior, and conflating them leads to bad capacity planning.&lt;/p&gt;

&lt;h2&gt;
  
  
  🧪 The Setup
&lt;/h2&gt;

&lt;p&gt;Here's the baseline harness — goroutines that block forever on a channel, so they stay alive and scheduled but don't do work:&lt;/p&gt;

&lt;p&gt;go&lt;br&gt;
package main&lt;/p&gt;

&lt;p&gt;import (&lt;br&gt;
    "fmt"&lt;br&gt;
    "os"&lt;br&gt;
    "runtime"&lt;br&gt;
    "runtime/pprof"&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;func main() {&lt;br&gt;
    const n = 1_000_000&lt;br&gt;
    block := make(chan struct{})&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;var readyWg = make(chan struct{}, n)
for i := 0; i &amp;lt; n; i++ {
    go func() {
        readyWg &amp;lt;- struct{}{}
        &amp;lt;-block
    }()
}
for i := 0; i &amp;lt; n; i++ {
    &amp;lt;-readyWg
}

var m runtime.MemStats
runtime.ReadMemStats(&amp;amp;m)
fmt.Printf("HeapAlloc: %d MB\n", m.HeapAlloc/1024/1024)
fmt.Printf("StackInuse: %d MB\n", m.StackInuse/1024/1024)
fmt.Printf("NumGoroutine: %d\n", runtime.NumGoroutine())

f, _ := os.Create("heap.pprof")
pprof.WriteHeapProfile(f)
f.Close()

&amp;lt;-make(chan struct{}) // hang so we can attach pprof live too
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;}&lt;/p&gt;

&lt;p&gt;Run with &lt;code&gt;GODEBUG=madvdontneed=1&lt;/code&gt; off (default) and &lt;code&gt;GOGC=400&lt;/code&gt; to reduce GC noise while we measure steady state, then:&lt;/p&gt;

&lt;p&gt;bash&lt;br&gt;
go run main.go &amp;amp;&lt;br&gt;
go tool pprof -top -alloc_space heap.pprof&lt;br&gt;
go tool pprof &lt;a href="http://localhost:6060/debug/pprof/goroutine" rel="noopener noreferrer"&gt;http://localhost:6060/debug/pprof/goroutine&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;(Yes, I added the standard &lt;code&gt;net/http/pprof&lt;/code&gt; import for the live endpoint — don't forget the underscore import or you'll wonder why &lt;code&gt;/debug/pprof/&lt;/code&gt; 404s.)&lt;/p&gt;

&lt;h2&gt;
  
  
  📊 What &lt;code&gt;MemStats&lt;/code&gt; Actually Shows
&lt;/h2&gt;

&lt;p&gt;With 1M blocked goroutines doing nothing but sitting on a channel receive:&lt;/p&gt;

&lt;p&gt;HeapAlloc: 366 MB&lt;br&gt;
StackInuse: 2147 MB&lt;br&gt;
NumGoroutine: 1000000&lt;/p&gt;

&lt;p&gt;That &lt;code&gt;StackInuse&lt;/code&gt; number is the headline: &lt;strong&gt;~2.1KB per goroutine&lt;/strong&gt;, even though the default initial stack is 2KB and these goroutines do almost nothing. That checks out — Go's runtime allocates the initial stack up front, and a goroutine parked on a channel receive still holds onto that stack because the scheduler needs somewhere to resume execution.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;HeapAlloc&lt;/code&gt; — 366MB — is the part people forget about. That's not stack, that's the &lt;code&gt;runtime.g&lt;/code&gt; structs, the &lt;code&gt;sudog&lt;/code&gt; entries used for channel waiters, and assorted bookkeeping. Divide it out: &lt;strong&gt;~384 bytes per goroutine&lt;/strong&gt; in pure scheduler/heap overhead, separate from the stack.&lt;/p&gt;

&lt;p&gt;So per goroutine, roughly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;2KB — initial stack (grows if the function needs more)&lt;/li&gt;
&lt;li&gt;~200-450 bytes — &lt;code&gt;g&lt;/code&gt; struct + scheduler metadata&lt;/li&gt;
&lt;li&gt;~48-100 bytes — &lt;code&gt;sudog&lt;/code&gt; if blocked on a channel/mutex/select&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That second bucket is the one that surprises people, because it scales with goroutine &lt;em&gt;count&lt;/em&gt;, not with what the goroutine is &lt;em&gt;doing&lt;/em&gt;. You can't shrink it by simplifying your goroutine's logic. It's the tax for existing.&lt;/p&gt;

&lt;h2&gt;
  
  
  📈 Where pprof Actually Helps
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;go tool pprof -alloc_space&lt;/code&gt; on the heap profile shows the scheduler overhead concretely:&lt;/p&gt;

&lt;p&gt;(pprof) top10&lt;br&gt;
Flat  Flat%   Sum%   Cum   Cum%   Name&lt;br&gt;
312MB 85.2%  85.2%  312MB 85.2%  runtime.malg&lt;br&gt;
38MB  10.4%  95.6%   38MB 10.4%  runtime.acquireSudog&lt;br&gt;
9MB    2.5%  98.1%    9MB  2.5%  runtime.newproc.func1&lt;/p&gt;

&lt;p&gt;&lt;code&gt;runtime.malg&lt;/code&gt; — goroutine struct allocation — dominates. This is the smoking gun for "scheduler overhead," separate entirely from the stack memory pprof doesn't even show you here (stack isn't heap-allocated in the traditional sense pprof tracks by default; you have to cross-reference &lt;code&gt;StackInuse&lt;/code&gt; from &lt;code&gt;MemStats&lt;/code&gt; to see it).&lt;/p&gt;

&lt;p&gt;That's the key methodological point: &lt;strong&gt;pprof's heap profile and &lt;code&gt;MemStats&lt;/code&gt;' &lt;code&gt;StackInuse&lt;/code&gt; are answering different questions&lt;/strong&gt;, and if you only look at one you'll misdiagnose the bottleneck. I've seen postmortems blame "goroutine leaks" on stack growth when the actual driver was thousands of goroutines each holding a &lt;code&gt;sudog&lt;/code&gt; because they were blocked on a busy mutex.&lt;/p&gt;

&lt;h2&gt;
  
  
  🐘 What Happens When Goroutines Actually Do Something
&lt;/h2&gt;

&lt;p&gt;Blocked-on-channel goroutines are the cheap case. Let's make it more realistic — recursive work that grows the stack:&lt;/p&gt;

&lt;p&gt;go&lt;br&gt;
func recurse(n int, block &amp;lt;-chan struct{}) {&lt;br&gt;
    if n == 0 {&lt;br&gt;
        &amp;lt;-block&lt;br&gt;
        return&lt;br&gt;
    }&lt;br&gt;
    var buf [64]byte&lt;br&gt;
    _ = buf&lt;br&gt;
    recurse(n-1, block)&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;Spawning 1M of these with &lt;code&gt;recurse(50, block)&lt;/code&gt; pushes &lt;code&gt;StackInuse&lt;/code&gt; from ~2.1GB to ~4.6GB — stacks grew from the 2KB default to accommodate the recursion depth, and &lt;strong&gt;Go doesn't shrink them back down eagerly&lt;/strong&gt; unless a GC cycle happens to trigger &lt;code&gt;shrinkstack&lt;/code&gt; on that goroutine (checked during stack scanning, roughly every other GC if the stack is &amp;lt;1/4 utilized). Under &lt;code&gt;GOGC=400&lt;/code&gt; that shrink check happens rarely, so stacks stay bloated far longer than you'd expect.&lt;/p&gt;

&lt;p&gt;This is the practical takeaway: if your workload has bursty deep call stacks (recursive JSON parsing, deep middleware chains, reflection-heavy code), the stack growth sticks around as a memory cost long after the burst ends, independent of the flat per-goroutine scheduler tax. Lowering &lt;code&gt;GOGC&lt;/code&gt; or calling &lt;code&gt;debug.FreeOSMemory()&lt;/code&gt; after a burst can reclaim it, at the cost of more frequent GC cycles overall — a real trade-off, not a free win.&lt;/p&gt;

&lt;h2&gt;
  
  
  🧮 The Cost Model, Summarized
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Cost center&lt;/th&gt;
&lt;th&gt;Scales with&lt;/th&gt;
&lt;th&gt;Reclaimed how&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Initial stack (2KB)&lt;/td&gt;
&lt;td&gt;goroutine count&lt;/td&gt;
&lt;td&gt;goroutine exit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stack growth&lt;/td&gt;
&lt;td&gt;max call depth reached&lt;/td&gt;
&lt;td&gt;GC-triggered &lt;code&gt;shrinkstack&lt;/code&gt;, rare under high GOGC&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;g&lt;/code&gt; struct / scheduler&lt;/td&gt;
&lt;td&gt;goroutine count&lt;/td&gt;
&lt;td&gt;goroutine exit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;sudog&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;count of &lt;em&gt;blocked&lt;/em&gt; goroutines&lt;/td&gt;
&lt;td&gt;unblock or exit&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If you're doing capacity planning for a goroutine-per-connection server, the number that matters isn't "goroutines are ~2KB each" — it's closer to 2.5-3KB steady state for idle blocked goroutines, and potentially multiples of that if your request handlers recurse or allocate large local buffers before blocking.&lt;/p&gt;

&lt;h2&gt;
  
  
  🙋 Over to You
&lt;/h2&gt;

&lt;p&gt;Have you actually hit a goroutine-count wall in production, or is 1M goroutines mostly an academic exercise for your workloads? I'd genuinely like to know where the real-world breakpoint sits — 100k? 10M? Drop your numbers (and your &lt;code&gt;GOMAXPROCS&lt;/code&gt;) in the comments.&lt;/p&gt;

&lt;p&gt;If you want to reproduce this, the full harness plus the recursive-stack variant is about 80 lines — worth running locally with your own &lt;code&gt;GOGC&lt;/code&gt; and &lt;code&gt;GOMAXPROCS&lt;/code&gt; settings, because scheduler contention at 1M goroutines on a 4-core box behaves noticeably differently than on 64 cores, and that's a whole separate profiling story.&lt;/p&gt;

</description>
      <category>go</category>
      <category>performance</category>
      <category>pprof</category>
      <category>concurrency</category>
    </item>
    <item>
      <title>Git Worktrees: The Missing Piece for Parallel AI Agents</title>
      <dc:creator>Renato Silva</dc:creator>
      <pubDate>Mon, 31 Aug 2026 16:38:50 +0000</pubDate>
      <link>https://dev.to/renato_silva_71eef0fc385f/git-worktrees-the-missing-piece-for-parallel-ai-agents-10lm</link>
      <guid>https://dev.to/renato_silva_71eef0fc385f/git-worktrees-the-missing-piece-for-parallel-ai-agents-10lm</guid>
      <description>&lt;h2&gt;
  
  
  🔧 The Problem
&lt;/h2&gt;

&lt;p&gt;If you're running more than one AI coding agent at a time — Claude Code in one terminal, Aider in another, maybe a Cursor background agent chewing on a refactor — you've probably hit the same wall: they all want to work on the same repo, but they can't share a working directory without stepping on each other.&lt;/p&gt;

&lt;p&gt;The usual workarounds are all bad:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;git stash&lt;/code&gt; juggling&lt;/strong&gt; — you stash, switch branches, let the agent work, unstash, repeat. Fine for one agent. A nightmare for three running concurrently, because stash is a single shared stack and agents don't know how to negotiate over it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cloning the repo N times&lt;/strong&gt; — works, but now you've got N full copies of &lt;code&gt;.git&lt;/code&gt;, N sets of dependencies to install, and N places for config drift to sneak in. On a large monorepo this is also just slow.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One giant branch with agents committing to subdirectories&lt;/strong&gt; — merge conflicts waiting to happen, and agents lose the ability to see a clean diff of just their own work.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What you actually want is N independent working directories, backed by &lt;em&gt;one&lt;/em&gt; &lt;code&gt;.git&lt;/code&gt;, so history, remotes, and object storage stay unified while the checked-out files stay isolated. That's exactly what &lt;code&gt;git worktree&lt;/code&gt; gives you, and it's been sitting in Git core since version 2.5 (2015), mostly ignored until "run four agents at once" became a normal Tuesday.&lt;/p&gt;

&lt;h2&gt;
  
  
  🌳 Worktrees in Practice
&lt;/h2&gt;

&lt;p&gt;The core workflow is boring in the best way:&lt;/p&gt;

&lt;p&gt;bash&lt;/p&gt;

&lt;h1&gt;
  
  
  from your main checkout
&lt;/h1&gt;

&lt;p&gt;git worktree add ../myapp-agent-a feature/agent-a&lt;br&gt;
git worktree add ../myapp-agent-b feature/agent-b&lt;br&gt;
git worktree add ../myapp-agent-c fix/flaky-test&lt;/p&gt;

&lt;h1&gt;
  
  
  see what's active
&lt;/h1&gt;

&lt;p&gt;git worktree list&lt;/p&gt;

&lt;h1&gt;
  
  
  /home/dev/myapp            abcd123 [main]
&lt;/h1&gt;

&lt;h1&gt;
  
  
  /home/dev/myapp-agent-a    ef01234 [feature/agent-a]
&lt;/h1&gt;

&lt;h1&gt;
  
  
  /home/dev/myapp-agent-b    5678aaa [feature/agent-b]
&lt;/h1&gt;

&lt;h1&gt;
  
  
  /home/dev/myapp-agent-c    9911bbb [fix/flaky-test]
&lt;/h1&gt;

&lt;p&gt;Each directory is a real, complete checkout — you can &lt;code&gt;cd&lt;/code&gt; into it, run tests, open it in an editor, point an agent at it — and none of them affect each other's index or working tree. Behind the scenes they all share &lt;code&gt;.git/objects&lt;/code&gt;, so you're not duplicating blobs, and any commit made in one worktree is immediately visible to &lt;code&gt;git log&lt;/code&gt; in the others (once you fetch/checkout).&lt;/p&gt;

&lt;p&gt;Cleaning up is just as direct:&lt;/p&gt;

&lt;p&gt;bash&lt;br&gt;
git worktree remove ../myapp-agent-c&lt;/p&gt;

&lt;h1&gt;
  
  
  or, if the agent left the directory dirty and you don't care:
&lt;/h1&gt;

&lt;p&gt;git worktree remove --force ../myapp-agent-c&lt;/p&gt;

&lt;p&gt;For an agent-driven workflow, I wrap this in a small script so I'm not hand-typing branch names every time I spin one up:&lt;/p&gt;

&lt;p&gt;bash&lt;/p&gt;

&lt;h1&gt;
  
  
  !/usr/bin/env bash
&lt;/h1&gt;

&lt;h1&gt;
  
  
  spawn-agent.sh 
&lt;/h1&gt;

&lt;p&gt;set -euo pipefail&lt;/p&gt;

&lt;p&gt;task="$1"&lt;br&gt;
branch="agent/${task}"&lt;br&gt;
worktree_path="../$(basename "$(pwd)")-${task}"&lt;/p&gt;

&lt;p&gt;git worktree add -b "$branch" "$worktree_path" main&lt;br&gt;
cd "$worktree_path"&lt;/p&gt;

&lt;h1&gt;
  
  
  per-worktree setup so agents don't fight over node_modules etc.
&lt;/h1&gt;

&lt;p&gt;cp ../.env.example .env&lt;br&gt;
npm install --prefer-offline&lt;/p&gt;

&lt;p&gt;echo "Worktree ready at $worktree_path on branch $branch"&lt;/p&gt;

&lt;p&gt;Now "give the agent a sandbox" is a single command, and tearing it down after review/merge is another single command. No stash stack, no second clone, no confusion about which branch is checked out where.&lt;/p&gt;

&lt;h2&gt;
  
  
  ⚠️ The Gotchas Nobody Mentions
&lt;/h2&gt;

&lt;p&gt;Worktrees are not a free lunch, and the failure modes are exactly the kind of thing that eats an afternoon if you don't know to look for them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Shared package manager caches can lie to you.&lt;/strong&gt; &lt;code&gt;node_modules&lt;/code&gt;, &lt;code&gt;.venv&lt;/code&gt;, and build caches are &lt;em&gt;not&lt;/em&gt; shared between worktrees by default — each one needs its own install. If your agents are installing dependencies in parallel across worktrees pointed at the same global npm/pip cache, you can get lock contention or, worse, a half-written cache entry that silently corrupts a build in a different worktree. Pin a per-worktree cache directory if you're running installs concurrently:&lt;/p&gt;

&lt;p&gt;bash&lt;br&gt;
npm install --cache "$(pwd)/.npm-cache"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;IDE indexing goes haywire.&lt;/strong&gt; VS Code, JetBrains IDEs, and language servers built with a single-repo assumption will happily index every worktree directory you open as if it's an unrelated project — which is technically correct but means you're running 4x the TypeScript server memory, 4x the file watchers, and sometimes 4x the "go to definition" confusion if symlinks or path aliases assume a fixed repo root. If you're not actively reading code in a worktree, don't leave it open in the IDE — close the window when the agent is just running headless.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Branches can't be checked out twice.&lt;/strong&gt; This one bites people immediately: Git will refuse to let two worktrees point at the same branch.&lt;/p&gt;

&lt;p&gt;fatal: 'feature/agent-a' is already checked out at '/home/dev/myapp-agent-a'&lt;/p&gt;

&lt;p&gt;This is a feature, not a bug — it's the mechanism that prevents two agents from independently committing to the same branch and creating divergent history in two places at once. But it does mean your orchestration script needs a real branch-per-agent naming scheme, not "reuse main for everything."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Detached HEAD surprises.&lt;/strong&gt; If an agent (or you) checks out a commit instead of a branch, you get a detached HEAD in that worktree — harmless, but if the agent then commits and you forget to create a branch before removing the worktree, &lt;code&gt;git worktree remove&lt;/code&gt; will happily let you lose those commits to garbage collection. Always &lt;code&gt;git branch tmp-recovery&lt;/code&gt; before tearing down a detached-HEAD worktree you're unsure about.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Submodules and &lt;code&gt;.git&lt;/code&gt; hooks need extra care.&lt;/strong&gt; Hooks live in the shared &lt;code&gt;.git&lt;/code&gt; directory by default (or &lt;code&gt;.git/worktrees/&amp;lt;name&amp;gt;&lt;/code&gt; for some internals), so a hook that assumes &lt;code&gt;$(pwd)&lt;/code&gt; is the repo root can misbehave across worktrees. If you use submodules, &lt;code&gt;git worktree add&lt;/code&gt; doesn't initialize them for you — add &lt;code&gt;--recurse-submodules&lt;/code&gt; or run &lt;code&gt;git submodule update --init&lt;/code&gt; explicitly per worktree.&lt;/p&gt;

&lt;h2&gt;
  
  
  🚀 Putting It Together
&lt;/h2&gt;

&lt;p&gt;The pattern that's worked well for me running 3-4 agents concurrently:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;One "orchestrator" checkout (your normal working directory) that never runs an agent directly — it's just for review and merging.&lt;/li&gt;
&lt;li&gt;One worktree per active task, named after the task, not the agent.&lt;/li&gt;
&lt;li&gt;A teardown step that force-removes the worktree &lt;em&gt;and&lt;/em&gt; deletes the branch once merged, so &lt;code&gt;git worktree list&lt;/code&gt; doesn't slowly fill up with zombie sandboxes.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;bash&lt;br&gt;
git worktree remove --force ../myapp-agent-a&lt;br&gt;
git branch -d agent/task-a   # or -D if the agent's commits got squashed on merge&lt;/p&gt;

&lt;p&gt;It's not a glamorous feature — worktrees have existed for a decade specifically for things like hotfix-while-mid-feature workflows — but it maps almost perfectly onto "isolated sandbox per autonomous process" once you swap the human for an agent. The shared object store keeps disk usage sane, and the isolated working trees keep agents from corrupting each other's in-progress edits.&lt;/p&gt;

&lt;p&gt;How are you isolating your parallel agents right now — worktrees, containers, or something else entirely? And if you've hit a worktree gotcha that isn't on this list, I'd genuinely like to hear about it in the comments.&lt;/p&gt;

</description>
      <category>git</category>
      <category>ai</category>
      <category>productivity</category>
      <category>cli</category>
    </item>
    <item>
      <title>SSE vs WebSockets vs Polling: Real-Time Sync From the Backend</title>
      <dc:creator>Renato Silva</dc:creator>
      <pubDate>Thu, 27 Aug 2026 19:23:10 +0000</pubDate>
      <link>https://dev.to/renato_silva_71eef0fc385f/sse-vs-websockets-vs-polling-real-time-sync-from-the-backend-4hgc</link>
      <guid>https://dev.to/renato_silva_71eef0fc385f/sse-vs-websockets-vs-polling-real-time-sync-from-the-backend-4hgc</guid>
      <description>&lt;p&gt;There's been a nice trick going around: use &lt;code&gt;BroadcastChannel&lt;/code&gt; in the browser to sync state across tabs without a server round-trip. It's elegant, but it only solves half the problem — it syncs tabs on &lt;em&gt;one&lt;/em&gt; device. The moment you have two different users, or one user on a phone and a laptop, you need the server to be the source of truth and push updates out.&lt;/p&gt;

&lt;p&gt;So let's flip it: how do you actually push consistent state to N connected clients from a Node API, and which transport should you reach for?&lt;/p&gt;

&lt;h2&gt;
  
  
  🔧 The Problem
&lt;/h2&gt;

&lt;p&gt;Say you're building something boring and real: a shared cart, a live dashboard, a "someone else is editing this" indicator. Multiple clients need to see the same state change at roughly the same time, without everyone hammering &lt;code&gt;GET /state&lt;/code&gt; every 500ms.&lt;/p&gt;

&lt;p&gt;You've got three realistic options:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Polling&lt;/strong&gt; — client asks, server answers, repeat&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SSE (Server-Sent Events)&lt;/strong&gt; — server pushes a one-way stream over plain HTTP&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;WebSockets&lt;/strong&gt; — full duplex, server and client both push&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Each one has a different cost model, and picking the "cool" one (WebSockets) is often the wrong call.&lt;/p&gt;

&lt;h2&gt;
  
  
  🐢 Polling: the boring baseline
&lt;/h2&gt;

&lt;p&gt;Polling gets a bad reputation it doesn't fully deserve. It's stateless, trivially horizontally scalable, works through every proxy and CDN ever built, and requires zero special infrastructure.&lt;/p&gt;

&lt;p&gt;javascript&lt;br&gt;
// client&lt;br&gt;
setInterval(async () =&amp;gt; {&lt;br&gt;
  const res = await fetch('/api/state');&lt;br&gt;
  const state = await res.json();&lt;br&gt;
  renderState(state);&lt;br&gt;
}, 2000);&lt;/p&gt;

&lt;p&gt;The honest trade-off: latency is bounded by your interval, and cost scales linearly with (clients × interval). Ten thousand clients polling every 2 seconds is 5,000 requests/sec hitting your server &lt;em&gt;even when nothing changed&lt;/em&gt;. Fine for a demo, painful at scale, and it never actually feels "real-time" — there's always a visible lag.&lt;/p&gt;

&lt;h2&gt;
  
  
  📡 SSE: push, but only one way
&lt;/h2&gt;

&lt;p&gt;SSE is the underrated option. It's just an HTTP response that never closes, with a text protocol on top. No new protocol, no special client library, works over regular HTTP/1.1 and HTTP/2, and reconnects automatically via &lt;code&gt;EventSource&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Here's a minimal Node/Express version that keeps a registry of connected clients and broadcasts state changes:&lt;/p&gt;

&lt;p&gt;javascript&lt;br&gt;
import express from 'express';&lt;br&gt;
const app = express();&lt;/p&gt;

&lt;p&gt;let state = { count: 0 };&lt;br&gt;
const clients = new Set();&lt;/p&gt;

&lt;p&gt;app.get('/events', (req, res) =&amp;gt; {&lt;br&gt;
  res.set({&lt;br&gt;
    'Content-Type': 'text/event-stream',&lt;br&gt;
    'Cache-Control': 'no-cache',&lt;br&gt;
    Connection: 'keep-alive',&lt;br&gt;
  });&lt;br&gt;
  res.flushHeaders();&lt;/p&gt;

&lt;p&gt;// send current state immediately so late joiners aren't out of sync&lt;br&gt;
  res.write(&lt;code&gt;data: ${JSON.stringify(state)}\n\n&lt;/code&gt;);&lt;/p&gt;

&lt;p&gt;clients.add(res);&lt;br&gt;
  req.on('close', () =&amp;gt; clients.delete(res));&lt;br&gt;
});&lt;/p&gt;

&lt;p&gt;function broadcast(newState) {&lt;br&gt;
  state = newState;&lt;br&gt;
  const payload = &lt;code&gt;data: ${JSON.stringify(state)}\n\n&lt;/code&gt;;&lt;br&gt;
  for (const res of clients) res.write(payload);&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;app.post('/increment', express.json(), (req, res) =&amp;gt; {&lt;br&gt;
  broadcast({ count: state.count + 1 });&lt;br&gt;
  res.sendStatus(204);&lt;br&gt;
});&lt;/p&gt;

&lt;p&gt;app.listen(3000);&lt;/p&gt;

&lt;p&gt;javascript&lt;br&gt;
// client&lt;br&gt;
const source = new EventSource('/events');&lt;br&gt;
source.onmessage = (e) =&amp;gt; renderState(JSON.parse(e.data));&lt;/p&gt;

&lt;p&gt;That's the whole system. No socket library, no handshake upgrade dance, no ping/pong heartbeat logic to babysit — the browser handles reconnects for you.&lt;/p&gt;

&lt;p&gt;The catch: SSE is one-directional. Clients still need a normal &lt;code&gt;POST&lt;/code&gt;/&lt;code&gt;fetch&lt;/code&gt; to send actions back. For a lot of real apps (dashboards, notifications, live scores, cart sync) that's not a limitation, it's a feature — you get a clean separation between "write path" (REST) and "read/subscribe path" (SSE).&lt;/p&gt;

&lt;p&gt;Also worth knowing: browsers cap concurrent &lt;code&gt;EventSource&lt;/code&gt; connections per origin (6 over HTTP/1.1), and some corporate proxies buffer streaming responses, which can delay delivery. HTTP/2 mostly fixes the connection-limit problem since it multiplexes over one TCP connection.&lt;/p&gt;

&lt;h2&gt;
  
  
  🔌 WebSockets: when you actually need two-way
&lt;/h2&gt;

&lt;p&gt;WebSockets are the right tool when the client needs to push frequently too — collaborative editing, multiplayer cursors, chat, game state. Otherwise they're often overkill: you now own a stateful, bidirectional connection with your own reconnect logic, your own heartbeat, and your own message framing.&lt;/p&gt;

&lt;p&gt;javascript&lt;br&gt;
import { WebSocketServer } from 'ws';&lt;br&gt;
const wss = new WebSocketServer({ port: 8080 });&lt;/p&gt;

&lt;p&gt;let state = { count: 0 };&lt;/p&gt;

&lt;p&gt;wss.on('connection', (ws) =&amp;gt; {&lt;br&gt;
  ws.send(JSON.stringify(state));&lt;/p&gt;

&lt;p&gt;ws.on('message', (raw) =&amp;gt; {&lt;br&gt;
    const msg = JSON.parse(raw);&lt;br&gt;
    if (msg.type === 'increment') {&lt;br&gt;
      state = { count: state.count + 1 };&lt;br&gt;
      const payload = JSON.stringify(state);&lt;br&gt;
      for (const client of wss.clients) {&lt;br&gt;
        if (client.readyState === client.OPEN) client.send(payload);&lt;br&gt;
      }&lt;br&gt;
    }&lt;br&gt;
  });&lt;br&gt;
});&lt;/p&gt;

&lt;p&gt;This works fine on one process. The real cost shows up when you scale horizontally: connections are pinned to whichever server instance accepted them, so a broadcast has to fan out across processes too — usually via Redis pub/sub, NATS, or a managed service like Pusher/Ably. That's infrastructure SSE and polling don't force on you nearly as early.&lt;/p&gt;

&lt;h2&gt;
  
  
  📊 Honest comparison
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Polling&lt;/th&gt;
&lt;th&gt;SSE&lt;/th&gt;
&lt;th&gt;WebSockets&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Direction&lt;/td&gt;
&lt;td&gt;client-pull&lt;/td&gt;
&lt;td&gt;server-push (one-way)&lt;/td&gt;
&lt;td&gt;bidirectional&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Transport&lt;/td&gt;
&lt;td&gt;plain HTTP&lt;/td&gt;
&lt;td&gt;plain HTTP (streamed)&lt;/td&gt;
&lt;td&gt;own protocol over TCP&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reconnect handling&lt;/td&gt;
&lt;td&gt;trivial (just retry)&lt;/td&gt;
&lt;td&gt;built into &lt;code&gt;EventSource&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;you build it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Horizontal scaling&lt;/td&gt;
&lt;td&gt;trivial (stateless)&lt;/td&gt;
&lt;td&gt;needs shared client registry&lt;/td&gt;
&lt;td&gt;needs pub/sub fan-out&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Proxy/firewall friendliness&lt;/td&gt;
&lt;td&gt;best&lt;/td&gt;
&lt;td&gt;good&lt;/td&gt;
&lt;td&gt;can be blocked/downgraded&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Good fit&lt;/td&gt;
&lt;td&gt;low-frequency, infrequent updates&lt;/td&gt;
&lt;td&gt;dashboards, notifications, live state&lt;/td&gt;
&lt;td&gt;chat, collab editing, multiplayer&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A pattern I keep coming back to: start with SSE for anything that's fundamentally "server tells clients what changed." Only reach for WebSockets once you have a genuine, frequent client-to-server-to-other-clients requirement that a &lt;code&gt;POST&lt;/code&gt; + SSE combo can't express cleanly. Polling is still the right call for admin dashboards or anything where a few seconds of staleness is genuinely fine and you'd rather not run a persistent-connection service at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  🧠 The part that actually matters: consistency, not transport
&lt;/h2&gt;

&lt;p&gt;Here's the thing none of the three options solve for you: what happens when two broadcasts race, or a client reconnects mid-update and misses a message? The transport is the easy 20%. The hard part is designing your broadcast payload so a client can always recover a consistent view — either by sending full state snapshots (like the example above) instead of deltas, or by including a version/sequence number so clients can detect gaps and request a resync.&lt;/p&gt;

&lt;p&gt;If you only ever broadcast diffs, a single dropped message means every client after it is silently wrong forever. That's the bug that doesn't show up in your demo and absolutely shows up in production three weeks later.&lt;/p&gt;

&lt;p&gt;What's your default pick for this kind of problem — do you reach for SSE first, or do you go straight to WebSockets out of habit? Curious how many people are still shipping raw polling in 2024 and just not talking about it.&lt;/p&gt;

</description>
      <category>node</category>
      <category>websocket</category>
      <category>realtime</category>
      <category>backend</category>
    </item>
    <item>
      <title>Graph Search Isn't Just a LeetCode Trick: BFS in Prod</title>
      <dc:creator>Renato Silva</dc:creator>
      <pubDate>Mon, 24 Aug 2026 09:40:56 +0000</pubDate>
      <link>https://dev.to/renato_silva_71eef0fc385f/graph-search-isnt-just-a-leetcode-trick-bfs-in-prod-5g2c</link>
      <guid>https://dev.to/renato_silva_71eef0fc385f/graph-search-isnt-just-a-leetcode-trick-bfs-in-prod-5g2c</guid>
      <description>&lt;p&gt;Every few months "six degrees of separation" and graph traversal puzzles trend again, and every time the comments split into two camps: people who think BFS/DFS are pure interview theater, and people quietly using them in production without telling anyone. I'm in the second camp. Here's a real feature — "related feedback threads" — built on a boring Node/Postgres stack, where breadth-first search turned out to be exactly the right tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  🎯 The Setup
&lt;/h2&gt;

&lt;p&gt;We run a support/feedback tool where users can link feedback items to each other: "this is related to that," "this duplicates that," "this was split off from that." Over time these links form a graph. Individually each link is trivial — a row in a join table. But support agents kept asking a very reasonable question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"If I'm looking at ticket #4521, what's the full cluster of stuff connected to it, even indirectly?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's not a JOIN. That's a graph traversal.&lt;/p&gt;

&lt;h2&gt;
  
  
  🔧 The Problem
&lt;/h2&gt;

&lt;p&gt;Our schema looks like this:&lt;/p&gt;

&lt;p&gt;sql&lt;br&gt;
CREATE TABLE feedback (&lt;br&gt;
  id SERIAL PRIMARY KEY,&lt;br&gt;
  title TEXT NOT NULL,&lt;br&gt;
  created_at TIMESTAMPTZ DEFAULT now()&lt;br&gt;
);&lt;/p&gt;

&lt;p&gt;CREATE TABLE feedback_links (&lt;br&gt;
  source_id INT REFERENCES feedback(id),&lt;br&gt;
  target_id INT REFERENCES feedback(id),&lt;br&gt;
  relation TEXT NOT NULL, -- 'related', 'duplicate', 'split_from'&lt;br&gt;
  PRIMARY KEY (source_id, target_id)&lt;br&gt;
);&lt;/p&gt;

&lt;p&gt;A single feedback item might link to 3 others, each of which links to 2 more, and so on. Agents don't just want direct neighbors — they want the whole connected component up to some reasonable depth, because context that's two or three hops away is often exactly what explains why a bug report and a feature request are secretly the same underlying issue.&lt;/p&gt;

&lt;p&gt;The naive fix — recursive JOINs pulled straight into application code with no depth limit — either times out on a dense cluster or returns way more noise than an agent can use in a support ticket sidebar.&lt;/p&gt;

&lt;h2&gt;
  
  
  🕸️ Modeling Feedback as a Graph
&lt;/h2&gt;

&lt;p&gt;Once you say "connected component up to N hops," you've already described BFS. Depth-first search would work too, but it explores one branch all the way down before backtracking, which is the wrong shape for "show me everything within 3 degrees, closest first." BFS naturally processes nodes in order of distance from the source, which maps directly onto the UI requirement: show closest-related items first, and stop expanding once you hit the depth cap.&lt;/p&gt;

&lt;h2&gt;
  
  
  🔍 The BFS Implementation
&lt;/h2&gt;

&lt;p&gt;We pull the edges relevant to the starting node's component lazily, level by level, straight from Postgres, and do the traversal logic in Node:&lt;/p&gt;

&lt;p&gt;javascript&lt;br&gt;
async function findRelatedFeedback(pool, startId, maxDepth = 3, maxResults = 50) {&lt;br&gt;
  const visited = new Set([startId]);&lt;br&gt;
  const result = [];&lt;br&gt;
  let frontier = [startId];&lt;br&gt;
  let depth = 0;&lt;/p&gt;

&lt;p&gt;while (frontier.length &amp;gt; 0 &amp;amp;&amp;amp; depth &amp;lt; maxDepth &amp;amp;&amp;amp; result.length &amp;lt; maxResults) {&lt;br&gt;
    const { rows } = await pool.query(&lt;br&gt;
      &lt;code&gt;SELECT source_id, target_id, relation&lt;br&gt;
       FROM feedback_links&lt;br&gt;
       WHERE source_id = ANY($1) OR target_id = ANY($1)&lt;/code&gt;,&lt;br&gt;
      [frontier]&lt;br&gt;
    );&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;const nextFrontier = [];

for (const row of rows) {
  const neighbor = frontier.includes(row.source_id) ? row.target_id : row.source_id;
  if (!visited.has(neighbor)) {
    visited.add(neighbor);
    nextFrontier.push(neighbor);
    result.push({ id: neighbor, depth: depth + 1, relation: row.relation });
  }
}

frontier = nextFrontier;
depth++;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;}&lt;/p&gt;

&lt;p&gt;return result;&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;This is textbook BFS with two production-shaped guardrails bolted on: &lt;code&gt;maxDepth&lt;/code&gt; so a densely connected cluster can't blow up the response, and &lt;code&gt;maxResults&lt;/code&gt; so a single mega-hub node (some "general feedback" catch-all ticket with 200 links) can't turn one API call into a full graph dump. Those two limits are doing more work for user experience than the algorithm itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  🐘 Doing It in Postgres Instead
&lt;/h2&gt;

&lt;p&gt;You can also push the whole traversal into the database with a recursive CTE, which is worth knowing even if you don't end up using it:&lt;/p&gt;

&lt;p&gt;sql&lt;br&gt;
WITH RECURSIVE related AS (&lt;br&gt;
  SELECT source_id AS id, 0 AS depth&lt;br&gt;
  FROM feedback WHERE id = $1&lt;br&gt;
  UNION&lt;br&gt;
  SELECT id, 0 FROM feedback WHERE id = $1&lt;/p&gt;

&lt;p&gt;UNION ALL&lt;/p&gt;

&lt;p&gt;SELECT&lt;br&gt;
    CASE WHEN fl.source_id = r.id THEN fl.target_id ELSE fl.source_id END,&lt;br&gt;
    r.depth + 1&lt;br&gt;
  FROM feedback_links fl&lt;br&gt;
  JOIN related r&lt;br&gt;
    ON fl.source_id = r.id OR fl.target_id = r.id&lt;br&gt;
  WHERE r.depth &amp;lt; 3&lt;br&gt;
)&lt;br&gt;
SELECT DISTINCT id, MIN(depth) AS depth&lt;br&gt;
FROM related&lt;br&gt;
WHERE id &amp;lt;&amp;gt; $1&lt;br&gt;
GROUP BY id&lt;br&gt;
ORDER BY depth;&lt;/p&gt;

&lt;p&gt;We tried this first. It's elegant and it's fewer round trips. But recursive CTEs don't cleanly enforce a "stop after N total nodes visited, regardless of depth" limit — you can cap depth, but capping total breadth requires awkward window functions or a hard row limit that can cut off a level halfway through and give you an inconsistent-looking result set. Doing BFS in application code, level by level, gave us a natural point to check "have I collected enough?" between each round trip. More queries, but more control.&lt;/p&gt;

&lt;h2&gt;
  
  
  ⚖️ Complexity Trade-offs at Small Scale
&lt;/h2&gt;

&lt;p&gt;Here's the part that actually matters for a "boring CRUD app with a graph feature" like ours: at our scale (tens of thousands of feedback items, average node degree under 4), textbook BFS complexity of O(V + E) is a non-issue. We're not traversing millions of edges. The real costs are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Round trips, not Big O.&lt;/strong&gt; Each BFS level is a network hop to Postgres. At depth 3 that's at most 3 queries, which is fine. If we ever needed depth 10, we'd batch differently or move to a native graph store.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fan-out nodes, not graph size.&lt;/strong&gt; The actual risk isn't "the graph is too big," it's "one node has too many neighbors." A single popular feedback thread with 80 links can make one BFS level return more rows than three normal traversals combined. This is why &lt;code&gt;maxResults&lt;/code&gt; matters more than &lt;code&gt;maxDepth&lt;/code&gt; in practice.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cycles are silent but real.&lt;/strong&gt; Feedback links can form loops (A relates to B relates to C relates back to A), and without the &lt;code&gt;visited&lt;/code&gt; set, plain recursive traversal would infinite-loop or duplicate work. It's the kind of bug that doesn't show up in a demo with 5 tickets and absolutely shows up once support agents start linking things liberally.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We explicitly did &lt;em&gt;not&lt;/em&gt; reach for Neo4j or any dedicated graph database. The join table plus application-level BFS handles our volume with room to spare, and it means one less piece of infrastructure to operate. If our average node degree climbed into the hundreds, or if we needed shortest-path-with-weighted-relations queries across millions of edges, that calculus would flip. Picking the graph database on day one for a feature that queries at most a few thousand edges is optimizing for a scale problem we don't have yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  🚀 Where This Goes Next
&lt;/h2&gt;

&lt;p&gt;The obvious next step is weighting edges by relation type — a "duplicate" link probably matters more than a "loosely related" one — which turns this from plain BFS into something closer to Dijkstra territory. We haven't needed it yet, but it's a good sign that starting with the simplest correct algorithm leaves you room to grow instead of boxing you in.&lt;/p&gt;

&lt;p&gt;Have you shipped a "basic" algorithm like BFS or DFS in a real feature and had people be surprised it wasn't over-engineered? I'd like to hear what problem it solved for you — and whether you eventually outgrew it.&lt;/p&gt;

</description>
      <category>postgres</category>
      <category>algorithms</category>
      <category>node</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Why I Started Rejecting My Own Giant PRs on a Solo Project</title>
      <dc:creator>Renato Silva</dc:creator>
      <pubDate>Thu, 20 Aug 2026 09:28:25 +0000</pubDate>
      <link>https://dev.to/renato_silva_71eef0fc385f/why-i-started-rejecting-my-own-giant-prs-on-a-solo-project-20ao</link>
      <guid>https://dev.to/renato_silva_71eef0fc385f/why-i-started-rejecting-my-own-giant-prs-on-a-solo-project-20ao</guid>
      <description>&lt;p&gt;I don't have teammates on this project. No one is waiting on my PRs, no one is blocked by my branch, and technically I could just push straight to &lt;code&gt;main&lt;/code&gt; and call it a day. For about eight months, that's exactly what I did. Then I started opening pull requests against myself, refusing to merge them until they passed a checklist, and my bug count dropped hard enough that I'm never going back.&lt;/p&gt;

&lt;p&gt;This isn't a productivity larp. It's a direct response to something I kept doing on a Node.js backend for a side project that grew into something people actually pay for: writing 1,200-line PRs that touched routing, database schema, auth middleware, and a new queue system all at once, then merging them at 1am because "it works locally."&lt;/p&gt;

&lt;h2&gt;
  
  
  🔧 The Problem
&lt;/h2&gt;

&lt;p&gt;Here's an actual PR title from my own history, from back when I didn't bother with PRs at all, just commits:&lt;/p&gt;

&lt;p&gt;commit 4a9f2c1&lt;br&gt;
Author: me&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;add subscription billing, refactor user model, switch to bullmq, fix cors bug
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;One commit. Four unrelated concerns. When something broke in production three weeks later — turned out the user model refactor silently changed how &lt;code&gt;email&lt;/code&gt; uniqueness was enforced — I had no way to bisect it cleanly. &lt;code&gt;git bisect&lt;/code&gt; pointed at a commit that also happened to introduce a queue system, so I spent an hour reading unrelated BullMQ code before I found the actual bug in a Mongoose schema change two files away.&lt;/p&gt;

&lt;p&gt;The mega-PR problem people are complaining about on GitHub right now — the 4,000-line diff nobody can meaningfully review — isn't really a GitHub problem. It's a batching problem. Solo devs get it too, we just don't call it a "review bottleneck" because there's no reviewer to bottleneck. The cost shows up later, as debugging tax instead of review tax.&lt;/p&gt;

&lt;h2&gt;
  
  
  🧩 What Changed
&lt;/h2&gt;

&lt;p&gt;I started treating my own future self as the reviewer. Concretely, that meant three habits, in order of how much they actually helped.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. One deployable concern per PR
&lt;/h3&gt;

&lt;p&gt;Not one &lt;em&gt;file&lt;/em&gt;. One &lt;em&gt;concern&lt;/em&gt;. A PR can touch six files if they all serve the same change. It cannot touch six unrelated changes even if it's technically "one file."&lt;/p&gt;

&lt;p&gt;Before:&lt;/p&gt;

&lt;p&gt;feat: subscription billing, user model refactor, bullmq, cors fix&lt;/p&gt;

&lt;p&gt;After, same work, split into four PRs merged over two days:&lt;/p&gt;

&lt;p&gt;fix: cors origin whitelist for staging subdomain&lt;br&gt;
refactor: normalize email field before uniqueness check&lt;br&gt;
feat: add BullMQ queue for email jobs (behind flag)&lt;br&gt;
feat: enable Stripe subscription billing on user model&lt;/p&gt;

&lt;p&gt;Each of those is independently revertible. When the queue system had a memory leak two weeks later, &lt;code&gt;git revert&lt;/code&gt; on one commit fixed it without touching billing.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Feature flags instead of long-lived branches
&lt;/h3&gt;

&lt;p&gt;The old instinct was to keep a branch alive for a week while I built something big, then merge it all at once — the exact mega-PR pattern. Now I merge small, working pieces behind a flag, even when the feature isn't done.&lt;/p&gt;

&lt;p&gt;javascript&lt;br&gt;
// config/flags.js&lt;br&gt;
const flags = {&lt;br&gt;
  QUEUE_EMAIL_JOBS: process.env.FLAG_QUEUE_EMAIL_JOBS === 'true',&lt;br&gt;
};&lt;/p&gt;

&lt;p&gt;module.exports = flags;&lt;/p&gt;

&lt;p&gt;javascript&lt;br&gt;
// services/emailService.js&lt;br&gt;
const { QUEUE_EMAIL_JOBS } = require('../config/flags');&lt;/p&gt;

&lt;p&gt;async function sendWelcomeEmail(user) {&lt;br&gt;
  if (QUEUE_EMAIL_JOBS) {&lt;br&gt;
    await emailQueue.add('welcome', { userId: user.id });&lt;br&gt;
  } else {&lt;br&gt;
    await mailer.sendNow(user.email, 'welcome');&lt;br&gt;
  }&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;This let me merge the BullMQ integration in small pieces — queue setup, worker process, retry logic — over four separate PRs, none of which changed production behavior until I flipped &lt;code&gt;FLAG_QUEUE_EMAIL_JOBS&lt;/code&gt; to &lt;code&gt;true&lt;/code&gt; in one final, tiny, easy-to-review PR:&lt;/p&gt;

&lt;p&gt;diff&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;FLAG_QUEUE_EMAIL_JOBS=false&lt;/li&gt;
&lt;li&gt;FLAG_QUEUE_EMAIL_JOBS=true&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If that broke something, the rollback was a one-line env change, not a &lt;code&gt;git revert&lt;/code&gt; across four commits with merge conflicts.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. A self-review checklist before I hit merge
&lt;/h3&gt;

&lt;p&gt;This is the part that actually changes behavior, because it forces a pause. Mine lives in &lt;code&gt;.github/pull_request_template.md&lt;/code&gt; and I fill it out even though I'm the only one who reads it:&lt;/p&gt;

&lt;p&gt;markdown&lt;/p&gt;

&lt;h2&gt;
  
  
  Self-review checklist
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;[ ] This PR does ONE thing. If I can't summarize it in one sentence, split it.&lt;/li&gt;
&lt;li&gt;[ ] No schema change and feature logic in the same PR.&lt;/li&gt;
&lt;li&gt;[ ] New code path is behind a flag if it touches billing, auth, or queues.&lt;/li&gt;
&lt;li&gt;[ ] I ran this against the staging DB dump, not just local seed data.&lt;/li&gt;
&lt;li&gt;[ ] Rollback plan: revert commit / flip flag / neither needed.&lt;/li&gt;
&lt;li&gt;[ ] Diff is under ~300 lines, or I have a good reason it isn't.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last checkbox alone killed most of my mega-PRs. "Under 300 lines" isn't a magic number — it's just small enough that I can actually reread the whole diff in one sitting and notice the thing I got wrong, instead of skimming because I already know what I meant to write.&lt;/p&gt;

&lt;h2&gt;
  
  
  📉 Before / After, With Real Numbers
&lt;/h2&gt;

&lt;p&gt;I pulled stats from my own git log across a 3-month window before and after adopting this.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;After&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Avg lines changed per PR&lt;/td&gt;
&lt;td&gt;640&lt;/td&gt;
&lt;td&gt;145&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Production incidents traced to a merge&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Time to &lt;code&gt;git bisect&lt;/code&gt; a regression&lt;/td&gt;
&lt;td&gt;~45 min avg&lt;/td&gt;
&lt;td&gt;~8 min avg&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PRs reverted in full&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;0 (partial reverts only)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The incident count matters most. Five of those six "before" incidents were bugs sitting quietly inside a large diff, unrelated to the actual thing I thought I was shipping. Smaller diffs didn't make me a better programmer overnight — they just made my mistakes smaller and easier to isolate.&lt;/p&gt;

&lt;h2&gt;
  
  
  🚦 Where I Still Cut Corners
&lt;/h2&gt;

&lt;p&gt;I'm not going to pretend this is pure discipline. Genuine one-off scripts, migrations I'll run exactly once, or throwaway debug endpoints still go straight to &lt;code&gt;main&lt;/code&gt; sometimes. The checklist is for anything touching auth, billing, data integrity, or anything a customer would notice if it broke. Applying full ceremony to a typo fix in a README would just be theater.&lt;/p&gt;

&lt;h2&gt;
  
  
  🙋 Your Turn
&lt;/h2&gt;

&lt;p&gt;If you're a solo dev or work on a small team with light review culture — do you actually PR your own work, or is &lt;code&gt;main&lt;/code&gt; still your review process? I'm curious whether feature flags feel like overhead to people working on smaller CRUD apps versus something like billing or queues where the blast radius of a bad merge is bigger.&lt;/p&gt;

&lt;p&gt;Drop your workflow in the comments, especially if you've got a better checklist item than mine — I'm always looking to steal a good one.&lt;/p&gt;

</description>
      <category>node</category>
      <category>git</category>
      <category>codereview</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Your API Doesn't Have an AI Problem, It Has a Design Problem</title>
      <dc:creator>Renato Silva</dc:creator>
      <pubDate>Wed, 19 Aug 2026 19:48:44 +0000</pubDate>
      <link>https://dev.to/renato_silva_71eef0fc385f/your-api-doesnt-have-an-ai-problem-it-has-a-design-problem-1l5f</link>
      <guid>https://dev.to/renato_silva_71eef0fc385f/your-api-doesnt-have-an-ai-problem-it-has-a-design-problem-1l5f</guid>
      <description>&lt;p&gt;Every week there's a new post about "adding AI to your API" — a chat endpoint, a summarization feature, an autocomplete widget. And every week, teams discover the same thing: the AI feature isn't the hard part. The hard part is that their API was never designed to answer real questions in the first place.&lt;/p&gt;

&lt;p&gt;AI doesn't create bad architecture. It just puts a spotlight on it and asks it to perform live, in front of an audience.&lt;/p&gt;

&lt;h2&gt;
  
  
  🔥 The Pattern Nobody Wants to Admit
&lt;/h2&gt;

&lt;p&gt;Here's the usual sequence:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Team builds a CRUD API around whatever tables were easiest to model.&lt;/li&gt;
&lt;li&gt;Product asks for an AI feature — "summarize customer sentiment," "suggest a response," "cluster similar feedback."&lt;/li&gt;
&lt;li&gt;Engineering discovers the API can't answer "similar to what?" or "sentiment over what time window, grouped how?" without a pile of N+1 queries, ad hoc joins, or a background job nobody wants to own.&lt;/li&gt;
&lt;li&gt;Someone ships a &lt;code&gt;/ai/summarize&lt;/code&gt; endpoint that quietly does three database round trips, a Python script, and a prayer.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The AI didn't break the system. The AI just needed the system to answer real, compositional questions — and it turns out the system was only ever designed to answer "give me row 42."&lt;/p&gt;

&lt;h2&gt;
  
  
  🧩 Case Study: minimalist-feedback-api
&lt;/h2&gt;

&lt;p&gt;Let's make this concrete with a small, honest example — a feedback API that looks totally reasonable at first glance.&lt;/p&gt;

&lt;p&gt;sql&lt;br&gt;
CREATE TABLE feedback (&lt;br&gt;
  id SERIAL PRIMARY KEY,&lt;br&gt;
  message TEXT NOT NULL,&lt;br&gt;
  rating INTEGER,&lt;br&gt;
  submitted_at TIMESTAMP DEFAULT now(),&lt;br&gt;
  user_email TEXT&lt;br&gt;
);&lt;/p&gt;

&lt;p&gt;And the API surface:&lt;/p&gt;

&lt;p&gt;http&lt;br&gt;
GET  /feedback&lt;br&gt;
GET  /feedback/:id&lt;br&gt;
POST /feedback&lt;br&gt;
DELETE /feedback/:id&lt;/p&gt;

&lt;p&gt;This is fine for a v1. It's minimal, it's CRUD, it ships fast. The problem is what it's missing: there's no concept of a &lt;em&gt;category&lt;/em&gt;, no &lt;em&gt;tags&lt;/em&gt;, no &lt;em&gt;source&lt;/em&gt; (web, mobile, support ticket), no &lt;em&gt;status&lt;/em&gt; (new, triaged, resolved), and no relationship to a product area or feature. &lt;code&gt;rating&lt;/code&gt; is a bare integer with no scale documented anywhere except a Slack message from eight months ago.&lt;/p&gt;

&lt;p&gt;Nobody complained, because the only client was an admin dashboard doing &lt;code&gt;SELECT * FROM feedback ORDER BY submitted_at DESC LIMIT 50&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  🤖 Where the AI Feature Broke Everything
&lt;/h2&gt;

&lt;p&gt;Then someone asks for: "Can we get an AI summary of feedback trends by feature area, this week vs. last week?"&lt;/p&gt;

&lt;p&gt;Suddenly every missing modeling decision becomes a blocking issue:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;There's no &lt;code&gt;feature_area&lt;/code&gt;, so the LLM prompt starts doing keyword matching on free text ("if message contains 'checkout'...") — which is just a worse, slower, non-deterministic version of a foreign key.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;rating&lt;/code&gt; isn't validated or scaled consistently, so "average sentiment" is comparing 1–5 stars against some rows where someone typed &lt;code&gt;-1&lt;/code&gt; two years ago and it never got caught.&lt;/li&gt;
&lt;li&gt;There's no &lt;code&gt;submitted_at&lt;/code&gt; index strategy for range queries, so "this week vs last week" becomes two full table scans through a text-heavy table, on every request, because there's no caching layer and no aggregation endpoint either.&lt;/li&gt;
&lt;li&gt;The endpoint that gets built to serve this, &lt;code&gt;/ai/summary&lt;/code&gt;, ends up doing the query, the grouping, the prompt construction, and the LLM call all inline, with no separation between "fetch relevant data" and "generate summary," which means you can't cache the first part or test it independently of the model.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;http&lt;br&gt;
GET /ai/summary?range=week&lt;/p&gt;

&lt;p&gt;{&lt;br&gt;
  "summary": "Feedback improved slightly...",&lt;br&gt;
  "note": "best effort, based on keyword matching, may be wrong"&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;That &lt;code&gt;note&lt;/code&gt; field is the tell. It's an apology baked into the response schema.&lt;/p&gt;

&lt;h2&gt;
  
  
  🛠 The Actual Fix: Model the Domain, Not the Table
&lt;/h2&gt;

&lt;p&gt;The fix has almost nothing to do with AI. It's the modeling work that should have happened before anyone typed &lt;code&gt;CREATE TABLE&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;sql&lt;br&gt;
CREATE TABLE feedback (&lt;br&gt;
  id SERIAL PRIMARY KEY,&lt;br&gt;
  message TEXT NOT NULL,&lt;br&gt;
  sentiment_score NUMERIC(3,2), -- normalized -1.0 to 1.0, computed once&lt;br&gt;
  source TEXT NOT NULL,          -- 'web', 'mobile', 'support'&lt;br&gt;
  feature_area_id INTEGER REFERENCES feature_areas(id),&lt;br&gt;
  status TEXT NOT NULL DEFAULT 'new',&lt;br&gt;
  submitted_at TIMESTAMP NOT NULL DEFAULT now(),&lt;br&gt;
  user_id INTEGER REFERENCES users(id)&lt;br&gt;
);&lt;/p&gt;

&lt;p&gt;CREATE INDEX idx_feedback_submitted_at ON feedback (submitted_at);&lt;br&gt;
CREATE INDEX idx_feedback_feature_area ON feedback (feature_area_id);&lt;/p&gt;

&lt;p&gt;And the endpoint set stops being pure CRUD and starts modeling actual questions people ask:&lt;/p&gt;

&lt;p&gt;http&lt;br&gt;
GET /feedback?feature_area=checkout&amp;amp;since=2024-05-01&amp;amp;until=2024-05-08&lt;br&gt;
GET /feedback/aggregate?group_by=feature_area&amp;amp;range=week&lt;br&gt;
GET /feature-areas/:id/trend?window=30d&lt;/p&gt;

&lt;p&gt;Notice what changed: the aggregation is a first-class resource (&lt;code&gt;/feedback/aggregate&lt;/code&gt;), not something invented inline inside an AI endpoint. Now the AI feature is almost boring:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
def generate_weekly_summary(feature_area_id: int) -&amp;gt; str:&lt;br&gt;
    trend = api.get(f"/feature-areas/{feature_area_id}/trend?window=7d")&lt;br&gt;
    prompt = build_summary_prompt(trend)  # deterministic, testable&lt;br&gt;
    return llm.complete(prompt)&lt;/p&gt;

&lt;p&gt;The LLM call is now the &lt;em&gt;last&lt;/em&gt; step, operating on well-shaped, pre-aggregated, already-correct data. If the summary is wrong, you can tell immediately whether it's a data problem or a prompting problem — because they're separated.&lt;/p&gt;

&lt;h2&gt;
  
  
  📐 What Good Looks Like
&lt;/h2&gt;

&lt;p&gt;A few concrete rules that fall out of this case study, not as abstract principles but as things you can check in a PR review:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;If an AI feature needs a join your API can't express, that join was always missing.&lt;/strong&gt; The AI request just made it visible faster than a human analyst would have.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Aggregation endpoints are not optional sugar.&lt;/strong&gt; &lt;code&gt;/resource/aggregate&lt;/code&gt; or &lt;code&gt;/resource/:id/trend&lt;/code&gt; should exist before anyone builds a summarization feature on top, not as a side effect of building one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Free text fields are where schema debt hides.&lt;/strong&gt; &lt;code&gt;message&lt;/code&gt; being a TEXT blob is fine; using string matching against it as a substitute for a &lt;code&gt;feature_area_id&lt;/code&gt; is a design smell wearing an AI costume.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Normalize before you summarize.&lt;/strong&gt; If &lt;code&gt;rating&lt;/code&gt; or &lt;code&gt;sentiment_score&lt;/code&gt; isn't validated at write time, no amount of prompt engineering downstream will make the aggregate trustworthy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep the retrieval and the generation separate, and testable separately.&lt;/strong&gt; If your only way to verify the LLM's output is to eyeball it, you've merged two very different failure modes into one.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this is AI-specific advice. It's just API design discipline that AI features are unusually good at exposing, because they demand compositional answers instead of row lookups.&lt;/p&gt;

&lt;h2&gt;
  
  
  💬 Over to You
&lt;/h2&gt;

&lt;p&gt;If you added an AI feature to an existing API recently — what actually broke first? Was it the schema, the missing aggregation layer, or something in how endpoints were shaped around CRUD instead of around the questions people actually ask?&lt;/p&gt;

&lt;p&gt;The uncomfortable version of this post is: if the AI feature made your API look bad, the API was already bad. AI is just an unusually blunt code reviewer.&lt;/p&gt;

</description>
      <category>api</category>
      <category>restapi</category>
      <category>softwaredesign</category>
      <category>ai</category>
    </item>
    <item>
      <title>Rate Limiting Lessons From a 100K-Request Meltdown</title>
      <dc:creator>Renato Silva</dc:creator>
      <pubDate>Fri, 14 Aug 2026 20:05:16 +0000</pubDate>
      <link>https://dev.to/renato_silva_71eef0fc385f/rate-limiting-lessons-from-a-100k-request-meltdown-7h0</link>
      <guid>https://dev.to/renato_silva_71eef0fc385f/rate-limiting-lessons-from-a-100k-request-meltdown-7h0</guid>
      <description>&lt;h2&gt;
  
  
  🔥 The Story That Made Every Backend Dev's Stomach Drop
&lt;/h2&gt;

&lt;p&gt;You probably saw it: a developer shipped a React component with a &lt;code&gt;useEffect&lt;/code&gt; that had a missing dependency array (or a state update that retriggered itself), and it quietly hammered their API with &lt;strong&gt;over 100,000 requests&lt;/strong&gt; before anyone noticed. No malicious actor, no botnet — just a bracket in the wrong place and a hook that fired on every render.&lt;/p&gt;

&lt;p&gt;The internet had a good laugh, but every backend dev reading that thread had the same intrusive thought: &lt;em&gt;"my API would've just... died."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That's the uncomfortable truth. A self-inflicted traffic spike from a buggy client is functionally indistinguishable from a DDoS if your server has no defenses. The fix isn't "tell frontend devs to be careful" — it's "assume they won't be, and build accordingly."&lt;/p&gt;

&lt;p&gt;This post walks through the three layers I now consider non-negotiable for any Node/Express API: &lt;strong&gt;token-bucket rate limiting&lt;/strong&gt;, &lt;strong&gt;circuit breakers&lt;/strong&gt;, and &lt;strong&gt;defensive defaults&lt;/strong&gt;. I'll also talk about a real (much smaller, thankfully) spike that hit my side project, &lt;code&gt;minimalist-feedback-api&lt;/code&gt;, and what actually saved it.&lt;/p&gt;

&lt;h2&gt;
  
  
  🪣 Why Token Bucket Beats Fixed Windows
&lt;/h2&gt;

&lt;p&gt;Most people's first rate limiter is a fixed window: "100 requests per minute per IP." It's easy to reason about and easy to implement badly. The problem is the boundary. If your window resets at :00, a client can send 100 requests at 11:59:59 and another 100 at 12:00:01 — 200 requests in two seconds, technically "within limits."&lt;/p&gt;

&lt;p&gt;Token bucket fixes this by modeling capacity as a continuously refilling resource instead of a hard reset:&lt;/p&gt;

&lt;p&gt;js&lt;br&gt;
// tokenBucket.js&lt;br&gt;
class TokenBucket {&lt;br&gt;
  constructor({ capacity, refillRatePerSec }) {&lt;br&gt;
    this.capacity = capacity;&lt;br&gt;
    this.tokens = capacity;&lt;br&gt;
    this.refillRate = refillRatePerSec;&lt;br&gt;
    this.lastRefill = Date.now();&lt;br&gt;
  }&lt;/p&gt;

&lt;p&gt;_refill() {&lt;br&gt;
    const now = Date.now();&lt;br&gt;
    const elapsedSec = (now - this.lastRefill) / 1000;&lt;br&gt;
    const refillAmount = elapsedSec * this.refillRate;&lt;br&gt;
    this.tokens = Math.min(this.capacity, this.tokens + refillAmount);&lt;br&gt;
    this.lastRefill = now;&lt;br&gt;
  }&lt;/p&gt;

&lt;p&gt;tryConsume(cost = 1) {&lt;br&gt;
    this._refill();&lt;br&gt;
    if (this.tokens &amp;gt;= cost) {&lt;br&gt;
      this.tokens -= cost;&lt;br&gt;
      return true;&lt;br&gt;
    }&lt;br&gt;
    return false;&lt;br&gt;
  }&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;module.exports = TokenBucket;&lt;/p&gt;

&lt;p&gt;Then the Express middleware, keyed per client (IP, API key, whatever identifies the caller):&lt;/p&gt;

&lt;p&gt;js&lt;br&gt;
// rateLimitMiddleware.js&lt;br&gt;
const TokenBucket = require('./tokenBucket');&lt;/p&gt;

&lt;p&gt;const buckets = new Map();&lt;/p&gt;

&lt;p&gt;function getBucket(key) {&lt;br&gt;
  if (!buckets.has(key)) {&lt;br&gt;
    buckets.set(key, new TokenBucket({ capacity: 20, refillRatePerSec: 2 }));&lt;br&gt;
  }&lt;br&gt;
  return buckets.get(key);&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;function rateLimit(req, res, next) {&lt;br&gt;
  const key = req.ip; // swap for API key if you have auth&lt;br&gt;
  const bucket = getBucket(key);&lt;/p&gt;

&lt;p&gt;if (bucket.tryConsume(1)) {&lt;br&gt;
    return next();&lt;br&gt;
  }&lt;/p&gt;

&lt;p&gt;res.status(429).set('Retry-After', '1').json({&lt;br&gt;
    error: 'Too many requests. Slow down.',&lt;br&gt;
  });&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;module.exports = rateLimit;&lt;/p&gt;

&lt;p&gt;The key insight: capacity 20 with a refill rate of 2/sec means a client gets a &lt;em&gt;burst allowance&lt;/em&gt; (handles legitimate rapid-fire usage like a form autosave) but can't sustain more than 2 requests/sec indefinitely. That's exactly the shape of a runaway &lt;code&gt;useEffect&lt;/code&gt; loop — it doesn't send 100 requests once, it sends them in a tight, sustained burst. Token bucket catches that pattern where a naive fixed window might not, depending on where the boundaries land.&lt;/p&gt;

&lt;p&gt;For anything beyond a single process, don't keep buckets in memory — use Redis (via something like &lt;code&gt;rate-limiter-flexible&lt;/code&gt;) so limits survive restarts and work across horizontally scaled instances. In-memory &lt;code&gt;Map&lt;/code&gt; is fine for a single-instance side project; it's a liability the moment you run two replicas behind a load balancer, because each instance tracks its own bucket and your effective limit doubles per replica.&lt;/p&gt;

&lt;h2&gt;
  
  
  🧯 Circuit Breakers: The Second Line of Defense
&lt;/h2&gt;

&lt;p&gt;Rate limiting protects your API from too many &lt;em&gt;incoming&lt;/em&gt; requests. Circuit breakers protect your API (and its downstream dependencies) from cascading failure once something's already struggling — usually a database, a third-party API, or an internal service call that's gone slow or unresponsive.&lt;/p&gt;

&lt;p&gt;Here's the pattern with &lt;code&gt;opossum&lt;/code&gt;, a solid circuit breaker library for Node:&lt;/p&gt;

&lt;p&gt;js&lt;br&gt;
const CircuitBreaker = require('opossum');&lt;br&gt;
const db = require('./db');&lt;/p&gt;

&lt;p&gt;async function fetchFeedback(id) {&lt;br&gt;
  return db.query('SELECT * FROM feedback WHERE id = $1', [id]);&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;const breakerOptions = {&lt;br&gt;
  timeout: 3000,              // fail fast after 3s&lt;br&gt;
  errorThresholdPercentage: 50, // trip if 50% of requests fail&lt;br&gt;
  resetTimeout: 10000,        // try again after 10s&lt;br&gt;
};&lt;/p&gt;

&lt;p&gt;const breaker = new CircuitBreaker(fetchFeedback, breakerOptions);&lt;/p&gt;

&lt;p&gt;breaker.fallback(() =&amp;gt; ({ error: 'Feedback service temporarily unavailable' }));&lt;/p&gt;

&lt;p&gt;app.get('/feedback/:id', async (req, res) =&amp;gt; {&lt;br&gt;
  const result = await breaker.fire(req.params.id);&lt;br&gt;
  res.json(result);&lt;br&gt;
});&lt;/p&gt;

&lt;p&gt;Without this, a slow database under load doesn't just cause slow responses — it causes &lt;em&gt;request pileup&lt;/em&gt;. Every incoming request holds a connection open waiting on a query that's never coming back fast enough, you exhaust your connection pool, and now healthy requests fail too. The circuit breaker trips, starts returning fast fallbacks immediately, and gives the database room to recover instead of getting buried under retries.&lt;/p&gt;

&lt;p&gt;Rate limiting stops the flood at the door. Circuit breakers stop one struggling dependency from taking the whole system down with it. You want both — they solve different failure modes.&lt;/p&gt;

&lt;h2&gt;
  
  
  🛡️ Defensive Defaults I Now Bake Into Every Express API
&lt;/h2&gt;

&lt;p&gt;Beyond the two big patterns above, there's a checklist of small things that cost nothing to add and save you on a bad day:&lt;/p&gt;

&lt;p&gt;js&lt;br&gt;
const express = require('express');&lt;br&gt;
const helmet = require('helmet');&lt;br&gt;
const compression = require('compression');&lt;/p&gt;

&lt;p&gt;const app = express();&lt;/p&gt;

&lt;p&gt;// Cap body size — don't let a malformed client send you a 500MB payload&lt;br&gt;
app.use(express.json({ limit: '100kb' }));&lt;/p&gt;

&lt;p&gt;// Basic security headers&lt;br&gt;
app.use(helmet());&lt;/p&gt;

&lt;p&gt;// Compress responses to reduce bandwidth under load&lt;br&gt;
app.use(compression());&lt;/p&gt;

&lt;p&gt;// Global request timeout so nothing hangs forever&lt;br&gt;
app.use((req, res, next) =&amp;gt; {&lt;br&gt;
  res.setTimeout(10000, () =&amp;gt; {&lt;br&gt;
    res.status(503).json({ error: 'Request timed out' });&lt;br&gt;
  });&lt;br&gt;
  next();&lt;br&gt;
});&lt;/p&gt;

&lt;p&gt;// Always have a catch-all error handler, even if it feels redundant&lt;br&gt;
app.use((err, req, res, next) =&amp;gt; {&lt;br&gt;
  console.error(err);&lt;br&gt;
  res.status(500).json({ error: 'Something went wrong' });&lt;br&gt;
});&lt;/p&gt;

&lt;p&gt;None of this is exciting. That's the point — defensive defaults are boring on purpose. The bracket-typo story went viral precisely because the API had no boring safety net, and 100,000 requests met zero resistance.&lt;/p&gt;

&lt;h2&gt;
  
  
  📈 The Day minimalist-feedback-api Got Hit
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;minimalist-feedback-api&lt;/code&gt; is a small feedback-collection service I built as a learning project — nothing fancy, just an endpoint for apps to POST feedback and a dashboard to read it. It's not built to handle enterprise traffic, but I treated it like production because that's where you actually learn this stuff.&lt;/p&gt;

&lt;p&gt;A few months in, one integrator's frontend had a retry loop with no backoff — every failed request immediately retried, and a brief blip in their own network turned into a sustained burst against my &lt;code&gt;/feedback&lt;/code&gt; endpoint. It wasn't 100,000 requests, but it was enough (a few thousand in under a minute) to be a real stress test.&lt;/p&gt;

&lt;p&gt;What actually saved it wasn't anything clever — it was the boring stuff: the token bucket limiter returned 429s immediately instead of letting requests queue up, the body size cap meant even the retries were cheap to reject, and the circuit breaker around my database call meant the brief connection pressure never turned into a full outage. The service degraded gracefully (some legitimate requests got 429'd too) instead of falling over entirely. That's the tradeoff you're signing up for: rate limiting means occasionally rejecting a request that would've been fine, in exchange for never going fully down.&lt;/p&gt;

&lt;h2&gt;
  
  
  🤔 What Would Your API Do?
&lt;/h2&gt;

&lt;p&gt;Honestly ask yourself: if a client-side bug sent your busiest endpoint 100,000 requests in five minutes right now, what would happen? Would it 429 gracefully, or would your database connection pool just... give up?&lt;/p&gt;

&lt;p&gt;If you're not sure, that uncertainty is the signal to add a rate limiter today — even a basic one. It's a couple hours of work that turns a viral "oops" story into a boring non-event. What's your go-to rate limiting setup, and have you ever had a spike (accidental or not) actually test it for real? I'd love to hear the war stories in the comments.&lt;/p&gt;

</description>
      <category>node</category>
      <category>express</category>
      <category>backend</category>
      <category>api</category>
    </item>
    <item>
      <title>I Stopped Trusting AI Agents With My API</title>
      <dc:creator>Renato Silva</dc:creator>
      <pubDate>Fri, 14 Aug 2026 19:51:29 +0000</pubDate>
      <link>https://dev.to/renato_silva_71eef0fc385f/i-stopped-trusting-ai-agents-with-my-api-117n</link>
      <guid>https://dev.to/renato_silva_71eef0fc385f/i-stopped-trusting-ai-agents-with-my-api-117n</guid>
      <description>&lt;h2&gt;
  
  
  🤖 The Problem
&lt;/h2&gt;

&lt;p&gt;A few weeks ago I wired an LLM agent up to &lt;code&gt;minimalist-feedback-api&lt;/code&gt;, my little side project for collecting product feedback. The pitch to myself was simple: let a support-bot agent read feedback threads and occasionally write a triage note or close a stale ticket, without me manually reviewing every call.&lt;/p&gt;

&lt;p&gt;It took about two days for the agent to do something I didn't ask for.&lt;/p&gt;

&lt;p&gt;Nothing catastrophic — it bulk-updated the status of a dozen feedback items because it decided, on its own, that they were "resolved" based on a fuzzy read of the conversation. Technically it used an endpoint I'd exposed to it. Technically the request was authenticated. But nobody had actually agreed that an agent should be allowed to do bulk writes, and there was no record of &lt;em&gt;why&lt;/em&gt; it thought that was a good idea.&lt;/p&gt;

&lt;p&gt;That's the part that got me. With a human client, a bad API call is a bug. With an agent, a bad API call is a &lt;em&gt;decision&lt;/em&gt;, made by a system that can also decide to make it again, faster, in a loop, at 3am.&lt;/p&gt;

&lt;p&gt;So I stopped trusting agents with the same trust model I give human-driven clients, and built a gatekeeper middleware specifically for tool-calling traffic.&lt;/p&gt;

&lt;h2&gt;
  
  
  🔐 What "Trust" Even Means for an Agent
&lt;/h2&gt;

&lt;p&gt;Before writing code, I had to get concrete about what I was actually worried about. It came down to three things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Scope&lt;/strong&gt; — an API key belonging to "the support agent" should not be able to call every write endpoint just because it's authenticated.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rate&lt;/strong&gt; — agents don't get bored or embarrassed. A misbehaving loop can hit your API way harder than a person ever would.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Audit&lt;/strong&gt; — when something weird happens, I need to reconstruct not just &lt;em&gt;what&lt;/em&gt; was called, but &lt;em&gt;which agent, with what identity, doing what it claimed to be doing.&lt;/em&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Regular auth middleware answers "who are you." This needed to answer "are you allowed to do &lt;em&gt;this specific thing&lt;/em&gt;, right now, at this rate, and is someone going to know about it."&lt;/p&gt;

&lt;h2&gt;
  
  
  🏗️ The Shape of the Middleware
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;minimalist-feedback-api&lt;/code&gt; has a handful of write endpoints: create feedback, update status, delete feedback, bulk operations. I treated agent access as a distinct concern from normal API auth — it sits &lt;em&gt;after&lt;/em&gt; authentication and &lt;em&gt;before&lt;/em&gt; the route handler.&lt;/p&gt;

&lt;p&gt;js&lt;br&gt;
// middleware/agentGatekeeper.js&lt;br&gt;
const agentScopes = {&lt;br&gt;
  'agent:support-triage': ['feedback:update-status', 'feedback:read'],&lt;br&gt;
  'agent:analytics-readonly': ['feedback:read'],&lt;br&gt;
};&lt;/p&gt;

&lt;p&gt;function requireAgentScope(action) {&lt;br&gt;
  return (req, res, next) =&amp;gt; {&lt;br&gt;
    const agentId = req.headers['x-agent-id'];&lt;br&gt;
    const agentToken = req.headers['x-agent-token'];&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;if (!agentId) {
  // Not an agent request, let normal auth handle it
  return next();
}

if (!verifyAgentToken(agentId, agentToken)) {
  return res.status(401).json({ error: 'invalid agent credentials' });
}

const allowed = agentScopes[agentId] || [];
if (!allowed.includes(action)) {
  return res.status(403).json({
    error: `agent '${agentId}' is not scoped for action '${action}'`,
  });
}

req.agent = { id: agentId, action };
next();
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;};&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;module.exports = { requireAgentScope };&lt;/p&gt;

&lt;p&gt;The key decision here: &lt;strong&gt;scopes are actions, not endpoints.&lt;/strong&gt; &lt;code&gt;feedback:update-status&lt;/code&gt; and &lt;code&gt;feedback:delete&lt;/code&gt; are separate permissions even though they might hit similar routes, because "update a status field" and "permanently delete a record" are very different risk levels. My support-triage agent gets the former, never the latter. No agent in this system currently has delete access, on purpose — if it needs to happen, a human does it.&lt;/p&gt;

&lt;h2&gt;
  
  
  🚦 Rate-Limiting Per Agent, Not Per IP
&lt;/h2&gt;

&lt;p&gt;Standard rate limiters key off IP address, which is close to useless for agents — they usually run from the same handful of server IPs as your other backend traffic. I keyed limiting off the agent identity instead, with tighter windows than I'd ever apply to a human-facing key:&lt;/p&gt;

&lt;p&gt;js&lt;br&gt;
const rateLimit = require('express-rate-limit');&lt;/p&gt;

&lt;p&gt;const agentWriteLimiter = rateLimit({&lt;br&gt;
  windowMs: 60 * 1000,&lt;br&gt;
  max: 5, // an agent doing &amp;gt;5 writes/min is suspicious, full stop&lt;br&gt;
  keyGenerator: (req) =&amp;gt; req.agent?.id || req.ip,&lt;br&gt;
  handler: (req, res) =&amp;gt; {&lt;br&gt;
    logAgentEvent({&lt;br&gt;
      agentId: req.agent?.id,&lt;br&gt;
      action: req.agent?.action,&lt;br&gt;
      outcome: 'rate_limited',&lt;br&gt;
    });&lt;br&gt;
    res.status(429).json({ error: 'agent rate limit exceeded' });&lt;br&gt;
  },&lt;br&gt;
});&lt;/p&gt;

&lt;p&gt;Five writes a minute felt aggressive when I set it, but it's forced something useful: if the agent legitimately needs to do more than that, it should be batching its reasoning into fewer, larger, more deliberate calls — not firing off a write per sentence of its own chain of thought.&lt;/p&gt;

&lt;h2&gt;
  
  
  📝 Auditing: The Part I Actually Use Every Day
&lt;/h2&gt;

&lt;p&gt;Scopes and rate limits prevent damage. The audit log is what lets me &lt;em&gt;trust&lt;/em&gt; the system incrementally instead of all-or-nothing. Every agent-originated write gets logged with enough context to answer "why did this happen" without me guessing:&lt;/p&gt;

&lt;p&gt;js&lt;br&gt;
function auditAgentWrite(req, res, next) {&lt;br&gt;
  const original = res.json.bind(res);&lt;br&gt;
  res.json = (body) =&amp;gt; {&lt;br&gt;
    if (req.agent) {&lt;br&gt;
      logAgentEvent({&lt;br&gt;
        agentId: req.agent.id,&lt;br&gt;
        action: req.agent.action,&lt;br&gt;
        method: req.method,&lt;br&gt;
        path: req.originalUrl,&lt;br&gt;
        requestBody: req.body,&lt;br&gt;
        statusCode: res.statusCode,&lt;br&gt;
        timestamp: new Date().toISOString(),&lt;br&gt;
      });&lt;br&gt;
    }&lt;br&gt;
    return original(body);&lt;br&gt;
  };&lt;br&gt;
  next();&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;Wiring it all together on a real route looks like this:&lt;/p&gt;

&lt;p&gt;js&lt;br&gt;
router.patch(&lt;br&gt;
  '/feedback/:id/status',&lt;br&gt;
  requireAgentScope('feedback:update-status'),&lt;br&gt;
  agentWriteLimiter,&lt;br&gt;
  auditAgentWrite,&lt;br&gt;
  updateFeedbackStatus,&lt;br&gt;
);&lt;/p&gt;

&lt;p&gt;The log entries go to a plain table (&lt;code&gt;agent_audit_log&lt;/code&gt;) rather than a generic app log stream, because I wanted to query it directly: "show me every write &lt;code&gt;agent:support-triage&lt;/code&gt; made in the last 24 hours" is a query I actually run now, especially after a prompt or model change.&lt;/p&gt;

&lt;h2&gt;
  
  
  ⚖️ Trade-offs I'm Consciously Accepting
&lt;/h2&gt;

&lt;p&gt;This isn't zero-cost. Static scope tables mean I have to redeploy to change what an agent can do — I'm fine with that friction on purpose, because "redeploy to expand agent permissions" is a feature, not a bug, at this stage. A more dynamic, database-backed scope system would remove the friction and also remove the forcing function that makes me think twice.&lt;/p&gt;

&lt;p&gt;I also don't do anything fancy with the audit data yet — no anomaly detection, no auto-revocation. It's a log I read. That's deliberately unglamorous; I'd rather have a boring, reliable trail than a clever system I don't fully understand when it fires.&lt;/p&gt;

&lt;h2&gt;
  
  
  🙋 Over to You
&lt;/h2&gt;

&lt;p&gt;If you're letting an agent call real write endpoints today, what's actually stopping it from doing something scoped, rate-limited access wouldn't have caught anyway? I'm curious whether people are seeing failure modes that permission systems can't touch — like an agent staying &lt;em&gt;within&lt;/em&gt; scope but still making bad judgment calls.&lt;/p&gt;

&lt;p&gt;Happy to share the full &lt;code&gt;minimalist-feedback-api&lt;/code&gt; gatekeeper module if there's interest — it's small enough to drop into most Express projects in an afternoon.&lt;/p&gt;

</description>
      <category>node</category>
      <category>express</category>
      <category>security</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
