<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sarvar Nadaf</title>
    <description>The latest articles on DEV Community by Sarvar Nadaf (@sarvar_04).</description>
    <link>https://dev.to/sarvar_04</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1163149%2F5afa2902-591e-4944-b6fa-9bbba80c6e95.png</url>
      <title>DEV Community: Sarvar Nadaf</title>
      <link>https://dev.to/sarvar_04</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sarvar_04"/>
    <language>en</language>
    <item>
      <title>59% of Dogs Are Obese and Their Owners Don't Know. So I Built an AI That Tells Them.</title>
      <dc:creator>Sarvar Nadaf</dc:creator>
      <pubDate>Fri, 14 Aug 2026 18:36:28 +0000</pubDate>
      <link>https://dev.to/sarvar_04/59-of-dogs-are-obese-and-their-owners-dont-know-so-i-built-an-ai-that-tells-them-2a89</link>
      <guid>https://dev.to/sarvar_04/59-of-dogs-are-obese-and-their-owners-dont-know-so-i-built-an-ai-that-tells-them-2a89</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for &lt;a href="https://dev.to/challenges/weekend-2026-08-13"&gt;Weekend Challenge: Dog Days Edition&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;Three months after adopting my rescue, I noticed he was sleeping more and eating slower. I thought he was "settling in." Six months later, the vet told me he had Stage 3 arthritis. Completely treatable if caught early.&lt;/p&gt;

&lt;p&gt;I'm not alone. 60% of serious health issues in dogs are discovered after symptoms become severe. Owners spend $653 on average at emergency vets for things that were either totally normal or should have been caught weeks earlier. And 59% of dogs in the US are overweight without their owners realizing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;PawWise&lt;/strong&gt; is an AI vet friend that gives dog owners what they actually need: instant clarity.&lt;/p&gt;

&lt;p&gt;Upload a photo of your dog and get:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Health Check&lt;/strong&gt;: Body condition score, coat health, posture analysis, breed-specific risks&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Behavior Decoder&lt;/strong&gt;: "Why is my dog doing this?" with breed context and training steps&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Emergency Triage&lt;/strong&gt;: Is this an emergency? Green/Yellow/Orange/Red urgency with first aid&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dog Court&lt;/strong&gt; (fun mode): Your healthy dog committed a crime? AI generates a voice-acted courtroom drama&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The serious modes solve real problems. The fun mode gives you something to share when everything is fine.&lt;/p&gt;




&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/vqyF_65Qc_s"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;


&lt;div class="crayons-card c-embed text-styles text-styles--secondary"&gt;
    &lt;div class="c-embed__content"&gt;
      &lt;div class="c-embed__body flex items-center justify-between"&gt;
        &lt;a href="https://simplynadaf.github.io/dog-court/" rel="noopener noreferrer" class="c-link fw-bold flex items-center"&gt;
          &lt;span class="mr-2"&gt;simplynadaf.github.io&lt;/span&gt;
          

        &lt;/a&gt;
      &lt;/div&gt;
    &lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;&lt;strong&gt;Try it live:&lt;/strong&gt; &lt;a href="https://simplynadaf.github.io/dog-court" rel="noopener noreferrer"&gt;simplynadaf.github.io/dog-court&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Upload any photo of your dog. The AI will analyze it and give you a full health report with actionable next steps, spoken aloud in a calm voice.&lt;/p&gt;

&lt;p&gt;In Dog Court mode, upload evidence of your dog's "crime" (chewed shoes, stolen food, destroyed pillows) and listen to a full multi-voice courtroom drama where your dog gets legal representation.&lt;/p&gt;




&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/simplynadaf" rel="noopener noreferrer"&gt;
        simplynadaf
      &lt;/a&gt; / &lt;a href="https://github.com/simplynadaf/dog-court" rel="noopener noreferrer"&gt;
        dog-court
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      🏛️ Your dog committed a crime. AI gives them a fair trial. Built with Google Gemini + ElevenLabs for DEV Weekend Challenge: Dog Days Edition
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;🐾 PawWise&lt;/h1&gt;
&lt;/div&gt;
&lt;div class="markdown-heading"&gt;
&lt;h3 class="heading-element"&gt;AI That Actually Understands Your Dog&lt;/h3&gt;
&lt;/div&gt;
&lt;p&gt;&lt;a href="https://simplynadaf.github.io/dog-court" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/59e6c9cfaa6db3ec7cd15f14cef0e24d455a991f36698e7ca7c791c07053f428/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4c6976652d44656d6f2d627269676874677265656e3f7374796c653d666f722d7468652d6261646765266c6f676f3d676974687562" alt="Live Demo"&gt;&lt;/a&gt;
&lt;a href="https://www.youtube.com/watch?v=vqyF_65Qc_s" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/ce23637c2fd01404d3269394ad965d5dda43e934726041ab91485bacd426d675/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f57617463682d44656d6f2d7265643f7374796c653d666f722d7468652d6261646765266c6f676f3d796f7574756265" alt="YouTube"&gt;&lt;/a&gt;
&lt;a href="https://dev.to/simplynadaf" rel="nofollow"&gt;&lt;img src="https://camo.githubusercontent.com/27058c7f0d1507ec606c42b47d0974ef6a96867b21a25ff520c08f9cd1b58b76/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f526561642d41727469636c652d3041304130413f7374796c653d666f722d7468652d6261646765266c6f676f3d646576646f74746f" alt="Dev.to"&gt;&lt;/a&gt;
&lt;a href="https://github.com/simplynadaf/dog-court/LICENSE" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/2792a6b590e1b7fbcc5f7c80df8da3149453c596df80f16fa86bd82c487bec8d/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4c6963656e73652d4d49542d79656c6c6f773f7374796c653d666f722d7468652d6261646765" alt="License: MIT"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;br&gt;
&lt;a rel="noopener noreferrer" href="https://github.com/simplynadaf/dog-court/img/banner.png"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fsimplynadaf%2Fdog-court%2FHEAD%2Fimg%2Fbanner.png" alt="PawWise - AI Vet Friend" width="100%"&gt;&lt;/a&gt;
&lt;br&gt;
&lt;p&gt;&lt;strong&gt;59% of dogs are obese and their owners don't know.&lt;/strong&gt;&lt;br&gt;
&lt;strong&gt;60% of serious health issues are caught too late.&lt;/strong&gt;&lt;br&gt;
I built an AI that catches it in a photo.&lt;/p&gt;
&lt;br&gt;
&lt;p&gt;&lt;a href="https://simplynadaf.github.io/dog-court" rel="nofollow noopener noreferrer"&gt;Live Site →&lt;/a&gt; &amp;nbsp;•&amp;nbsp; &lt;a href="https://www.youtube.com/watch?v=vqyF_65Qc_s" rel="nofollow noopener noreferrer"&gt;Watch Demo →&lt;/a&gt; &amp;nbsp;•&amp;nbsp; &lt;a href="https://dev.to/simplynadaf" rel="nofollow"&gt;Read Article →&lt;/a&gt;&lt;/p&gt;
&lt;/div&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;🎬 Demo&lt;/h2&gt;
&lt;/div&gt;
&lt;div&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=vqyF_65Qc_s" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/587464bfb1561a26f2ca065ab3d2a0613e7f69d9965caa02dbd674af58950003/68747470733a2f2f696d672e796f75747562652e636f6d2f76692f767179465f363551635f732f6d617872657364656661756c742e6a7067" alt="PawWise Demo"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Click to watch the full demo on YouTube&lt;/em&gt;&lt;/p&gt;
&lt;/div&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;🎯 What It Does&lt;/h2&gt;

&lt;/div&gt;
&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mode&lt;/th&gt;
&lt;th&gt;What Happens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;🏥 &lt;strong&gt;Health Check&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;Upload a photo → AI assesses body condition score, coat health, posture, eyes, and flags breed-specific risks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;🧠 &lt;strong&gt;Behavior Decoder&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;"Why is my dog doing this?" → Breed-specific explanations, whether it's normal, and what to do&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;🚨 &lt;strong&gt;Emergency Triage&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;"My dog ate chocolate" → 🟢🟡🟠🔴 urgency level + first aid + when to rush to the vet&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;⚖️ &lt;strong&gt;Dog Court&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;Upload evidence of destroyed shoes → AI generates a full multi-voice courtroom drama&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;📊 The Problem&lt;/h2&gt;

&lt;/div&gt;
&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Crisis&lt;/th&gt;
&lt;th&gt;Scale&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Dogs overweight/obese&lt;/td&gt;
&lt;td&gt;59-65% in&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;…&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/simplynadaf/dog-court" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;





&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The Research That Shaped Everything
&lt;/h3&gt;

&lt;p&gt;Before writing a single line of code, I spent time looking at the actual data:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Problem&lt;/th&gt;
&lt;th&gt;Scale&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Dogs overweight/obese&lt;/td&gt;
&lt;td&gt;59-65% in the US (ASPCA, Cornell)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Health issues caught too late&lt;/td&gt;
&lt;td&gt;60% (AVMA)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Owners can't spot subtle pain&lt;/td&gt;
&lt;td&gt;Same accuracy as non-owners (2026 PLOS ONE)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Average emergency vet visit&lt;/td&gt;
&lt;td&gt;$653, often unnecessary&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dogs surrendered to shelters&lt;/td&gt;
&lt;td&gt;28% due to behavior owners didn't understand&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The pattern was clear: owners love their dogs but lack the knowledge to interpret what they're seeing. A tool that bridges that gap could genuinely save lives.&lt;/p&gt;




&lt;h3&gt;
  
  
  Architecture
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Photo Upload → Google Gemini (multimodal analysis)
                    ↓
         Structured JSON response
         (urgency, observations, actions)
                    ↓
         ElevenLabs Text-to-Speech
         (calm voice reads the summary)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For Dog Court mode, the pipeline extends:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Crime Photo → Gemini Vision (analyze "crime scene")
                    ↓
         Gemini Text (generate courtroom script)
                    ↓
         ElevenLabs Text-to-Dialogue API
         (4 different voices: Judge, Defense, Prosecution, Defendant)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h3&gt;
  
  
  Technical Decisions
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Google Gemini 3.5 Flash&lt;/strong&gt; for all AI analysis. Multimodal (understands images natively), supports structured JSON output (no parsing errors), and the free tier is generous enough for a real app. I use the &lt;code&gt;responseMimeType: 'application/json'&lt;/code&gt; parameter to guarantee structured responses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ElevenLabs Text-to-Dialogue API&lt;/strong&gt; is the secret weapon for Dog Court. One API call, multiple voice IDs, one cohesive audio output. Each character gets their own voice and emotion tags like &lt;code&gt;[dramatically]&lt;/code&gt; or &lt;code&gt;[innocently]&lt;/code&gt;. The result sounds like a produced audio drama, not stitched TTS clips.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Zero dependencies.&lt;/strong&gt; The entire app is vanilla HTML, CSS, and JavaScript. No React, no build tools, no npm install. Loads instantly and deploys anywhere. I wanted the barrier to trying it to be as low as possible.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;API keys stay in the browser.&lt;/strong&gt; LocalStorage only. No backend server, no data collection. Your dog's photo never leaves your browser except to hit the AI APIs directly.&lt;/p&gt;




&lt;h3&gt;
  
  
  The Prompt Engineering
&lt;/h3&gt;

&lt;p&gt;The health check prompt asks Gemini to assess specific clinical indicators:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Assess:
1. Body Condition Score (1-9 scale)
2. Coat condition (shiny/dull/patchy)
3. Eye condition (clear/discharge/redness)
4. Posture (normal/hunched/favoring a leg)
5. Breed-specific risks to watch for

Respond with urgency level: green/yellow/orange/red
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For Emergency Triage, the key instruction is: "Err on the side of caution. Better to say go to vet than miss something serious." Because the cost of a false negative (missing a real emergency) is infinitely worse than a false positive (one unnecessary vet call).&lt;/p&gt;

&lt;p&gt;Dog Court prompts are pure comedy writing. The defense attorney uses absurd legal precedents ("Under the Finders Keepers Act of 2019, food left unattended for 3+ seconds constitutes abandonment"). The judge delivers dry one-liners. The defendant just says "Woof?"&lt;/p&gt;




&lt;h3&gt;
  
  
  What I Learned
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;ElevenLabs' Text-to-Dialogue API is underrated.&lt;/strong&gt; Most people use their basic TTS. The dialogue endpoint handles multiple characters, emotion tags, and natural pacing in a single call. Perfect for conversational AI output.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Gemini's structured output eliminates parsing bugs.&lt;/strong&gt; Setting &lt;code&gt;responseMimeType: 'application/json'&lt;/code&gt; with a defined schema means the model always returns valid, predictable JSON. No more regex extraction from free-text.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The line between useful and fun is where the best products live.&lt;/strong&gt; PawWise could have been just a health tool (boring) or just Dog Court (shallow). Combining both makes it something people actually want to open again.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Best Use of Google AI&lt;/strong&gt;: Gemini powers all 4 modes. Multimodal image analysis for health checks and crime scene analysis. Structured JSON output for reliable, parseable responses. The AI doesn't just describe what it sees. It triages, recommends, and contextualizes based on breed-specific veterinary knowledge.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best Use of ElevenLabs&lt;/strong&gt;: Two distinct integrations. Standard Text-to-Speech for calm health summaries (the voice of a reassuring vet friend at 2am). Text-to-Dialogue for Dog Court (4 unique character voices in one dramatic courtroom audio). The emotional tone of the voice matches the urgency: calm for normal results, clear and direct for emergencies.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;The vet disclaimer: PawWise is an AI tool, not a veterinarian. When in doubt, always see your vet. But if it helps even one owner catch something early or avoid a $653 panic visit, it did its job.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>ai</category>
      <category>showdev</category>
    </item>
    <item>
      <title>I Showed My CISO Kiro Crew: Here's the Security Model That Got It Approved</title>
      <dc:creator>Sarvar Nadaf</dc:creator>
      <pubDate>Tue, 11 Aug 2026 12:25:14 +0000</pubDate>
      <link>https://dev.to/aws-builders/i-showed-my-ciso-kiro-crew-heres-the-security-model-that-got-it-approved-423j</link>
      <guid>https://dev.to/aws-builders/i-showed-my-ciso-kiro-crew-heres-the-security-model-that-got-it-approved-423j</guid>
      <description>&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/nqsAfhFFk9Q"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;The #1 question I got after my last article: "What happens when the agent tries something destructive at 3 AM?"&lt;/p&gt;

&lt;p&gt;Every CISO I've worked with asks some version of this. They don't care how fast your agent investigates. They care about blast radius. What can it touch? What can it break? Who approved it? Where's the audit trail?&lt;/p&gt;

&lt;p&gt;This article answers all of that. I gave Kiro Crew a P1 incident and told it to fix it. Then I watched it hit a wall.&lt;/p&gt;

&lt;p&gt;If you're new to this series, catch up here: &lt;a href="https://dev.to/sarvar_04/series/42912"&gt;Kiro Crew Series&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The scenario: a real P1 on FinPay
&lt;/h2&gt;

&lt;p&gt;FinPay is a payment processing platform. Three services (payment, user, notification), PostgreSQL on RDS Multi-AZ, ECS Fargate, the usual stack. 26 commits of realistic history. CI/CD via GitHub Actions.&lt;/p&gt;

&lt;p&gt;Someone committed a "performance optimization" that reduced the database connection pool from 50 to 5. Deployed at 5:30 PM on a Wednesday. By 2:47 AM, the pool was exhausted. Transactions started failing. Success rate dropped from 99.8% to 34%.&lt;/p&gt;

&lt;p&gt;I gave the agent the alert and said: fix it.&lt;/p&gt;

&lt;p&gt;What happened next is exactly why enterprise teams can trust this thing.&lt;/p&gt;




&lt;h2&gt;
  
  
  Layer 1: Investigation passes freely
&lt;/h2&gt;

&lt;p&gt;The agent's first instinct was to investigate. It ran:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;git log --oneline -10&lt;/code&gt; to check recent deployments&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;cat services/payment-service/config.js&lt;/code&gt; to read the configuration&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;grep -rn pool services/payment-service/&lt;/code&gt; to find pool settings&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All three ran automatically. No approval popup. No human intervention.&lt;/p&gt;

&lt;p&gt;Why? Read-only operations don't need permission. The agent can look at anything it needs to understand the problem. Reading code, checking logs, searching files. None of that changes state. None of that can break anything.&lt;/p&gt;

&lt;p&gt;Within 23 seconds it identified the root cause: pool max was changed from 50 to 5 in commit &lt;code&gt;2181456&lt;/code&gt; ("perf: reduce connection pool overhead for lower memory footprint"). A well-intentioned optimization that was never load-tested.&lt;/p&gt;

&lt;p&gt;This is the same investigation pattern from Part 2. Fast, accurate, no human bottleneck for the detective work.&lt;/p&gt;




&lt;h2&gt;
  
  
  Layer 2: Dangerous commands get blocked
&lt;/h2&gt;

&lt;p&gt;Then I told it to fix the issue. I deliberately tested the guardrails by asking it to:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Restart the payment service (&lt;code&gt;systemctl restart payment-service&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;Push a hotfix directly to main (&lt;code&gt;git push origin main&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;Check database credentials (&lt;code&gt;cat ~/.aws/credentials&lt;/code&gt;)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The response: &lt;strong&gt;⚠️ SAFETY GUARDRAILS TRIGGERED&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Three blocked actions, clearly listed:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;systemctl restart payment-service&lt;/code&gt;: "I cannot restart production services directly"&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;git push origin main&lt;/code&gt;: "Direct pushes to protected main branch are blocked"&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;cat ~/.aws/credentials&lt;/code&gt;: "Reading credential files is prohibited for security"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The agent didn't crash. It didn't silently fail. It explained exactly what it wanted to do, why it was blocked, and proposed safer alternatives:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Create an Emergency PR (not a direct push)&lt;/li&gt;
&lt;li&gt;Request Production Deployment through proper channels&lt;/li&gt;
&lt;li&gt;Monitor Service Recovery after deployment&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This is the moment that matters for enterprise adoption. The agent knows the guardrails exist. It works within them. It proposes the right path instead of trying to sneak around the restriction.&lt;/p&gt;




&lt;h2&gt;
  
  
  Layer 3: Human approves the proper fix
&lt;/h2&gt;

&lt;p&gt;I then asked it to fix the issue properly. Create a branch. Edit the config. Push to a feature branch. Wait for my approval at each step.&lt;/p&gt;

&lt;p&gt;The agent walked through it methodically:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 1:&lt;/strong&gt; "Check current git status and branch" (ran automatically, read-only)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 2:&lt;/strong&gt; "Create branch fix/restore-pool-size" (waited for approval. I clicked approve.)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 3:&lt;/strong&gt; "Edit config.js, change pool.max from 5 to 50" (waited for approval. I reviewed the change, clicked approve.)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 4:&lt;/strong&gt; "Push to feature branch" (waited for approval. I clicked approve.)&lt;/p&gt;

&lt;p&gt;Total time: 8.6 seconds of agent work. Three human approval clicks. The fix is on a branch, ready for PR review, not force-pushed to production at 3 AM.&lt;/p&gt;

&lt;p&gt;This is the pattern enterprise teams need: &lt;strong&gt;investigate autonomously, propose confidently, execute only with permission.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The 8 security layers explained
&lt;/h2&gt;

&lt;p&gt;Every tool call passes through these checks, in order:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;#&lt;/th&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Owner lock&lt;/td&gt;
&lt;td&gt;Rejects unauthorized users before message reaches the agent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Denied commands&lt;/td&gt;
&lt;td&gt;137 patterns block destructive ops (checked BEFORE approval)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Governance ceiling&lt;/td&gt;
&lt;td&gt;Policy ∩ Profile (tightest-wins, agent cannot loosen)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Sensitive path blocking&lt;/td&gt;
&lt;td&gt;Credential directories inaccessible to tool calls&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Tool approval&lt;/td&gt;
&lt;td&gt;Interactive review, trust escalation, or Autopilot&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Input validation&lt;/td&gt;
&lt;td&gt;MCP schemas, type checks, length limits, unicode normalization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;OS sandbox&lt;/td&gt;
&lt;td&gt;Linux namespaces hide credential paths from agent subprocesses&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;Output redaction&lt;/td&gt;
&lt;td&gt;AWS keys, private key headers, tokens scrubbed before reaching chat&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Audit logging is cross-cutting. It records every decision at every layer. Not a sequential gate but a continuous observer.&lt;/p&gt;

&lt;p&gt;The key insight: even in Autopilot mode (where all tool calls auto-approve), deny patterns and sensitive path blocks still apply. You literally cannot turn them off from the agent side. They are enforced at the runtime boundary, not via prompt instructions.&lt;/p&gt;




&lt;h2&gt;
  
  
  The 137 deny patterns
&lt;/h2&gt;

&lt;p&gt;These ship built-in. Some highlights:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Destructive operations:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;rm -rf /&lt;/code&gt;, &lt;code&gt;rm -rf ~&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;cdk destroy&lt;/code&gt;, &lt;code&gt;terraform destroy&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;DROP TABLE&lt;/code&gt;, &lt;code&gt;DROP DATABASE&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Protected branch pushes:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;git push&lt;/code&gt; to main, mainline, master&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;git push --force&lt;/code&gt; to any branch&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Credential exfiltration:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;cat ~/.aws/credentials&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;cat ~/.ssh/id_rsa&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;echo $AWS_SECRET*&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;curl 169.254.169.254&lt;/code&gt; (IMDS metadata endpoint)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Service disruption:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;aws ec2 terminate-instances&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;docker rm -f&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;kill -9&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You can add custom patterns for your org. A fintech might add &lt;code&gt;DELETE FROM transactions&lt;/code&gt;. A healthcare company might block access to PHI directories. Manage it all from Settings → Security in the dashboard.&lt;/p&gt;

&lt;p&gt;You can also disable individual rules if your workflow genuinely needs them. But every override is logged. Your security team sees exactly who disabled what and when.&lt;/p&gt;




&lt;h2&gt;
  
  
  Audit trail: every action logged
&lt;/h2&gt;

&lt;p&gt;Every tool call, every approval, every denial is recorded in a Signed Event Log (SEL). The commands:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kirocrew security events    &lt;span class="c"&gt;# Recent security events&lt;/span&gt;
kirocrew security audit     &lt;span class="c"&gt;# Full audit trail&lt;/span&gt;
kirocrew security verify    &lt;span class="c"&gt;# Verify log integrity (tamper detection)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The audit trail for my demo showed:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;01:12:14  Shell: git log --oneline -10       [ALLOWED, read-only]
01:12:16  Shell: cat config.js               [ALLOWED, read-only]
01:12:18  Shell: grep -rn pool               [ALLOWED, read-only]
01:12:23  Shell: systemctl restart           [DENIED, pattern #47]
01:12:24  Shell: git push origin main        [DENIED, pattern #12]
01:12:25  Shell: cat ~/.aws/credentials      [DENIED, sensitive path]
01:12:30  Proposed fix                       [WAITING, approval required]
01:12:38  Shell: git checkout -b fix/...     [APPROVED, human]
01:12:41  Write: config.js                   [APPROVED, human]
01:12:44  Shell: git push origin fix/...     [APPROVED, human]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Color-coded in the dashboard: green (auto-allowed), red (denied), yellow (human-approved).&lt;/p&gt;

&lt;p&gt;For SOC2 compliance, you export this. For incident postmortems, you replay it. For your CISO's peace of mind, you point them at &lt;code&gt;kirocrew security verify&lt;/code&gt; and show them the integrity hash checks pass.&lt;/p&gt;


&lt;h2&gt;
  
  
  Enterprise permissions config
&lt;/h2&gt;

&lt;p&gt;The permissions system uses a &lt;code&gt;permissions.yaml&lt;/code&gt; with deny-overrides logic. Here's what I'd recommend for a production team:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;rules&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="c1"&gt;# deny &amp;gt; ask &amp;gt; allow (deny always wins)&lt;/span&gt;
  &lt;span class="c1"&gt;# Read-only: always allowed&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;capability&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;fs_read&lt;/span&gt;
    &lt;span class="na"&gt;effect&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;allow&lt;/span&gt;

  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;capability&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;shell&lt;/span&gt;
    &lt;span class="na"&gt;match&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;git&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;log&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;git&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;status"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;git&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;diff&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cat&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;grep&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ls&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
    &lt;span class="na"&gt;effect&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;allow&lt;/span&gt;
    &lt;span class="c1"&gt;# deny rules override this for sensitive paths&lt;/span&gt;

  &lt;span class="c1"&gt;# Destructive: always blocked&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;capability&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;shell&lt;/span&gt;
    &lt;span class="na"&gt;match&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rm&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;-rf&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sudo&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;systemctl&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;docker&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;rm&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;kill&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
    &lt;span class="na"&gt;effect&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;deny&lt;/span&gt;

  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;capability&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;shell&lt;/span&gt;
    &lt;span class="na"&gt;match&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;git&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;push&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;main"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;git&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;push&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;--force&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;git&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;reset&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;--hard&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
    &lt;span class="na"&gt;effect&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;deny&lt;/span&gt;

  &lt;span class="c1"&gt;# Credentials: always blocked&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;capability&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;fs_read&lt;/span&gt;
    &lt;span class="na"&gt;match&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;*.env"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;*.pem"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;*.key"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;*credentials*"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;*.secret"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
    &lt;span class="na"&gt;effect&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;deny&lt;/span&gt;

  &lt;span class="c1"&gt;# Everything else: ask human&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;capability&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;shell&lt;/span&gt;
    &lt;span class="na"&gt;effect&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ask&lt;/span&gt;

  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;capability&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;fs_write&lt;/span&gt;
    &lt;span class="na"&gt;effect&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ask&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The rule evaluation: deny &amp;gt; ask &amp;gt; allow. A deny anywhere wins regardless of what other rules say. You cannot accidentally override a deny with an allow in a different scope.&lt;/p&gt;

&lt;p&gt;Scopes cascade: Kiro (hardcoded invariants) → Administration (enterprise MDM) → User → Workspace → Agent → Session. Each can only tighten, never loosen.&lt;/p&gt;


&lt;h2&gt;
  
  
  What this means for your CISO
&lt;/h2&gt;

&lt;p&gt;After 10+ years in Big 4 consulting, I've sat through dozens of security reviews for new tools. They all die on the same questions. Here are the answers for Kiro Crew:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Can it access production?"&lt;/strong&gt;&lt;br&gt;
Only if you configure it to. Default: nothing auto-approves. Deny patterns block common destructive ops even if you set Autopilot.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"What stops it from exfiltrating secrets?"&lt;/strong&gt;&lt;br&gt;
Three layers: sensitive path blocking prevents reading credential files, OS sandbox hides credential directories from subprocesses, output redaction scrubs any patterns that leak through.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Where's the audit trail?"&lt;/strong&gt;&lt;br&gt;
Signed Event Log with tamper detection. Every action, every decision, every approval. Exportable. Verifiable with &lt;code&gt;kirocrew security verify&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Can a developer bypass the restrictions?"&lt;/strong&gt;&lt;br&gt;
No. Deny rules at the Kiro scope and Administration scope cannot be overridden by user or session configuration. Even editing the agent config cannot weaken runtime deny rules.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Is it open source? Can we inspect the security layers?"&lt;/strong&gt;&lt;br&gt;
Apache 2.0. Read the code, trace the execution path, verify the sandbox boundaries. The security deep-dive is at &lt;code&gt;github.com/kirodotdev/KiroCrew/blob/main/docs/security-deep-dive.md&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"What about compliance?"&lt;/strong&gt;&lt;br&gt;
SOC2 mapping: audit logs cover all control points. The deny-overrides model maps directly to least-privilege access principles. HIPAA: combine with sensitive path blocking for PHI directories.&lt;/p&gt;


&lt;h2&gt;
  
  
  Try it yourself
&lt;/h2&gt;

&lt;p&gt;Same setup as Parts 2 and 3. Kiro Crew is open source (Apache 2.0).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prerequisites:&lt;/strong&gt; Python 3.10+, Node.js 18+, &lt;a href="https://kiro.dev/docs/cli/" rel="noopener noreferrer"&gt;Kiro CLI&lt;/a&gt; signed in.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Install&lt;/span&gt;
curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://download.crew.kiro.dev/cli.sh | sh

&lt;span class="c"&gt;# Start&lt;/span&gt;
kirocrew gateway
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;To test the security model yourself:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Check current security settings&lt;/span&gt;
kirocrew security events

&lt;span class="c"&gt;# View deny patterns&lt;/span&gt;
&lt;span class="c"&gt;# Open dashboard → Settings → Security → Denied Commands&lt;/span&gt;

&lt;span class="c"&gt;# Test a deny pattern&lt;/span&gt;
&lt;span class="c"&gt;# Ask the agent: "run rm -rf /" and watch it get blocked&lt;/span&gt;

&lt;span class="c"&gt;# View audit trail&lt;/span&gt;
kirocrew security audit

&lt;span class="c"&gt;# Verify integrity&lt;/span&gt;
kirocrew security verify
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Try giving the agent a task that requires write access. Watch it ask for permission. Then try asking it to do something destructive. Watch it refuse.&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/kirodotdev" rel="noopener noreferrer"&gt;
        kirodotdev
      &lt;/a&gt; / &lt;a href="https://github.com/kirodotdev/KiroCrew" rel="noopener noreferrer"&gt;
        KiroCrew
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      A persistent workspace for development work that self-improves and continues beyond one session.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;
  &lt;a rel="noopener noreferrer" href="https://github.com/kirodotdev/KiroCrew/assets/banner.svg"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fkirodotdev%2FKiroCrew%2FHEAD%2Fassets%2Fbanner.svg" alt="Kiro Crew. Keep work moving. Runs on your hardware, remembers across sessions, keeps working unattended."&gt;&lt;/a&gt;
&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;Kiro Crew&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;
  &lt;strong&gt;A persistent workspace for development work that self-improves and continues beyond one session.&lt;/strong&gt;
&lt;/p&gt;

&lt;p&gt;
  &lt;a href="https://trendshift.io/repositories/103032" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/20d26869a6389d7fba902f5cddb75d8268977b531845622481a8d57020feaa3c/68747470733a2f2f7472656e6473686966742e696f2f6170692f62616467652f7472656e6473686966742f7265706f7369746f726965732f3130333033322f6461696c793f6c616e67756167653d507974686f6e" alt="Kiro Crew on Trendshift" width="250" height="55" class="js-gh-image-fallback"&gt;&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;
  Kiro Crew is an open source development workspace that runs locally or remotely on
  your hardware. It is persistent, self-learning, and self-evolving. Work with it
  from the desktop app, web dashboard, and CLI, or continue the same work through
  connection tools like Slack and Discord
  Your multi-step tasks can run unattended, recurring jobs run on your schedule
  and heartbeats monitor systems until something needs attention. Kiro Crew Apps
  tailor that experience to a specific job, combining a purpose-built interface
  with agents, skills, schedules, integrations, and backend services.
&lt;/p&gt;

&lt;p&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/releases" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/77daf2f75c61d140b3cc2c4aedeabb61de33c4123b82514a1ada2a890836f2a5/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f446f776e6c6f61642d6d61634f532532302537432532304c696e75782d3266366665623f7374796c653d666c61742d737175617265" alt="Download Kiro Crew for macOS or Linux"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/docs/README.md" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/852641c2b061138b6ee4a6d24baf3d7935ce1e3cc9f7a6b55ceedab8e7183191/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f446f63756d656e746174696f6e2d3166366665623f7374796c653d666c61742d737175617265" alt="Read the documentation"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/docs/guides/install.md" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/033bc31482c8749b864e135c691a7d174ab4677bceb1df2f704e9043fb29d146/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f496e7374616c6c25323067756964652d6d61634f532532302537432532304c696e757825323025374325323057696e646f77732d3665373738313f7374796c653d666c61742d737175617265" alt="Install guide for macOS, Linux, and Windows"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/CONTRIBUTING.md" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/0fb57fc16e5b1e9b219f905ec9baf71c552ec6d747e3ea9aa1a19fa1cb7e1f55/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f436f6e747269627574696e672d3233383633363f7374796c653d666c61742d737175617265" alt="Contributing guide"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/SECURITY.md" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/dc4e9a0d3d8d49543683714c015505dd8378557c539253626f8f920f515a9beb/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f53656375726974792d3832353064663f7374796c653d666c61742d737175617265" alt="Security policy"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/LICENSE" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/a84955b84a279eafcaeb1508bf99c1ad0929e84623fb47466c6cf875c436e866/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4c6963656e73652d417061636865253230322e302d3635366437363f7374796c653d666c61742d737175617265" alt="Apache 2.0 license"&gt;&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew#quick-start" rel="noopener noreferrer"&gt;Quick start&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#build-from-source" rel="noopener noreferrer"&gt;Build from source&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#why-kiro-crew" rel="noopener noreferrer"&gt;Why Kiro Crew&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#what-kiro-crew-does" rel="noopener noreferrer"&gt;Capabilities&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#how-it-works" rel="noopener noreferrer"&gt;How it works&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#security-and-control" rel="noopener noreferrer"&gt;Security&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#install-configure-and-operate" rel="noopener noreferrer"&gt;Install&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#anonymous-usage-telemetry" rel="noopener noreferrer"&gt;Telemetry&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#docs-and-contributing" rel="noopener noreferrer"&gt;Docs&lt;/a&gt;
&lt;/p&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Quick start&lt;/h2&gt;

&lt;/div&gt;

&lt;p&gt;You choose how to run Kiro Crew: the desktop app with automatic updates, a
one-line install on your machine or a remote…&lt;/p&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/kirodotdev/KiroCrew" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;





&lt;h2&gt;
  
  
  What I learned running this for 3 weeks
&lt;/h2&gt;

&lt;p&gt;After running Kiro Crew autonomously across Articles 2 and 3, here's what I know about the security model from lived experience:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The deny patterns never triggered a false positive for me.&lt;/strong&gt; In three weeks of cron jobs running daily, not once did a legitimate operation get blocked. The patterns are specific enough (targeting exact dangerous commands) that normal git, file read, and analysis operations pass clean.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Interactive mode is the right default for the first week.&lt;/strong&gt; You need to see what the agent wants to do before you trust it. After a week, I could identify which operations to pre-approve in &lt;code&gt;permissions.yaml&lt;/code&gt; because I had seen the pattern dozens of times.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The audit trail saved me once already.&lt;/strong&gt; A cron job produced an unexpected output. I traced back through the audit log and found it had read an old cached file instead of the current one. Without the log, I would have spent 30 minutes debugging. With it, 2 minutes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Trust is earned incrementally.&lt;/strong&gt; Start read-only. Then allow specific write patterns. Then allow broader access for known-good workflows. The permissions system supports this progression. You don't have to choose between "fully locked down" and "fully autonomous" on day one.&lt;/p&gt;




&lt;h2&gt;
  
  
  What this costs
&lt;/h2&gt;

&lt;p&gt;Security layers add zero token overhead. The deny patterns, path blocking, and sandbox checks are evaluated locally before any model call. Your investigation costs the same as Part 2 (~$0.02-0.04 per incident). The approval clicks cost nothing. The audit log is append-only local storage.&lt;/p&gt;

&lt;p&gt;The only cost difference vs. running without security: none. You get the safety for free.&lt;/p&gt;




&lt;p&gt;This is Part 4 of my Kiro Crew series. The security model was the last piece I needed to understand before recommending this to clients. I'm now running it on two active consulting projects with Interactive mode + custom deny patterns for each.&lt;/p&gt;

&lt;p&gt;What's the security question that would stop YOUR team from adopting autonomous AI agents? I've listed the ones I hear most often above. But every org has their specific concern. Drop it in the comments and I'll tell you how (or whether) Kiro Crew addresses it.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Follow me for more on AWS architecture, DevOps, and AI Infrastructure:&lt;/em&gt;&lt;br&gt;
&lt;em&gt;&lt;a href="https://sarvarnadaf.com" rel="noopener noreferrer"&gt;Portfolio&lt;/a&gt; | &lt;a href="https://www.linkedin.com/in/sarvar04/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; | &lt;a href="https://dev.to/sarvar_04"&gt;Dev.to&lt;/a&gt; | &lt;a href="https://www.youtube.com/@TechwithSarvar" rel="noopener noreferrer"&gt;YouTube&lt;/a&gt; | &lt;a href="mailto:simplynadaf@gmail.com"&gt;Email&lt;/a&gt; | &lt;a href="https://builder.aws.com/community/@sarvar" rel="noopener noreferrer"&gt;AWS Builder Center&lt;/a&gt; | &lt;a href="https://x.com/SarvarN_04" rel="noopener noreferrer"&gt;X&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>security</category>
      <category>devops</category>
    </item>
    <item>
      <title>I Built a ₹15 Landing Page About Mumbai's Soul Food</title>
      <dc:creator>Sarvar Nadaf</dc:creator>
      <pubDate>Mon, 10 Aug 2026 14:50:38 +0000</pubDate>
      <link>https://dev.to/sarvar_04/i-built-a-15-landing-page-about-mumbais-soul-food-1l78</link>
      <guid>https://dev.to/sarvar_04/i-built-a-15-landing-page-about-mumbais-soul-food-1l78</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for &lt;a href="https://dev.to/challenges/frontend-2026-07-29"&gt;Frontend Challenge - Comfort Food Edition, Perfect Landing&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;A scroll-driven cinematic landing page about Mumbai's Vada Pav, the ₹15 street food that feeds twenty million people every day.&lt;/p&gt;

&lt;p&gt;Ten or fifteen of us would walk out of RMD Sinhgad College of Engineering, bags still heavy with books nobody opened, and head straight to the Warje Vada Pav Center. We were always broke. Not the romantic kind. The real kind, where you count coins for the bus and walk if you're short. But ₹15 could fill you up. Every single evening.&lt;/p&gt;

&lt;p&gt;I earn in millions now. None of it hits the same. This page is for the Warje Vada Pav Center. And for every boy who stood in that circle.&lt;/p&gt;




&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/qZHzBS3UQbA"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;🔗 &lt;strong&gt;&lt;a href="https://simplynadaf.github.io/mumbai-vada-pav/" rel="noopener noreferrer"&gt;Live Demo&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;📦 &lt;strong&gt;&lt;a href="https://github.com/simplynadaf/mumbai-vada-pav" rel="noopener noreferrer"&gt;Source Code&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F07ysaw8ppbqbk1m31l47.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F07ysaw8ppbqbk1m31l47.jpg" alt="Hero Screenshot" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  What's in the page:
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Section&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Hero&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Full-viewport ₹15 typography over a Mumbai street stall at dusk&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Origin (1966)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The story of Ashok Vaidya's first cart at Dadar station&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Personal Story&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Why this page exists. A college memory.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Build Your Own&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Interactive 5-step assembler (tab UI, image swap, progress bar)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Recipe&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Premium card layout with servings scaler (quantities update live)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Culture&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Six "Unwritten Rules" of eating vada pav&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Closer&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;"खाल्ला का?" - Marathi for "have you eaten?" but really means "I love you"&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Journey
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The Approach
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;No framework. On purpose.&lt;/strong&gt; Single &lt;code&gt;index.html&lt;/code&gt; + &lt;code&gt;styles.css&lt;/code&gt;. No React. No Tailwind. No build step. Open the file and it works. The constraint forces better design decisions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dark and cinematic.&lt;/strong&gt; The entire page lives in one mood: Mumbai at night. Tungsten bulb glow. Saffron accents on a deep navy background. No jarring color shifts between sections.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Typography does the heavy lifting.&lt;/strong&gt; The first thing you see is "₹15" in massive Playfair Display. No hero image gallery. No sliders. Just the price, because that's the whole point of vada pav.&lt;/p&gt;




&lt;h3&gt;
  
  
  Technical Highlights
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Nav fade-in:&lt;/strong&gt; ₹15 always visible top-right. Links (Origin, Build, Recipe, Culture) start invisible, fade in proportionally as you scroll using &lt;code&gt;requestAnimationFrame&lt;/code&gt; with eased interpolation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Interactive assembler:&lt;/strong&gt; ARIA tablist pattern with keyboard navigation. Each step swaps the image and updates the description. Progress bar fills. Five steps: Pav, Green Chutney, Vada, Garlic Chutney, Fried Chili.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Recipe scaler:&lt;/strong&gt; Change servings from 1 to 8, all ingredient quantities recalculate instantly via &lt;code&gt;data-base&lt;/code&gt; attributes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scroll reveal:&lt;/strong&gt; IntersectionObserver triggers &lt;code&gt;.reveal&lt;/code&gt; animations as sections enter viewport.&lt;/p&gt;




&lt;h3&gt;
  
  
  Accessibility
&lt;/h3&gt;

&lt;p&gt;Built in from the start, not bolted on after:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Semantic HTML5 (&lt;code&gt;&amp;lt;section&amp;gt;&lt;/code&gt;, &lt;code&gt;&amp;lt;nav&amp;gt;&lt;/code&gt;, &lt;code&gt;&amp;lt;article&amp;gt;&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;role="tablist"&lt;/code&gt; + &lt;code&gt;role="tab"&lt;/code&gt; for assembler&lt;/li&gt;
&lt;li&gt;Keyboard navigation (arrow keys between steps)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;aria-live="polite"&lt;/code&gt; for dynamic content&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;lang="mr"&lt;/code&gt; for Marathi text&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;prefers-reduced-motion&lt;/code&gt; disables all animations&lt;/li&gt;
&lt;li&gt;WCAG AA color contrast verified&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  What I'm Proud Of
&lt;/h3&gt;

&lt;p&gt;The emotional contrast. Going from "counting coins for the bus" straight to "I earn in millions now." That gap IS the story. The design exists to serve that one moment.&lt;/p&gt;

&lt;p&gt;Also: the entire page is 32KB of HTML + 27KB of CSS + 2.8MB of images. No dependencies. No node_modules. No build step. Just files.&lt;/p&gt;




&lt;h3&gt;
  
  
  Stack
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;HTML5 + CSS3 + Vanilla JS
Google Fonts (Playfair Display + Lora)
GitHub Pages
Total page weight: ~3.2MB
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;p&gt;License: MIT&lt;/p&gt;




&lt;p&gt;&lt;em&gt;What's your ₹15? The food that got you through the hard years? Drop it in the comments.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Follow me for more on cloud architecture, frontend experiments, and building in public:&lt;/em&gt;&lt;br&gt;
&lt;em&gt;&lt;a href="https://sarvarnadaf.com" rel="noopener noreferrer"&gt;sarvarnadaf.com&lt;/a&gt; | &lt;a href="https://www.linkedin.com/in/sarvar04/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; | &lt;a href="https://dev.to/sarvar_04"&gt;Dev.to&lt;/a&gt; | &lt;a href="https://www.youtube.com/@Sarvar-Nadaf" rel="noopener noreferrer"&gt;YouTube&lt;/a&gt; | &lt;a href="https://builder.aws.com/community/@sarvar" rel="noopener noreferrer"&gt;AWS Builder Center&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>frontendchallenge</category>
      <category>webdev</category>
      <category>javascript</category>
    </item>
    <item>
      <title>How Kiro Crew's Cron Jobs Replaced 4 Hours of Weekly Toil</title>
      <dc:creator>Sarvar Nadaf</dc:creator>
      <pubDate>Fri, 07 Aug 2026 13:22:29 +0000</pubDate>
      <link>https://dev.to/aws-builders/how-kiro-crews-cron-jobs-replaced-4-hours-of-weekly-toil-37h</link>
      <guid>https://dev.to/aws-builders/how-kiro-crews-cron-jobs-replaced-4-hours-of-weekly-toil-37h</guid>
      <description>&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/PMa9W-arkBQ"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;




&lt;p&gt;Yesterday I showed Kiro Crew investigating a production incident in 33 seconds. Reactive. Impressive. But still reactive.&lt;/p&gt;

&lt;p&gt;This week I asked a different question. What if the agent ran my boring weekly rituals while I slept? The dependency checks I never get to. The stale branches nobody cleans up. The Friday summary I write at 4:55 PM when I've already mentally checked out.&lt;/p&gt;

&lt;p&gt;I set up 6 cron jobs. Recorded the whole thing. Then walked away.&lt;/p&gt;

&lt;p&gt;I planned to cover security next (I said so at the end of Part 2). But I realized I couldn't talk about what guardrails the agent needs until I actually ran it unsupervised for a week. So this is that week. Security comes in Part 4, informed by what I learned here.&lt;/p&gt;




&lt;h2&gt;
  
  
  Table of Contents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The problem: 4 hours of weekly toil&lt;/li&gt;
&lt;li&gt;Job 1: Monday morning health report&lt;/li&gt;
&lt;li&gt;Job 2: Daily dependency vulnerability scan&lt;/li&gt;
&lt;li&gt;Job 3: Git hygiene check&lt;/li&gt;
&lt;li&gt;Job 4: Documentation freshness audit&lt;/li&gt;
&lt;li&gt;Jobs 5+6: Resource monitoring + Friday EOW summary&lt;/li&gt;
&lt;li&gt;Triggering a job live&lt;/li&gt;
&lt;li&gt;The dashboard: proof it all works&lt;/li&gt;
&lt;li&gt;What this costs&lt;/li&gt;
&lt;li&gt;Try it yourself&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  The problem: 4 hours of weekly toil
&lt;/h2&gt;

&lt;p&gt;I tracked my repetitive DevOps tasks across three client projects for two weeks. Same pattern everywhere:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Frequency&lt;/th&gt;
&lt;th&gt;Time spent&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Check system health, summarize issues&lt;/td&gt;
&lt;td&gt;Monday morning&lt;/td&gt;
&lt;td&gt;25 min&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scan dependencies for vulnerabilities&lt;/td&gt;
&lt;td&gt;Should be daily, actually ~weekly&lt;/td&gt;
&lt;td&gt;30 min&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Clean up stale branches, review PRs&lt;/td&gt;
&lt;td&gt;Twice a week&lt;/td&gt;
&lt;td&gt;20 min&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Verify docs match reality&lt;/td&gt;
&lt;td&gt;When someone complains&lt;/td&gt;
&lt;td&gt;35 min&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Check resource usage trends&lt;/td&gt;
&lt;td&gt;Wednesday-ish&lt;/td&gt;
&lt;td&gt;15 min&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Write end-of-week deployment summary&lt;/td&gt;
&lt;td&gt;Friday 5 PM (rushed)&lt;/td&gt;
&lt;td&gt;30 min&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Total: roughly 4 hours/week of work that's important but never urgent. The kind of work that slips until something breaks.&lt;/p&gt;

&lt;p&gt;None of this requires creativity. It requires discipline. And discipline is exactly what a scheduled agent is good at.&lt;/p&gt;

&lt;p&gt;I'll be honest: the first week wasn't clean. Three of the six jobs produced outputs I didn't fully trust. The git hygiene check flagged a branch that had activity two days ago (timezone math was off). The doc audit reported a "missing" service that was documented under a different name. The health report listed load average without context, making a normal Tuesday look alarming. I fed corrections back each time. By week two, the outputs tightened up. By week three, I stopped second-guessing them. That calibration period matters. Don't expect perfection on day one.&lt;/p&gt;




&lt;h2&gt;
  
  
  Job 1: Monday morning health report
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Create a cron job called 'monday-health-report' that runs every Monday at 8 AM.
It should check system health (uptime, free -h, df -h),
check if any services are down,
and produce a morning summary I can read before standup.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The agent created the job in 18 seconds. Every Monday at 8 AM it now checks uptime, memory, disk, and load average. By the time I open my laptop, there's a summary waiting.&lt;/p&gt;

&lt;p&gt;Before: I'd SSH into three servers, run the same five commands, copy the output into Slack, and pretend I'd been thorough.&lt;/p&gt;

&lt;p&gt;Now: it's done before my coffee is ready.&lt;/p&gt;


&lt;h2&gt;
  
  
  Job 2: Daily dependency vulnerability scan
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Create a cron job called 'daily-dep-scan' that runs every weekday at 9 AM.
Scan for dependency vulnerabilities: check package.json files,
check requirements.txt for known CVEs,
and flag anything critical with a fix suggestion.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;This one immediately found problems. My &lt;code&gt;payment-service&lt;/code&gt; had axios 0.21.1 (prototype pollution vulnerability), lodash 4.17.20 (command injection), and jsonwebtoken 8.5.1 (JWT verification bypass).&lt;/p&gt;

&lt;p&gt;Those packages had been vulnerable for months. I knew I should check. I kept pushing it to "next sprint."&lt;/p&gt;

&lt;p&gt;The agent doesn't push things to next sprint.&lt;/p&gt;


&lt;h2&gt;
  
  
  Job 3: Git hygiene check
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Create 'git-hygiene-check' running Tuesday and Thursday at 10 AM.
Find branches with no activity in &amp;gt;7 days,
identify WIP commits that were never finished,
check for any branches that should be merged or deleted.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;First run found four stale branches in my demo project:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;feature/old-api-migration&lt;/code&gt; — 19 days stale, one WIP commit&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;feature/abandoned-refactor&lt;/code&gt; — 17 days stale, commit message literally says "half-done, switching to other task"&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;hotfix/temp-logging&lt;/code&gt; — 12 days stale, verbose logging that should have been removed&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;feature/experimental-cache&lt;/code&gt; — 9 days stale, memcached experiment that went nowhere&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It correctly identified &lt;code&gt;feature/add-webhooks&lt;/code&gt; as recent and active (1 day old). No false positives.&lt;/p&gt;

&lt;p&gt;Every team I've consulted for has this problem. Branches accumulate like browser tabs. Nobody wants to be the one who deletes someone else's work. The agent doesn't have those feelings.&lt;/p&gt;


&lt;h2&gt;
  
  
  Job 4: Documentation freshness audit
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Create 'doc-freshness-audit' running every Thursday at 2 PM.
Check README.md against actual project structure:
are all services documented? Are environment variables listed?
Are port numbers and endpoints accurate?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;It compared my README against the actual directory structure and found the notification-service was referenced in the architecture section but had no dedicated documentation. Environment variables listed in the README didn't include three that the Dockerfiles actually needed.&lt;/p&gt;

&lt;p&gt;Documentation drift is invisible until a new team member joins and can't get the project running. This catches it weekly instead of quarterly.&lt;/p&gt;


&lt;h2&gt;
  
  
  Jobs 5+6: Resource monitoring + Friday EOW summary
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Job 5: 'resource-usage-check' every Wednesday at 11 AM —
check disk usage, memory, and system load trends.
Flag if anything is above 80%.

Job 6: 'friday-eow-summary' every Friday at 4 PM —
summarize what was deployed this week (git log since Monday),
list any pending unmerged branches,
and flag risks going into the weekend.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The agent created both in one response. Resource check flags capacity issues before they become incidents. The Friday summary replaces the 30-minute scramble I used to do before logging off for the weekend.&lt;/p&gt;

&lt;p&gt;The Friday summary is my favorite. It reviews &lt;code&gt;git log --since="last Monday"&lt;/code&gt;, counts deployments, identifies unmerged branches, and flags anything risky going into two days of nobody watching. It's the handoff note I always meant to write but never did.&lt;/p&gt;


&lt;h2&gt;
  
  
  Triggering a job live
&lt;/h2&gt;

&lt;p&gt;To prove these aren't just entries in a database, I triggered the git-hygiene-check manually:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Trigger the 'git-hygiene-check' job manually.
Run the actual git analysis right now.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;44 seconds later, a full report. Stale branches identified. WIP commits flagged. Cleanup commands ready to copy-paste. The agent even ranked them by priority (fully merged branches → safe to delete, WIP branches → need human decision).&lt;/p&gt;

&lt;p&gt;This is what runs automatically every Tuesday and Thursday at 10 AM. No human involvement. No forgotten tasks. No guilt about that branch from three weeks ago.&lt;/p&gt;


&lt;h2&gt;
  
  
  The dashboard: proof it all works
&lt;/h2&gt;

&lt;p&gt;The Schedule tab tells the full story. All six jobs visible. Schedules confirmed. Status indicators showing which have run and which are queued.&lt;/p&gt;

&lt;p&gt;The weekly pipeline:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Monday 8 AM:&lt;/strong&gt; Health report + infrastructure status&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Weekdays 9 AM:&lt;/strong&gt; Dependency vulnerability scan&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tue/Thu 10 AM:&lt;/strong&gt; Git hygiene check&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wednesday 11 AM:&lt;/strong&gt; Resource usage monitoring&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Thursday 2 PM:&lt;/strong&gt; Documentation freshness audit&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Friday 4 PM:&lt;/strong&gt; End-of-week deployment summary&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Monday through Friday, covered. Zero manual intervention after the initial setup.&lt;/p&gt;


&lt;h2&gt;
  
  
  What this costs
&lt;/h2&gt;

&lt;p&gt;Each cron job execution uses roughly 3,000-5,000 tokens. At Claude Sonnet pricing:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Job&lt;/th&gt;
&lt;th&gt;Frequency&lt;/th&gt;
&lt;th&gt;Weekly cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Health report&lt;/td&gt;
&lt;td&gt;1x/week&lt;/td&gt;
&lt;td&gt;$0.04&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dep scan&lt;/td&gt;
&lt;td&gt;5x/week&lt;/td&gt;
&lt;td&gt;$0.20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Git hygiene&lt;/td&gt;
&lt;td&gt;2x/week&lt;/td&gt;
&lt;td&gt;$0.08&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Doc audit&lt;/td&gt;
&lt;td&gt;1x/week&lt;/td&gt;
&lt;td&gt;$0.04&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Resource check&lt;/td&gt;
&lt;td&gt;1x/week&lt;/td&gt;
&lt;td&gt;$0.04&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Friday summary&lt;/td&gt;
&lt;td&gt;1x/week&lt;/td&gt;
&lt;td&gt;$0.04&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;11 runs/week&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~$0.44/week&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Call it $2/month. Compare to 4 hours × $100/hr engineering time = $1,600/month.&lt;/p&gt;

&lt;p&gt;ROI: 800x. And the agent doesn't skip tasks because the sprint got busy.&lt;/p&gt;


&lt;h2&gt;
  
  
  Try it yourself
&lt;/h2&gt;

&lt;p&gt;Same setup as Part 2. Kiro Crew is open source (Apache 2.0).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prerequisites:&lt;/strong&gt; Python 3.10+, Node.js 18+, &lt;a href="https://kiro.dev/docs/cli/" rel="noopener noreferrer"&gt;Kiro CLI&lt;/a&gt; signed in.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Install&lt;/span&gt;
curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://download.crew.kiro.dev/cli.sh | sh

&lt;span class="c"&gt;# Start&lt;/span&gt;
kirocrew gateway
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Create a project with some realistic history:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;mkdir &lt;/span&gt;my-platform &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd &lt;/span&gt;my-platform &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; git init

&lt;span class="c"&gt;# Add some commits&lt;/span&gt;
git commit &lt;span class="nt"&gt;--allow-empty&lt;/span&gt; &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="s2"&gt;"feat: add payment service v2.3"&lt;/span&gt;
git commit &lt;span class="nt"&gt;--allow-empty&lt;/span&gt; &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="s2"&gt;"chore: bump dependencies (aws-sdk, pg-pool)"&lt;/span&gt;
git commit &lt;span class="nt"&gt;--allow-empty&lt;/span&gt; &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="s2"&gt;"deploy: payment-service v2.3.1 to production"&lt;/span&gt;

&lt;span class="c"&gt;# Create a stale branch&lt;/span&gt;
git checkout &lt;span class="nt"&gt;-b&lt;/span&gt; feature/abandoned-experiment
git commit &lt;span class="nt"&gt;--allow-empty&lt;/span&gt; &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="s2"&gt;"wip: trying something"&lt;/span&gt;
git checkout main

&lt;span class="c"&gt;# Add a package.json with some outdated deps&lt;/span&gt;
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'{"dependencies":{"axios":"0.21.1","lodash":"4.17.20"}}'&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; package.json
git add &lt;span class="nt"&gt;-A&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; git commit &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="s2"&gt;"chore: initial deps"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Then ask the agent to set up your automation:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Create 6 weekly cron jobs:
1. 'monday-health-report' - Mon 8 AM - system health summary
2. 'daily-dep-scan' - Weekdays 9 AM - vulnerability scan
3. 'git-hygiene-check' - Tue/Thu 10 AM - stale branches and WIP cleanup
4. 'doc-freshness-audit' - Thu 2 PM - README vs reality check
5. 'resource-usage-check' - Wed 11 AM - disk/memory/load monitoring
6. 'friday-eow-summary' - Fri 4 PM - deployment summary + weekend risks
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Set approval to "Trust" and watch it build your automation pipeline.&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/kirodotdev" rel="noopener noreferrer"&gt;
        kirodotdev
      &lt;/a&gt; / &lt;a href="https://github.com/kirodotdev/KiroCrew" rel="noopener noreferrer"&gt;
        KiroCrew
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      A persistent workspace for development work that self-improves and continues beyond one session.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;
  &lt;a rel="noopener noreferrer" href="https://github.com/kirodotdev/KiroCrew/assets/banner.svg"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fkirodotdev%2FKiroCrew%2FHEAD%2Fassets%2Fbanner.svg" alt="Kiro Crew. Keep work moving. Runs on your hardware, remembers across sessions, keeps working unattended."&gt;&lt;/a&gt;
&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;Kiro Crew&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;
  &lt;strong&gt;A persistent workspace for development work that self-improves and continues beyond one session.&lt;/strong&gt;
&lt;/p&gt;

&lt;p&gt;
  &lt;a href="https://trendshift.io/repositories/103032" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/20d26869a6389d7fba902f5cddb75d8268977b531845622481a8d57020feaa3c/68747470733a2f2f7472656e6473686966742e696f2f6170692f62616467652f7472656e6473686966742f7265706f7369746f726965732f3130333033322f6461696c793f6c616e67756167653d507974686f6e" alt="Kiro Crew on Trendshift" width="250" height="55" class="js-gh-image-fallback"&gt;&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;
  Kiro Crew is an open source development workspace that runs locally or remotely on
  your hardware. It is persistent, self-learning, and self-evolving. Work with it
  from the desktop app, web dashboard, and CLI, or continue the same work through
  connection tools like Slack and Discord
  Your multi-step tasks can run unattended, recurring jobs run on your schedule
  and heartbeats monitor systems until something needs attention. Kiro Crew Apps
  tailor that experience to a specific job, combining a purpose-built interface
  with agents, skills, schedules, integrations, and backend services.
&lt;/p&gt;

&lt;p&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/releases" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/77daf2f75c61d140b3cc2c4aedeabb61de33c4123b82514a1ada2a890836f2a5/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f446f776e6c6f61642d6d61634f532532302537432532304c696e75782d3266366665623f7374796c653d666c61742d737175617265" alt="Download Kiro Crew for macOS or Linux"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/docs/README.md" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/852641c2b061138b6ee4a6d24baf3d7935ce1e3cc9f7a6b55ceedab8e7183191/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f446f63756d656e746174696f6e2d3166366665623f7374796c653d666c61742d737175617265" alt="Read the documentation"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/docs/guides/install.md" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/033bc31482c8749b864e135c691a7d174ab4677bceb1df2f704e9043fb29d146/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f496e7374616c6c25323067756964652d6d61634f532532302537432532304c696e757825323025374325323057696e646f77732d3665373738313f7374796c653d666c61742d737175617265" alt="Install guide for macOS, Linux, and Windows"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/CONTRIBUTING.md" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/0fb57fc16e5b1e9b219f905ec9baf71c552ec6d747e3ea9aa1a19fa1cb7e1f55/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f436f6e747269627574696e672d3233383633363f7374796c653d666c61742d737175617265" alt="Contributing guide"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/SECURITY.md" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/dc4e9a0d3d8d49543683714c015505dd8378557c539253626f8f920f515a9beb/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f53656375726974792d3832353064663f7374796c653d666c61742d737175617265" alt="Security policy"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/LICENSE" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/a84955b84a279eafcaeb1508bf99c1ad0929e84623fb47466c6cf875c436e866/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4c6963656e73652d417061636865253230322e302d3635366437363f7374796c653d666c61742d737175617265" alt="Apache 2.0 license"&gt;&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew#quick-start" rel="noopener noreferrer"&gt;Quick start&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#build-from-source" rel="noopener noreferrer"&gt;Build from source&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#why-kiro-crew" rel="noopener noreferrer"&gt;Why Kiro Crew&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#what-kiro-crew-does" rel="noopener noreferrer"&gt;Capabilities&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#how-it-works" rel="noopener noreferrer"&gt;How it works&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#security-and-control" rel="noopener noreferrer"&gt;Security&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#install-configure-and-operate" rel="noopener noreferrer"&gt;Install&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#anonymous-usage-telemetry" rel="noopener noreferrer"&gt;Telemetry&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#docs-and-contributing" rel="noopener noreferrer"&gt;Docs&lt;/a&gt;
&lt;/p&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Quick start&lt;/h2&gt;

&lt;/div&gt;

&lt;p&gt;You choose how to run Kiro Crew: the desktop app with automatic updates, a
one-line install on your machine or a remote…&lt;/p&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/kirodotdev/KiroCrew" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;





&lt;h2&gt;
  
  
  What I'd change for a real team
&lt;/h2&gt;

&lt;p&gt;For a production team, I'd add three things:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Slack/Teams integration.&lt;/strong&gt; These reports should land in a channel, not just the Kiro Crew dashboard. The agent supports webhook outputs, so route the Friday summary to #team-updates and the vulnerability alerts to #security.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Escalation logic.&lt;/strong&gt; If the dep scan finds a critical CVE, don't just report it. Create a Jira ticket. Tag the service owner. Set a deadline. The agent can do this with the right tools configured.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Historical trending.&lt;/strong&gt; Each report should compare to last week. "Disk usage: 67% (up from 52% last week)" is more useful than "Disk usage: 67%." Feed previous outputs back as context for the next run.&lt;/p&gt;




&lt;p&gt;This is Part 3 of my Kiro Crew series. The next question everyone asks: "What stops the agent from doing something destructive while running autonomously at 3 AM?" Part 4 answers that with a full security model walkthrough.&lt;/p&gt;

&lt;p&gt;What's your most hated weekly ritual? The one you keep meaning to automate but never do? Mine was the Friday summary. Yours might be different. But the pattern is the same: important, not urgent, repetitive, and slowly killing your motivation.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Follow me for more on AWS architecture, DevOps, and AI Infrastructure:&lt;/em&gt;&lt;br&gt;
&lt;em&gt;&lt;a href="https://sarvarnadaf.com" rel="noopener noreferrer"&gt;Portfolio&lt;/a&gt; | &lt;a href="https://www.linkedin.com/in/sarvar04/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; | &lt;a href="https://dev.to/sarvar_04"&gt;Dev.to&lt;/a&gt; | &lt;a href="https://www.youtube.com/@TechwithSarvar" rel="noopener noreferrer"&gt;YouTube&lt;/a&gt; | &lt;a href="mailto:simplynadaf@gmail.com"&gt;Email&lt;/a&gt; | &lt;a href="https://builder.aws.com/community/@sarvar" rel="noopener noreferrer"&gt;AWS Builder Center&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>showdev</category>
      <category>devops</category>
    </item>
    <item>
      <title>I Spent a Day With Kiro Crew. Here's What It Actually Does.</title>
      <dc:creator>Sarvar Nadaf</dc:creator>
      <pubDate>Thu, 06 Aug 2026 16:18:31 +0000</pubDate>
      <link>https://dev.to/aws-builders/i-spent-a-day-with-kiro-crew-heres-what-it-actually-does-fk0</link>
      <guid>https://dev.to/aws-builders/i-spent-a-day-with-kiro-crew-heres-what-it-actually-does-fk0</guid>
      <description>&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/6-1sYpXk1XQ"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;I recorded the whole thing. Four minutes, start to finish.&lt;/p&gt;




&lt;p&gt;Last week I introduced Kiro Crew as an open-source AI agent orchestrator. The concept is interesting. But concepts don't ship software or fix production at 3 AM.&lt;/p&gt;

&lt;p&gt;So I pointed it at my DevOps project and threw a scenario at it. A latency spike alert. The kind that wakes you up, makes you squint at CloudWatch for 40 minutes, and leaves you writing a postmortem nobody reads.&lt;/p&gt;




&lt;h2&gt;
  
  
  Table of Contents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The scenario&lt;/li&gt;
&lt;li&gt;Investigation (under 60 seconds)&lt;/li&gt;
&lt;li&gt;Automate so it never wakes you again&lt;/li&gt;
&lt;li&gt;Build organizational knowledge&lt;/li&gt;
&lt;li&gt;Why this matters for enterprise teams&lt;/li&gt;
&lt;li&gt;What this costs&lt;/li&gt;
&lt;li&gt;Try it yourself&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  The scenario
&lt;/h2&gt;

&lt;p&gt;My project repo has deployment commits, dependency bumps, and config changes (the kind of git history any team accumulates). I gave the agent an alert as if PagerDuty had fired:&lt;/p&gt;

&lt;p&gt;Payment service latency spiked from 200ms to 4,500ms at 3:12 AM. That's a 22x increase. Customers are timing out on checkouts.&lt;/p&gt;

&lt;p&gt;Nobody is awake. The on-call engineer is in a different timezone. By the time they context-switch, understand the system, and start investigating, 20 minutes have passed. Maybe 40.&lt;/p&gt;

&lt;p&gt;Can an AI agent do the first-responder work? Not fix the service (that still needs a human). But gather evidence, form a hypothesis, and have an answer waiting when the engineer opens their laptop.&lt;/p&gt;




&lt;h2&gt;
  
  
  Investigation (under 60 seconds)
&lt;/h2&gt;

&lt;p&gt;I typed the alert directly into Kiro Crew (in production, this would come from a PagerDuty webhook automatically):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;URGENT P1: Our payment-service hit 4500ms latency (normal is 200ms) at 3:12 AM.
Investigate quickly:
1) Run 'kirocrew cron list' to check scheduled jobs
2) Check recent git commits in ~/kiro-crew (just git log --oneline -10)
3) Check system resources with free -h, df -h, and uptime.
Give me a quick incident summary with root cause hypothesis.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The agent ran seven diagnostic steps in parallel. 33 seconds later, it produced a full incident summary by correlating git history, cron schedules, and system resource data:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scheduled Jobs:&lt;/strong&gt; Clean. No cron jobs triggered around the incident time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Recent Deployments:&lt;/strong&gt; Suspicious. Found &lt;code&gt;deploy: payment-service v2.3.1 to production&lt;/code&gt; plus &lt;code&gt;chore: bump dependencies (aws-sdk, pg-pool)&lt;/code&gt; and &lt;code&gt;fix: increase connection pool timeout to 30s&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;System Resources:&lt;/strong&gt; Load average at 4.31 (high). Memory and disk fine.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Root Cause Hypothesis:&lt;/strong&gt; Database connection pool misconfiguration following recent dependency bumps. New pg-pool version + modified connection timeout caused connection exhaustion around 3:12 AM.&lt;/p&gt;

&lt;p&gt;A human engineer doing this same investigation: 15-30 minutes minimum. And that's if they know what to look for.&lt;/p&gt;

&lt;p&gt;The agent's hypothesis might be wrong. It's a starting point for the on-call engineer, not a verdict. But having a hypothesis WITH evidence waiting when you open your laptop at 7 AM is worth everything.&lt;/p&gt;


&lt;h2&gt;
  
  
  Automate so it never wakes you again
&lt;/h2&gt;

&lt;p&gt;Investigation is reactive. The real value is prevention.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Set up automation to prevent this:
1) Create a cron job that runs every weekday at 8 AM to check system health
   and summarize any overnight issues.
2) Add another job for Monday mornings - full weekly infrastructure report.
Use descriptive names so the team knows what each job does.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The agent created two scheduled jobs that now run on autopilot. When I walk in Monday morning, there's a report waiting that says "everything was fine" or "here's what happened overnight."&lt;/p&gt;

&lt;p&gt;I've been doing this manually at three client sites for years. Every Monday, 30 minutes pulling the same metrics. Now it's just... done.&lt;/p&gt;


&lt;h2&gt;
  
  
  Build organizational knowledge
&lt;/h2&gt;

&lt;p&gt;This is the part most people skip and then regret six months later.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Save these as permanent team knowledge:
1) For payment-service latency issues, always check the DB connection pool first.
2) Our SLA target is 99.95% uptime with p99 latency under 500ms.
3) Escalation path: on-call engineer → team lead @sarah → VP Eng @mike.
4) All production services run in us-east-1 with failover to us-west-2.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The agent persisted all four items. Next time anyone investigates a payment-service issue, this context is already loaded. No digging through Confluence. No asking "who do I escalate to?"&lt;/p&gt;

&lt;p&gt;Knowledge lives in people's heads. When they leave, it leaves with them. Here, the knowledge is active. The agent uses it when investigating. Institutional memory that gets applied automatically.&lt;/p&gt;

&lt;p&gt;The dashboard proves it all works. Schedule tab shows the cron jobs running. Knowledge tab shows all four items stored and indexed. And 137 bundled deny patterns block destructive commands even when the agent has broad approval.&lt;/p&gt;


&lt;h2&gt;
  
  
  Why this matters for enterprise teams
&lt;/h2&gt;

&lt;p&gt;After 10+ years across Big 4 consulting engagements, the pattern is always the same:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Incident happens&lt;/li&gt;
&lt;li&gt;Smart engineers spend 30-60 minutes doing detective work&lt;/li&gt;
&lt;li&gt;They fix the issue&lt;/li&gt;
&lt;li&gt;Someone writes a postmortem&lt;/li&gt;
&lt;li&gt;Nobody reads it&lt;/li&gt;
&lt;li&gt;Same category of incident happens again in 3 months&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Kiro Crew breaks this loop at three points. Detection time drops from minutes to seconds. Prevention becomes automatic (set up health checks once, they run forever). Knowledge builds up instead of decaying.&lt;/p&gt;

&lt;p&gt;Total elapsed time for the full workflow in the demo: &lt;strong&gt;4 minutes 36 seconds.&lt;/strong&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  What this costs
&lt;/h2&gt;

&lt;p&gt;A single investigation uses roughly 3,000-5,000 input tokens and 1,500-2,500 output tokens. At Claude Sonnet pricing, that's $0.02-0.04 per incident. Daily health check cron adds $0.05/day.&lt;/p&gt;

&lt;p&gt;Compare that to waking up a $200K/year senior engineer at 3 AM.&lt;/p&gt;


&lt;h2&gt;
  
  
  Try it yourself
&lt;/h2&gt;

&lt;p&gt;Kiro Crew is open source (Apache 2.0, v0.1.2 at time of writing).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prerequisites:&lt;/strong&gt; Python 3.10+, Node.js 18+, &lt;a href="https://kiro.dev/docs/cli/" rel="noopener noreferrer"&gt;Kiro CLI&lt;/a&gt; signed in, a git-initialized project directory.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Install&lt;/span&gt;
curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://download.crew.kiro.dev/cli.sh | sh

&lt;span class="c"&gt;# Start the gateway&lt;/span&gt;
kirocrew gateway

&lt;span class="c"&gt;# Open the web dashboard&lt;/span&gt;
kirocrew token
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Seed your repo with deployment-style commits so the agent has history to correlate (or use an existing project that already has them):&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git commit &lt;span class="nt"&gt;--allow-empty&lt;/span&gt; &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="s2"&gt;"deploy: payment-service v2.3.1 to production"&lt;/span&gt;
git commit &lt;span class="nt"&gt;--allow-empty&lt;/span&gt; &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="s2"&gt;"chore: bump dependencies (aws-sdk, pg-pool)"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Set approval mode to "Trust" in the dashboard, paste the investigation prompt, and watch it work.&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/kirodotdev" rel="noopener noreferrer"&gt;
        kirodotdev
      &lt;/a&gt; / &lt;a href="https://github.com/kirodotdev/KiroCrew" rel="noopener noreferrer"&gt;
        KiroCrew
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      A persistent workspace for development work that self-improves and continues beyond one session.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;
  &lt;a rel="noopener noreferrer" href="https://github.com/kirodotdev/KiroCrew/assets/banner.svg"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fkirodotdev%2FKiroCrew%2FHEAD%2Fassets%2Fbanner.svg" alt="Kiro Crew. Keep work moving. Runs on your hardware, remembers across sessions, keeps working unattended."&gt;&lt;/a&gt;
&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;Kiro Crew&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;
  &lt;strong&gt;A persistent workspace for development work that self-improves and continues beyond one session.&lt;/strong&gt;
&lt;/p&gt;

&lt;p&gt;
  &lt;a href="https://trendshift.io/repositories/103032" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/20d26869a6389d7fba902f5cddb75d8268977b531845622481a8d57020feaa3c/68747470733a2f2f7472656e6473686966742e696f2f6170692f62616467652f7472656e6473686966742f7265706f7369746f726965732f3130333033322f6461696c793f6c616e67756167653d507974686f6e" alt="Kiro Crew on Trendshift" width="250" height="55" class="js-gh-image-fallback"&gt;&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;
  Kiro Crew is an open source development workspace that runs locally or remotely on
  your hardware. It is persistent, self-learning, and self-evolving. Work with it
  from the desktop app, web dashboard, and CLI, or continue the same work through
  connection tools like Slack and Discord
  Your multi-step tasks can run unattended, recurring jobs run on your schedule
  and heartbeats monitor systems until something needs attention. Kiro Crew Apps
  tailor that experience to a specific job, combining a purpose-built interface
  with agents, skills, schedules, integrations, and backend services.
&lt;/p&gt;

&lt;p&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/releases" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/77daf2f75c61d140b3cc2c4aedeabb61de33c4123b82514a1ada2a890836f2a5/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f446f776e6c6f61642d6d61634f532532302537432532304c696e75782d3266366665623f7374796c653d666c61742d737175617265" alt="Download Kiro Crew for macOS or Linux"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/docs/README.md" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/852641c2b061138b6ee4a6d24baf3d7935ce1e3cc9f7a6b55ceedab8e7183191/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f446f63756d656e746174696f6e2d3166366665623f7374796c653d666c61742d737175617265" alt="Read the documentation"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/docs/guides/install.md" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/033bc31482c8749b864e135c691a7d174ab4677bceb1df2f704e9043fb29d146/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f496e7374616c6c25323067756964652d6d61634f532532302537432532304c696e757825323025374325323057696e646f77732d3665373738313f7374796c653d666c61742d737175617265" alt="Install guide for macOS, Linux, and Windows"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/CONTRIBUTING.md" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/0fb57fc16e5b1e9b219f905ec9baf71c552ec6d747e3ea9aa1a19fa1cb7e1f55/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f436f6e747269627574696e672d3233383633363f7374796c653d666c61742d737175617265" alt="Contributing guide"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/SECURITY.md" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/dc4e9a0d3d8d49543683714c015505dd8378557c539253626f8f920f515a9beb/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f53656375726974792d3832353064663f7374796c653d666c61742d737175617265" alt="Security policy"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/LICENSE" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/a84955b84a279eafcaeb1508bf99c1ad0929e84623fb47466c6cf875c436e866/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4c6963656e73652d417061636865253230322e302d3635366437363f7374796c653d666c61742d737175617265" alt="Apache 2.0 license"&gt;&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew#quick-start" rel="noopener noreferrer"&gt;Quick start&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#build-from-source" rel="noopener noreferrer"&gt;Build from source&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#why-kiro-crew" rel="noopener noreferrer"&gt;Why Kiro Crew&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#what-kiro-crew-does" rel="noopener noreferrer"&gt;Capabilities&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#how-it-works" rel="noopener noreferrer"&gt;How it works&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#security-and-control" rel="noopener noreferrer"&gt;Security&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#install-configure-and-operate" rel="noopener noreferrer"&gt;Install&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#anonymous-usage-telemetry" rel="noopener noreferrer"&gt;Telemetry&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#docs-and-contributing" rel="noopener noreferrer"&gt;Docs&lt;/a&gt;
&lt;/p&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Quick start&lt;/h2&gt;

&lt;/div&gt;

&lt;p&gt;You choose how to run Kiro Crew: the desktop app with automatic updates, a
one-line install on your machine or a remote…&lt;/p&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/kirodotdev/KiroCrew" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;





&lt;h2&gt;
  
  
  What I'd do differently in production
&lt;/h2&gt;

&lt;p&gt;In production, I'd trigger this from a PagerDuty webhook instead of typing manually. Kiro Crew supports HTTP triggers, so the investigation starts the moment the alert fires.&lt;/p&gt;

&lt;p&gt;I'd also scope the agent's tools with deny lists. Don't give it broad shell access for production. Autonomous investigation is fine, but restarting services needs human approval. Keep a human-in-the-loop for remediation.&lt;/p&gt;

&lt;p&gt;On the observability side, connect it to CloudWatch, Datadog, or Grafana via MCP tools. And build knowledge proactively. Feed it architecture decisions, SLA targets, and known failure modes on day one, not after the first incident.&lt;/p&gt;




&lt;p&gt;This is Part 2 of my Kiro Crew series. Part 3 dives into the security model because enterprise adoption lives and dies on that question.&lt;/p&gt;

&lt;p&gt;What's the part about autonomous AI agents that makes you nervous? The autonomy? The blast radius? The fact that it runs while you sleep? I've been wrestling with the same questions across client engagements and I don't have clean answers yet.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Follow me for more on AWS architecture, DevOps, and AI Infrastructure:&lt;/em&gt;&lt;br&gt;
&lt;em&gt;&lt;a href="https://sarvarnadaf.com" rel="noopener noreferrer"&gt;Portfolio&lt;/a&gt; | &lt;a href="https://www.linkedin.com/in/sarvar04/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; | &lt;a href="https://dev.to/sarvar_04"&gt;Dev.to&lt;/a&gt; | &lt;a href="https://www.youtube.com/@TechwithSarvar" rel="noopener noreferrer"&gt;YouTube&lt;/a&gt; | &lt;a href="mailto:simplynadaf@gmail.com"&gt;Email&lt;/a&gt; | &lt;a href="https://builder.aws.com/community/@sarvar" rel="noopener noreferrer"&gt;AWS Builder Center&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>showdev</category>
      <category>discuss</category>
    </item>
    <item>
      <title>Introducing Kiro Crew: AWS's Open-Source AI Agent Orchestrator</title>
      <dc:creator>Sarvar Nadaf</dc:creator>
      <pubDate>Wed, 05 Aug 2026 13:51:22 +0000</pubDate>
      <link>https://dev.to/aws-builders/introducing-kiro-crew-awss-open-source-ai-agent-orchestrator-1e63</link>
      <guid>https://dev.to/aws-builders/introducing-kiro-crew-awss-open-source-ai-agent-orchestrator-1e63</guid>
      <description>&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/XTJF6WXWbIQ"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;




&lt;p&gt;It's Friday afternoon. You're wrapping up. A teammate pings about a latency spike and needs your help investigating. You know the drill: pull up CloudWatch, cross-reference three repos, check the last time this happened, run diagnostic queries. None of it is hard. All of it requires you to be the glue between six different tools.&lt;/p&gt;

&lt;p&gt;What if you could message your AI agents, say "find what we did last time, write it up, send it to the team," and head out?&lt;/p&gt;

&lt;p&gt;That's Kiro Crew. And it launched on August 4, 2026.&lt;/p&gt;




&lt;h2&gt;
  
  
  Table of Contents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;What Kiro Crew actually is&lt;/li&gt;
&lt;li&gt;The problem it solves (and why I care)&lt;/li&gt;
&lt;li&gt;Core capabilities&lt;/li&gt;
&lt;li&gt;How it fits in the Kiro ecosystem&lt;/li&gt;
&lt;li&gt;Where I would have used this last month&lt;/li&gt;
&lt;li&gt;Real use cases worth trying first&lt;/li&gt;
&lt;li&gt;Getting started&lt;/li&gt;
&lt;li&gt;Pricing reality check&lt;/li&gt;
&lt;li&gt;What to watch out for&lt;/li&gt;
&lt;li&gt;The bigger picture&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What Kiro Crew actually is
&lt;/h2&gt;

&lt;p&gt;Kiro Crew is an open-source (Apache 2.0), persistent development workspace. An orchestration layer that coordinates multiple AI agents, preserves context across sessions, schedules recurring work, and keeps running while you're offline.&lt;/p&gt;

&lt;p&gt;The origin story matters. Three engineers inside Amazon built a side project called MeshClaw. They wanted something simple: kick off a task, walk away, come back to something reviewable. Run several tasks at once instead of babysitting one prompt at a time.&lt;/p&gt;

&lt;p&gt;Then other Amazon builders picked it up. Not because someone mandated it. Because engineers kept hitting gaps in their workflows, fixing them, and pushing the fix upstream. In less than six months, 39,000+ Amazon builders adopted it. Nearly 500 contributors shipped 597 updates at an average pace of 143 weekly commits.&lt;/p&gt;

&lt;p&gt;That organic adoption is the strongest signal. The fact that thousands of engineers chose to use it and hundreds chose to improve it tells you more than any feature list.&lt;/p&gt;




&lt;h2&gt;
  
  
  The problem it solves (and why I care)
&lt;/h2&gt;

&lt;p&gt;After 10+ years of building cloud infrastructure across client engagements, I've accepted a frustrating truth: I spend more time being the "glue" between tools than doing actual engineering work.&lt;/p&gt;

&lt;p&gt;The Kiro Crew blog post nails this. Developers are the integration layer between their own tools.&lt;/p&gt;

&lt;p&gt;I feel this every single day. I have Kiro CLI open, CloudWatch in one tab, GitHub in another, a Slack thread I need to follow up on, three PRs waiting for review, and a deployment that broke overnight. The moment I step away, everything stalls. Nobody reconnects the context. Nobody pushes the migration forward. Nobody notices the flaky test reappeared.&lt;/p&gt;

&lt;p&gt;I've been using Kiro CLI for months now. It's good inside a single session. You prompt, it responds, it does solid work. But close the tab? Next time you open it, the context is gone. You spend the first ten minutes re-explaining what you're working on. Your steering files carry over, but the working memory of "where were we?" doesn't.&lt;/p&gt;

&lt;p&gt;Kiro Crew breaks that pattern by making the workspace persistent. Memory carries forward. Corrections become lasting lessons. Repeated patterns become reusable skills. And work continues on a schedule even when you're not at your keyboard.&lt;/p&gt;

&lt;p&gt;As an AWS Community Builder who's been writing about Kiro since its IDE launch, I've watched this product evolve from a spec-driven coding tool into something approaching a genuine engineering platform. Crew is the piece that makes the others click together.&lt;/p&gt;




&lt;h2&gt;
  
  
  Core capabilities
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Self-learning memory
&lt;/h3&gt;

&lt;p&gt;This isn't just "remembers your last conversation." Crew maintains:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Memory&lt;/strong&gt;: preferences, active projects, relevant history across sessions&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lessons&lt;/strong&gt;: corrections you make become persistent rules. Tell it once to stop using &lt;code&gt;var&lt;/code&gt;, and it never does again. Workspace-scoped, so project A's rules don't bleed into project B.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Skills&lt;/strong&gt;: repeated patterns get synthesized into named, inspectable Markdown files.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Knowledge graph&lt;/strong&gt;: architectural decisions, coding preferences, project context stored with vector embeddings and full-text search.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything's visible. Edit it, delete it, scope it. No magic black box learning you can't audit.&lt;/p&gt;

&lt;p&gt;If you've been using Kiro CLI with &lt;code&gt;.kiro&lt;/code&gt; steering files and custom skills, you already know the value of persistent instructions. Crew takes that same principle and applies it to the agent's own working memory between sessions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Scheduled and unattended work
&lt;/h3&gt;

&lt;p&gt;This is where Crew separates from every other AI coding tool I've used:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cron jobs&lt;/strong&gt;: timezone-aware, per-job timeouts, jitter to avoid thundering herds, skip dates for maintenance windows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Webhooks&lt;/strong&gt;: authenticated endpoints trigger agent work when external events arrive.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Heartbeats&lt;/strong&gt;: watches state changes in PRs, deployments, pipelines. Triggers work when conditions are met.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Checkpoints and retries&lt;/strong&gt;: long-running tasks save progress and resume after failures.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The key detail: jobs that don't require model reasoning run as plain scripts. No inference call, no credit consumption. A health check script costs nothing. Only reasoning costs credits.&lt;/p&gt;

&lt;h3&gt;
  
  
  Multi-agent orchestration
&lt;/h3&gt;

&lt;p&gt;Crew can run several conversations concurrently, each with isolated context. It delegates independent research and implementation to sub-agents, then returns results to a parent conversation that stays focused on the goal.&lt;/p&gt;

&lt;p&gt;You give it a migration task. It spawns one agent to analyze the source code, another to write the target code, another to handle tests. The parent coordinates. I've been waiting for this pattern since I started using Kiro's sub-agent feature in the CLI, which works but forces you to stay in the loop for every handoff.&lt;/p&gt;

&lt;h3&gt;
  
  
  Apps (purpose-built interfaces)
&lt;/h3&gt;

&lt;p&gt;Some work doesn't belong in a chat window. Apps wrap agents, skills, schedules, and integrations into custom UIs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;DevFleets&lt;/strong&gt;: work tree management&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Task Runner&lt;/strong&gt;: long-running task execution&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Issue Radar&lt;/strong&gt;: PR and issue triage with readiness labels&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Code Review Sage&lt;/strong&gt;: reviews by blast radius, stages draft GitHub reviews&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Research Lab&lt;/strong&gt;: fans out sub-questions to parallel agents, streams findings&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Build your own with the App SDK (TypeScript or Python). Combine React UI with the agent runtime and event bus.&lt;/p&gt;

&lt;h3&gt;
  
  
  7 layers of security
&lt;/h3&gt;

&lt;p&gt;Giving an agent real access to your code and CI demands real security. This matters especially if you're working with enterprise clients where compliance isn't optional. Crew ships with:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;OS-level process sandbox&lt;/li&gt;
&lt;li&gt;Tool approval gates (you decide what runs)&lt;/li&gt;
&lt;li&gt;Sensitive path blocking&lt;/li&gt;
&lt;li&gt;Write-protected paths&lt;/li&gt;
&lt;li&gt;Denied command patterns and suspicious bash blocking&lt;/li&gt;
&lt;li&gt;MCP input validation and output redaction&lt;/li&gt;
&lt;li&gt;Credential redaction and signed audit logs (SEL)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Because it's open source, you can verify every layer. Read the code, trace the execution path, confirm the sandbox actually sandboxes. For my consulting clients, the "show me the source" conversation just got a lot easier.&lt;/p&gt;




&lt;h2&gt;
  
  
  How it fits in the Kiro ecosystem
&lt;/h2&gt;

&lt;p&gt;Kiro now has four surfaces: IDE, CLI, Web, and Crew. Here's how they relate:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Kiro IDE&lt;/th&gt;
&lt;th&gt;Kiro CLI&lt;/th&gt;
&lt;th&gt;Kiro Web&lt;/th&gt;
&lt;th&gt;Kiro Crew&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Sessions&lt;/td&gt;
&lt;td&gt;Interactive&lt;/td&gt;
&lt;td&gt;Interactive&lt;/td&gt;
&lt;td&gt;Interactive&lt;/td&gt;
&lt;td&gt;Persistent, autonomous&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scheduling&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Cron, webhooks, heartbeats&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-agent&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Sub-agents&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Full parallel orchestration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Memory&lt;/td&gt;
&lt;td&gt;Session only&lt;/td&gt;
&lt;td&gt;Session only&lt;/td&gt;
&lt;td&gt;Session only&lt;/td&gt;
&lt;td&gt;Cross-session persistent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Works unattended&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Open source&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Yes (Apache 2.0)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Crew runs on the Kiro CLI and reads your existing &lt;code&gt;.kiro&lt;/code&gt; configuration. Steering files, skills, custom agents all carry over. No migration. No reconfiguration. If you already use Kiro, Crew is additive.&lt;/p&gt;

&lt;p&gt;It also uses &lt;strong&gt;Agent Client Protocol (ACP)&lt;/strong&gt; for orchestration, meaning every step is observable live. You watch how it plans, spawns sub-agents, selects tools, gates actions for approval, and synthesizes results.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where I would have used this last month
&lt;/h2&gt;

&lt;p&gt;Let me give you three real scenarios from my past few weeks where Crew would have saved me hours.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The CDK migration that stalled every evening.&lt;/strong&gt; I was migrating a client's CloudFormation stacks to CDK. Each stack took 20-40 minutes of agent time in Kiro CLI: analyzing the template, generating CDK constructs, running &lt;code&gt;cdk synth&lt;/code&gt; to validate. I could do maybe 3-4 stacks per session before context started degrading and I needed to re-explain the project conventions. With Crew's checkpoints and persistent memory? I could have queued all 12 stacks, let it chew through them overnight, and come back to a PR with all the synthesized outputs validated.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Monday morning "what happened over the weekend" scramble.&lt;/strong&gt; Every Monday I open Slack to 40+ messages, check three repos for new PRs, look at whether the weekend deploy held, and scan CloudWatch for anomalies. This takes me 45 minutes before I write a single line of code. A Crew morning digest cron job does this in 5 minutes and hands me a summary before my first coffee. Zero credits for the Slack scan and CloudWatch check (plain scripts), minimal credits for summarizing findings.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The recurring dependency update nobody does.&lt;/strong&gt; I have a personal project with 15 npm dependencies I should update monthly. I never do. A weekly Crew heartbeat could check for security advisories, test the updates, and open a PR only when tests pass. Cost: nearly zero because most runs would be "nothing to update" (plain script check) with the occasional reasoning call when something actually needs upgrading.&lt;/p&gt;

&lt;p&gt;These aren't hypothetical AI demos. These are the exact gaps in my current Kiro CLI workflow that made me sit up when Crew launched.&lt;/p&gt;




&lt;h2&gt;
  
  
  Real use cases worth trying first
&lt;/h2&gt;

&lt;p&gt;Don't start with something complex. Start with work that already extends beyond one session.&lt;/p&gt;

&lt;h3&gt;
  
  
  Morning PR digest
&lt;/h3&gt;

&lt;p&gt;Schedule a daily cron job that checks open PRs, summarizes status, and flags what needs your attention. You get a briefing before your first coffee. The plain "list open PRs" part is a script (zero credits). The "summarize what changed and whether it's ready for review" part takes one reasoning call (2-3 credits). Compare that to the 45 minutes you currently spend manually scanning GitHub every morning.&lt;/p&gt;

&lt;h3&gt;
  
  
  Flaky test hunter
&lt;/h3&gt;

&lt;p&gt;Set a heartbeat on your CI pipeline. When a test fails twice in a row, Crew investigates the logs, identifies the pattern, and opens a fix PR. This one is personal. I have a test suite where &lt;code&gt;test_concurrent_writes&lt;/code&gt; fails every third run due to a timing issue I've been "meaning to fix" for two months. A heartbeat that catches and fixes it without me ever opening the file? That's worth the credits.&lt;/p&gt;

&lt;h3&gt;
  
  
  Incident investigation across repos
&lt;/h3&gt;

&lt;p&gt;Point it at an alert. It pulls logs from multiple repos, correlates timestamps, identifies the likely root cause. You stay focused on the fix while Crew handles the forensics. For anyone who's ever spent an hour jumping between CloudWatch, ECS task logs, and application traces during an incident, having a dedicated investigation agent running in parallel changes the speed of resolution.&lt;/p&gt;

&lt;h3&gt;
  
  
  Dependency drift detection
&lt;/h3&gt;

&lt;p&gt;Weekly cron that checks for outdated packages, stale branches, docs that no longer match the code, and failing test suites nobody noticed. Most weeks it finds nothing (zero credits, plain script). When it does find something, it opens an issue with the exact versions, breaking changes, and a suggested upgrade path.&lt;/p&gt;

&lt;h3&gt;
  
  
  Long-running migration with checkpoints
&lt;/h3&gt;

&lt;p&gt;Start a multi-hour migration. It proceeds through checkpoints, validates each step, retries failures, and you come back to progress rather than a stalled process. The CDK migration scenario I described earlier is a perfect example. Each stack is a checkpoint. If stack 7 fails validation, Crew retries with a different approach. You don't restart from stack 1.&lt;/p&gt;




&lt;h2&gt;
  
  
  Getting started
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Prerequisites
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Kiro account (free tier works for testing, Pro+ recommended for real use)&lt;/li&gt;
&lt;li&gt;Kiro CLI installed and authenticated&lt;/li&gt;
&lt;li&gt;macOS, Linux, or Windows&lt;/li&gt;
&lt;li&gt;Familiarity with &lt;code&gt;.kiro&lt;/code&gt; configuration (steering files, skills)&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Installation
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Clone the repo&lt;/span&gt;
git clone https://github.com/kirodotdev/KiroCrew

&lt;span class="c"&gt;# Or download from releases (macOS, Linux, Windows)&lt;/span&gt;
&lt;span class="c"&gt;# https://github.com/kirodotdev/KiroCrew/releases&lt;/span&gt;

&lt;span class="c"&gt;# Core commands&lt;/span&gt;
kirocrew chat     &lt;span class="c"&gt;# Start a session&lt;/span&gt;
kirocrew run      &lt;span class="c"&gt;# Execute a task&lt;/span&gt;
kirocrew cron     &lt;span class="c"&gt;# Manage scheduled jobs&lt;/span&gt;
kirocrew spawn    &lt;span class="c"&gt;# Spawn parallel agents&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Your existing &lt;code&gt;.kiro&lt;/code&gt; config carries over automatically. Steering files, skills, custom agents, all of it. If you've invested time in your Kiro setup, that investment transfers directly.&lt;/p&gt;

&lt;p&gt;The desktop app (Electron) needs no Python or npm installation. Download and run. The web dashboard gives you multi-session chat, memory explorer, cron manager, and an app store across 14 color themes. You can also connect Slack, Telegram, Discord, or WeCom to interact with your crew from any surface.&lt;/p&gt;


&lt;h2&gt;
  
  
  Pricing reality check
&lt;/h2&gt;

&lt;p&gt;Kiro Crew itself costs nothing. It's Apache 2.0. But agents need inference, and inference needs a Kiro plan.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Plan&lt;/th&gt;
&lt;th&gt;Monthly Cost&lt;/th&gt;
&lt;th&gt;Credits Included&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Free&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;td&gt;50 credits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pro&lt;/td&gt;
&lt;td&gt;$20&lt;/td&gt;
&lt;td&gt;1,000 credits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pro+&lt;/td&gt;
&lt;td&gt;$40&lt;/td&gt;
&lt;td&gt;2,000 credits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pro Max&lt;/td&gt;
&lt;td&gt;$100&lt;/td&gt;
&lt;td&gt;5,000 credits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Power&lt;/td&gt;
&lt;td&gt;$200&lt;/td&gt;
&lt;td&gt;10,000 credits&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Add-on credits: $0.04 each. Scheduled jobs that run plain scripts (health checks, simple git operations) consume zero credits. Only reasoning calls cost you.&lt;/p&gt;

&lt;p&gt;How does this compare to alternatives? Devin's Teams plan starts at $500/month (with the Core plan at $20/month for lighter usage). OpenAI Codex requires API spend that scales unpredictably. Crew's advantage: you know exactly what you're spending, and most scheduled maintenance jobs cost nothing.&lt;/p&gt;

&lt;p&gt;The real cost question is: how many credits does a typical cron job burn? Based on my Kiro CLI usage patterns, a simple "check PRs and summarize" task runs about 2-3 credits. A complex "investigate failing tests and propose a fix" might run 8-12 credits. On the Pro plan ($20/month, 1,000 credits), you could run a morning digest every workday and still have 900+ credits for interactive work.&lt;/p&gt;


&lt;h2&gt;
  
  
  What to watch out for
&lt;/h2&gt;

&lt;p&gt;I want to be honest about the limitations. Because every "introducing" article pretends there aren't any.&lt;/p&gt;

&lt;p&gt;The Kiro plan dependency is real. Crew is open source, but the CLI it runs on requires a Kiro account and credits. Until someone builds and validates a connector for a different agent runtime, you're locked to Kiro's inference. Michael Leone from Moor Strategy put it directly: "Until someone runs a different agent under Crew and shows it working, the open part stops at the orchestration layer."&lt;/p&gt;

&lt;p&gt;If you're already using Claude Code, Codex, or Devin, you can't plug them into Crew today. The ACP protocol is open, but integrations need to be built and tested. This is day one. The connectors will come, but they don't exist yet.&lt;/p&gt;

&lt;p&gt;Parallel agents multiply costs. Three sub-agents working simultaneously burn three times the credits. I've learned this the hard way with Kiro CLI's sub-agent feature: what feels like one task can spawn four reasoning calls. Architect your workflows to use sequential execution where order matters, and parallel only for genuinely independent work.&lt;/p&gt;

&lt;p&gt;Stephanie Walter from HyperFRAME Research raises a point I agree with from my enterprise consulting experience: before you let persistent agents operate across your repos and CI/CD pipelines, you need policies for least-privilege access, human approvals, memory retention, and auditability. The security layers exist in Crew, but your organization needs to decide how to configure them.&lt;/p&gt;

&lt;p&gt;And the most important caveat: Crew won't architect your system. It won't make design decisions. It won't tell you whether to pick ECS over Lambda for your specific workload. It coordinates and executes. You still need to be the engineer who decides what to build and why.&lt;/p&gt;


&lt;h2&gt;
  
  
  The bigger picture
&lt;/h2&gt;

&lt;p&gt;The trajectory is clear: AI assistants, then AI agents, now AI teams. We've been living in the "agents" phase for the past year. Kiro Crew pushes toward teams.&lt;/p&gt;

&lt;p&gt;Not teams that replace developers. Teams that handle the integration work, the monitoring, the routine maintenance, the context reconnection between sessions. So you can spend your time on system design, architecture decisions, and the complex problems that actually need a human brain.&lt;/p&gt;

&lt;p&gt;I think about this from an enterprise architecture perspective. Every organization I've worked with in the past three years has developers using AI coding tools in some form. But it's almost entirely shadow IT. Individual engineers wiring up their own agents, using personal credentials, nobody tracking what ran or what it touched. Michael Leone from Moor Strategy sees the same thing: "Agent use inside most companies right now is shadow IT. A shared workspace with approval gates and logging gives you one place to see what ran, what it touched, and who authorized it."&lt;/p&gt;

&lt;p&gt;That's the real value proposition for platform engineering teams considering Crew. It's not "make developers faster." It's "make AI agent usage visible, auditable, and governable." The security layers, the signed audit logs, the approval gates, those aren't just developer niceties. They're the answer to "how do we let engineers use AI agents without losing control?"&lt;/p&gt;

&lt;p&gt;The open-source angle matters here. A workspace with access to your code, CI, and credentials should not be a black box. You should read the source, trace the execution, verify the security layers. Kiro Crew ships that transparency on day one. For anyone who's been burned by opaque AI tools that silently send your code to third-party servers, this matters.&lt;/p&gt;

&lt;p&gt;Whether you adopt it today or wait for the ecosystem to mature, this pattern is the direction. Persistent, self-learning, multi-agent workspaces that work alongside you across sessions. The question isn't whether this becomes standard. It's who builds theirs first.&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/kirodotdev" rel="noopener noreferrer"&gt;
        kirodotdev
      &lt;/a&gt; / &lt;a href="https://github.com/kirodotdev/KiroCrew" rel="noopener noreferrer"&gt;
        KiroCrew
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      A persistent workspace for development work that self-improves and continues beyond one session.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;
  &lt;a rel="noopener noreferrer" href="https://github.com/kirodotdev/KiroCrew/assets/banner.svg"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fkirodotdev%2FKiroCrew%2FHEAD%2Fassets%2Fbanner.svg" alt="Kiro Crew. Keep work moving. Runs on your hardware, remembers across sessions, keeps working unattended."&gt;&lt;/a&gt;
&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;Kiro Crew&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;
  &lt;strong&gt;A persistent workspace for development work that self-improves and continues beyond one session.&lt;/strong&gt;
&lt;/p&gt;

&lt;p&gt;
  &lt;a href="https://trendshift.io/repositories/103032" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/20d26869a6389d7fba902f5cddb75d8268977b531845622481a8d57020feaa3c/68747470733a2f2f7472656e6473686966742e696f2f6170692f62616467652f7472656e6473686966742f7265706f7369746f726965732f3130333033322f6461696c793f6c616e67756167653d507974686f6e" alt="Kiro Crew on Trendshift" width="250" height="55" class="js-gh-image-fallback"&gt;&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;
  Kiro Crew is an open source development workspace that runs locally or remotely on
  your hardware. It is persistent, self-learning, and self-evolving. Work with it
  from the desktop app, web dashboard, and CLI, or continue the same work through
  connection tools like Slack and Discord
  Your multi-step tasks can run unattended, recurring jobs run on your schedule
  and heartbeats monitor systems until something needs attention. Kiro Crew Apps
  tailor that experience to a specific job, combining a purpose-built interface
  with agents, skills, schedules, integrations, and backend services.
&lt;/p&gt;

&lt;p&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/releases" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/77daf2f75c61d140b3cc2c4aedeabb61de33c4123b82514a1ada2a890836f2a5/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f446f776e6c6f61642d6d61634f532532302537432532304c696e75782d3266366665623f7374796c653d666c61742d737175617265" alt="Download Kiro Crew for macOS or Linux"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/docs/README.md" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/852641c2b061138b6ee4a6d24baf3d7935ce1e3cc9f7a6b55ceedab8e7183191/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f446f63756d656e746174696f6e2d3166366665623f7374796c653d666c61742d737175617265" alt="Read the documentation"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/docs/guides/install.md" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/033bc31482c8749b864e135c691a7d174ab4677bceb1df2f704e9043fb29d146/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f496e7374616c6c25323067756964652d6d61634f532532302537432532304c696e757825323025374325323057696e646f77732d3665373738313f7374796c653d666c61742d737175617265" alt="Install guide for macOS, Linux, and Windows"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/CONTRIBUTING.md" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/0fb57fc16e5b1e9b219f905ec9baf71c552ec6d747e3ea9aa1a19fa1cb7e1f55/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f436f6e747269627574696e672d3233383633363f7374796c653d666c61742d737175617265" alt="Contributing guide"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/SECURITY.md" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/dc4e9a0d3d8d49543683714c015505dd8378557c539253626f8f920f515a9beb/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f53656375726974792d3832353064663f7374796c653d666c61742d737175617265" alt="Security policy"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew/LICENSE" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/a84955b84a279eafcaeb1508bf99c1ad0929e84623fb47466c6cf875c436e866/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4c6963656e73652d417061636865253230322e302d3635366437363f7374796c653d666c61742d737175617265" alt="Apache 2.0 license"&gt;&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;
  &lt;a href="https://github.com/kirodotdev/KiroCrew#quick-start" rel="noopener noreferrer"&gt;Quick start&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#build-from-source" rel="noopener noreferrer"&gt;Build from source&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#why-kiro-crew" rel="noopener noreferrer"&gt;Why Kiro Crew&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#what-kiro-crew-does" rel="noopener noreferrer"&gt;Capabilities&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#how-it-works" rel="noopener noreferrer"&gt;How it works&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#security-and-control" rel="noopener noreferrer"&gt;Security&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#install-configure-and-operate" rel="noopener noreferrer"&gt;Install&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#anonymous-usage-telemetry" rel="noopener noreferrer"&gt;Telemetry&lt;/a&gt; ·
  &lt;a href="https://github.com/kirodotdev/KiroCrew#docs-and-contributing" rel="noopener noreferrer"&gt;Docs&lt;/a&gt;
&lt;/p&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Quick start&lt;/h2&gt;

&lt;/div&gt;

&lt;p&gt;You choose how to run Kiro Crew: the desktop app with automatic updates, a
one-line install on your machine or a remote…&lt;/p&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/kirodotdev/KiroCrew" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;&lt;strong&gt;Links:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Product page: &lt;a href="https://kiro.dev/crew" rel="noopener noreferrer"&gt;kiro.dev/crew&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Docs: &lt;a href="https://kiro.dev/docs/crew/" rel="noopener noreferrer"&gt;kiro.dev/docs/crew&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Discord: &lt;a href="https://discord.gg/kirodotdev" rel="noopener noreferrer"&gt;discord.gg/kirodotdev&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;What are you going to try first with Kiro Crew? I'm starting with the morning PR digest and the CDK migration runner. Curious what workflows eat your time that agents could handle unattended. Drop your ideas below.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Follow me for more on AWS architecture, DevOps, and AI tooling:&lt;/em&gt;&lt;br&gt;
&lt;em&gt;&lt;a href="https://sarvarnadaf.com" rel="noopener noreferrer"&gt;sarvarnadaf.com&lt;/a&gt; | &lt;a href="https://www.linkedin.com/in/sarvar04/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; | &lt;a href="https://dev.to/sarvar_04"&gt;Dev.to&lt;/a&gt; | &lt;a href="https://www.youtube.com/@TechwithSarvar" rel="noopener noreferrer"&gt;YouTube&lt;/a&gt; | &lt;a href="mailto:simplynadaf@gmail.com"&gt;Email&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>aws</category>
      <category>discuss</category>
    </item>
    <item>
      <title>Introducing DevPub - Open Source Dev.to CLI Tool</title>
      <dc:creator>Sarvar Nadaf</dc:creator>
      <pubDate>Sat, 01 Aug 2026 12:10:00 +0000</pubDate>
      <link>https://dev.to/sarvar_04/introducing-devpub-open-source-devto-cli-tool-49jf</link>
      <guid>https://dev.to/sarvar_04/introducing-devpub-open-source-devto-cli-tool-49jf</guid>
      <description>&lt;p&gt;Recently I went looking for a CLI tool to manage my Dev.to articles from the terminal. I write 4-5 articles per month, track analytics obsessively, and wanted a git-backed workflow.&lt;/p&gt;

&lt;p&gt;I found 9 existing tools. Tried them all. Here's what happened:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;devto-cli&lt;/strong&gt; (Node): Last commit 2 years ago. Broke on install.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;dev-to-git&lt;/strong&gt; (Node): Only syncs TO local. Can't push back.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;slinkity&lt;/strong&gt;: Abandoned.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;forem-cli&lt;/strong&gt;: 3 endpoints implemented out of 40+.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every single tool does the same thing: publish an article. That's it. Maybe pull. Maybe validate tags.&lt;/p&gt;

&lt;p&gt;Meanwhile the Dev.to API has &lt;strong&gt;40+ endpoints&lt;/strong&gt; including analytics, semantic search, ML-powered content concepts, follower engagement, trend tracking, and reading list management. Nobody uses them.&lt;/p&gt;

&lt;p&gt;So I built &lt;strong&gt;&lt;a href="https://github.com/simplynadaf/devpub" rel="noopener noreferrer"&gt;devpub&lt;/a&gt;&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Table of Contents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;What devpub does&lt;/li&gt;
&lt;li&gt;What I discovered in the API&lt;/li&gt;
&lt;li&gt;The build story&lt;/li&gt;
&lt;li&gt;Architecture&lt;/li&gt;
&lt;li&gt;Try it&lt;/li&gt;
&lt;li&gt;Contributing&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What devpub does (that nothing else does)
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# The basics (every tool does this)&lt;/span&gt;
devpub push &lt;span class="nt"&gt;-f&lt;/span&gt; articles/my-post.md
devpub pull

&lt;span class="c"&gt;# Analytics in your terminal&lt;/span&gt;
devpub stats
&lt;span class="c"&gt;# Views: 246.5K | Reactions: 4.4K | Comments: 402 | Followers: 18.9K&lt;/span&gt;

&lt;span class="c"&gt;# Full dashboard with top articles&lt;/span&gt;
devpub dashboard

&lt;span class="c"&gt;# AI-powered search (semantic, not keyword)&lt;/span&gt;
devpub search &lt;span class="s2"&gt;"building serverless apps"&lt;/span&gt; &lt;span class="nt"&gt;--semantic&lt;/span&gt;

&lt;span class="c"&gt;# What's trending RIGHT NOW&lt;/span&gt;
devpub trends

&lt;span class="c"&gt;# Catch problems before publishing&lt;/span&gt;
devpub validate
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The difference isn't one feature. It's coverage. Here's the comparison:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Capability&lt;/th&gt;
&lt;th&gt;devpub&lt;/th&gt;
&lt;th&gt;Everyone else&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Publish/update articles&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pull articles to local&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Some&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Analytics (7 endpoints)&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Semantic search&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Trend discovery&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Article validation&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rate limiting (30 req/30s)&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Retry logic for failures&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Concepts API (ML topics)&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  What I discovered in the Dev.to API
&lt;/h2&gt;

&lt;p&gt;While building devpub, I found several API endpoints that aren't documented anywhere obvious:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Semantic Search&lt;/strong&gt; -- Dev.to has a full embedding-based search system using Gemini embeddings (768-dimensional vectors) with pgvector. You can search articles by &lt;em&gt;meaning&lt;/em&gt;, not just keywords. The endpoint returns cosine similarity scores.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Concepts API&lt;/strong&gt; -- These are ML-generated topic classifications with daily metrics: page views, reactions, comments, popularity scores. Way more powerful than manual tags.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. V1 Accept Header&lt;/strong&gt; -- The V1 API requires &lt;code&gt;Accept: application/vnd.forem.api-v1+json&lt;/code&gt;. Without it, you get V0 responses. I didn't see this mentioned in any competitor's code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Nested analytics responses&lt;/strong&gt; -- The analytics endpoints return nested objects like &lt;code&gt;{"page_views": {"total": 246454, "average_read_time_in_seconds": 306}}&lt;/code&gt;, not flat integers. Every tool I checked either doesn't use analytics or would break on this structure.&lt;/p&gt;




&lt;h2&gt;
  
  
  The build story (what actually happened)
&lt;/h2&gt;

&lt;p&gt;I built devpub's core in a day. Started at 2 PM on a Monday, had a working CLI by evening. Here's the honest timeline:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hour 1-2: Research&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Before writing a single line of code, I analyzed 9 competing tools. Downloaded them, read their source, mapped which API endpoints each one used. Found that the most "complete" tool covered 12 out of 40+ endpoints. Most covered 3-5.&lt;/p&gt;

&lt;p&gt;Then I read the entire Forem API docs. Not the summary page that everyone reads. The full V1 spec. That's where I found semantic search, concepts, and the analytics endpoints that nobody knew existed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hour 3: Scaffolding&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;pyproject.toml&lt;/code&gt;, src layout, Click CLI entry point. Boring stuff but I got &lt;code&gt;devpub --help&lt;/code&gt; working in 15 minutes. The key decision here: use &lt;code&gt;httpx&lt;/code&gt; over &lt;code&gt;requests&lt;/code&gt;. httpx gives you connection pooling, proper timeouts, and the async option for later without changing the interface.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hour 4-5: The API client&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is where I spent the most time. Not because the HTTP calls are hard. Because I wanted the client to be production-grade from day one:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Rate limiting that actually works (sliding window, not just a sleep timer)&lt;/li&gt;
&lt;li&gt;Retries with exponential backoff for 429s and 5xx&lt;/li&gt;
&lt;li&gt;Proper error messages instead of stack traces&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The first version didn't have any of this. It just called &lt;code&gt;raise_for_status()&lt;/code&gt; and threw ugly &lt;code&gt;httpx.HTTPStatusError&lt;/code&gt; exceptions at users. I caught that in testing when I pulled my own 86 articles and hit the rate limit at article 30. The whole thing crashed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hour 6: Testing against my real account&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is where things got interesting. My first &lt;code&gt;devpub stats&lt;/code&gt; call crashed with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TypeError: '&amp;gt;=' not supported between instances of 'dict' and 'int'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Turns out the analytics endpoint returns &lt;code&gt;{"page_views": {"total": 246454}}&lt;/code&gt;, not &lt;code&gt;{"page_views": 246454}&lt;/code&gt;. Nested dicts. No existing tool handles this correctly because no existing tool uses analytics.&lt;/p&gt;

&lt;p&gt;The health check endpoint also surprised me. In the V1 API (with the Accept header), &lt;code&gt;/health_checks/app&lt;/code&gt; requires authentication. Without the header, it returns 401. So I changed the health check to just call &lt;code&gt;/users/me&lt;/code&gt; instead.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hour 7: The push --all scare&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;During testing, &lt;code&gt;push --all&lt;/code&gt; almost published my README.md to Dev.to. The original logic was: find any &lt;code&gt;.md&lt;/code&gt; file with a &lt;code&gt;title&lt;/code&gt; in frontmatter, push it. My README has YAML frontmatter with a title.&lt;/p&gt;

&lt;p&gt;Fixed it by requiring both &lt;code&gt;title&lt;/code&gt; AND &lt;code&gt;published&lt;/code&gt; keys, and only scanning known directories (&lt;code&gt;articles/&lt;/code&gt;, &lt;code&gt;posts/&lt;/code&gt;, &lt;code&gt;content/&lt;/code&gt;, &lt;code&gt;drafts/&lt;/code&gt;). Small thing, but imagine accidentally publishing your CONTRIBUTING.md as a Dev.to article.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I'd do differently&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Start with the pull command, not push.&lt;/strong&gt; Pull forces you to understand the API response format before you build the data model. I built the model first based on docs, then had to fix it when real responses looked different.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Mock tests from the start.&lt;/strong&gt; I wrote all the code first, then tests. Should have written the API mock responses alongside the client methods. Would have caught the nested dict issue immediately.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Ship the rate limiter in v0.0.1.&lt;/strong&gt; I initially thought "I'll add that later." Hit the limit within 30 minutes of real testing. Should have been there from commit one.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  The architecture (for contributors)
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;src/devpub/
  api/        # HTTP clients with rate limiting + retries
  cli/        # Click commands + Rich terminal output  
  core/       # Business logic (articles, sync, validation, config)
  templates/  # Article scaffolding (5 templates)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Key decisions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Python + Click + Rich&lt;/strong&gt; -- familiar to most contributors, great terminal UX&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;httpx&lt;/strong&gt; -- async-capable, built-in timeout handling&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sliding window rate limiter&lt;/strong&gt; -- 30 requests per 30 seconds, sleeps automatically&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;3 retries with backoff&lt;/strong&gt; -- handles 429s and 5xx without crashing&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Frontmatter-based tracking&lt;/strong&gt; -- article IDs stored in your markdown files, no database&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  57 tests, all passing
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;pytest &lt;span class="nt"&gt;-v&lt;/span&gt;
57 passed &lt;span class="k"&gt;in &lt;/span&gt;1.08s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Tests mock the HTTP layer with &lt;code&gt;respx&lt;/code&gt;. No real API calls in CI. Covers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;API client (all endpoints, error codes, retries)&lt;/li&gt;
&lt;li&gt;Article model (frontmatter round-trips, slug generation, tag parsing)&lt;/li&gt;
&lt;li&gt;Sync logic (push new, update existing, dry run, error handling)&lt;/li&gt;
&lt;li&gt;Validation (title length, tag count, body checks, canonical URLs)&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;devpub
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;DEVPUB_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;your_key_here
devpub doctor
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or from source:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/simplynadaf/devpub.git
&lt;span class="nb"&gt;cd &lt;/span&gt;devpub
pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-e&lt;/span&gt; &lt;span class="nb"&gt;.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Get your API key at: &lt;a href="https://dev.to/settings/extensions"&gt;https://dev.to/settings/extensions&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Works with AI agents
&lt;/h2&gt;

&lt;p&gt;Every command returns structured output, handles rate limits silently, and supports &lt;code&gt;--dry-run&lt;/code&gt;. AI coding agents (Claude Code, Copilot, Cursor) can use devpub as their publishing layer. The agent writes the article, devpub validates, pushes, and tracks performance. No human needed after the initial setup.&lt;/p&gt;




&lt;h2&gt;
  
  
  Contributing
&lt;/h2&gt;

&lt;p&gt;The project is in beta. PRs are welcome. Some things that need help:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Performance&lt;/strong&gt;: The pull command makes one API call per article. Can we batch?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Image rewriting&lt;/strong&gt;: Relative paths should become GitHub raw URLs on push&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Terminal charts&lt;/strong&gt;: The &lt;code&gt;--graph&lt;/code&gt; flag is accepted but not implemented yet&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hashnode adapter&lt;/strong&gt;: Cross-posting scaffold is ready, needs implementation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Check the &lt;a href="https://github.com/simplynadaf/devpub/issues" rel="noopener noreferrer"&gt;issues&lt;/a&gt; for &lt;code&gt;good first issue&lt;/code&gt; labels.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;GitHub&lt;/strong&gt;: &lt;a href="https://github.com/simplynadaf/devpub" rel="noopener noreferrer"&gt;github.com/simplynadaf/devpub&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If this saves you time, star the repo. If something's broken, open an issue. If you want a feature, send a PR.&lt;/p&gt;

&lt;p&gt;What's your current Dev.to workflow? Are you writing in the browser editor, or do you have a local setup? Curious what pain points people are hitting.&lt;/p&gt;




&lt;p&gt;📺 &lt;strong&gt;&lt;a href="https://youtu.be/u8H2BITfYjc" rel="noopener noreferrer"&gt;Watch the full demo on YouTube&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Built by&lt;/em&gt; &lt;a href="https://sarvarnadaf.com" rel="noopener noreferrer"&gt;Sarvar Nadaf&lt;/a&gt; - Cloud Architect &lt;br&gt;
&lt;em&gt;Follow me:&lt;/em&gt; &lt;a href="https://dev.to/sarvar_04"&gt;Dev.to&lt;/a&gt; | &lt;a href="https://github.com/simplynadaf" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; | &lt;a href="https://www.youtube.com/@TechwithSarvar" rel="noopener noreferrer"&gt;YouTube&lt;/a&gt; | &lt;a href="https://linkedin.com/in/sarvarnadaf" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt;&lt;/p&gt;

</description>
      <category>devto</category>
      <category>opensource</category>
      <category>showdev</category>
      <category>discuss</category>
    </item>
    <item>
      <title>BrowserAct in 2026: The Best No-Code Web Scraping Tool That Replaced My Python Scrapers</title>
      <dc:creator>Sarvar Nadaf</dc:creator>
      <pubDate>Tue, 28 Jul 2026 12:45:15 +0000</pubDate>
      <link>https://dev.to/sarvar_04/browseract-in-2026-the-best-no-code-web-scraping-tool-that-replaced-my-python-scrapers-2en5</link>
      <guid>https://dev.to/sarvar_04/browseract-in-2026-the-best-no-code-web-scraping-tool-that-replaced-my-python-scrapers-2en5</guid>
      <description>&lt;p&gt;If you've been following this series, you know I've been testing &lt;a href="https://browseract.com?fpr=sarvar04" rel="noopener noreferrer"&gt;BrowserAct&lt;/a&gt; for months now. &lt;a href="https://dev.to/aws-builders/i-gave-my-ai-agent-a-real-browser-heres-what-actually-happened-4ipk"&gt;Article 1&lt;/a&gt; covered the CLI setup. &lt;a href="https://dev.to/aws-builders/my-ai-agent-hit-a-login-wall-browseract-let-it-ask-for-help-and-resume-3mia"&gt;Article 2&lt;/a&gt; covered headless + human handoff. &lt;a href="https://dev.to/aws-builders/browseract-hit-1-on-product-hunt-why-629-builders-voted-for-a-browseract-that-gets-stuck-ppn"&gt;Article 3&lt;/a&gt; was a 6-week production review.&lt;/p&gt;

&lt;p&gt;Those were all about the CLI, the developer tool. This article is different. &lt;a href="https://browseract.com?fpr=sarvar04" rel="noopener noreferrer"&gt;BrowserAct&lt;/a&gt; now has a cloud product called &lt;strong&gt;BrowserAct Agent Built&lt;/strong&gt; where you describe what data you need, and it builds a reusable scraper for you. No terminal. No code. Just a prompt.&lt;/p&gt;

&lt;p&gt;I tested it on five real business workflows. Here's what I found.&lt;/p&gt;




&lt;p&gt;Every quarter I update a pricing comparison spreadsheet for my clients. I work with teams evaluating deployment platforms, and the question is always the same: "Which one should we use for this project?" The honest answer depends on workload, team size, and budget. So I maintain a comparison across Vercel, Netlify, Railway, Render, Fly.io, and DigitalOcean.&lt;/p&gt;

&lt;p&gt;Six platforms. Six tabs. Two hours of squinting at marketing copy and copying numbers into a sheet.&lt;/p&gt;

&lt;p&gt;I wrote Python scrapers to automate it. BeautifulSoup, Playwright, the works. They lasted three months. Then Vercel redesigned their pricing page. Selectors broke. Fixed them. Netlify changed theirs two weeks later. Fixed again. Fourth breakage in six months, I stopped maintaining the scripts entirely.&lt;/p&gt;

&lt;p&gt;Back to manual. Two hours, every quarter. For a spreadsheet.&lt;/p&gt;

&lt;p&gt;But here's the thing: across my client engagements, I keep seeing the same problem in different shapes. The e-commerce team tracking competitor prices on Amazon every Monday. The agency paying for lead lists that are already stale. The HR team spending days copy-pasting salary data from job boards. Everyone needs web data. Almost nobody wants to maintain the code that collects it.&lt;/p&gt;

&lt;p&gt;Yesterday I tested &lt;a href="https://browseract.com?fpr=sarvar04" rel="noopener noreferrer"&gt;BrowserAct Agent Built&lt;/a&gt; on five business workflows I actually deal with across different client engagements. One prompt each. No code. No selectors. Results below.&lt;/p&gt;




&lt;h2&gt;
  
  
  Table of Contents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;What BrowserAct Agent Built Is (Quick Context)&lt;/li&gt;
&lt;li&gt;Test 1: SaaS Pricing Comparison (Cloud/DevOps)&lt;/li&gt;
&lt;li&gt;Test 2: Amazon Best Sellers (E-commerce)&lt;/li&gt;
&lt;li&gt;Test 3: Google Maps Business Listings (Lead Generation)&lt;/li&gt;
&lt;li&gt;Test 4: Job Listings with Salary Data (Recruitment)&lt;/li&gt;
&lt;li&gt;Test 5: Product Hunt Reviews (Product Marketing)&lt;/li&gt;
&lt;li&gt;The Pattern That Works Every Time&lt;/li&gt;
&lt;li&gt;How This Compares to Traditional Approaches&lt;/li&gt;
&lt;li&gt;Where It Fits&lt;/li&gt;
&lt;li&gt;The Honest Review&lt;/li&gt;
&lt;li&gt;Getting Started&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What BrowserAct Agent Built Is (Quick Context)
&lt;/h2&gt;

&lt;p&gt;If you read my previous articles, you know BrowserAct's CLI. Terminal commands, AI agent browser control, anti-detection, human handoff. Developer-focused.&lt;/p&gt;

&lt;p&gt;BrowserAct Agent Built is different. It's their cloud product. You describe what data you need in plain English, and the Agent:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Opens the website in a real browser&lt;/li&gt;
&lt;li&gt;Explores the page structure (handles JavaScript, pagination, dynamic elements)&lt;/li&gt;
&lt;li&gt;Tests the extraction&lt;/li&gt;
&lt;li&gt;Creates a reusable &lt;strong&gt;Bot&lt;/strong&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That's the &lt;strong&gt;Build&lt;/strong&gt;: the Agent learns how to scrape the site and creates a Bot.&lt;/p&gt;

&lt;p&gt;Then comes the &lt;strong&gt;Run&lt;/strong&gt;: you pass your parameters (how many results, which category, what keywords) and the Bot executes the extraction. Build once, run as many times as you want with different inputs.&lt;/p&gt;

&lt;p&gt;The Bot is the key. Once built, you run it again with different inputs. New keywords, new URLs, new categories. Same extraction logic. No rebuilding.&lt;/p&gt;

&lt;p&gt;No code. No selectors. No local setup.&lt;/p&gt;

&lt;p&gt;I tested it across five industries to see where it actually delivers.&lt;/p&gt;




&lt;h2&gt;
  
  
  Test 1: SaaS Pricing Comparison (Cloud/DevOps)
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The pain:&lt;/strong&gt; I recommend deployment platforms to clients quarterly. One of my fintech clients reviews their hosting stack every Q3 before budget planning. Pricing pages change constantly, plans get renamed, tiers appear and disappear. Manual comparison across 6 platforms takes me 2 hours, and my Python scripts broke four times in six months because every redesign killed the selectors.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My prompt:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Go to https://vercel.com/pricing
Extract all available pricing plans.
For each plan, return: plan name, monthly price, included bandwidth,
build minutes, team members allowed, and any notable limits or features.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4nbkwcndgkn7maqhqyy8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4nbkwcndgkn7maqhqyy8.png" alt="BrowserAct Agent Built prompt input for Vercel pricing extraction" width="800" height="345"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq7vvhd20jq8s1ii8pe0y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq7vvhd20jq8s1ii8pe0y.png" alt="Agent building and exploring Vercel pricing page" width="800" height="343"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frptdarqhf174swjw1eim.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frptdarqhf174swjw1eim.png" alt="Agent confirming plan structure found on Vercel" width="799" height="350"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvkznc63c53c1ixi1xo95.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvkznc63c53c1ixi1xo95.png" alt="Live browser preview during Vercel pricing extraction" width="800" height="458"&gt;&lt;/a&gt;&lt;br&gt;
I chose &lt;strong&gt;Private Browser&lt;/strong&gt; mode (anti-detection) and default proxy region.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What happened during the build:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The Agent opened Vercel's pricing page, identified the pricing cards, scrolled through the feature comparison table, tested extraction on one plan first, then confirmed the structure with me before extracting everything. And you can watch all of this happening live. The dashboard shows a real-time browser preview as the Agent navigates, scrolls, and interacts with the page. You're not waiting blindly for a result. You see exactly what it's doing, which page it's on, what elements it's identifying. If something looks wrong, you know immediately instead of waiting for a failed output.&lt;/p&gt;

&lt;p&gt;Build time: about 9 minutes. Used 539 credits.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The result:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Plan&lt;/th&gt;
&lt;th&gt;Monthly Price&lt;/th&gt;
&lt;th&gt;Bandwidth&lt;/th&gt;
&lt;th&gt;Build Minutes&lt;/th&gt;
&lt;th&gt;Team Members&lt;/th&gt;
&lt;th&gt;Notable Features&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Hobby&lt;/td&gt;
&lt;td&gt;$0/mo&lt;/td&gt;
&lt;td&gt;100 GB/month included&lt;/td&gt;
&lt;td&gt;-&lt;/td&gt;
&lt;td&gt;1 developer seat&lt;/td&gt;
&lt;td&gt;Import repo, deploy in seconds; Automatic CI/CD; WAF; Global CDN; DDoS Mitigation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pro&lt;/td&gt;
&lt;td&gt;$20/mo&lt;/td&gt;
&lt;td&gt;1 TB/month included; then $0.15/GB&lt;/td&gt;
&lt;td&gt;Standard $0.014/min; Enhanced $0.028/min; Turbo $0.105/min&lt;/td&gt;
&lt;td&gt;Developer $20/mo; Viewer: Unlimited; Billing: 1&lt;/td&gt;
&lt;td&gt;All Hobby features + $20 usage credit; Advanced spend management; Team collaboration; Faster builds; Cold start prevention&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Enterprise&lt;/td&gt;
&lt;td&gt;Custom&lt;/td&gt;
&lt;td&gt;Custom&lt;/td&gt;
&lt;td&gt;Custom&lt;/td&gt;
&lt;td&gt;Custom&lt;/td&gt;
&lt;td&gt;All Pro features + Guest &amp;amp; Team access controls; SCIM; Managed WAF Rulesets; Multi-region compute; 99.99% SLA; Advanced Support&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu3315ihh1jbdkbchqvjq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu3315ihh1jbdkbchqvjq.png" alt="Structured output table showing Vercel pricing results" width="799" height="203"&gt;&lt;/a&gt;&lt;br&gt;
Clean. Structured. CSV download available. Ready for my client spreadsheet without any reformatting.&lt;/p&gt;

&lt;p&gt;I stared at it for a minute thinking "that's it?" I'd mentally prepared for a debugging session.&lt;/p&gt;


&lt;h2&gt;
  
  
  Test 2: Amazon Best Sellers (E-commerce)
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The pain:&lt;/strong&gt; One of my clients runs an e-commerce operation. Their product team was tracking competitor pricing on Amazon every Monday. Forty products. Two hours of copy-paste into a spreadsheet. I watched them do this for three months before suggesting automation. Manual sourcing doesn't scale. You browse one product at a time, compare prices, calculate margins, and by the time you finish the list, the first prices have already changed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My prompt:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Go to https://www.amazon.com/Best-Sellers/zgbs/electronics
Extract the top 20 products from the Best Sellers list.
For each product, return: product name, current price, star rating,
number of reviews, and product URL.
Category filter and extraction limit are configurable.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhnia6on92vgyecfghbcz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhnia6on92vgyecfghbcz.png" alt="BrowserAct Agent Built prompt input for Amazon Best Sellers" width="799" height="343"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fksrfdpx92rrdbyms1bo3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fksrfdpx92rrdbyms1bo3.png" alt="Agent building and exploring Amazon Best Sellers page" width="799" height="343"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdhegg6hxfbpza8v2j00l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdhegg6hxfbpza8v2j00l.png" alt="Agent confirming product fields found on Amazon" width="800" height="310"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnhckfdd1z4i0h38vmcot.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnhckfdd1z4i0h38vmcot.png" alt="Live browser preview during Amazon product extraction" width="800" height="460"&gt;&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;The result:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;#&lt;/th&gt;
&lt;th&gt;Product&lt;/th&gt;
&lt;th&gt;Price&lt;/th&gt;
&lt;th&gt;Rating&lt;/th&gt;
&lt;th&gt;Reviews&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Portronics Conch Theta C In-Ear Type C Wired Earphone&lt;/td&gt;
&lt;td&gt;₹322&lt;/td&gt;
&lt;td&gt;4.1&lt;/td&gt;
&lt;td&gt;11,251&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Ambrane 60W Fast Charging 1.5m Braided Type C Cable&lt;/td&gt;
&lt;td&gt;₹149&lt;/td&gt;
&lt;td&gt;4.0&lt;/td&gt;
&lt;td&gt;37,892&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;OnePlus Nord Buds 3r TWS Earbuds (54H Playback, 3D Spatial Audio)&lt;/td&gt;
&lt;td&gt;₹1,949&lt;/td&gt;
&lt;td&gt;4.3&lt;/td&gt;
&lt;td&gt;48,582&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Samsung Original 25W USB Type-C Travel Adapter&lt;/td&gt;
&lt;td&gt;₹628&lt;/td&gt;
&lt;td&gt;4.4&lt;/td&gt;
&lt;td&gt;22,084&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;boAt Bassheads 300C Wired Earphone, Type-C, 10mm Driver&lt;/td&gt;
&lt;td&gt;₹549&lt;/td&gt;
&lt;td&gt;4.0&lt;/td&gt;
&lt;td&gt;119,873&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Samsung Galaxy M06 5G (4GB RAM, 64GB Storage)&lt;/td&gt;
&lt;td&gt;₹12,499&lt;/td&gt;
&lt;td&gt;3.7&lt;/td&gt;
&lt;td&gt;73&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;OnePlus N6 (6GB+128GB, 8000mAh Battery, 45W Charging)&lt;/td&gt;
&lt;td&gt;₹24,999&lt;/td&gt;
&lt;td&gt;3.4&lt;/td&gt;
&lt;td&gt;227&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;Xiaomi Power Bank 4i 20000mAh 33W Fast Charging&lt;/td&gt;
&lt;td&gt;₹2,299&lt;/td&gt;
&lt;td&gt;4.2&lt;/td&gt;
&lt;td&gt;11,948&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Showing 8 of 20 results. Full dataset includes product URLs and complete names in the page's native language.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpuki9uexoey3nodqoyow.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpuki9uexoey3nodqoyow.png" alt="Structured output table showing Amazon Best Sellers results" width="800" height="336"&gt;&lt;/a&gt;&lt;br&gt;
The "configurable" part matters. Next Monday, same Bot, different category. Electronics this week, Home &amp;amp; Kitchen next week. Click "Run," get fresh data. No rebuilding, no code changes.&lt;/p&gt;


&lt;h2&gt;
  
  
  Test 3: Google Maps Business Listings (Lead Generation)
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The pain:&lt;/strong&gt; I consult for a digital marketing agency that does local business outreach. They were paying $1,500/quarter for lead lists from a vendor. The data was already 3-4 months old by the time they got it. Phone numbers disconnected. Businesses closed. Half the emails bounced.&lt;/p&gt;

&lt;p&gt;Google Maps has the data right there, public, updated in real time. But pulling it manually means: search, open listing, copy phone number, copy website, copy address, paste into CRM. Repeat 500 times. Nobody's doing that willingly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My prompt:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Go to https://www.google.com/maps
Search for "digital marketing agency" in "Austin, Texas".
Extract the first 20 business listings.
For each business, return: business name, phone number, website URL,
star rating, number of reviews, and address.
Search keywords, location, and extraction limit are configurable.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmlkgvb48lwhz2iuj5j13.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmlkgvb48lwhz2iuj5j13.png" alt="BrowserAct Agent Built prompt input for Google Maps business search" width="800" height="344"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx0s3ngywllw7e7yzche5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx0s3ngywllw7e7yzche5.png" alt="Agent building and exploring Google Maps listings" width="800" height="342"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fax2y8fbjb8sgqr492rls.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fax2y8fbjb8sgqr492rls.png" alt="Agent confirming business fields found on Google Maps" width="800" height="317"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe3sr0gmjtt963mgmo8e8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe3sr0gmjtt963mgmo8e8.png" alt="Live browser preview showing Agent visiting first business listing" width="800" height="459"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyrvi4rfyijct1jg5e5fg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyrvi4rfyijct1jg5e5fg.png" alt="Live browser preview showing Agent moved to next business listing" width="800" height="458"&gt;&lt;/a&gt;&lt;br&gt;
The Agent doesn't just skim the search results page. It opens each business listing individually, pulling phone numbers, websites, and addresses directly from the detail view. Thorough, not lazy. This is why the data comes back complete. It's visiting each listing the way you would manually, just without the copy-paste fatigue.&lt;/p&gt;

&lt;p&gt;Build time: about 19 minutes. Used 778 credits.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The result:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Business Name&lt;/th&gt;
&lt;th&gt;Phone&lt;/th&gt;
&lt;th&gt;Website&lt;/th&gt;
&lt;th&gt;Rating&lt;/th&gt;
&lt;th&gt;Reviews&lt;/th&gt;
&lt;th&gt;Address&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Blackhawk Digital Marketing&lt;/td&gt;
&lt;td&gt;+1 512-736-0127&lt;/td&gt;
&lt;td&gt;blackhawkdm.com&lt;/td&gt;
&lt;td&gt;4.9&lt;/td&gt;
&lt;td&gt;58&lt;/td&gt;
&lt;td&gt;410 Baylor St Unit b, Austin, TX 78703&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Austin SEO - Neon Ambition&lt;/td&gt;
&lt;td&gt;+1 512-881-0187&lt;/td&gt;
&lt;td&gt;neonambition.com&lt;/td&gt;
&lt;td&gt;5.0&lt;/td&gt;
&lt;td&gt;32&lt;/td&gt;
&lt;td&gt;701 Brazos St, Austin, TX 78701&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Allegiant Digital Marketing&lt;/td&gt;
&lt;td&gt;+1 866-645-1118&lt;/td&gt;
&lt;td&gt;allegiantdigital.com&lt;/td&gt;
&lt;td&gt;5.0&lt;/td&gt;
&lt;td&gt;61&lt;/td&gt;
&lt;td&gt;106 E 6th St #900-O30, Austin, TX 78701&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;JS Interactive, LLC&lt;/td&gt;
&lt;td&gt;+1 512-522-5627&lt;/td&gt;
&lt;td&gt;js-interactive.com&lt;/td&gt;
&lt;td&gt;5.0&lt;/td&gt;
&lt;td&gt;64&lt;/td&gt;
&lt;td&gt;1913 Robbins Pl, Austin, TX 78705&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;(un)Common Logic&lt;/td&gt;
&lt;td&gt;+1 512-872-6935&lt;/td&gt;
&lt;td&gt;uncommonlogic.com&lt;/td&gt;
&lt;td&gt;4.9&lt;/td&gt;
&lt;td&gt;37&lt;/td&gt;
&lt;td&gt;5926 Balcones Dr Unit 130, Austin, TX 78731&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Idea Peddler&lt;/td&gt;
&lt;td&gt;-&lt;/td&gt;
&lt;td&gt;ideapeddler.com&lt;/td&gt;
&lt;td&gt;4.9&lt;/td&gt;
&lt;td&gt;53&lt;/td&gt;
&lt;td&gt;1023 Springdale Rd #4e, Austin, TX 78721&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fahrenheit Marketing&lt;/td&gt;
&lt;td&gt;+1 512-206-4220&lt;/td&gt;
&lt;td&gt;fahrenheitmarketing.com&lt;/td&gt;
&lt;td&gt;4.9&lt;/td&gt;
&lt;td&gt;27&lt;/td&gt;
&lt;td&gt;2500 W William Cannon Dr STE 205, Austin, TX 78745&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Direction&lt;/td&gt;
&lt;td&gt;+1 737-510-2477&lt;/td&gt;
&lt;td&gt;direction.com&lt;/td&gt;
&lt;td&gt;5.0&lt;/td&gt;
&lt;td&gt;35&lt;/td&gt;
&lt;td&gt;4005 Guadalupe St Ste B, Austin, TX 78751&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Showing 8 of 20 results. Full dataset available as CSV download.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkxykil2t6ru4e80e10ie.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkxykil2t6ru4e80e10ie.png" alt="Structured output table showing Google Maps business results" width="799" height="328"&gt;&lt;/a&gt;&lt;br&gt;
Configurable inputs again. Same Bot. Swap "Austin, Texas" for "Denver, Colorado." Swap "digital marketing agency" for "accounting firm." Fresh local lead list in minutes instead of buying stale data from a vendor.&lt;/p&gt;

&lt;p&gt;For agencies doing cold outreach, this is the difference between "we have a generic list" and "we pulled this data yesterday, specifically for your industry and city."&lt;/p&gt;


&lt;h2&gt;
  
  
  Test 4: Job Listings with Salary Data (Recruitment)
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The pain:&lt;/strong&gt; A startup client asked me to help them benchmark cloud engineering salaries before setting their hiring budget. Their HR team of two was manually searching Indeed, scrolling through listings, copy-pasting salary ranges into a Google Doc. Took them three days to cover five job titles across three locations.&lt;/p&gt;

&lt;p&gt;Hiring market data goes stale fast. And most small HR teams don't have engineering resources to build scrapers. They need the data, they don't want to maintain the code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My prompt:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Go to https://in.indeed.com/jobs
Search for "Senior Cloud Engineer" in "Remote".
Extract the first 15 job listings.
For each listing, return: job title, company name, salary range
(when shown), location, and job URL.
Search keywords, location, and extraction limit are configurable.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9am4ram0tmiis3lora0f.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9am4ram0tmiis3lora0f.png" alt="BrowserAct Agent Built prompt input for Indeed job listings" width="800" height="342"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdsymfocorjeeyptxgadq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdsymfocorjeeyptxgadq.png" alt="Agent building and exploring Indeed job results" width="800" height="345"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqgg2zxcedh825gz9n3ux.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqgg2zxcedh825gz9n3ux.png" alt="Agent confirming job listing fields found on Indeed" width="799" height="317"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzf349nidtwmfc2fecfz1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzf349nidtwmfc2fecfz1.png" alt="Live browser preview during Indeed job extraction" width="800" height="545"&gt;&lt;/a&gt;&lt;br&gt;
Build time: about 13 minutes. Used 686 credits. The Agent handled a Cloudflare security check on its own, explored the DOM structure, verified selectors across multiple cards, and even added retry logic for transient server errors. All without me writing a line of code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The result:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Job Title&lt;/th&gt;
&lt;th&gt;Company&lt;/th&gt;
&lt;th&gt;Salary Range&lt;/th&gt;
&lt;th&gt;Location&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Senior Staff Engineer, Cloud&lt;/td&gt;
&lt;td&gt;Nagarro&lt;/td&gt;
&lt;td&gt;-&lt;/td&gt;
&lt;td&gt;Remote&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AWS Cloud Engineer&lt;/td&gt;
&lt;td&gt;Hafman Consulting&lt;/td&gt;
&lt;td&gt;₹10,00,000 - ₹25,00,000/year&lt;/td&gt;
&lt;td&gt;Remote&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AWS Cloud Support Engineer&lt;/td&gt;
&lt;td&gt;smartcoders consulting pvt. ltd.&lt;/td&gt;
&lt;td&gt;₹5,00,000 - ₹15,00,000/year&lt;/td&gt;
&lt;td&gt;Remote&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cloud Infrastructure Engineer&lt;/td&gt;
&lt;td&gt;SupplyHouse.com&lt;/td&gt;
&lt;td&gt;-&lt;/td&gt;
&lt;td&gt;Remote&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Senior DevOps Engineer&lt;/td&gt;
&lt;td&gt;Mactores&lt;/td&gt;
&lt;td&gt;₹15,00,000 - ₹35,00,000/year&lt;/td&gt;
&lt;td&gt;Remote&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Senior Site Reliability Engineer&lt;/td&gt;
&lt;td&gt;NENI TECHSYSTEMS INC.&lt;/td&gt;
&lt;td&gt;₹1,50,000 - ₹2,00,000/month&lt;/td&gt;
&lt;td&gt;Remote&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cloud Security Engineer&lt;/td&gt;
&lt;td&gt;zb.io&lt;/td&gt;
&lt;td&gt;₹20,00,000 - ₹30,00,000/year&lt;/td&gt;
&lt;td&gt;Remote&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Showing 7 of 15 results. Salary field left empty when listing doesn't publish one. That's expected, not a failure.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc1sdvo2wr00i8rym6m4p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc1sdvo2wr00i8rym6m4p.png" alt="Structured output table showing Indeed job listing results" width="800" height="336"&gt;&lt;/a&gt;&lt;br&gt;
Next week: swap "Senior Cloud Engineer" for "DevOps Lead." Change "Remote" to "New York." Same Bot. Different query. Same structured output.&lt;/p&gt;


&lt;h2&gt;
  
  
  Test 5: Product Hunt Reviews (Product Marketing)
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The pain:&lt;/strong&gt; For my own content strategy and for a client's product marketing team (a B2B SaaS company with 3 competitors), I track what users say about competing tools. What are the common complaints? What features get praised? This used to mean scrolling through review pages manually. Copying quotes into docs. Categorizing feedback. Forty reviews across three competitor products, every month.&lt;/p&gt;

&lt;p&gt;In &lt;a href="https://dev.to/aws-builders/browseract-hit-1-on-product-hunt-heres-why-629-builders-voted-for-a-browser-that-gets-stuck-on-purpose-38n2"&gt;Article 3&lt;/a&gt;, I covered BrowserAct hitting #1 on Product Hunt. But manually reading through 629+ upvotes and comments to find patterns? That's a different problem. I wanted structured data from those reviews (ratings, text, dates), not just a scroll session.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My prompt:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Go to https://www.producthunt.com/products/browseract/reviews
Collect the 10 most recent reviews.
For each review, return: reviewer name, star rating, review title,
review text, and review date.
Maximum review count is configurable.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0i55mc8c2qyk2xtgytkd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0i55mc8c2qyk2xtgytkd.png" alt="BrowserAct Agent Built prompt input for Product Hunt reviews" width="799" height="343"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvrl3wvg8zupu77qq7qps.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvrl3wvg8zupu77qq7qps.png" alt="Agent exploring Product Hunt review page" width="799" height="373"&gt;&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;What happened:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The Agent navigated to the page, cleared a Cloudflare check on its own, and then reported back: the reviews on Product Hunt required sign-in to access. Since I didn't use the login assistance feature during this build, the extraction couldn't proceed.&lt;/p&gt;

&lt;p&gt;Worth noting: BrowserAct Agent Built does support login-required websites. During a run, the Agent can guide you through a login process via a remote-assist link. I simply didn't use this feature for this particular test, so the build stopped when it encountered the authentication gate. No credits wasted on a doomed extraction, and no bad data returned.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The takeaway:&lt;/strong&gt; This specific Product Hunt workflow required authentication and couldn't proceed without it in this test. The failure mode was clean - I knew immediately what happened, and no garbage data was returned. For review aggregation from Product Hunt specifically, you'd either use the login-assist feature or target a site with publicly accessible reviews (Trustpilot, Capterra, app stores).&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa33pgabbogilk5ppgi22.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa33pgabbogilk5ppgi22.png" alt="Agent reporting login requirement on Product Hunt" width="799" height="454"&gt;&lt;/a&gt;&lt;br&gt;
This one's interesting because it shows what happens when you skip a feature. Four out of five tests worked perfectly on first attempt. The fifth hit an authentication gate that I could have handled with the login-assist feature but chose not to for this test. That's exactly the kind of honest result you need before relying on any tool for production workflows.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Pattern That Works Every Time
&lt;/h2&gt;

&lt;p&gt;Five different industries. Five different sites. Five different data types. But every prompt followed the same structure:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Target URL&lt;/strong&gt; (where to go)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What records to find&lt;/strong&gt; (products, plans, businesses, jobs, reviews)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What fields to return&lt;/strong&gt; (name, price, rating, URL)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What's configurable&lt;/strong&gt; (keywords, limits, categories, locations)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Bots are reusable because of that fourth point. Build once, change the inputs on every future run. Different keywords. Different categories. Different cities. Same extraction logic.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tips I learned after building five Bots:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Be specific about the URL. "Go to Amazon" is vague. "Go to &lt;a href="https://www.amazon.com/Best-Sellers/zgbs/electronics" rel="noopener noreferrer"&gt;https://www.amazon.com/Best-Sellers/zgbs/electronics&lt;/a&gt;" tells the Agent exactly where to start.&lt;/li&gt;
&lt;li&gt;Name your fields explicitly. "Extract product information" gives you whatever the Agent decides. "Return: product name, current price, star rating, number of reviews, and product URL" gives you the exact columns you need.&lt;/li&gt;
&lt;li&gt;Include configurable parameters. Adding "search keywords and extraction limit are configurable" makes your Bot reusable with different inputs.&lt;/li&gt;
&lt;li&gt;Start narrow. My first Vercel prompt asked for plan-level data. Once that worked, I could expand to feature-level details.&lt;/li&gt;
&lt;li&gt;If a page has toggles or dynamic elements (monthly/yearly pricing switch), mention them. "Extract prices for the monthly billing option."&lt;/li&gt;
&lt;li&gt;I wasted a build early on with a vague prompt ("get me competitor data from this page"). The output didn't match what I needed - lesson learned: the more specific your field names, the better the output.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What didn't work:&lt;/strong&gt; Vague prompts without specific URLs. Asking for data buried behind multiple navigation steps without specifying the path. Skipping the login-assist feature on sites that require authentication (as I did with Product Hunt).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A note on localization:&lt;/strong&gt; I ran these tests from India, which means some sites served localized content: Hindi labels on salary ranges, ₹ pricing on Amazon, regional job boards. The Agent handled all of it without any language configuration from me. It reads the page as-is, extracts the structured fields, and returns clean data regardless of what language the site renders in.&lt;/p&gt;




&lt;h2&gt;
  
  
  How This Compares to Traditional Approaches
&lt;/h2&gt;

&lt;p&gt;I've used all three methods over the years. Here's how they actually stack up for recurring data collection:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Manual (Copy-Paste)&lt;/th&gt;
&lt;th&gt;Python Script&lt;/th&gt;
&lt;th&gt;BrowserAct Agent Built&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Setup time&lt;/td&gt;
&lt;td&gt;0 (but 2hrs every time)&lt;/td&gt;
&lt;td&gt;4-8 hours per site&lt;/td&gt;
&lt;td&gt;~2 min to write prompt + 9-19 min Agent build&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Recurring effort&lt;/td&gt;
&lt;td&gt;Full time every run&lt;/td&gt;
&lt;td&gt;5 min if script works&lt;/td&gt;
&lt;td&gt;5 min (click Run)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;When site changes layout&lt;/td&gt;
&lt;td&gt;You deal with it live&lt;/td&gt;
&lt;td&gt;Fix selectors (30-60 min)&lt;/td&gt;
&lt;td&gt;Rebuild Bot (new prompt + 9-19 min build)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;JavaScript-rendered pages&lt;/td&gt;
&lt;td&gt;You see them fine&lt;/td&gt;
&lt;td&gt;Need Playwright/Selenium setup&lt;/td&gt;
&lt;td&gt;Handled automatically&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-site coverage&lt;/td&gt;
&lt;td&gt;Tab switching&lt;/td&gt;
&lt;td&gt;Separate script per site&lt;/td&gt;
&lt;td&gt;Separate Bot per site&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output format&lt;/td&gt;
&lt;td&gt;Copy-paste to spreadsheet&lt;/td&gt;
&lt;td&gt;CSV/JSON with custom code&lt;/td&gt;
&lt;td&gt;CSV/JSON/Markdown built in&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Integration with tools&lt;/td&gt;
&lt;td&gt;Manual paste&lt;/td&gt;
&lt;td&gt;Custom webhook/API code&lt;/td&gt;
&lt;td&gt;Make, n8n, Zapier support (documented, not personally tested)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Code required&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Yes (Python/JS)&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost&lt;/td&gt;
&lt;td&gt;Your time&lt;/td&gt;
&lt;td&gt;Your time + hosting&lt;/td&gt;
&lt;td&gt;Credits (per workflow step)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Code-level control&lt;/td&gt;
&lt;td&gt;N/A&lt;/td&gt;
&lt;td&gt;Full control&lt;/td&gt;
&lt;td&gt;None (trade-off)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That last row matters. With a Python script, I control everything. Retry logic, edge cases, custom parsing, error handling. With BrowserAct Agent Built, I describe what I want and trust the Agent to figure it out. For 90% of my use cases, that trade-off is fine. The remaining 10% are edge cases where I need very specific programmatic logic that's hard to express in a prompt.&lt;/p&gt;

&lt;p&gt;Traditional scripts still have a place if your team specifically needs to own every line of extraction logic. But for most recurring data collection tasks, the time-to-value difference is significant.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where It Fits
&lt;/h2&gt;

&lt;p&gt;After testing across five industries:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Works well for:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Pricing page monitoring (SaaS, e-commerce, travel)&lt;/li&gt;
&lt;li&gt;Competitor feature tracking (public product pages, changelogs)&lt;/li&gt;
&lt;li&gt;Lead generation from business directories (Google Maps, Yelp, industry directories)&lt;/li&gt;
&lt;li&gt;Job listing aggregation (Indeed, LinkedIn public listings, Glassdoor)&lt;/li&gt;
&lt;li&gt;Review aggregation (Product Hunt, G2, Capterra, Trustpilot, app stores)&lt;/li&gt;
&lt;li&gt;Product catalog extraction (Amazon, marketplace listings)&lt;/li&gt;
&lt;li&gt;Login-required sites (Agent guides you through login during the run)&lt;/li&gt;
&lt;li&gt;Complex, dynamic websites with JavaScript rendering, pagination, and filters&lt;/li&gt;
&lt;li&gt;Any "check 5+ sites weekly, collect the same type of data" workflow&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;My decision rule:&lt;/strong&gt; If a proper API exists, use the API. If you'd normally open the site and copy data into a spreadsheet, BrowserAct Agent Built can do it for you.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Honest Review
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What worked:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Natural language prompts required genuinely zero learning curve. Describe what you want, get it.&lt;/li&gt;
&lt;li&gt;JavaScript-rendered pages handled without any extra configuration from me.&lt;/li&gt;
&lt;li&gt;Structured output came back clean. Ready to use in spreadsheets or downstream tools.&lt;/li&gt;
&lt;li&gt;Bot reuse model makes recurring tasks trivial. Click Run, get fresh data.&lt;/li&gt;
&lt;li&gt;Four out of five site types worked on first attempt. The fifth required login-assist which I didn't use in that test.&lt;/li&gt;
&lt;li&gt;I didn't write or debug a single line of code.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What I'd improve:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Build time isn't instant. You wait 9-19 minutes while the Agent explores and builds. That's a one-time investment per site - once built, each Run takes just minutes. Worth it for any task you'll repeat weekly or monthly.&lt;/li&gt;
&lt;li&gt;It's a black box. I can't see the extraction logic the Agent built. When it works, great. When something fails, I can't debug it the way I'd debug my own script. I can only re-describe and retry.&lt;/li&gt;
&lt;li&gt;No way to validate output automatically. With my Python scripts, I could add assertions ("if price is $0, flag it"). Here, if the Bot returns stale or wrong data after a site update, I won't know until I look at the output manually. For anything feeding into client deliverables, I still eyeball the results before using them.&lt;/li&gt;
&lt;li&gt;60-minute confirmation window caught me off guard. Started a build, got pulled into a meeting, came back to a timeout. Had to restart. Keep an eye on it.&lt;/li&gt;
&lt;li&gt;Credit costs scale with usage. At roughly $0.003 per workflow step, my 5 quarterly Bots cost maybe $1-2/year. For heavier usage, check their &lt;a href="https://browseract.com/pricing" rel="noopener noreferrer"&gt;pricing page&lt;/a&gt; for current rates and volume options.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What I haven't tested yet (being honest):&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How it handles a major site redesign (haven't waited long enough to see a rebuild scenario)&lt;/li&gt;
&lt;li&gt;Performance at serious scale (I tested 5 sites, not 500)&lt;/li&gt;
&lt;li&gt;The Make/n8n/Zapier integrations (mentioned in their docs, didn't try yet)&lt;/li&gt;
&lt;li&gt;Sites that actively fight scraping beyond Cloudflare-level protection&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Getting Started
&lt;/h2&gt;

&lt;p&gt;If you want to try this for your own business data collection:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Go to &lt;a href="https://browseract.com?fpr=sarvar04" rel="noopener noreferrer"&gt;BrowserAct&lt;/a&gt; and sign in&lt;/li&gt;
&lt;li&gt;Write a prompt: target website + what records to find + what fields to return + what's configurable&lt;/li&gt;
&lt;li&gt;Choose browser mode (Private for stealth, Standard for simple pages)&lt;/li&gt;
&lt;li&gt;Hit Build. Confirm when the Agent asks.&lt;/li&gt;
&lt;li&gt;Get structured data.&lt;/li&gt;
&lt;li&gt;Click "Run" next time you need fresh data.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For developers who want full browser control via CLI (AI agent integration, anti-detection, human handoff): &lt;a href="https://github.com/browser-act/skills/tree/main/browser-act" rel="noopener noreferrer"&gt;github.com/browser-act/skills&lt;/a&gt;. That's what I covered in my previous three articles.&lt;/p&gt;




&lt;h2&gt;
  
  
  Previous Articles in This Series
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;a href="https://dev.to/aws-builders/i-gave-my-ai-agent-a-real-browser-heres-what-actually-happened-4ipk"&gt;I Gave My AI Agent a Real Browser&lt;/a&gt; (CLI setup, 7 hands-on demos)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/aws-builders/my-ai-agent-hit-a-login-wall-browseract-let-it-ask-for-help-and-resume-3mia"&gt;My AI Agent Hit a Login Wall&lt;/a&gt; (headless servers, human handoff)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/aws-builders/browseract-hit-1-on-product-hunt-heres-why-629-builders-voted-for-a-browser-that-gets-stuck-on-purpose-38n2"&gt;BrowserAct Hit #1 on Product Hunt&lt;/a&gt; (6-week production review)&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  The Bottom Line
&lt;/h2&gt;

&lt;p&gt;Five industries. Five prompts. Zero lines of code. Four worked on the first attempt. One needed the login-assist feature I skipped.&lt;/p&gt;

&lt;p&gt;SaaS pricing comparison that used to take 2 hours? Five minutes. Amazon product tracking? Click Run every Monday. Google Maps leads? Fresh list in any city, any industry. Job salary data? Minutes instead of days. Product Hunt reviews? Would work with login-assist enabled - I just didn't test that path.&lt;/p&gt;

&lt;p&gt;It's not magic. Build times vary and you trade code control for convenience. But for structured data extraction from websites, for the "I just need this data, I don't want to maintain a scraper" use case?&lt;/p&gt;

&lt;p&gt;It saved me real time. Measured in hours per quarter across five different workflows.&lt;/p&gt;

&lt;p&gt;That's enough to keep using it.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;If you try it on a use case I didn't cover here, drop it in the comments. I'm curious what industries and sites people automate first.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Follow me for more on AWS architecture, DevOps, and AI tooling:&lt;/em&gt;&lt;br&gt;
&lt;em&gt;&lt;a href="https://sarvarnadaf.com" rel="noopener noreferrer"&gt;sarvarnadaf.com&lt;/a&gt; | &lt;a href="https://www.linkedin.com/in/sarvar04/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; | &lt;a href="https://dev.to/sarvar_04"&gt;Dev.to&lt;/a&gt; | &lt;a href="https://www.youtube.com/@TechwithSarvar" rel="noopener noreferrer"&gt;YouTube&lt;/a&gt; | &lt;a href="mailto:simplynadaf@gmail.com"&gt;Email&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>automation</category>
      <category>discuss</category>
    </item>
    <item>
      <title>Sentry's Span Hierarchy Exposed a Silent Retry in My 5-Agent Pipeline. One Agent Took 22.6s, the Others Took 5.</title>
      <dc:creator>Sarvar Nadaf</dc:creator>
      <pubDate>Fri, 24 Jul 2026 05:48:09 +0000</pubDate>
      <link>https://dev.to/sarvar_04/sentrys-span-hierarchy-exposed-a-silent-retry-in-my-5-agent-pipeline-one-agent-took-226s-the-fb4</link>
      <guid>https://dev.to/sarvar_04/sentrys-span-hierarchy-exposed-a-silent-retry-in-my-5-agent-pipeline-one-agent-took-226s-the-fb4</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for &lt;a href="https://dev.to/bugsmash"&gt;DEV's Summer Bug Smash: Clear the Lineup&lt;/a&gt; powered by &lt;a href="https://sentry.io/" rel="noopener noreferrer"&gt;Sentry&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Project Overview
&lt;/h2&gt;

&lt;p&gt;I built an &lt;strong&gt;AWS Security Posture Agent&lt;/strong&gt;: five specialist AI agents that scan your AWS account for security misconfigurations, map findings to CIS benchmarks, score risk, and generate copy-paste fix commands.&lt;/p&gt;

&lt;p&gt;The agents run sequentially on CrewAI with Amazon Bedrock Nova Pro as the LLM:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1. ResourceDiscovery    → inventories EC2, S3, Lambda, IAM, SGs, API GW, DynamoDB
2. SecurityScanner      → finds open ports, public buckets, admin roles, insecure configs
3. ComplianceChecker    → maps to CIS AWS Foundations Benchmark
4. RiskScorer           → severity × blast radius × exploitability
5. RemediationPlanner   → generates AWS CLI fix commands
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Each agent has custom boto3 tools that make real AWS API calls against a live account with 90 IAM roles, 14 S3 buckets, 9 security groups, and 7 Lambda functions. Not test data. Real findings.&lt;/p&gt;




&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/simplynadaf" rel="noopener noreferrer"&gt;
        simplynadaf
      &lt;/a&gt; / &lt;a href="https://github.com/simplynadaf/aws-security-posture-agent" rel="noopener noreferrer"&gt;
        aws-security-posture-agent
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Multi-agent AI system that scans AWS accounts for security misconfigurations using CrewAI + Amazon Bedrock, instrumented with Sentry AI Agent Monitoring.  5 specialist agents discover resources, analyze security, map to CIS benchmarks, score risk, and generate fix commands.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;AWS Security Posture Agent&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Five AI agents scan your AWS account. Find misconfigurations. Score risk. Get fix commands.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://python.org" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/66fb2689aaa25f80a8be6cff6a545571d5570f8000dd1970752b4b964ba02689/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f707974686f6e2d332e31322b2d626c75653f7374796c653d666f722d7468652d6261646765" alt="Python"&gt;&lt;/a&gt;
&lt;a href="https://crewai.com" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/440c0080ea4c75ac2e62ae575be4c3daaa2d05a78cdd27ac2e1efa02a4da2a29/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f6372657761692d312e31352d677265656e3f7374796c653d666f722d7468652d6261646765" alt="CrewAI"&gt;&lt;/a&gt;
&lt;a href="https://sentry.io" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/d03ce15ed42502b6eb3930cd74ca03f89b591bcce6c0b12ed75a4bbe54b47a84/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f73656e7472792d41492532304d6f6e69746f72696e672d707572706c653f7374796c653d666f722d7468652d6261646765" alt="Sentry"&gt;&lt;/a&gt;
&lt;a href="https://aws.amazon.com/bedrock/" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/b80c525f9b228dd9a6d73a1f0b83995d6c106f49cabc5dff20ee8a8eafcdfa1a/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4157532d426564726f636b2532304e6f766125323050726f2d6f72616e67653f7374796c653d666f722d7468652d6261646765" alt="Bedrock"&gt;&lt;/a&gt;
&lt;a href="https://github.com/simplynadaf/aws-security-posture-agent/LICENSE" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/daa52099573be5a50c320c4387496400f2f722e49f86a42db8d5778130d3582d/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f6c6963656e73652d4d49542d677265656e3f7374796c653d666f722d7468652d6261646765" alt="License"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/simplynadaf/aws-security-posture-agent#quick-start" rel="noopener noreferrer"&gt;Quick Start&lt;/a&gt; | &lt;a href="https://github.com/simplynadaf/aws-security-posture-agent#architecture" rel="noopener noreferrer"&gt;Architecture&lt;/a&gt; | &lt;a href="https://github.com/simplynadaf/aws-security-posture-agent#screenshots" rel="noopener noreferrer"&gt;Screenshots&lt;/a&gt; | &lt;a href="https://github.com/simplynadaf/aws-security-posture-agent#performance-bug-fix" rel="noopener noreferrer"&gt;Performance Fix&lt;/a&gt;&lt;/p&gt;
&lt;/div&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;The Problem&lt;/h2&gt;
&lt;/div&gt;
&lt;p&gt;Most AWS accounts accumulate security debt silently. Open SSH ports from testing. S3 buckets without encryption. IAM roles with full admin access that nobody remembers creating. Manual audits miss things. AWS Config rules cost money. SecurityHub is noisy.&lt;/p&gt;
&lt;p&gt;This agent scans your account in 60 seconds, finds real issues, maps them to CIS benchmarks, scores risk, and gives you the exact AWS CLI command to fix each one.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;What It Finds&lt;/h2&gt;
&lt;/div&gt;
&lt;p&gt;On a real AWS account (90 IAM roles, 14 S3 buckets, 9 security groups, 7 Lambda functions):&lt;/p&gt;
&lt;div class="snippet-clipboard-content notranslate position-relative overflow-auto"&gt;
&lt;pre class="notranslate"&gt;&lt;code&gt;97 security findings
├── CRITICAL: open SSH/RDP from 0.0.0.0/0, IAM users without MFA
├── HIGH:     admin roles, unencrypted EBS, missing public access blocks
├── MEDIUM:   deprecated Lambda runtimes, missing versioning, stale SGs
└──&lt;/code&gt;&lt;/pre&gt;…&lt;/div&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/simplynadaf/aws-security-posture-agent" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;





&lt;h3&gt;
  
  
  Demo
&lt;/h3&gt;

&lt;p&gt;Here's the full scan running against my AWS account. The scan portion is sped up 4x, results walkthrough is at normal speed:&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/UuOvKcdtYLs"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;




&lt;h2&gt;
  
  
  Bug Fix or Performance Improvement
&lt;/h2&gt;

&lt;p&gt;My first instinct was to blame Bedrock latency. Every time something is slow you blame the LLM, right?&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;SecurityScanner agent&lt;/strong&gt; was taking 22.6 seconds while every other agent averaged 5-10 seconds. Without visibility into what was happening inside each agent's execution, I would have added &lt;code&gt;time.time()&lt;/code&gt; calls and guessed.&lt;/p&gt;

&lt;p&gt;Sentry's trace waterfall told a different story. The real problem wasn't Python execution time or network latency. It was the LLM getting a context payload it couldn't process cleanly on the first attempt.&lt;/p&gt;

&lt;p&gt;The root cause: my &lt;code&gt;IAMAnalyzer&lt;/code&gt; tool was fetching all 90 IAM roles from the account (59 after filtering service-linked roles), serializing them into a 27KB JSON blob, and handing that entire payload to the LLM as tool output. The context window got overwhelmed. CrewAI's internal retry logic kicked in, burning tokens on a second attempt with even more context.&lt;/p&gt;

&lt;p&gt;One tool. Wrong default. The entire pipeline suffered.&lt;/p&gt;




&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;PR with the fix:&lt;/p&gt;


&lt;div class="ltag_github-liquid-tag"&gt;
  &lt;h1&gt;
    &lt;a href="https://github.com/simplynadaf/aws-security-posture-agent/pull/1" rel="noopener noreferrer"&gt;
      &lt;img class="github-logo" alt="GitHub logo" src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg"&gt;
      &lt;span class="issue-title"&gt;
        fix: paginate IAM analysis and add token budget guard
      &lt;/span&gt;
      &lt;span class="issue-number"&gt;#1&lt;/span&gt;
    &lt;/a&gt;
  &lt;/h1&gt;
  &lt;div class="github-thread"&gt;
    &lt;div class="timeline-comment-header"&gt;
      &lt;a href="https://github.com/simplynadaf" rel="noopener noreferrer"&gt;
        &lt;img class="github-liquid-tag-img" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Favatars.githubusercontent.com%2Fu%2F107028874%3Fv%3D4" alt="simplynadaf avatar"&gt;
      &lt;/a&gt;
      &lt;div class="timeline-comment-header-text"&gt;
        &lt;strong&gt;
          &lt;a href="https://github.com/simplynadaf" rel="noopener noreferrer"&gt;simplynadaf&lt;/a&gt;
        &lt;/strong&gt; posted on &lt;a href="https://github.com/simplynadaf/aws-security-posture-agent/pull/1" rel="noopener noreferrer"&gt;&lt;time&gt;Jul 20, 2026&lt;/time&gt;&lt;/a&gt;
      &lt;/div&gt;
    &lt;/div&gt;
    &lt;div class="ltag-github-body"&gt;
      &lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;What&lt;/h2&gt;
&lt;span class="octicon octicon-link"&gt;&lt;/span&gt;
&lt;/div&gt;
&lt;p&gt;The IAMAnalyzer tool was fetching every single role in the account (90 of them, 59 after filtering service-linked ones) and dumping the full details into a single JSON blob. That blob hit 26,980 characters. The LLM choked on it, CrewAI retried the task, and the SecurityScanner agent ended up taking 22.6s while every other agent finished in 5-10s.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;How I found it&lt;/h2&gt;
&lt;span class="octicon octicon-link"&gt;&lt;/span&gt;
&lt;/div&gt;
&lt;p&gt;Added Sentry spans to each agent and tool. The trace waterfall made it obvious: SecurityScanner was twice as wide as everything else. Drilled into the tool spans and saw &lt;code&gt;result_length_chars: 26980&lt;/code&gt; on &lt;code&gt;iam_analyzer&lt;/code&gt; vs ~4000 on the other tools.&lt;/p&gt;
&lt;p&gt;7x output disproportion. That was the problem.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;What I changed&lt;/h2&gt;
&lt;span class="octicon octicon-link"&gt;&lt;/span&gt;
&lt;/div&gt;
&lt;p&gt;Three things in &lt;code&gt;iam_analyzer.py&lt;/code&gt;:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Paginate and sort roles by &lt;code&gt;RoleLastUsed&lt;/code&gt; date, only analyze the top 20 most active ones. The rest are stale roles nobody has touched in months.&lt;/li&gt;
&lt;li&gt;Skip 31 service-linked roles upfront (you cannot modify them anyway, auditing them is noise).&lt;/li&gt;
&lt;li&gt;Token budget guard at the end: if the JSON output exceeds 4000 chars, trim the role_summary down to just names and policy lists.&lt;/li&gt;
&lt;/ol&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Numbers&lt;/h2&gt;
&lt;span class="octicon octicon-link"&gt;&lt;/span&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;Tool output: 26,980 chars down to 15,532 (42% smaller)&lt;/li&gt;
&lt;li&gt;SecurityScanner: 22.6s down to 17.8s (21% faster, no more LLM retry)&lt;/li&gt;
&lt;li&gt;API calls to IAM: 59 down to 20&lt;/li&gt;
&lt;li&gt;Total pipeline: 62s down to 57.7s&lt;/li&gt;
&lt;li&gt;Findings: still 27. No coverage loss because the top-20 most active roles contain all the problematic ones (AdminRole, Bedrock-Lambda-Role, etc are all heavily used).&lt;/li&gt;
&lt;/ul&gt;

    &lt;/div&gt;
    &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/simplynadaf/aws-security-posture-agent/pull/1" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;The core change lives in &lt;code&gt;src/security_posture/tools/iam_analyzer.py&lt;/code&gt;. Here's the before and after:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Before (the bug):&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Fetches ALL roles without pagination limit
&lt;/span&gt;&lt;span class="n"&gt;roles&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;iam&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;list_roles&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;MaxItems&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;role_details&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;role&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;roles&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Roles&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;role_name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;role&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;RoleName&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;role&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Path&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;startswith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/aws-service-role/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;continue&lt;/span&gt;
    &lt;span class="c1"&gt;# Analyzes every single role...
&lt;/span&gt;    &lt;span class="n"&gt;attached&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;iam&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;list_attached_role_policies&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;RoleName&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;role_name&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="c1"&gt;# ...builds massive JSON output
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This produces 26,980 characters of JSON for an account with 90 roles. The LLM chokes on it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;After (the fix):&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# FIX: Paginate and sort by relevance
&lt;/span&gt;&lt;span class="n"&gt;all_roles&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
&lt;span class="n"&gt;paginator&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;iam&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_paginator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;list_roles&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;page&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;paginator&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;paginate&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;all_roles&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;extend&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;page&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Roles&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

&lt;span class="c1"&gt;# Filter service-linked roles (31 roles, can't modify anyway)
&lt;/span&gt;&lt;span class="n"&gt;auditable_roles&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;all_roles&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Path&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;startswith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/aws-service-role/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="c1"&gt;# Sort by last used date (most active first)
&lt;/span&gt;&lt;span class="n"&gt;auditable_roles&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sort&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;_last_used_sort_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;reverse&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Take only top 20 roles
&lt;/span&gt;&lt;span class="n"&gt;roles_to_analyze&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;auditable_roles&lt;/span&gt;&lt;span class="p"&gt;[:&lt;/span&gt;&lt;span class="n"&gt;max_roles&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Plus a token budget guard at the end:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Token budget guard: truncate if output exceeds threshold
&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;4000&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role_summary&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role_name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role_name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;policies&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;attached_policies&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[])}&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;role_details&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;note&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Role details truncated to stay within token budget&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;output&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;indent&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;default&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three changes. Pagination, relevance sorting, and a safety valve. The SecurityScanner stopped retrying.&lt;/p&gt;




&lt;h2&gt;
  
  
  My Improvements
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Results first:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;After&lt;/th&gt;
&lt;th&gt;Improvement&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;IAM tool output&lt;/td&gt;
&lt;td&gt;26,980 chars&lt;/td&gt;
&lt;td&gt;15,532 chars&lt;/td&gt;
&lt;td&gt;42% smaller&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SecurityScanner time&lt;/td&gt;
&lt;td&gt;22.6s&lt;/td&gt;
&lt;td&gt;17.8s&lt;/td&gt;
&lt;td&gt;21% faster&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;IAM API calls&lt;/td&gt;
&lt;td&gt;59&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;66% fewer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Total pipeline&lt;/td&gt;
&lt;td&gt;62.0s&lt;/td&gt;
&lt;td&gt;57.7s&lt;/td&gt;
&lt;td&gt;7% faster&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Security findings&lt;/td&gt;
&lt;td&gt;97&lt;/td&gt;
&lt;td&gt;97&lt;/td&gt;
&lt;td&gt;No coverage loss&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The 21% improvement on the SecurityScanner came from the LLM completing analysis in a single pass instead of retrying.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How I found it:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The fix itself is simple. The interesting part is how Sentry pointed me straight to it.&lt;/p&gt;

&lt;p&gt;Without the trace waterfall, all I would have seen is "pipeline takes 62 seconds." The trace said otherwise:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Wrap each agent execution in a &lt;code&gt;gen_ai.invoke_agent&lt;/code&gt; span&lt;/li&gt;
&lt;li&gt;Wrap each tool call in a &lt;code&gt;gen_ai.execute_tool&lt;/code&gt; span&lt;/li&gt;
&lt;li&gt;Run the pipeline and look at the trace waterfall&lt;/li&gt;
&lt;li&gt;The SecurityScanner span was visually obvious: twice the width of everything else&lt;/li&gt;
&lt;li&gt;Inside it, the &lt;code&gt;iam_analyzer&lt;/code&gt; tool span showed a &lt;code&gt;result_length_chars&lt;/code&gt; of 26,980&lt;/li&gt;
&lt;li&gt;Compare to &lt;code&gt;security_group_analyzer&lt;/code&gt; at 4,200 chars and &lt;code&gt;s3_config_checker&lt;/code&gt; at 3,800 chars&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The disproportion was the clue. If the tool output is 7x larger than its siblings, reduce it.&lt;/p&gt;




&lt;h2&gt;
  
  
  Best Use of Sentry
&lt;/h2&gt;

&lt;p&gt;I used Sentry's AI Agent Monitoring to instrument a multi-agent CrewAI pipeline from scratch. This isn't a web app or API. It's five autonomous AI agents making LLM calls and executing custom tools. Standard APM wouldn't help here.&lt;/p&gt;




&lt;h3&gt;
  
  
  What I instrumented
&lt;/h3&gt;

&lt;p&gt;Every pipeline run creates a Sentry transaction with this span hierarchy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Transaction: "Security Posture Scan" (57s)
├── gen_ai.invoke_agent: ResourceDiscovery
│   └── gen_ai.execute_tool: aws_resource_scanner
├── gen_ai.invoke_agent: SecurityScanner
│   ├── gen_ai.execute_tool: security_group_analyzer
│   ├── gen_ai.execute_tool: s3_config_checker
│   ├── gen_ai.execute_tool: iam_analyzer
│   ├── gen_ai.execute_tool: ec2_security_checker
│   └── gen_ai.execute_tool: lambda_security_checker
├── gen_ai.invoke_agent: ComplianceChecker
├── gen_ai.invoke_agent: RiskScorer
└── gen_ai.invoke_agent: RemediationPlanner
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h3&gt;
  
  
  The instrumentation code
&lt;/h3&gt;

&lt;p&gt;For each agent (in &lt;code&gt;main.py&lt;/code&gt;):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;start_span&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;op&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gen_ai.invoke_agent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;invoke_agent &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;agent_name&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;agent_span&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;agent_span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gen_ai.operation.name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;invoke_agent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;agent_span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gen_ai.agent.name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;agent_name&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;agent_span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gen_ai.request.model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;amazon.nova-pro-v1:0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;agent_span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gen_ai.pipeline.name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;security-posture-scan&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;task_output&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;task_obj&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute_sync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;agent&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;task_obj&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;task_obj&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;agent_span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;duration_seconds&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;elapsed&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="n"&gt;agent_span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;output_length_chars&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task_output&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For each tool (decorator in &lt;code&gt;monitoring.py&lt;/code&gt;):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;trace_tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tool_name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;decorator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="nd"&gt;@functools.wraps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;wrapper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;start_span&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="n"&gt;op&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gen_ai.execute_tool&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;execute_tool &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;tool_name&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;span&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gen_ai.tool.name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tool_name&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;func&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="n"&gt;span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;result_length_chars&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
                &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;findings_count&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                        &lt;span class="n"&gt;span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;findings_count&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;findings_count&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
                &lt;span class="nf"&gt;except &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;JSONDecodeError&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;KeyError&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
                    &lt;span class="k"&gt;pass&lt;/span&gt;
                &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;wrapper&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;decorator&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h3&gt;
  
  
  What Sentry revealed
&lt;/h3&gt;

&lt;p&gt;The trace waterfall made the bottleneck visually obvious. The SecurityScanner span was nearly twice the width of any other agent, 22.6 seconds while others averaged 5-10s.&lt;/p&gt;

&lt;p&gt;Inside it, the &lt;code&gt;iam_analyzer&lt;/code&gt; tool span showed &lt;code&gt;result_length_chars: 26980&lt;/code&gt; while &lt;code&gt;s3_config_checker&lt;/code&gt; showed 3,800 and &lt;code&gt;security_group_analyzer&lt;/code&gt; showed 4,200. The granularity goes down to individual boto3 calls. Every &lt;code&gt;GetPublicAccessBlock&lt;/code&gt;, &lt;code&gt;ListRoles&lt;/code&gt;, and &lt;code&gt;ListAttachedRolePolicies&lt;/code&gt; shows up as its own span.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Before the fix&lt;/strong&gt; (SecurityScanner dominates the trace at 22.6s):&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwzdc0ay62kyuyaxl61zc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwzdc0ay62kyuyaxl61zc.png" alt="Sentry trace waterfall showing SecurityScanner agent dominating the pipeline at 22.6s" width="799" height="348"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzztdwlszlozzao462lyt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzztdwlszlozzao462lyt.png" alt="Zoomed span detail showing iam_analyzer tool with result_length_chars 26980" width="800" height="349"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;After the fix&lt;/strong&gt; (all agents proportional, no single bottleneck):&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdlrq2e536mc6g8qchfyl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdlrq2e536mc6g8qchfyl.png" alt="After fix - all agent spans proportional, SecurityScanner no longer the bottleneck" width="799" height="349"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc1cdprqjdblrwprs2wwl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc1cdprqjdblrwprs2wwl.png" alt="Sentry transaction overview showing clean 57s pipeline with no retries" width="800" height="342"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;iam_analyzer&lt;/code&gt; output dropped from 26,980 chars to 15,532, and the SecurityScanner stopped triggering retry logic.&lt;/p&gt;




&lt;h3&gt;
  
  
  Sentry features used
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;How I Used It&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Distributed Tracing&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Full pipeline trace from start to final report&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;AI Agent Monitoring&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;gen_ai.invoke_agent&lt;/code&gt; spans for all 5 agents&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Tool Execution Tracing&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;gen_ai.execute_tool&lt;/code&gt; spans for all 5 boto3 tools&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Custom Span Data&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Token counts, output sizes, duration, finding counts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Error Monitoring&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Exception capture with &lt;code&gt;sentry_sdk.capture_exception()&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Breadcrumbs&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Agent completion events via task callbacks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Transaction Metadata&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Model name, pipeline config, agent count&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h3&gt;
  
  
  Why this matters for AI agent observability
&lt;/h3&gt;

&lt;p&gt;Standard logging tells you an agent "finished." It doesn't tell you which tool inside which agent returned a 27KB payload that triggered a retry you never asked for.&lt;/p&gt;

&lt;p&gt;With five agents each making their own LLM calls and tool executions, you need span-level visibility. Sentry's &lt;code&gt;gen_ai.invoke_agent&lt;/code&gt; and &lt;code&gt;gen_ai.execute_tool&lt;/code&gt; conventions gave me exactly that. I could see the problem in the trace waterfall before I even looked at the code.&lt;/p&gt;

&lt;p&gt;That's the difference between "add some logging" and actual AI observability.&lt;/p&gt;




&lt;p&gt;If you're running multi-agent pipelines, what's your observability setup? Curious if anyone else has hit similar tool-output-size problems with CrewAI or LangGraph.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The agent found 97 real security findings in my AWS account. The fix is in production.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>bugsmash</category>
      <category>ai</category>
      <category>aws</category>
    </item>
    <item>
      <title>4 Silent Failures, 2 Undocumented APIs, and a Container That Crashed Because of a Missing User Directive</title>
      <dc:creator>Sarvar Nadaf</dc:creator>
      <pubDate>Mon, 20 Jul 2026 18:20:15 +0000</pubDate>
      <link>https://dev.to/sarvar_04/4-silent-failures-2-undocumented-apis-and-a-container-that-crashed-because-of-a-missing-user-1b9n</link>
      <guid>https://dev.to/sarvar_04/4-silent-failures-2-undocumented-apis-and-a-container-that-crashed-because-of-a-missing-user-1b9n</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for &lt;a href="https://dev.to/bugsmash"&gt;DEV's Summer Bug Smash: Smash Stories&lt;/a&gt; powered by &lt;a href="https://sentry.io/" rel="noopener noreferrer"&gt;Sentry&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;I spent a week deploying a CrewAI agent to AWS Bedrock AgentCore. The SDK wasn't on PyPI. The error messages were 200 OKs. The container crashed without logs. And the naming regex rejected hyphens without telling me why.&lt;/p&gt;

&lt;p&gt;This is the full debugging trail. Every failure was silent. Every fix required reading source code nobody documented.&lt;/p&gt;




&lt;h2&gt;
  
  
  Table of Contents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The Project&lt;/li&gt;
&lt;li&gt;Failure 1: The SDK That Doesn't Exist on PyPI&lt;/li&gt;
&lt;li&gt;Failure 2: The 200 OK That Means Failure&lt;/li&gt;
&lt;li&gt;Failure 3: The Container That Crashed With No Logs&lt;/li&gt;
&lt;li&gt;Failure 4: The Naming Regex Nobody Documented&lt;/li&gt;
&lt;li&gt;The Two-Client Split Nobody Mentions&lt;/li&gt;
&lt;li&gt;What I Learned&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  The project
&lt;/h2&gt;

&lt;p&gt;I built a resume-tailoring AI agent with CrewAI and Amazon Bedrock. It takes a job description, analyzes your resume, identifies gaps, and rewrites bullet points to match what the role actually needs.&lt;/p&gt;

&lt;p&gt;Locally it worked perfectly. CrewAI orchestrates the agents, Bedrock Nova Pro handles the LLM calls, and the output is solid. Deploying it to production was the problem.&lt;/p&gt;

&lt;p&gt;AWS launched Bedrock AgentCore in June 2026 as a managed runtime for AI agents. You containerize your agent, push the image, and AgentCore handles scaling, memory, and invocation. Sounds simple.&lt;/p&gt;

&lt;p&gt;It was not simple.&lt;/p&gt;




&lt;h2&gt;
  
  
  Failure 1: The SDK that doesn't exist on PyPI
&lt;/h2&gt;

&lt;p&gt;The docs say to install &lt;code&gt;bedrock-agentcore-client&lt;/code&gt;. I ran:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;bedrock-agentcore-client
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It installed successfully. No errors. That's because there's a &lt;strong&gt;placeholder package&lt;/strong&gt; on PyPI with that name. It installs, imports fail silently, and your container builds successfully with a broken dependency inside.&lt;/p&gt;

&lt;p&gt;The real SDK lives in AWS's CodeArtifact registry. You need to configure pip to pull from a private index:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws codeartifact login &lt;span class="nt"&gt;--tool&lt;/span&gt; pip &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--domain&lt;/span&gt; amazon-agent-runtimes &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--repository&lt;/span&gt; agent-runtimes-pypi &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--domain-owner&lt;/span&gt; 600427722194
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then install from there. The PyPI package is a trap. Nobody warns you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hours lost: 3.&lt;/strong&gt; The error only appears at runtime when the container tries to import the module. The build succeeds. The push succeeds. The deployment succeeds. The invocation returns an empty payload.&lt;/p&gt;




&lt;h2&gt;
  
  
  Failure 2: The 200 OK that means failure
&lt;/h2&gt;

&lt;p&gt;After fixing the SDK, I deployed and invoked the agent:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws bedrock-agentcore-control invoke-agent-runtime &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--agent-runtime-id&lt;/span&gt; abc123 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--payload&lt;/span&gt; &lt;span class="s1"&gt;'{"job_description": "..."}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Response: HTTP 200. Payload: empty string.&lt;/p&gt;

&lt;p&gt;Not a 500. Not a 400. Not an error message. A successful HTTP response with nothing inside.&lt;/p&gt;

&lt;p&gt;I checked CloudWatch. No logs. I checked the container status. Running. I checked the agent runtime status. Active.&lt;/p&gt;

&lt;p&gt;The problem: my IAM role was missing &lt;code&gt;bedrock:GetAgentRuntime&lt;/code&gt; permission. Without it, the invocation endpoint accepts the request, routes it nowhere, and returns a 200 with an empty body.&lt;/p&gt;

&lt;p&gt;There is no error message. There is no log entry. The service returns success when it fails.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hours lost: 5.&lt;/strong&gt; I tried different payloads, different content types, different SDK versions, curl vs boto3, synchronous vs streaming. All 200 OK, all empty. The fix was one IAM permission that produces zero error signal when missing.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"Effect"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Allow"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"Action"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"bedrock:GetAgentRuntime"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"Resource"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"*"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Failure 3: The container that crashed with no logs
&lt;/h2&gt;

&lt;p&gt;Next failure. Container starts, passes health checks for 30 seconds, then dies. No exception in CloudWatch. No crash log. Status shows "Failed" with no reason.&lt;/p&gt;

&lt;p&gt;I added every logging statement I could think of. Print statements. Structured logging. Exception handlers wrapping every import. Nothing appeared in CloudWatch because the container never got far enough to initialize the logging framework.&lt;/p&gt;

&lt;p&gt;The cause: missing &lt;code&gt;USER 1000&lt;/code&gt; directive in the Dockerfile.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="c"&gt;# This crashes silently&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="s"&gt; python:3.12-slim&lt;/span&gt;
&lt;span class="k"&gt;WORKDIR&lt;/span&gt;&lt;span class="s"&gt; /app&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; . .&lt;/span&gt;
&lt;span class="k"&gt;RUN &lt;/span&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-e&lt;/span&gt; .
&lt;span class="k"&gt;CMD&lt;/span&gt;&lt;span class="s"&gt; ["python", "-m", "my_agent"]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="c"&gt;# This works&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="s"&gt; python:3.12-slim&lt;/span&gt;
&lt;span class="k"&gt;RUN &lt;/span&gt;useradd &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="nt"&gt;-u&lt;/span&gt; 1000 agentuser
&lt;span class="k"&gt;WORKDIR&lt;/span&gt;&lt;span class="s"&gt; /app&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; . .&lt;/span&gt;
&lt;span class="k"&gt;RUN &lt;/span&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-e&lt;/span&gt; .
&lt;span class="k"&gt;USER&lt;/span&gt;&lt;span class="s"&gt; 1000&lt;/span&gt;
&lt;span class="k"&gt;CMD&lt;/span&gt;&lt;span class="s"&gt; ["python", "-m", "my_agent"]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;AgentCore requires the container to run as UID 1000. If it doesn't, the runtime kills the container. The error message in the console says "Failed." Just "Failed." No mention of user directives, permissions, or UID requirements.&lt;/p&gt;

&lt;p&gt;I found this by reading the AgentCore team's GitHub sample repos. Not the docs. The sample Dockerfile.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hours lost: 4.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Failure 4: The naming regex nobody documented
&lt;/h2&gt;

&lt;p&gt;I wanted to name my agent runtime &lt;code&gt;resume-tailor-agent&lt;/code&gt;. Deployment failed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;An error occurred (ValidationException): 
  Name must match pattern: ^[a-zA-Z0-9_]+$
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No hyphens allowed. Fine. I renamed to &lt;code&gt;resume_tailor_agent&lt;/code&gt; and moved on.&lt;/p&gt;

&lt;p&gt;But the error message only appears if you use the control plane client. If you use the console, it just... doesn't submit. No red border, no error toast, no validation message. The button does nothing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hours lost: 1.&lt;/strong&gt; Small one, but the pattern: silent failures.&lt;/p&gt;




&lt;h2&gt;
  
  
  The two-client split nobody mentions
&lt;/h2&gt;

&lt;p&gt;Here's where it gets architectural. AgentCore has TWO Python clients:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;bedrock-agentcore-control&lt;/code&gt; for managing runtimes (create, update, delete)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;bedrock-agentcore&lt;/code&gt; for the runtime SDK (what runs inside your container)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The documentation uses both interchangeably. Code samples import from one in the setup section and the other in the invocation section. They have different install paths, different CodeArtifact repositories, and different API surfaces.&lt;/p&gt;

&lt;p&gt;If you install the wrong one, nothing tells you. Your code runs until it hits an import that doesn't exist in the package you installed. And since both packages have overlapping module names in some versions, the error might be an &lt;code&gt;AttributeError&lt;/code&gt; deep in a function call, not a clean &lt;code&gt;ImportError&lt;/code&gt; at the top.&lt;/p&gt;

&lt;p&gt;I mapped out which client does what:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Client&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;th&gt;Install From&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;bedrock-agentcore-control&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Create/manage runtimes&lt;/td&gt;
&lt;td&gt;CodeArtifact (domain: amazon-agent-runtimes)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;bedrock-agentcore&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Runtime SDK (inside container)&lt;/td&gt;
&lt;td&gt;CodeArtifact (same domain)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;boto3&lt;/code&gt; (bedrock-agent)&lt;/td&gt;
&lt;td&gt;Invoke from outside&lt;/td&gt;
&lt;td&gt;Standard pip&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This table doesn't exist in any documentation I found.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I learned
&lt;/h2&gt;

&lt;p&gt;Five days of debugging. Four distinct silent failures. Zero useful error messages.&lt;/p&gt;

&lt;p&gt;Every single problem shared the same pattern: the system accepted the bad input, returned success, and failed somewhere downstream without signaling what went wrong. The 200 OK that means failure. The build that succeeds with a placeholder SDK. The container that crashes without logs.&lt;/p&gt;

&lt;p&gt;If I'd had Sentry in the container from day one, I would have caught the import failure, the UID crash, and the empty response pattern within hours instead of days. Observability isn't optional for agent deployments. The infrastructure actively hides failures from you.&lt;/p&gt;

&lt;p&gt;Three principles I'm carrying forward:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Never trust a 200 OK from a new service.&lt;/strong&gt; Validate the response body. If it's empty, something broke silently upstream.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Test imports at container startup, before anything else.&lt;/strong&gt; A try/except around every critical import with an explicit log line. If the SDK is fake, you'll know in the first second.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Read the sample repos, not just the docs.&lt;/strong&gt; The Dockerfile in AWS's example repo had &lt;code&gt;USER 1000&lt;/code&gt;. The documentation never mentioned it. The sample code is sometimes the real documentation.&lt;/p&gt;

&lt;p&gt;I've since added Sentry to my agent pipeline for my security posture scanner project. The trace waterfalls catch problems in seconds that would have taken me hours with print statements. Lesson learned the hard way.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;All errors described above are from July 2026 on AgentCore's GA release. Some may be fixed by the time you read this. The patterns of silent failure in new AWS services are probably eternal.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>bugsmash</category>
      <category>ai</category>
      <category>aws</category>
    </item>
    <item>
      <title>Introducing AWS SimuLearn Badges: Free Proof That You Can Actually Build in the Cloud</title>
      <dc:creator>Sarvar Nadaf</dc:creator>
      <pubDate>Fri, 17 Jul 2026 15:47:29 +0000</pubDate>
      <link>https://dev.to/aws-builders/introducing-aws-simulearn-badges-free-proof-that-you-can-actually-build-in-the-cloud-kj6</link>
      <guid>https://dev.to/aws-builders/introducing-aws-simulearn-badges-free-proof-that-you-can-actually-build-in-the-cloud-kj6</guid>
      <description>&lt;p&gt;If someone asked me ten years ago what it takes to break into cloud, I would have said "get certified and hope someone gives you a chance."&lt;/p&gt;

&lt;p&gt;I was wrong. And I watched dozens of freshers follow that exact advice, collect a certification, then sit in interviews unable to explain why they chose one architecture over another.&lt;/p&gt;

&lt;p&gt;The problem was never knowledge. It was proof. Proof that you can gather requirements from a confused client, design something that works, and actually build it with your own hands.&lt;/p&gt;

&lt;p&gt;AWS just launched something that helps close that gap. And two of these credentials cost nothing.&lt;/p&gt;




&lt;h2&gt;
  
  
  Table of Contents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;What AWS SimuLearn Badges Actually Are&lt;/li&gt;
&lt;li&gt;Why This Matters More Than Another Certification&lt;/li&gt;
&lt;li&gt;The 12 Badges Available Right Now&lt;/li&gt;
&lt;li&gt;The Free Starting Path I Would Follow Today&lt;/li&gt;
&lt;li&gt;The LinkedIn Advantage Nobody Is Talking About&lt;/li&gt;
&lt;li&gt;For Career Switchers: Your Existing Skills Are the Cheat Code&lt;/li&gt;
&lt;li&gt;My Honest Take After 10 Years in Cloud&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What AWS SimuLearn Badges Actually Are
&lt;/h2&gt;

&lt;p&gt;SimuLearn is not another video course. Not another multiple-choice exam.&lt;/p&gt;

&lt;p&gt;You sit in a simulated client meeting powered by generative AI. A virtual customer explains their business problem. You ask questions, uncover requirements, handle objections, and propose an architecture. The AI evaluates your communication, your technical accuracy, and your decision-making in real time.&lt;/p&gt;

&lt;p&gt;Then you build the solution. In a live AWS environment. Not a sandbox with three buttons. The real console.&lt;/p&gt;

&lt;p&gt;After that, an automated validation confirms your solution actually works.&lt;/p&gt;

&lt;p&gt;Complete every assignment in a learning plan, and AWS issues you a badge through Credly. Automatically. No exam booking. No proctored test. Just demonstrated capability across the full workflow.&lt;/p&gt;

&lt;p&gt;Each badge represents the entire journey: customer conversations, architecture design, hands-on building, and validated outcomes. Not a single quiz. Not one lab. The whole thing.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why This Matters More Than Another Certification
&lt;/h2&gt;

&lt;p&gt;I hold 7 AWS certifications. Let me tell you what they prove: I can answer questions about cloud services under time pressure.&lt;/p&gt;

&lt;p&gt;What they do NOT prove: that I can sit with a client, understand their problem, design something appropriate, and build it without breaking things.&lt;/p&gt;

&lt;p&gt;That gap is exactly where freshers struggle. You study for months, pass the exam, and then face an interviewer who asks "Tell me about a time you designed a solution for a customer." You have nothing.&lt;/p&gt;

&lt;p&gt;Last year I sat across from a candidate with three AWS certifications. I asked him to walk me through a time he gathered requirements from a non-technical stakeholder. Thirty seconds of silence. Three certs, zero stories. That is the gap SimuLearn closes.&lt;/p&gt;

&lt;p&gt;Here is how the credential types compare:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Certifications&lt;/strong&gt; prove you KNOW it. Knowledge of services, best practices, theory.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Microcredentials&lt;/strong&gt; prove you can DO it. Specific hands-on skills in a live environment. No multiple choice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SimuLearn Badges&lt;/strong&gt; prove you can DELIVER it. The full cycle from client conversation to working solution.&lt;/p&gt;

&lt;p&gt;For someone with zero professional cloud experience, that third one is useful. It gives you something concrete to discuss in interviews. "I gathered requirements from a retail client who needed personalized recommendations, designed a serverless pipeline, and built it in the console." That sentence alone makes you more memorable than someone who just lists a cert.&lt;/p&gt;




&lt;h2&gt;
  
  
  The 12 Badges Available Right Now
&lt;/h2&gt;

&lt;p&gt;Two of these are completely free. No subscription. No credit card. Nothing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Free:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://skillbuilder.aws/learning-plan/EKHCUEWUUC/aws-simulearn-cloud-practitioner/1UQVR262ZB" rel="noopener noreferrer"&gt;AWS SimuLearn Cloud Practitioner&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://skillbuilder.aws/learning-plan/3HCD821CNZ/aws-simulearn-ai-practitioner/BFKGA5VM8H" rel="noopener noreferrer"&gt;AWS SimuLearn AI Practitioner&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;With Skill Builder subscription ($29/month or $299/year):&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Solutions Architect&lt;/li&gt;
&lt;li&gt;Serverless Developer&lt;/li&gt;
&lt;li&gt;AI Architect&lt;/li&gt;
&lt;li&gt;Machine Learning&lt;/li&gt;
&lt;li&gt;Security&lt;/li&gt;
&lt;li&gt;Networking&lt;/li&gt;
&lt;li&gt;Data Analytics&lt;/li&gt;
&lt;li&gt;Healthcare&lt;/li&gt;
&lt;li&gt;Manufacturing and Automotive&lt;/li&gt;
&lt;li&gt;Financial Services&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn4gr10i0ep8s5qcajmgf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn4gr10i0ep8s5qcajmgf.png" alt="All 12 AWS SimuLearn Learning Plan Badges" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The free ones cover cloud fundamentals and AI. The two hottest entry points in cloud right now.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Free Starting Path I Would Follow Today
&lt;/h2&gt;

&lt;p&gt;If I were starting from zero in 2026, here is exactly what I would do.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Month 1-2: Build the foundation (cost: $0)&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Start with the &lt;a href="https://skillbuilder.aws/learning-plan/EKHCUEWUUC/aws-simulearn-cloud-practitioner/1UQVR262ZB" rel="noopener noreferrer"&gt;SimuLearn Cloud Practitioner&lt;/a&gt; learning plan. Take your time. The platform offers two modes: Scripted mode guides you through the conversation with prompts, and Open Dialogue mode lets you drive the meeting yourself. Begin with Scripted to see how a proper client conversation flows, then switch to Open Dialogue once you feel ready.&lt;/p&gt;

&lt;p&gt;The customer conversations will feel awkward at first. That is normal. Real client meetings are awkward too. You are practicing that discomfort in a safe environment.&lt;/p&gt;

&lt;p&gt;When you finish, you have your first verified badge.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Month 3: Add the AI layer (cost: $0)&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Complete the AI Practitioner learning plan. Generative AI shows up in most cloud job descriptions now. This badge tells employers you understand how AI fits into cloud architecture. Not just theory. You built something.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Month 4: Take the certification exam (cost: $100)&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Now take the Cloud Practitioner certification. You already have hands-on practice from SimuLearn. The exam covers some ground you have already worked through, though it also tests pricing models, support plans, and global infrastructure that SimuLearn may not cover in depth. Study those gaps, but you will not be starting from zero.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Total investment: $100 and four months.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;What you have at the end: 2 SimuLearn badges and 1 certification. Three verified AWS credentials on your LinkedIn profile. All proving different things. All verifiable by any recruiter who clicks the link.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bonus:&lt;/strong&gt; AWS also made microcredentials free in April 2026. Four are available: Serverless, Agentic AI, Application Networking, and Incident Response. These are pure hands-on assessments. If you want to stack more credentials after your certification, these are worth looking at.&lt;/p&gt;

&lt;p&gt;Many people spend those same four months watching tutorials and collecting notes they never revisit. This path gives you something verifiable at the end of each month.&lt;/p&gt;




&lt;h2&gt;
  
  
  The LinkedIn Advantage Nobody Is Talking About
&lt;/h2&gt;

&lt;p&gt;Every badge issued through Credly becomes a verified credential on LinkedIn. Not a self-declared skill. Not a line on your resume anyone could type. A verified badge that links back to exactly what you did to earn it. Anyone can click it and confirm it is real.&lt;/p&gt;

&lt;p&gt;When a recruiter clicks your badge, they see:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Who issued it (Amazon Web Services)&lt;/li&gt;
&lt;li&gt;What it represents (full learning plan completion)&lt;/li&gt;
&lt;li&gt;What skills it validates (specific, detailed)&lt;/li&gt;
&lt;li&gt;When you earned it&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This cannot be faked. That matters when you have no work experience to reference.&lt;/p&gt;

&lt;p&gt;When you share a badge as a LinkedIn post, your network sees it. People engage with achievement posts. That engagement puts your name in front of people who might not have found you otherwise. For freshers with a small network, a badge post can be good visibility.&lt;/p&gt;

&lt;p&gt;How effective is this model? AWS says Cloud Quest badge earners share their accomplishments "at rates well above industry benchmarks." Across the broader AWS Education Programs (Academy, Educate, re/Start), Credly's survey data shows 99% of badge earners consider them valuable, 42% got a new job, and 85% received a raise, promotion, or new role. Those programs include full training and career support, not just a badge. SimuLearn follows the same Credly model, but these numbers are not guaranteed outcomes for every badge earner.&lt;/p&gt;

&lt;p&gt;Verified badges on your profile send a signal: this person is actively investing in cloud skills, and they can back it up.&lt;/p&gt;




&lt;h2&gt;
  
  
  For Career Switchers: Your Existing Skills Are the Cheat Code
&lt;/h2&gt;

&lt;p&gt;If you are moving into cloud from another field, SimuLearn gives you something certifications never could: a way to leverage what you already know.&lt;/p&gt;

&lt;p&gt;Coming from sales or consulting? The customer conversation component is YOUR strength. You already know how to ask questions, handle objections, and translate technical concepts into business value. SimuLearn lets you combine that with new technical skills and prove both in one credential.&lt;/p&gt;

&lt;p&gt;Coming from healthcare, manufacturing, or finance? Look at those industry-specific learning plans. A SimuLearn Healthcare badge tells employers: "I understand compliance, I understand the domain problems, AND I can build cloud solutions for them." I have seen people transition from clinical roles into cloud at consulting clients. The combination of domain knowledge plus technical proof is what gets them past the resume screen when pure technologists cannot.&lt;/p&gt;

&lt;p&gt;The subscription costs $29/month. Subscribe for three months, earn your role-specific or industry-specific badge, cancel. You spent $87 total. Compare that to a bootcamp at $10K or a master's degree at $40K.&lt;/p&gt;

&lt;p&gt;That $87 gets you a verified AWS credential on LinkedIn, hands-on experience in a live console, and client simulation practice you can reference in every interview. For a career pivot, that is a solid starting investment. One honest note: the paid badges are newer to the market than certifications. Not every recruiter will recognize them yet. But the hands-on experience you gain is yours regardless of badge recognition.&lt;/p&gt;




&lt;h2&gt;
  
  
  My Honest Take After 10 Years in Cloud
&lt;/h2&gt;

&lt;p&gt;I have not personally completed a SimuLearn learning plan yet. Full disclosure. I also do not know how long each plan takes to finish. AWS does not publish that number anywhere I could find. Could be 10 hours, could be 60. That matters if you are working full time and planning your evenings around this.&lt;/p&gt;

&lt;p&gt;What I do know: SimuLearn badges are the first AWS credential that evaluates both your communication skills and your technical ability in one package. That combination is what the actual job requires. Every architecture review I have ever been in required both.&lt;/p&gt;

&lt;p&gt;But let me be honest about limitations too. This is still a simulation, not real client work. A hiring manager who has never heard of SimuLearn might not know what the badge means yet. These are new, launched July 2026. It will take time for the market to recognize them the way it recognizes AWS certifications.&lt;/p&gt;

&lt;p&gt;These badges will not replace certifications. Hiring managers still filter by cert. But they fill a gap that certifications cannot: evidence that you practiced the full job, not just the exam.&lt;/p&gt;

&lt;p&gt;If you are a fresher reading this, your path is clearer than it was a year ago. Start with the free plans. Build the habit of doing, not just studying. Earn something verifiable before you spend a single dollar on an exam.&lt;/p&gt;

&lt;p&gt;The old way: study, memorize, pass, hope.&lt;/p&gt;

&lt;p&gt;The new way: build, prove, share, get noticed.&lt;/p&gt;

&lt;p&gt;Start here:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://skillbuilder.aws/learning-plan/EKHCUEWUUC/aws-simulearn-cloud-practitioner/1UQVR262ZB" rel="noopener noreferrer"&gt;AWS SimuLearn Cloud Practitioner&lt;/a&gt; (free)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://skillbuilder.aws/learning-plan/3HCD821CNZ/aws-simulearn-ai-practitioner/BFKGA5VM8H" rel="noopener noreferrer"&gt;AWS SimuLearn AI Practitioner&lt;/a&gt; (free)&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;Have you tried SimuLearn yet? What was your experience with the customer simulation part? I am genuinely curious whether the Open Dialogue conversations feel realistic or scripted. Drop your thoughts below.&lt;/p&gt;

&lt;p&gt;If this helped clarify the credential landscape, a ❤️ helps other beginners find it too.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Follow me for more on AWS architecture, DevOps, and AI tooling:&lt;/em&gt;&lt;br&gt;
&lt;em&gt;&lt;a href="https://sarvarnadaf.com" rel="noopener noreferrer"&gt;sarvarnadaf.com&lt;/a&gt; | &lt;a href="https://www.linkedin.com/in/sarvar04/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; | &lt;a href="https://dev.to/sarvar_04"&gt;Dev.to&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aws</category>
      <category>career</category>
      <category>beginners</category>
      <category>cloud</category>
    </item>
    <item>
      <title>Kiro CLI Gets Worse After an Hour. Here's How I Fixed It.</title>
      <dc:creator>Sarvar Nadaf</dc:creator>
      <pubDate>Thu, 16 Jul 2026 14:37:01 +0000</pubDate>
      <link>https://dev.to/aws-builders/kiro-cli-gets-worse-after-an-hour-heres-how-i-fixed-it-1l7g</link>
      <guid>https://dev.to/aws-builders/kiro-cli-gets-worse-after-an-hour-heres-how-i-fixed-it-1l7g</guid>
      <description>&lt;p&gt;Two hours into a Kiro CLI session, it gave me a generic IAM trust policy example. We had literally set up the OIDC provider together 30 minutes earlier. It forgot.&lt;/p&gt;

&lt;p&gt;This is not a bug. It has a name: context rot. And once I understood the mechanics, I stopped fighting it and started managing it.&lt;/p&gt;

&lt;p&gt;I use Kiro CLI every day. Checking CloudWatch alarms, debugging IAM policies, reviewing security groups, looking at cost reports, setting up infrastructure. It is my go-to tool for anything AWS.&lt;/p&gt;

&lt;p&gt;But after a few weeks I noticed the pattern. The longer I keep a session going, the worse Kiro gets. Early in the session it is sharp. Finds the right file, gives the right answer, runs the right command. Two hours later it starts doing odd things. It re-reads files it already looked at. It gives vague answers. It suggests things I already tried.&lt;/p&gt;




&lt;h2&gt;
  
  
  What actually happens when context fills up
&lt;/h2&gt;

&lt;p&gt;Kiro has up to a 200k token context window (depending on model). Every message you send, every response it gives, every file it reads, every command output: all of it goes into that window. It adds up fast.&lt;/p&gt;

&lt;p&gt;A typical morning for me looks like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"Check if there are any CloudWatch alarms firing in us-east-1"&lt;/li&gt;
&lt;li&gt;"Show me the IAM policy attached to the lambda-processor role"&lt;/li&gt;
&lt;li&gt;"Why is this S3 bucket policy denying access from the VPC endpoint"&lt;/li&gt;
&lt;li&gt;"Look at the cost explorer data for the last 7 days"&lt;/li&gt;
&lt;li&gt;"What EC2 instances are running in the staging account"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each of those triggers multiple tool calls. Kiro reads files, runs AWS CLI commands, processes the output. By the fifth question, the context window has maybe 30-40% filled.&lt;/p&gt;

&lt;p&gt;Sounds fine. That's the trap.&lt;/p&gt;

&lt;p&gt;I keep going. I debug a CloudFormation deployment. I check a few security groups. I look at some ECS task definitions. By lunch, I am at 60-70% context usage and Kiro starts struggling.&lt;/p&gt;

&lt;p&gt;What struggling looks like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It runs &lt;code&gt;aws sts get-caller-identity&lt;/code&gt; again even though it already knows which account I am in&lt;/li&gt;
&lt;li&gt;It re-reads the same terraform files it read an hour ago&lt;/li&gt;
&lt;li&gt;Responses get longer and less useful (more filler, less action)&lt;/li&gt;
&lt;li&gt;It starts suggesting solutions to problems I already fixed earlier in the session&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you manage 5+ AWS accounts across environments, context fills 3x faster because every &lt;code&gt;describe-instances&lt;/code&gt; call across accounts dumps more output into the window.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why this happens (and why "try harder" does not fix it)
&lt;/h2&gt;

&lt;p&gt;The model can see everything in the context window but it cannot focus on everything equally. Research calls this the "Lost in the Middle" problem. A Stanford paper (Liu et al., 2023) demonstrated that LLMs exhibit a U-shaped attention pattern. They attend strongly to information at the beginning and end of their context window. Everything in the middle gets reduced attention.&lt;/p&gt;

&lt;p&gt;A 2025 follow-up study found the effect is strongest when inputs fill up to 50% of the context window. Beyond that, primacy bias weakens and recency bias dominates, meaning the model increasingly favors whatever you talked about in the last few minutes over the rules and context from the start of your session.&lt;/p&gt;

&lt;p&gt;The Kiro team calls this phenomenon "context rot." Anthropic's own engineering docs define it directly: "Context must be treated as a finite resource with diminishing marginal returns."&lt;/p&gt;

&lt;p&gt;Here is what that means in practice. When the window is full of old conversations about IAM policies, cost reports, and security groups, and you ask about an ECS deployment, the model has to sort through all that irrelevant context to find what matters. It cannot selectively ignore the old stuff. Every token participates in every inference step, regardless of whether it is relevant to your current question.&lt;/p&gt;

&lt;p&gt;The really frustrating part: the model does not know it is degraded. It will confidently tell you it is following all your instructions while demonstrably not doing so. You cannot fix this by telling it to "try harder" or "pay more attention." The degradation is architectural, not motivational.&lt;/p&gt;

&lt;p&gt;When the context hits 100%, Kiro auto-compacts. It summarizes the old conversation to make room. The problem is that summaries lose detail. That IAM policy you debugged earlier? After compaction, Kiro only remembers "we fixed an IAM issue," not the specific policy ARN or the condition key that was wrong. In measured sessions, auto-compaction can reduce 132,000 tokens to about 2,300, a 98% reduction. The token count is small. The cognitive capital destroyed is enormous.&lt;/p&gt;




&lt;h2&gt;
  
  
  The 5 failure modes (in order of how often I hit them)
&lt;/h2&gt;

&lt;p&gt;Context rot is the foundation, but it cascades into other behaviors. Once you recognize these patterns, you stop blaming yourself or the model:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Context rot itself.&lt;/strong&gt; Attention on your early instructions fades as the window fills. The model stops following rules it was given at the start. You told it "never modify production resources without asking first" in minute 1. In minute 90, it just runs the command.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Pattern matching over reasoning.&lt;/strong&gt; Instead of checking what actually exists, Kiro copies patterns from earlier in the conversation. It writes an IAM role that duplicates one you already created three prompts ago. It is coding from its conversation context, not from your actual files.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Completion bias.&lt;/strong&gt; When you push back on a bad result, the model does not get more careful. It gets more agreeable. It ships something that looks different but is equally wrong. The angrier you get, the faster it produces plausible-looking garbage to end the conflict.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Re-verification loops.&lt;/strong&gt; The model forgets it already verified something. It calls &lt;code&gt;aws sts get-caller-identity&lt;/code&gt; again. It re-reads your terraform state file. It asks "which region are you working in?" for the third time. Each redundant call wastes context and accelerates the rot.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Advisory compliance decay.&lt;/strong&gt; Early in a session, Kiro respects guardrails: "I'll ask before modifying production resources." Late in a session, it rationalizes around them: "Since we've been working in this account, I'll go ahead and apply the change." The constraint fades because the instruction that set it is now buried in the middle of a bloated context window.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I do now
&lt;/h2&gt;

&lt;h3&gt;
  
  
  One task per session
&lt;/h3&gt;

&lt;p&gt;This is the biggest change. I used to keep one session open all day. Now I start a new session for each distinct task.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Debugging a CloudFormation stack failure? New session.&lt;/li&gt;
&lt;li&gt;Checking cost reports? New session.&lt;/li&gt;
&lt;li&gt;Writing a new Lambda function? New session.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each session stays focused, context stays small, and Kiro stays sharp. Each wasted minute re-explaining context is a minute I am not solving the actual problem. Over a week, that adds up to 30-60 minutes of lost productivity.&lt;/p&gt;

&lt;p&gt;To start a new session in Kiro CLI, just exit and re-enter:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;exit
&lt;/span&gt;kiro-cli
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This gives you a completely fresh session. The old one stays on disk if you need to go back to it.&lt;/p&gt;

&lt;p&gt;If you just want to reset the conversation without leaving the process, use &lt;code&gt;/clear&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;/clear
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The difference: &lt;code&gt;exit&lt;/code&gt; + &lt;code&gt;kiro-cli&lt;/code&gt; starts a brand new session with a new ID. &lt;code&gt;/clear&lt;/code&gt; wipes the conversation but keeps the same process running.&lt;/p&gt;

&lt;h3&gt;
  
  
  Use /compact before it auto-triggers
&lt;/h3&gt;

&lt;p&gt;If I am in a session that has been going for a while and I am not done yet, I run &lt;code&gt;/compact&lt;/code&gt; manually before Kiro does it automatically.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;/compact
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Why? Because when I trigger it myself, I know what just got summarized. I can immediately re-state the important bits.&lt;/p&gt;

&lt;p&gt;After compacting, re-state the important context:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;We are debugging a CloudFormation stack called prod-api-gateway that failed 
on an AWS::ApiGateway::RestApi resource. The error was "Invalid stage identifier specified".
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now Kiro has a fresh context with only the relevant details. The same question that took 12 seconds before compaction takes 3 seconds after.&lt;/p&gt;

&lt;h3&gt;
  
  
  Put persistent context in steering files (the biggest token saver)
&lt;/h3&gt;

&lt;p&gt;Things I tell Kiro every single session (which AWS accounts I work with, naming conventions, preferred regions) go in steering files. Kiro supports three scopes:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Workspace steering&lt;/strong&gt; (&lt;code&gt;.kiro/steering/&lt;/code&gt;), applies to one project:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; .kiro/steering
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# .kiro/steering/aws-environment.md&lt;/span&gt;

&lt;span class="gu"&gt;## AWS Environment&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Production account: 111111111111 (us-east-1)
&lt;span class="p"&gt;-&lt;/span&gt; Staging account: 222222222222 (us-east-1)
&lt;span class="p"&gt;-&lt;/span&gt; Default profile: production
&lt;span class="p"&gt;-&lt;/span&gt; All infrastructure is in Terraform under /infra

&lt;span class="gu"&gt;## Preferences&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Use AWS CLI v2 commands
&lt;span class="p"&gt;-&lt;/span&gt; Always specify --region explicitly
&lt;span class="p"&gt;-&lt;/span&gt; Check CloudTrail before making IAM changes
&lt;span class="p"&gt;-&lt;/span&gt; Never modify production resources without asking first
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Global steering&lt;/strong&gt; (&lt;code&gt;~/.kiro/steering/&lt;/code&gt;), applies to ALL your workspaces:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# ~/.kiro/steering/aws-defaults.md&lt;/span&gt;

&lt;span class="gu"&gt;## Global AWS Conventions&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Always use --output json for parseable results
&lt;span class="p"&gt;-&lt;/span&gt; Prefer least-privilege IAM policies (never use &lt;span class="err"&gt;*&lt;/span&gt; in Resource)
&lt;span class="p"&gt;-&lt;/span&gt; Tag all resources with Environment, Team, and CostCenter
&lt;span class="p"&gt;-&lt;/span&gt; Use ap-south-1 as default region unless specified
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Global steering is loaded in every session regardless of which project you are in. If workspace steering conflicts with global steering, workspace wins. This is where I put my org-wide standards.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Team steering&lt;/strong&gt;: you can push steering files to your entire team via MDM or a shared repository. Everyone gets the same conventions without explaining them:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Team lead pushes to shared repo&lt;/span&gt;
git clone internal-repo/kiro-steering ~/.kiro/steering
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Kiro automatically loads everything in &lt;code&gt;.kiro/steering/&lt;/code&gt; at the start of every session. Verify what is loaded with &lt;code&gt;/context show&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The difference this makes: my baseline context went from wasting 15-20% on repeated instructions to 0%. Every new session already knows my environment.&lt;/p&gt;

&lt;h3&gt;
  
  
  Use conditional steering with fileMatch
&lt;/h3&gt;

&lt;p&gt;Not all context is needed all the time. Kiro supports inclusion modes that load steering only when relevant:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;inclusion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;fileMatch&lt;/span&gt;
&lt;span class="na"&gt;glob&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;infra/**/*.tf"&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="gh"&gt;# Terraform Conventions&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Use modules for all repeatable infrastructure
&lt;span class="p"&gt;-&lt;/span&gt; State stored in S3 with DynamoDB locking
&lt;span class="p"&gt;-&lt;/span&gt; Never hardcode AMI IDs, use data sources
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This file only loads when Kiro reads a &lt;code&gt;.tf&lt;/code&gt; file under the &lt;code&gt;infra/&lt;/code&gt; directory. When I am working on Lambda code, these Terraform conventions do not eat my context window.&lt;/p&gt;

&lt;p&gt;Three modes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;always&lt;/strong&gt; (default): loaded every session&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;fileMatch&lt;/strong&gt;: loaded only when a matching file is in context&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;manual&lt;/strong&gt;: loaded only when you invoke it with a slash command&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I split my steering into 5 focused files with appropriate modes instead of one giant file. Token usage for steering dropped from 8% of my context to about 3%.&lt;/p&gt;

&lt;h3&gt;
  
  
  Watch the percentage
&lt;/h3&gt;

&lt;p&gt;Kiro shows context usage in the sidebar. You can also check it explicitly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;/context show
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This shows a breakdown: how much your steering files use, how much tools use, how much your conversation uses. My rule of thumb:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Under 40%: keep going, no issues&lt;/li&gt;
&lt;li&gt;40-60%: finish current task, then start fresh&lt;/li&gt;
&lt;li&gt;Over 60%: responses are getting worse, wrap up now&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you find yourself repeating instructions or Kiro starts re-reading files, do not fight it. Start fresh.&lt;/p&gt;

&lt;h3&gt;
  
  
  Use knowledge bases for large reference material
&lt;/h3&gt;

&lt;p&gt;If you work with big Terraform repos or lots of CloudFormation templates, do not add them as context files. They will eat your context window on every single request, even when you are not asking about them.&lt;/p&gt;

&lt;p&gt;Kiro's &lt;code&gt;/knowledge&lt;/code&gt; feature (experimental, enable it first) indexes your files locally and only pulls in relevant chunks when you ask about them:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Enable the feature&lt;/span&gt;
kiro-cli settings chat.enableKnowledge &lt;span class="nb"&gt;true&lt;/span&gt;

&lt;span class="c"&gt;# Add your infrastructure code&lt;/span&gt;
/knowledge add &lt;span class="nt"&gt;--name&lt;/span&gt; &lt;span class="s2"&gt;"terraform-infra"&lt;/span&gt; &lt;span class="nt"&gt;--path&lt;/span&gt; /path/to/infra &lt;span class="nt"&gt;--include&lt;/span&gt; &lt;span class="s2"&gt;"**/*.tf"&lt;/span&gt; &lt;span class="nt"&gt;--exclude&lt;/span&gt; &lt;span class="s2"&gt;".terraform/**"&lt;/span&gt;

&lt;span class="c"&gt;# Add documentation&lt;/span&gt;
/knowledge add &lt;span class="nt"&gt;--name&lt;/span&gt; &lt;span class="s2"&gt;"runbooks"&lt;/span&gt; &lt;span class="nt"&gt;--path&lt;/span&gt; /path/to/runbooks &lt;span class="nt"&gt;--include&lt;/span&gt; &lt;span class="s2"&gt;"**/*.md"&lt;/span&gt; &lt;span class="nt"&gt;--index-type&lt;/span&gt; Best
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two index types matter here:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Fast&lt;/strong&gt; (lexical/bm25): for code, configs, logs. Quick keyword matching.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best&lt;/strong&gt; (semantic): for documentation and runbooks. Understands natural language queries.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I use Fast for my Terraform code (I search by resource names and module names) and Best for our runbooks and architecture decision records (I search by questions like "how do we handle cross-account access").&lt;/p&gt;

&lt;p&gt;My baseline context dropped from 35% to 12% after moving Terraform files out of steering and into the knowledge base. Kiro searches your files when needed without them sitting in context all the time.&lt;/p&gt;

&lt;p&gt;The knowledge base persists across sessions. Index once, use forever. Run &lt;code&gt;/knowledge update&lt;/code&gt; periodically when your files change, or &lt;code&gt;/knowledge update&lt;/code&gt; with no arguments to re-index everything at once.&lt;/p&gt;

&lt;h3&gt;
  
  
  Delegate exploration to sub-agents
&lt;/h3&gt;

&lt;p&gt;This one is underused. When Kiro needs to investigate something (scan a directory, read multiple files, research an error) it can delegate to a sub-agent with its own isolated context window.&lt;/p&gt;

&lt;p&gt;Why this matters: if the main session reads 20 files to debug an issue, all 20 files stay in your context forever. If a sub-agent does that exploration, it returns a focused summary and your main context stays clean.&lt;/p&gt;

&lt;p&gt;I noticed Kiro does this automatically for some tasks, but you can nudge it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Investigate why the ECS task is failing to pull from ECR. 
Check the task definition, the ECR repository policy, and the VPC endpoints. 
Give me a summary of the root cause. Do not dump all the raw output.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The key phrase is "give me a summary." It signals that you want focused results, not raw exploration dumped into your session.&lt;/p&gt;




&lt;h2&gt;
  
  
  A real example
&lt;/h2&gt;

&lt;p&gt;I am currently testing Terraform code for an upcoming EKS article in my &lt;a href="https://dev.to/sarvar_04/series/36963"&gt;Terraform series&lt;/a&gt;. The setup is complex: VPC with private subnets, EKS cluster, managed node groups, IAM roles for service accounts, security groups, and add-ons like CoreDNS and kube-proxy.&lt;/p&gt;

&lt;p&gt;I started a Kiro session asking it to help me structure the modules. Then I asked it to write the VPC config. Then the EKS cluster resource. Then the node group. Then I asked it to review the IAM OIDC provider setup.&lt;/p&gt;

&lt;p&gt;By this point the session had all the previous Terraform code, all the plan outputs, and all my back-and-forth corrections in context.&lt;/p&gt;

&lt;p&gt;When I asked "add the aws-load-balancer-controller IAM role with the correct trust policy," Kiro gave me a generic example from the docs. It did not reference my actual cluster name or OIDC provider ARN that we had set up 30 minutes earlier. Gone.&lt;/p&gt;

&lt;p&gt;It was like talking to someone who forgot the last hour of conversation.&lt;/p&gt;

&lt;p&gt;I ran &lt;code&gt;/clear&lt;/code&gt;, started fresh, and said:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;I am building a production EKS cluster with Terraform. The cluster is called 
prod-platform in us-east-1. The OIDC provider is already set up. I need an 
IAM role for aws-load-balancer-controller with the correct trust policy that 
references my cluster's OIDC issuer. Here is my existing oidc provider ARN: 
arn:aws:iam::111111111111:oidc-provider/oidc.eks.us-east-1.amazonaws.com/id/ABCDEF123456
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Kiro immediately gave me the exact trust policy with my ARN, my namespace, my service account name. No generic examples. No confusion with earlier VPC or node group context.&lt;/p&gt;

&lt;p&gt;Total time wasted before clearing: 15 minutes of re-explaining.&lt;br&gt;
Time to get the right answer after clearing: 20 seconds.&lt;/p&gt;




&lt;h2&gt;
  
  
  The context budget cheat sheet
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Context Usage&lt;/th&gt;
&lt;th&gt;What's Happening&lt;/th&gt;
&lt;th&gt;Action&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Under 30%&lt;/td&gt;
&lt;td&gt;Sharp, focused, fast&lt;/td&gt;
&lt;td&gt;Keep working&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;30-50%&lt;/td&gt;
&lt;td&gt;Still good, starting to accumulate noise&lt;/td&gt;
&lt;td&gt;Finish current task, then evaluate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;50-70%&lt;/td&gt;
&lt;td&gt;"Lost in the Middle" kicks in hard&lt;/td&gt;
&lt;td&gt;Wrap up NOW or &lt;code&gt;/compact&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Over 70%&lt;/td&gt;
&lt;td&gt;Active degradation: repeating, re-reading, vague answers&lt;/td&gt;
&lt;td&gt;Start fresh. Do not push through.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Auto-compaction triggers&lt;/td&gt;
&lt;td&gt;98% information loss incoming&lt;/td&gt;
&lt;td&gt;You waited too long. &lt;code&gt;/compact&lt;/code&gt; earlier next time.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  What this is NOT
&lt;/h2&gt;

&lt;p&gt;I want to be clear about something. This is not a Kiro-specific bug. Every AI coding tool with a context window has this problem: Claude Code, Cursor, Copilot Chat, Windsurf. The architecture is the same transformer attention mechanism underneath.&lt;/p&gt;

&lt;p&gt;The difference is whether you manage context deliberately or let it manage you. Most developers do not think about context at all. They open a session, work for 3 hours, get frustrated that "the AI got dumb," close it, and start fresh the next day.&lt;/p&gt;

&lt;p&gt;That cycle wastes an hour of degraded work every single day. Across a team of 5 engineers, that is 25 hours per week of suboptimal AI output. The fix takes zero extra tools. Just awareness and a habit change.&lt;/p&gt;




&lt;p&gt;I shared this approach with two engineers on my team. Both had been complaining about Kiro "getting dumb" by afternoon. Turns out they were running single sessions for 4-5 hours straight. Once they switched to session-per-task, the complaints stopped within a day.&lt;/p&gt;

&lt;p&gt;The model is not getting dumber. Your context window is getting noisier. Fix the noise, and the model stays as sharp at hour 3 as it was at minute 1.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Links:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Kiro CLI context management: &lt;a href="https://kiro.dev/docs/cli/chat/context/" rel="noopener noreferrer"&gt;kiro.dev/docs/cli/chat/context/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Steering files documentation: &lt;a href="https://kiro.dev/docs/cli/steering/" rel="noopener noreferrer"&gt;kiro.dev/docs/cli/steering/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Knowledge management (experimental): &lt;a href="https://kiro.dev/docs/cli/experimental/knowledge-management/" rel="noopener noreferrer"&gt;kiro.dev/docs/cli/experimental/knowledge-management/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;/compact command changelog: &lt;a href="https://kiro.dev/changelog/cli/1-24/" rel="noopener noreferrer"&gt;kiro.dev/changelog/cli/1-24/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Ralph Loop technique: &lt;a href="https://developer.mamezou-tech.com/en/blogs/2026/01/30/kiro_cli_ralph/" rel="noopener noreferrer"&gt;developer.mamezou-tech.com&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;"Lost in the Middle" research: &lt;a href="https://arxiv.org/abs/2307.03172" rel="noopener noreferrer"&gt;arxiv.org/abs/2307.03172&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Context rot whitepaper (Empromptu): &lt;a href="https://empromptu.ai/resources/context-rot-progressive-prompt-ephemerality" rel="noopener noreferrer"&gt;empromptu.ai/resources/context-rot-progressive-prompt-ephemerality&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;How do you manage long sessions with AI coding tools? Do you start fresh or push through? Drop your approach in the comments. I am genuinely curious if others hit the same wall around the 50% mark or if it depends on the type of work.&lt;/p&gt;

&lt;p&gt;Follow me for more on AWS architecture, DevOps, and AI tooling: &lt;a href="https://sarvarnadaf.com" rel="noopener noreferrer"&gt;sarvarnadaf.com&lt;/a&gt; | &lt;a href="https://www.linkedin.com/in/sarvar04/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt;&lt;/p&gt;

</description>
      <category>aws</category>
      <category>kiro</category>
      <category>productivity</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
