<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Kyalo460</title>
    <description>The latest articles on DEV Community by Kyalo460 (@kyalo460).</description>
    <link>https://dev.to/kyalo460</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4158746%2Faee68f84-8f8e-417a-b684-13e5f9649626.jpg</url>
      <title>DEV Community: Kyalo460</title>
      <link>https://dev.to/kyalo460</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/kyalo460"/>
    <language>en</language>
    <item>
      <title>AiWatchDog, monitor your agents, increase productivity</title>
      <dc:creator>Kyalo460</dc:creator>
      <pubDate>Sat, 03 Oct 2026 01:50:17 +0000</pubDate>
      <link>https://dev.to/kyalo460/aiwatchdog-monitor-your-agents-increase-productivity-213o</link>
      <guid>https://dev.to/kyalo460/aiwatchdog-monitor-your-agents-increase-productivity-213o</guid>
      <description>&lt;p&gt;I Built an Open Source Watchdog for AI Coding Agents&lt;br&gt;
Live demo → &lt;a href="https://aiwatchdog.vercel.app/" rel="noopener noreferrer"&gt;https://aiwatchdog.vercel.app/&lt;/a&gt; · Open source (MIT)&lt;/p&gt;

&lt;p&gt;The problem&lt;br&gt;
I run a lot of AI coding agents — Kilo, Claude Code, Codex, Aider, my own scripts. And I kept hitting the same wall:&lt;/p&gt;

&lt;p&gt;I start a task, walk away, come back 20 minutes later, and stare at a terminal with no idea whether the agent is crushing it, waiting for me, or wedged in an infinite retry loop burning tokens.&lt;/p&gt;

&lt;p&gt;So I built the thing I kept wishing existed.&lt;/p&gt;

&lt;p&gt;AI Watchdog — real-time visibility and intervention for AI coding agents. It ingests everything your agent does, decides with a multi-signal confidence model whether it's working / waiting / genuinely stuck, and notifies you only when you actually matter to the outcome.&lt;/p&gt;

&lt;p&gt;Why "stuck detection" is the hard part&lt;br&gt;
The naive rule — no output for 5 minutes → STUCK — is wrong, and I refused to ship it. Agents legitimately go silent while compiling, installing dependencies, running Docker, or grinding through a long test suite. Alert on that and you'll mute the tool within a day.&lt;/p&gt;

&lt;p&gt;So Watchdog runs a weighted ensemble of independent detectors, each producing an explainable signal:&lt;/p&gt;

&lt;p&gt;Signal                   Weight   Suppressed by&lt;br&gt;
No heartbeat               20     explicit waiting_for_user&lt;br&gt;
No tool activity       20     known long phases (build/test/docker)&lt;br&gt;
Repeated similar errors    25     one-off failures&lt;br&gt;
Repeated commands      15     parameterized commands&lt;br&gt;
No filesystem changes      10     non-coding executions&lt;br&gt;
Process looks idle     10     returns null if processes aren't visible&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbrjfs3zfxgc4m45zqrvb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbrjfs3zfxgc4m45zqrvb.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
Confidence is the clamped sum of triggered weights — and every contributing weight is traceable, so the UI can always answer "why does it think you're stuck?"&lt;/p&gt;

&lt;p&gt;Here's what I'm proudest of: the arithmetic works in your favor. Total silence scores 55 — five points below the 60 threshold. A coding execution with no file changes hits 65 and crosses. Silence is not evidence. That was a design rule, not a tuning accident.&lt;/p&gt;

&lt;p&gt;What's in it&lt;br&gt;
Live Command Center — every execution streaming over WebSocket&lt;br&gt;
Trace view — the full event tree, so you see the exact moment it derailed&lt;br&gt;
One command connects any agent: watchdog wrap -- kilo run "fix the failing test" Kilo, Claude Code, Codex, Aider, OpenCode, npm test, your own script. stdin passes through, exit code propagates, heartbeats every 15s.&lt;br&gt;
TypeScript + Python SDKs — five lines to instrument an existing agent&lt;br&gt;
VS Code extension — terminal commands and file changes&lt;br&gt;
A stuck-agent simulator with 5 modes, so you can prove the detector before trusting it&lt;br&gt;
Multi-tenant by construction — organizationId is required in every repo method, so a method that forgets it doesn't compile&lt;br&gt;
Fail-closed redaction — if the redaction module can't load, ingest throws rather than persisting raw secrets&lt;br&gt;
~1,500 tests, nothing mocked in e2e — real API process, real Postgres, real WebSocket, real child CLI&lt;br&gt;
The honest hard parts&lt;br&gt;
I shipped a cross-tenant read oracle and then found it. A global UNIQUE(dedupe_key) with an unscoped lookup meant anyone guessing stuck: could read another tenant's alert. Rescoped to (organization_id, dedupe_key).&lt;/p&gt;

&lt;p&gt;My partition helper could delete the dev database. pg_class matched across all schemas, so a parallel test schema made a test drop production partitions. I observed this rather than theorized about it. Now pinned to current_schema().&lt;/p&gt;

&lt;p&gt;State writes weren't compare-and-set — a read-then-write race. Now UPDATE ... WHERE id=$1 AND state=$2.&lt;/p&gt;

&lt;p&gt;Deliberate deviations: no Next.js (SSR buys nothing for an auth-gated realtime SPA), no ORM (hand-written SQL, deterministic, full control of monthly RANGE partitioning + CAS), and the detection engine is pure — no DB, no network, no process.env — which is exactly what makes it exhaustively testable.&lt;/p&gt;

&lt;p&gt;The load-bearing detail nobody sees: persisting hysteresis state is the write that matters. Skip it and STUCK detection silently never fires.&lt;/p&gt;

&lt;p&gt;What's next&lt;br&gt;
Extracting the sweeper to a real BullMQ worker with advisory locks, so multiple API replicas stop double-sweeping&lt;br&gt;
Turning on real adapter capabilities — supportsPauseResume and supportsProcessMetrics are false for every adapter today. I'd rather ship an honest capability matrix than pretend&lt;br&gt;
A real observability loop for editor-based agents — the VS Code extension can't see chat traffic inside another extension (Kilo, Cline, Roo); no API exposes it, so the localhost bridge plus terminal/file signals is the honest ceiling right now&lt;br&gt;
SSO/MFA, audit logs, richer notification rules, deeper CI integrations&lt;br&gt;
Why open source&lt;br&gt;
The interesting part — the detection methodology, the weight table, the suppression rules, the hysteresis design — is more valuable shared than defended.&lt;/p&gt;

&lt;p&gt;LLM tracing tools see model calls but not the shell, file, and git activity where agents actually get stuck. APM tools don't model agent execution at all. Agent platforms are closed and single-vendor. Nobody owns the neutral, execution-centric layer. I'd like this to be it.&lt;/p&gt;

&lt;p&gt;→ Try it at &lt;a href="https://aiwatchdog.vercel.app/" rel="noopener noreferrer"&gt;https://aiwatchdog.vercel.app/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you've ever lost 20 minutes wondering whether your agent was alive, I'd love to hear about it. And if your agent does something Watchdog doesn't understand yet, tell me — those reports are literally the roadmap.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>opensource</category>
      <category>showdev</category>
    </item>
  </channel>
</rss>
