<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ícaro Galvão do Nascimento</title>
    <description>The latest articles on DEV Community by Ícaro Galvão do Nascimento (@icaro0310).</description>
    <link>https://dev.to/icaro0310</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4166945%2F53aea24e-0267-48bc-a9cc-b90b9cc7af2f.jpg</url>
      <title>DEV Community: Ícaro Galvão do Nascimento</title>
      <link>https://dev.to/icaro0310</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/icaro0310"/>
    <language>en</language>
    <item>
      <title>Star for star, but I actually read your code first</title>
      <dc:creator>Ícaro Galvão do Nascimento</dc:creator>
      <pubDate>Tue, 06 Oct 2026 18:22:32 +0000</pubDate>
      <link>https://dev.to/icaro0310/star-for-star-but-i-actually-read-your-code-first-had</link>
      <guid>https://dev.to/icaro0310/star-for-star-but-i-actually-read-your-code-first-had</guid>
      <description>&lt;p&gt;Most star-for-star threads are hollow — people star without ever opening the repo. I'd rather do the honest version.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The deal:&lt;/strong&gt; drop your repo in the comments with one line about what it does. I'll read the code (not just the README), star it if it's real, and tell you what I actually think. Star back whichever of mine you find interesting — no obligation if nothing clicks.&lt;/p&gt;

&lt;p&gt;What I maintain, all open source, AI coding agent tooling:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://github.com/Icaro0310/poordjaevin" rel="noopener noreferrer"&gt;poordjaevin&lt;/a&gt;&lt;/strong&gt; — an MCP server that answers "should I trust this agent's output?" with calibrated probabilities (temperature-scaled NLI + conformal abstention, fully local, no API key)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://github.com/Icaro0310/devin-doctor" rel="noopener noreferrer"&gt;devin-doctor&lt;/a&gt;&lt;/strong&gt; — &lt;code&gt;brew doctor&lt;/code&gt; for a Devin install; detects stale tools across uv/pipx/npm/PATH wrappers&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://github.com/Icaro0310/devin-qa-pack" rel="noopener noreferrer"&gt;devin-qa-pack&lt;/a&gt;&lt;/strong&gt; — audits whether an agent's claims are backed by its real tool calls&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://github.com/Icaro0310/devin-devkit" rel="noopener noreferrer"&gt;devin-devkit&lt;/a&gt;&lt;/strong&gt; — one installer for all 19 tools, with &lt;code&gt;outdated&lt;/code&gt;/&lt;code&gt;update&lt;/code&gt; against a live registry&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything else: &lt;a href="https://github.com/Icaro0310" rel="noopener noreferrer"&gt;github.com/Icaro0310&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;What do you build?&lt;/p&gt;

</description>
      <category>discuss</category>
      <category>watercooler</category>
      <category>github</category>
      <category>opensource</category>
    </item>
    <item>
      <title>I built tools to verify what my AI coding agent actually did</title>
      <dc:creator>Ícaro Galvão do Nascimento</dc:creator>
      <pubDate>Tue, 06 Oct 2026 17:12:11 +0000</pubDate>
      <link>https://dev.to/icaro0310/i-built-tools-to-verify-what-my-ai-coding-agent-actually-did-2a6f</link>
      <guid>https://dev.to/icaro0310/i-built-tools-to-verify-what-my-ai-coding-agent-actually-did-2a6f</guid>
      <description>&lt;p&gt;AI coding agents are great at one thing that isn't writing code: &lt;em&gt;asserting&lt;/em&gt;. "Tests pass." "The file was updated." "I pushed the fix." And if you've run agents for real work, you know these claims are sometimes... optimistic.&lt;/p&gt;

&lt;p&gt;I'm a QA engineer who runs &lt;a href="https://devin.ai" rel="noopener noreferrer"&gt;Devin&lt;/a&gt; daily. At some point I got tired of manually checking whether the agent's claims matched reality, so I did what a QA engineer does: I built a test harness. Then an eval suite. Then a policy layer. Then a decision layer. Twenty local-first tools later, the whole stack runs on a RAM-constrained laptop with zero telemetry leaving the machine.&lt;/p&gt;

&lt;p&gt;This is the short version of what exists, why, and what I learned.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gap: agents assert, QA verifies
&lt;/h2&gt;

&lt;p&gt;The trigger was simple. An agent session reported "fixed, all green" — and the file on disk disagreed. No malice; agents report from their own narrative, not from ground truth. Classic QA problem: the &lt;em&gt;claim&lt;/em&gt; and the &lt;em&gt;state&lt;/em&gt; are different artifacts, and only one of them is evidence.&lt;/p&gt;

&lt;p&gt;The insight that made everything else possible: agent CLIs already write structured telemetry locally — session files, tool-call state, token usage. The evidence was sitting on disk the whole time. I just had to stop trusting the story and start reading the log.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five tools, one pipeline
&lt;/h2&gt;

&lt;p&gt;The tools compose into a sequence: understand → verify → measure → control → judge.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/Icaro0310/devin-internals-spec" rel="noopener noreferrer"&gt;devin-internals-spec&lt;/a&gt;&lt;/strong&gt; — &lt;em&gt;understand.&lt;/em&gt; Before you can trust tool output you need to know the contracts: the file formats, exit codes, and behavioral rules of the runtime itself. This is the spec layer — everything downstream reads against it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/Icaro0310/devin-qa-pack" rel="noopener noreferrer"&gt;devin-qa-pack&lt;/a&gt;&lt;/strong&gt; — &lt;em&gt;verify.&lt;/em&gt; Runs a QA audit over actual agent work: file diffs present, tests run, commits exist, pushes landed, verification commands executed. Claims are checked against &lt;code&gt;tool_call_state&lt;/code&gt;, not against the agent's narrative. 47 tests, CI on Ubuntu + Windows. This is the tool that started it all.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/Icaro0310/devin-evals" rel="noopener noreferrer"&gt;devin-evals&lt;/a&gt;&lt;/strong&gt; — &lt;em&gt;measure.&lt;/em&gt; Once verification works, you can ask the harder question: how good is the agent on &lt;em&gt;this&lt;/em&gt; kind of task? Golden tasks, rubric scoring, regression tracking across sessions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/Icaro0310/devin-bridge" rel="noopener noreferrer"&gt;devin-bridge&lt;/a&gt;&lt;/strong&gt; — &lt;em&gt;control.&lt;/em&gt; A policy gate between intent and execution — ACP-based control with a &lt;code&gt;--devin-only&lt;/code&gt; mode that enforces "this session does exactly what it was scoped to do."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/Icaro0310/poordjaevin" rel="noopener noreferrer"&gt;poordjaevin&lt;/a&gt;&lt;/strong&gt; — &lt;em&gt;judge.&lt;/em&gt; The piece I'm proudest of. Takes a task description and produces a calibrated confidence score: should I trust this delegation? The calibration story is the interesting part — the judge went from &lt;strong&gt;ECE 0.170 to 0.071&lt;/strong&gt; through iterative refinement on real session data. (ECE = expected calibration error: when the judge says "80% confident," it should be right ~80% of the time. Most confidence scores don't do this. Now mine roughly does.) It's also a real MCP server — &lt;code&gt;poordjaevin serve&lt;/code&gt; — so agents can consult it mid-flight.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "local-first" actually buys you
&lt;/h2&gt;

&lt;p&gt;Every tool follows the same contract: read local agent state, write local artifacts, expose a stable CLI, degrade gracefully when optional infrastructure is absent. No required daemons, no SaaS dashboard, no telemetry.&lt;/p&gt;

&lt;p&gt;Concretely:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;devin-history&lt;/code&gt; — cross-session memory: searchable SQLite over all past sessions (7+ commands of grep-able agent archaeology)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;devin-metrics&lt;/code&gt; — telemetry aggregation for quality signals over time&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;devin-memory&lt;/code&gt; + &lt;code&gt;devin-search&lt;/code&gt; + &lt;code&gt;devin-graph&lt;/code&gt; — memory store, retrieval, and a knowledge graph over sessions, projects and decisions&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;devin-doctor&lt;/code&gt; — environment diagnostics across Windows/Linux&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;devin-backup&lt;/code&gt; — snapshot + verify + restore of agent state&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;devin-janitor&lt;/code&gt; + &lt;code&gt;devin-redact&lt;/code&gt; — cleanup and secret-redaction so state can move safely&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;devin-office&lt;/code&gt; — the fun one: a live circuit-board dashboard that renders real sessions, subagents and tool calls from the local store. Purely visual, read-only, zero telemetry — and it makes for a great demo GIF.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Three lessons from building this
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. The agent's own telemetry is an untapped QA datasource.&lt;/strong&gt; Session files and tool-call state are structured evidence most people ignore. Reading them turns "trust the agent" into "verify the agent" — and verification is what makes delegation safe at scale.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Calibration &amp;gt; confidence.&lt;/strong&gt; A judge that's confidently wrong is worse than no judge. Iterating on ECE (0.170 → 0.071) took real labeled outcomes, not prompt tweaks. If you build any kind of AI decision layer, measure calibration explicitly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Devin-only mode matters more than integrations.&lt;/strong&gt; Every tool works standalone with just the agent's CLI present. Obsidian, Slack, MCP — all optional. The ecosystem must survive a locked-down corporate box with nothing but Devin installed, because that's where a lot of real work happens.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it / tear it apart
&lt;/h2&gt;

&lt;p&gt;Everything is MIT-licensed, cross-platform, and documented in English + PT-BR:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Profile/catalog:&lt;/strong&gt; &lt;a href="https://github.com/Icaro0310" rel="noopener noreferrer"&gt;github.com/Icaro0310&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Website:&lt;/strong&gt; &lt;a href="https://icaro0310.github.io" rel="noopener noreferrer"&gt;icaro0310.github.io&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Start here:&lt;/strong&gt; &lt;code&gt;devin-qa-pack&lt;/code&gt; (the verifier) or &lt;code&gt;poordjaevin&lt;/code&gt; (the calibrated judge — &lt;code&gt;pip install poordjaevin&lt;/code&gt; / &lt;code&gt;uv tool install poordjaevin&lt;/code&gt;)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Honest question for the comments: if you run AI agents on real work — Copilot, Devin, Cursor, Claude Code — &lt;strong&gt;how do you verify their claims today?&lt;/strong&gt; Manual spot-checks? CI gates? Nothing? I suspect "nothing" is the most common answer, and I built this stack partly because it scared me that it was mine.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>devops</category>
      <category>tools</category>
    </item>
  </channel>
</rss>
