<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Maryan K</title>
    <description>The latest articles on DEV Community by Maryan K (@maryan_k_bef6cf83fa64e809).</description>
    <link>https://dev.to/maryan_k_bef6cf83fa64e809</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4036373%2F3a4f192b-4cd0-4f57-a285-283d8f8515dc.png</url>
      <title>DEV Community: Maryan K</title>
      <link>https://dev.to/maryan_k_bef6cf83fa64e809</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/maryan_k_bef6cf83fa64e809"/>
    <language>en</language>
    <item>
      <title>Weekly Startup Engineering Signal Report: 2026-08-20</title>
      <dc:creator>Maryan K</dc:creator>
      <pubDate>Mon, 31 Aug 2026 20:10:30 +0000</pubDate>
      <link>https://dev.to/maryan_k_bef6cf83fa64e809/weekly-startup-engineering-signal-report-2026-08-20-8ki</link>
      <guid>https://dev.to/maryan_k_bef6cf83fa64e809/weekly-startup-engineering-signal-report-2026-08-20-8ki</guid>
      <description>&lt;p&gt;Originally published on &lt;a href="https://signals.gitdealflow.com/blog/weekly-signal-report-2026-08-20" rel="noopener noreferrer"&gt;VC Deal Flow Signal&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This is the weekly engineering-signal report from &lt;a href="https://signals.gitdealflow.com" rel="noopener noreferrer"&gt;VC Deal Flow Signal (signals.gitdealflow.com)&lt;/a&gt;. It ranks startup engineering acceleration from public GitHub activity, a leading indicator of fundraise announcements by 3-6 weeks.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Top movers this week
&lt;/h2&gt;

&lt;p&gt;The fastest-accelerating startups in the 350+ tracked panel:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Startup&lt;/th&gt;
&lt;th&gt;Stage&lt;/th&gt;
&lt;th&gt;Geography&lt;/th&gt;
&lt;th&gt;14d velocity change&lt;/th&gt;
&lt;th&gt;Signal&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;liberusoftware&lt;/td&gt;
&lt;td&gt;Seed&lt;/td&gt;
&lt;td&gt;UK&lt;/td&gt;
&lt;td&gt;+3275%&lt;/td&gt;
&lt;td&gt;Engineering hiring burst&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;untrustedmodders&lt;/td&gt;
&lt;td&gt;Pre-seed&lt;/td&gt;
&lt;td&gt;Unknown&lt;/td&gt;
&lt;td&gt;+2900%&lt;/td&gt;
&lt;td&gt;Deploy frequency spike&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;fleetbase&lt;/td&gt;
&lt;td&gt;Pre-seed&lt;/td&gt;
&lt;td&gt;APAC&lt;/td&gt;
&lt;td&gt;+1600%&lt;/td&gt;
&lt;td&gt;Engineering hiring burst&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;aphp&lt;/td&gt;
&lt;td&gt;Pre-seed&lt;/td&gt;
&lt;td&gt;EU&lt;/td&gt;
&lt;td&gt;+999%&lt;/td&gt;
&lt;td&gt;Infrastructure buildout&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;StanfordSpezi&lt;/td&gt;
&lt;td&gt;Pre-seed&lt;/td&gt;
&lt;td&gt;US&lt;/td&gt;
&lt;td&gt;+999%&lt;/td&gt;
&lt;td&gt;Deploy frequency spike&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ere-health&lt;/td&gt;
&lt;td&gt;Series A/B&lt;/td&gt;
&lt;td&gt;EU&lt;/td&gt;
&lt;td&gt;+999%&lt;/td&gt;
&lt;td&gt;Deploy frequency spike&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;epic-open-source&lt;/td&gt;
&lt;td&gt;Seed&lt;/td&gt;
&lt;td&gt;Unknown&lt;/td&gt;
&lt;td&gt;+999%&lt;/td&gt;
&lt;td&gt;Engineering hiring burst&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;insightsengineering&lt;/td&gt;
&lt;td&gt;Seed&lt;/td&gt;
&lt;td&gt;EU&lt;/td&gt;
&lt;td&gt;+999%&lt;/td&gt;
&lt;td&gt;Deploy frequency spike&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Sector activity
&lt;/h2&gt;

&lt;p&gt;The 15 tracked sectors, ranked by startup count:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Web3&lt;/strong&gt; - 42 startups&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gaming&lt;/strong&gt; - 39 startups&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;EdTech&lt;/strong&gt; - 37 startups&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data Infrastructure&lt;/strong&gt; - 35 startups&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enterprise SaaS&lt;/strong&gt; - 31 startups&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Robotics&lt;/strong&gt; - 28 startups&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  About the data
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Panel: 350+ startups, refreshed every Monday.&lt;/li&gt;
&lt;li&gt;Period: Q3 2026.&lt;/li&gt;
&lt;li&gt;Methodology: &lt;a href="https://signals.gitdealflow.com/methodology" rel="noopener noreferrer"&gt;https://signals.gitdealflow.com/methodology&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Free JSON and CSV: &lt;a href="https://signals.gitdealflow.com/api/signals.json" rel="noopener noreferrer"&gt;https://signals.gitdealflow.com/api/signals.json&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Citation: VC Deal Flow Signal (signals.gitdealflow.com), Q3 2026 data.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Use the live data without a dashboard
&lt;/h2&gt;

&lt;p&gt;The weekly table above is generated from the same public endpoint. Here is a direct shell check:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; https://signals.gitdealflow.com/api/signals.json &lt;span class="se"&gt;\&lt;/span&gt;
  | jq &lt;span class="s1"&gt;'.trending[:3] | map({name, signalType, commitVelocityChange})'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For an AI client, run the MCP server locally:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:{&lt;/span&gt;&lt;span class="nl"&gt;"vc-deal-flow-signal"&lt;/span&gt;&lt;span class="p"&gt;:{&lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"npx"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"args"&lt;/span&gt;&lt;span class="p"&gt;:[&lt;/span&gt;&lt;span class="s2"&gt;"-y"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s2"&gt;"@gitdealflow/mcp-signal"&lt;/span&gt;&lt;span class="p"&gt;]}}}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then ask: "Which startups are accelerating fastest this week?" Both examples read the live public dataset. They do not make a funding prediction.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>I Made Valet Damage Provable With GPS, Dual Timestamps, and SHA-256</title>
      <dc:creator>Maryan K</dc:creator>
      <pubDate>Fri, 14 Aug 2026 11:07:10 +0000</pubDate>
      <link>https://dev.to/maryan_k_bef6cf83fa64e809/i-made-valet-damage-provable-with-gps-dual-timestamps-and-sha-256-3l06</link>
      <guid>https://dev.to/maryan_k_bef6cf83fa64e809/i-made-valet-damage-provable-with-gps-dual-timestamps-and-sha-256-3l06</guid>
      <description>&lt;h1&gt;
  
  
  I Made Valet Damage Provable With GPS, Dual Timestamps, and SHA-256
&lt;/h1&gt;

&lt;p&gt;When you hand your keys to a valet and the car comes back with a scratch, the burden of proof is on you. A photo you took on your phone is almost useless: there is no proof of when or where it was taken. I built CarShake to turn a before-and-after photo into a record that is hard to argue with.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three layers of evidence
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;GPS.&lt;/strong&gt; Every photo is stamped with where it was taken, so "that scratch was already there" meets "here is the geotag at the valet stand".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dual timestamps.&lt;/strong&gt; One from the device, one from the server. Timing that does not depend on a clock you could have set.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SHA-256 hashing.&lt;/strong&gt; Each photo record is hashed and stored immutably, so once a record exists it cannot be edited later, by you or anyone else.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The before-and-after flow: photograph all four sides, wheels, windshield, and any existing damage before you hand over the keys. Repeat when you get the car back. A computer-vision pass flags new dents and scratches you might miss in bad light.&lt;/p&gt;

&lt;h2&gt;
  
  
  The "Instant Proof" trick that gets people in the door
&lt;/h2&gt;

&lt;p&gt;The full product is a guided before-and-after capture. But there is also a free Instant Proof tool that just stamps date, time, and location on any photo in about thirty seconds. That single feature is the distribution: it is useful on its own, with no account.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pricing is deliberately boring
&lt;/h2&gt;

&lt;p&gt;The core capture flow is free. A $7 premium kit adds a printable checklist, a twelve-month photo log, and a shareable verification report. One-time, not a subscription. 200+ drivers have signed up for the free checklist.&lt;/p&gt;

&lt;p&gt;This is a small, honest product for a specific fear: being charged for damage you did not cause. If you build in the evidence-integrity space, the lesson is that the boring parts (timestamp, geotag, hash) are the whole product.&lt;/p&gt;

</description>
      <category>showdev</category>
      <category>saas</category>
      <category>security</category>
      <category>ai</category>
    </item>
    <item>
      <title>I Bundled 5 AI Tools Into a $0.97 System for Building a Side Income</title>
      <dc:creator>Maryan K</dc:creator>
      <pubDate>Fri, 14 Aug 2026 11:06:12 +0000</pubDate>
      <link>https://dev.to/maryan_k_bef6cf83fa64e809/i-bundled-5-ai-tools-into-a-097-system-for-building-a-side-income-1hec</link>
      <guid>https://dev.to/maryan_k_bef6cf83fa64e809/i-bundled-5-ai-tools-into-a-097-system-for-building-a-side-income-1hec</guid>
      <description>&lt;h1&gt;
  
  
  I Bundled 5 AI Tools Into a $0.97 System for Building a Side Income
&lt;/h1&gt;

&lt;p&gt;The founder who wants a side income has a stack of problems that are all the same shape: how much do I actually need, what do I build, is this allowed under my contract, how do I launch, and how do I market it without it becoming a full-time job. I built Invisible Exit as five connected tools for those five questions.&lt;/p&gt;

&lt;h2&gt;
  
  
  The five tools
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;A freedom-number dashboard.&lt;/strong&gt; Your recurring revenue, churn, and growth against the number you actually need to reach.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An idea pipeline.&lt;/strong&gt; Hundreds of micro-SaaS ideas scored by fit, time investment, and revenue potential.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A compliance hub.&lt;/strong&gt; Entity separation and a check against common contract clauses, non-compete and IP assignment.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Launch control.&lt;/strong&gt; Stripe, landing page, and a launch sequence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A brand builder.&lt;/strong&gt; Content and audience playbooks that do not require your face or your real name.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;One subscription, all five. The founding price is $0.97 a month.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part I get asked about: "is this allowed?"
&lt;/h2&gt;

&lt;p&gt;The honest answer is that it depends on your contract, and the product is built around that fact. The compliance tool is there specifically to check your own agreement before you start, and to keep the business in a separate structure. I am not going to tell you to break a non-compete. I built a tool that flags the risk so you can decide with your eyes open.&lt;/p&gt;

&lt;h2&gt;
  
  
  The math that anchors it
&lt;/h2&gt;

&lt;p&gt;The target is concrete: a $29/month micro-SaaS with about 138 customers is $4,000 a month in recurring revenue. That number is what the whole system is organized around. If you do not reach $4K/month within twelve months, there is a refund.&lt;/p&gt;

&lt;p&gt;This is a niche product for people who want an exit ramp, not a get-rich promise. The honest version is "five tools, one low price, and the work is still yours to do."&lt;/p&gt;

</description>
      <category>saas</category>
      <category>startup</category>
      <category>showdev</category>
      <category>productivity</category>
    </item>
    <item>
      <title>I Shipped 12 Products Nobody Paid For. The Fix Was a Playbook, Not More Code.</title>
      <dc:creator>Maryan K</dc:creator>
      <pubDate>Fri, 14 Aug 2026 11:06:06 +0000</pubDate>
      <link>https://dev.to/maryan_k_bef6cf83fa64e809/i-shipped-12-products-nobody-paid-for-the-fix-was-a-playbook-not-more-code-3bgp</link>
      <guid>https://dev.to/maryan_k_bef6cf83fa64e809/i-shipped-12-products-nobody-paid-for-the-fix-was-a-playbook-not-more-code-3bgp</guid>
      <description>&lt;h1&gt;
  
  
  I Shipped 12 Products Nobody Paid For. The Fix Was a Playbook, Not More Code.
&lt;/h1&gt;

&lt;p&gt;This is a build-in-public confession. I shipped twelve products that nobody paid for, and spent a year convinced the problem was traffic. It was not. The problem was that I never did the uncomfortable work: name one real person, make one real promise, and send one real message.&lt;/p&gt;

&lt;h2&gt;
  
  
  The product is the playbook, not another course
&lt;/h2&gt;

&lt;p&gt;Unlock SaaS is a web app, not a Notion template, and not a course. It is a seven-step playbook engine for the post-launch, pre-revenue founder. It pushes back on vague answers until three things exist: one named real person, one real offer, one real outreach message. It tracks the work in software instead of in your willpower.&lt;/p&gt;

&lt;p&gt;The reason it is software and not a PDF is specific: the avoidance option has to be removed. A template you can close is not a plan. A tool that shows you step four is unfinished, and step one was never done, is.&lt;/p&gt;

&lt;h2&gt;
  
  
  The guarantee is enforced by code
&lt;/h2&gt;

&lt;p&gt;Day 60 is when the guarantee fires. If you did the work (steps one through five done in product, twenty outreach actions logged) and your Stripe line is still zero, the refund runs automatically. No support ticket, no "we will review your case". I built the refund into the engine because a guarantee that requires a human to approve it is not a guarantee.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who this is for, exactly
&lt;/h2&gt;

&lt;p&gt;The founder who shipped something real with Lovable, Claude, Replit, or Bolt. Users signed up. Nobody paid. They are starting to think the product is the problem. It is not. The problem is the work nobody taught us to do.&lt;/p&gt;

&lt;p&gt;The founding price is $49/month for life, and it closes at 100 builders. If you are in that post-launch flatline, the honest question is not "how do I get more traffic". It is "have I named one real person and made one real promise". Most of us have not.&lt;/p&gt;

</description>
      <category>saas</category>
      <category>startup</category>
      <category>showdev</category>
      <category>productivity</category>
    </item>
    <item>
      <title>I Built an AI Employee That Files Invoices and Refuses to Guess</title>
      <dc:creator>Maryan K</dc:creator>
      <pubDate>Fri, 14 Aug 2026 11:04:43 +0000</pubDate>
      <link>https://dev.to/maryan_k_bef6cf83fa64e809/i-built-an-ai-employee-that-files-invoices-and-refuses-to-guess-2dgm</link>
      <guid>https://dev.to/maryan_k_bef6cf83fa64e809/i-built-an-ai-employee-that-files-invoices-and-refuses-to-guess-2dgm</guid>
      <description>&lt;h1&gt;
  
  
  I Built an AI Employee That Files Invoices and Refuses to Guess
&lt;/h1&gt;

&lt;p&gt;The "AI agent" promise got way ahead of the product. Nika is the opposite: a deliberately narrow AI employee that does one job, files invoices, and is designed to stop and ask rather than guess.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why an "employee" and not an app
&lt;/h2&gt;

&lt;p&gt;The founder who still does their own bookkeeping does not need another dashboard to learn. They need the invoices handled. So Nika is positioned as an employee you hire in five minutes. She watches a mailbox, pulls every invoice, files it into your records, and asks before she ever guesses.&lt;/p&gt;

&lt;p&gt;You pay per finished invoice, $0.40, and $5 gets her started. No subscription, no card required to look around.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one design decision that matters: ask, do not guess
&lt;/h2&gt;

&lt;p&gt;Every AI bookkeeping demo shows a model confidently extracting a supplier name and a VAT number and being wrong. The interesting engineering problem is not extraction, it is confidence. Nika routes anything she is not sure about to a human with a specific question, instead of filing a guess.&lt;/p&gt;

&lt;p&gt;A filed error is expensive to unwind. A held item waiting for an OK is cheap. The copy on the site says it plainly: "Anything she isn't sure about waits for your OK." That sentence is the product.&lt;/p&gt;

&lt;h2&gt;
  
  
  Per-task pricing instead of seats
&lt;/h2&gt;

&lt;p&gt;A part-time bookkeeper for four hours of weekly invoice work does not justify a salary, holiday, and training. Per-finished-task pricing matches the actual job: it is four hours a week, not a role. No work done, no charge.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is live and what is not
&lt;/h2&gt;

&lt;p&gt;Nika is the only live employee today. The site teases more employees (a speed-to-lead caller, a win-back agent, a review manager), but those are clearly marked "coming soon", and I will not pretend they are shipped. One real employee doing one job well is the honest version of this product.&lt;/p&gt;

&lt;p&gt;The stack is a Next.js app with a three-language site (English, Greek, Ukrainian). If you are building an "AI agent" product, my advice is to shrink the scope until the agent does one verifiable job and nothing else. That is where the trust is.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>showdev</category>
      <category>saas</category>
      <category>automation</category>
    </item>
    <item>
      <title>I Turned Spoken Construction Reports Into Court-Ready Evidence</title>
      <dc:creator>Maryan K</dc:creator>
      <pubDate>Fri, 14 Aug 2026 11:04:40 +0000</pubDate>
      <link>https://dev.to/maryan_k_bef6cf83fa64e809/i-turned-spoken-construction-reports-into-court-ready-evidence-1124</link>
      <guid>https://dev.to/maryan_k_bef6cf83fa64e809/i-turned-spoken-construction-reports-into-court-ready-evidence-1124</guid>
      <description>&lt;h1&gt;
  
  
  I Turned Spoken Construction Reports Into Court-Ready Evidence
&lt;/h1&gt;

&lt;p&gt;Subcontractors lose money because their documentation is paper. A general contractor asks "prove you were on site Tuesday", and the sub is holding a crumpled notepad. I built VoiceLogPro to fix that: you talk through your day on site, and it becomes a timestamped PDF that holds up when a payment is disputed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem is evidence, not convenience
&lt;/h2&gt;

&lt;p&gt;A daily construction report matters for exactly one reason: it is contemporaneous evidence. When a mechanic's lien or a delay claim reaches a hearing, the question is never "did you keep nice notes". It is "can you prove, with a dated record, that you did the work and flagged the delay when it happened".&lt;/p&gt;

&lt;p&gt;Paper fails that test. So do most apps, which are built for project managers behind a desk, not for an electrician holding a phone in one hand and a drill in the other.&lt;/p&gt;

&lt;p&gt;VoiceLogPro is voice-first on purpose. You speak the report: weather, crew, work performed, materials, any delay or change order. The app turns it into a structured PDF with those fields filled in. No typing on a jobsite.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "court-admissible" actually means in the code
&lt;/h2&gt;

&lt;p&gt;Three things, and they are boring on purpose.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;A timestamp.&lt;/strong&gt; The report is stamped when you record it, not when you export it later. That distinction is the first thing a lawyer asks about.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fixed fields.&lt;/strong&gt; Weather, crew, work performed, materials, delays. Every report has the same shape, which is what makes it comparable across days and defensible as a series.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A PDF you can hand over.&lt;/strong&gt; Not a link, not an app someone has to log into. A file that can be attached to a claim.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I did not try to build another Procore or Raken. Those are project management platforms. VoiceLogPro is a documentation defense for the trade subcontractor who needs to get paid.&lt;/p&gt;

&lt;h2&gt;
  
  
  The $40,000 story behind it
&lt;/h2&gt;

&lt;p&gt;The dream customer is an electrician running three crews. Last year he lost about $40,000 because his documentation was paper and a general contractor knew it. That is the whole pitch. I am not selling productivity. I am selling proof.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it is now
&lt;/h2&gt;

&lt;p&gt;It is a $49/month crew plan, live at voicelogpro.com, with a waitlist for beta. The site also carries state lien-law guides (Texas Chapter 53, California 20-day preliminary notices, and others) because the documentation and the legal deadline are the same problem. A report filed a day late is a report that does not protect your lien rights.&lt;/p&gt;

&lt;p&gt;If you build anything for the trades, my one piece of advice: build for the phone in one hand, not the desk. The jobsite decides what actually gets used.&lt;/p&gt;

</description>
      <category>showdev</category>
      <category>webdev</category>
      <category>saas</category>
      <category>programming</category>
    </item>
    <item>
      <title>The 52-Point GitHub Due Diligence Checklist I Use Before Every Seed Investment</title>
      <dc:creator>Maryan K</dc:creator>
      <pubDate>Fri, 14 Aug 2026 08:51:22 +0000</pubDate>
      <link>https://dev.to/maryan_k_bef6cf83fa64e809/the-52-point-github-due-diligence-checklist-i-use-before-every-seed-investment-1kjf</link>
      <guid>https://dev.to/maryan_k_bef6cf83fa64e809/the-52-point-github-due-diligence-checklist-i-use-before-every-seed-investment-1kjf</guid>
      <description>&lt;p&gt;Most VC due diligence checklists focus on financials, legal, and market size. They miss the single most honest signal a startup has: its GitHub.&lt;/p&gt;

&lt;p&gt;A pitch deck tells you what the founders &lt;em&gt;want&lt;/em&gt; you to know. Their Git history tells you what they actually &lt;em&gt;did&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;After analyzing GitHub data from 219 funded startups (paper &lt;a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6606558" rel="noopener noreferrer"&gt;here&lt;/a&gt;), I built a 52-point checklist. Here it is, free.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why GitHub matters in due diligence
&lt;/h2&gt;

&lt;p&gt;Three things show up in code that never appear in a deck:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Actual shipping velocity&lt;/strong&gt; (not claimed velocity)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Team health&lt;/strong&gt; (are contributors leaving silently?)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Infrastructure maturity&lt;/strong&gt; (are they building for scale or winging it?)&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The checklist (abridged)
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Repository health (12 checks)
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Commits in last 30 days?&lt;/li&gt;
&lt;li&gt;No abandoned repos (&amp;gt;90 days inactive)?&lt;/li&gt;
&lt;li&gt;README and LICENSE present?&lt;/li&gt;
&lt;li&gt;CI/CD workflows configured?&lt;/li&gt;
&lt;li&gt;Branch protection enabled?&lt;/li&gt;
&lt;li&gt;PR merge rate &amp;gt;50%?&lt;/li&gt;
&lt;li&gt;No critical CVEs in dependencies?&lt;/li&gt;
&lt;li&gt;SECURITY.md present?&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Commit velocity (8 checks)
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Weekly commits stable or increasing over 90 days?&lt;/li&gt;
&lt;li&gt;No sudden drops &amp;gt;60% week-over-week?&lt;/li&gt;
&lt;li&gt;Commit messages follow a consistent pattern?&lt;/li&gt;
&lt;li&gt;Night/weekend commits (genuine engagement)?&lt;/li&gt;
&lt;li&gt;No suspicious bot-generated patterns?&lt;/li&gt;
&lt;li&gt;Velocity correlates with claimed headcount?&lt;/li&gt;
&lt;li&gt;Spikes align with claimed milestones?&lt;/li&gt;
&lt;li&gt;Compare to &lt;a href="https://signals.gitdealflow.com" rel="noopener noreferrer"&gt;sector benchmarks&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Team signals (10 checks)
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;More than 3 active contributors in last 90 days?&lt;/li&gt;
&lt;li&gt;Contributor count growing quarter-over-quarter?&lt;/li&gt;
&lt;li&gt;No mass contributor exodus (&amp;gt;40% drop)?&lt;/li&gt;
&lt;li&gt;Key engineers have LinkedIn profiles matching GitHub?&lt;/li&gt;
&lt;li&gt;CTO/founder still commits code?&lt;/li&gt;
&lt;li&gt;No ghost contributors (empty profiles)?&lt;/li&gt;
&lt;li&gt;Bot accounts labeled and scoped?&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Infrastructure buildout (8 checks)
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;CI/CD pipeline configured (GitHub Actions, etc.)?&lt;/li&gt;
&lt;li&gt;Containerization present (Dockerfile)?&lt;/li&gt;
&lt;li&gt;Infrastructure-as-code (Terraform, Pulumi)?&lt;/li&gt;
&lt;li&gt;Monitoring/observability configs present?&lt;/li&gt;
&lt;li&gt;Staging environment exists?&lt;/li&gt;
&lt;li&gt;Feature flags configured?&lt;/li&gt;
&lt;li&gt;API documentation published?&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Security posture (6 checks)
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Dependabot or equivalent enabled?&lt;/li&gt;
&lt;li&gt;No public secrets/keys in commit history?&lt;/li&gt;
&lt;li&gt;2FA enforcement visible?&lt;/li&gt;
&lt;li&gt;Code scanning (CodeQL/Semgrep) enabled?&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Product-market fit signals from code (8 checks)
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Rate of new feature branches increasing?&lt;/li&gt;
&lt;li&gt;Customer-facing repos (SDKs, integrations) active?&lt;/li&gt;
&lt;li&gt;Localization efforts in code?&lt;/li&gt;
&lt;li&gt;Billing integration present and current?&lt;/li&gt;
&lt;li&gt;Database migration scripts (real product evolution)?&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How to use it
&lt;/h2&gt;

&lt;p&gt;Run this checklist against any startup's public GitHub &lt;em&gt;before&lt;/em&gt; the pitch meeting. Most of it can be automated using the &lt;a href="https://docs.github.com/en/rest" rel="noopener noreferrer"&gt;GitHub API&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;I track 350+ startups across 15 sectors automatically at &lt;a href="https://signals.gitdealflow.com" rel="noopener noreferrer"&gt;VC Deal Flow Signal&lt;/a&gt;, which scores each startup on engineering acceleration and flags when signals spike before a funding round.&lt;/p&gt;

&lt;h2&gt;
  
  
  The sector benchmark trap
&lt;/h2&gt;

&lt;p&gt;Don't compare a Healthcare startup to a Developer Tools startup. Developer Tools companies ship 2.3x more than Healthcare. Always compare within the same sector.&lt;/p&gt;

&lt;p&gt;Here are median weekly commits by sector (Q2 2026):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Sector&lt;/th&gt;
&lt;th&gt;Median weekly commits&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Developer Tools&lt;/td&gt;
&lt;td&gt;89&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data Infrastructure&lt;/td&gt;
&lt;td&gt;52&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Web3&lt;/td&gt;
&lt;td&gt;56&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AI/ML&lt;/td&gt;
&lt;td&gt;47&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Enterprise SaaS&lt;/td&gt;
&lt;td&gt;38&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fintech&lt;/td&gt;
&lt;td&gt;34&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Healthcare&lt;/td&gt;
&lt;td&gt;22&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Full sector breakdown with 350+ startups is at &lt;a href="https://signals.gitdealflow.com" rel="noopener noreferrer"&gt;signals.gitdealflow.com&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest part
&lt;/h2&gt;

&lt;p&gt;My first 10 pre-registered predictions went 0-for-10. I published that result anyway. The methodology improved. The current backtest covers 219 fundraises.&lt;/p&gt;

&lt;p&gt;The point isn't that GitHub signals are perfect. They're a &lt;em&gt;leading&lt;/em&gt; indicator that gives you 30-47 days of runway before a round is announced. Use it to get in early, then verify with traditional diligence.&lt;/p&gt;

&lt;p&gt;Full methodology and the 0-for-10 transparency ledger: &lt;a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6606558" rel="noopener noreferrer"&gt;SSRN paper&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The full 52-point checklist is also on GitHub: &lt;a href="https://github.com/kindrat86/vc-due-diligence-checklist" rel="noopener noreferrer"&gt;vc-due-diligence-checklist&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;The author runs &lt;a href="https://gitdealflow.com" rel="noopener noreferrer"&gt;GitDealFlow&lt;/a&gt;, a GitHub-powered deal flow intelligence tool. All data and methodology are open. CC BY 4.0.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>vc</category>
      <category>startup</category>
      <category>github</category>
      <category>duediligence</category>
    </item>
    <item>
      <title>I Tracked 350 Startups' GitHub Activity for 6 Months. Here's What Predicts Fundraises.</title>
      <dc:creator>Maryan K</dc:creator>
      <pubDate>Fri, 14 Aug 2026 08:48:18 +0000</pubDate>
      <link>https://dev.to/maryan_k_bef6cf83fa64e809/i-tracked-350-startups-github-activity-for-6-months-heres-what-predicts-fundraises-4e11</link>
      <guid>https://dev.to/maryan_k_bef6cf83fa64e809/i-tracked-350-startups-github-activity-for-6-months-heres-what-predicts-fundraises-4e11</guid>
      <description>&lt;p&gt;Six months ago I started tracking the GitHub activity of 350+ startups across 15 sectors. The goal: find public signals that predict which startups are about to raise.&lt;/p&gt;

&lt;p&gt;I backtested against 219 actual fundraises. Here's what I found.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;Every week, my system pulls data from the GitHub API for 350+ startup organizations:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Commit counts per repository&lt;/li&gt;
&lt;li&gt;Active contributor counts&lt;/li&gt;
&lt;li&gt;New repository creation&lt;/li&gt;
&lt;li&gt;Infrastructure file changes (CI/CD, Dockerfiles, monitoring configs)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I normalize by sector (a Healthcare startup ships differently than a Dev Tools startup) and score each startup on an Engineering Acceleration Score.&lt;/p&gt;

&lt;p&gt;Live data is at &lt;a href="https://signals.gitdealflow.com" rel="noopener noreferrer"&gt;signals.gitdealflow.com&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three signals that predict fundraises
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Commit velocity acceleration (the strongest signal)
&lt;/h3&gt;

&lt;p&gt;Startups that raised within 90 days showed a median &lt;strong&gt;34% week-over-week increase&lt;/strong&gt; in commit velocity. The pattern is distinctive: steady baseline for months, then a sharp ramp.&lt;/p&gt;

&lt;p&gt;This makes intuitive sense. Teams push hard to ship milestones before pitching investors. The code acceleration starts before the deck is finished.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Contributor count spike
&lt;/h3&gt;

&lt;p&gt;Fundraising startups added a median of &lt;strong&gt;2.3 new active contributors&lt;/strong&gt; in the 60 days before their round. This reflects pre-round hiring pushes. You're building the team before you ask for money.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Infrastructure buildout
&lt;/h3&gt;

&lt;p&gt;Startups about to raise were &lt;strong&gt;3.2x more likely&lt;/strong&gt; to add new CI/CD workflows, Dockerfiles, or monitoring configs in the 90 days pre-round. This signals preparing for scale (and investor demos).&lt;/p&gt;

&lt;h2&gt;
  
  
  Sector benchmarks (Q2 2026)
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Sector&lt;/th&gt;
&lt;th&gt;Startups&lt;/th&gt;
&lt;th&gt;Median weekly commits&lt;/th&gt;
&lt;th&gt;Top quartile&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;AI &amp;amp; ML&lt;/td&gt;
&lt;td&gt;67&lt;/td&gt;
&lt;td&gt;47&lt;/td&gt;
&lt;td&gt;112&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Developer Tools&lt;/td&gt;
&lt;td&gt;34&lt;/td&gt;
&lt;td&gt;89&lt;/td&gt;
&lt;td&gt;203&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fintech&lt;/td&gt;
&lt;td&gt;42&lt;/td&gt;
&lt;td&gt;34&lt;/td&gt;
&lt;td&gt;78&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Web3&lt;/td&gt;
&lt;td&gt;28&lt;/td&gt;
&lt;td&gt;56&lt;/td&gt;
&lt;td&gt;134&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Healthcare&lt;/td&gt;
&lt;td&gt;19&lt;/td&gt;
&lt;td&gt;22&lt;/td&gt;
&lt;td&gt;51&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Enterprise SaaS&lt;/td&gt;
&lt;td&gt;45&lt;/td&gt;
&lt;td&gt;38&lt;/td&gt;
&lt;td&gt;84&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Full table across all 15 sectors at &lt;a href="https://signals.gitdealflow.com/data/vc-signal-benchmarks" rel="noopener noreferrer"&gt;signals.gitdealflow.com&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest result
&lt;/h2&gt;

&lt;p&gt;My first pre-registered cohort went &lt;strong&gt;0-for-10&lt;/strong&gt;. Zero correct predictions out of 10 startups I said would raise within 90 days.&lt;/p&gt;

&lt;p&gt;I published that result. Then I refined the methodology, expanded the dataset, and re-ran the backtest against 219 actual fundraises.&lt;/p&gt;

&lt;p&gt;The improved signals now identify acceleration patterns that precede fundraises. But I lead with the failure because that's what makes the data trustworthy.&lt;/p&gt;

&lt;p&gt;Full paper with the 0-for-10 transparency ledger: &lt;a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6606558" rel="noopener noreferrer"&gt;SSRN abstract 6606558&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to use this
&lt;/h2&gt;

&lt;p&gt;If you're an angel or VC:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Track GitHub activity for startups in your thesis sectors&lt;/li&gt;
&lt;li&gt;Watch for the three-signal composite (velocity + contributors + infra)&lt;/li&gt;
&lt;li&gt;When all three fire, the startup is likely 30-47 days from announcing a round&lt;/li&gt;
&lt;li&gt;Reach out &lt;em&gt;before&lt;/em&gt; the round is announced&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The automated version of this is at &lt;a href="https://gitdealflow.com" rel="noopener noreferrer"&gt;GitDealFlow&lt;/a&gt;. You can also browse the open data at &lt;a href="https://signals.gitdealflow.com" rel="noopener noreferrer"&gt;signals.gitdealflow.com&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Open source everything
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Methodology: &lt;a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6606558" rel="noopener noreferrer"&gt;SSRN paper&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Due diligence checklist: &lt;a href="https://github.com/kindrat86/vc-due-diligence-checklist" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Live data: &lt;a href="https://signals.gitdealflow.com" rel="noopener noreferrer"&gt;signals.gitdealflow.com&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Signal interpretation guide: &lt;a href="https://github.com/kindrat86/github-startup-signals-guide" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;The author runs &lt;a href="https://gitdealflow.com" rel="noopener noreferrer"&gt;GitDealFlow&lt;/a&gt;. All methodology is open. CC BY 4.0.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>startups</category>
      <category>data</category>
      <category>vc</category>
      <category>github</category>
    </item>
    <item>
      <title>Why Your Agent Burned $2,800 at 3 AM: And How to See It</title>
      <dc:creator>Maryan K</dc:creator>
      <pubDate>Thu, 13 Aug 2026 13:36:02 +0000</pubDate>
      <link>https://dev.to/maryan_k_bef6cf83fa64e809/why-your-agent-burned-2800-at-3-am-and-how-to-see-it-21l1</link>
      <guid>https://dev.to/maryan_k_bef6cf83fa64e809/why-your-agent-burned-2800-at-3-am-and-how-to-see-it-21l1</guid>
      <description>&lt;p&gt;&lt;em&gt;Co-authored with Jacopo (&lt;a href="https://github.com/Jacopos311/Agent-Devtools" rel="noopener noreferrer"&gt;Agent-Devtools&lt;/a&gt;)&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;At 3:07 AM, my agent made 21 API calls to a premium LLM endpoint. Each cost $133. Total time: 60 seconds. Total bill: &lt;strong&gt;$2,800&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I was asleep. The budget alert email, the thing I'd set up so this couldn't happen, arrived 4 minutes later. Also while I was asleep. It was a very informative email about money that was already gone.&lt;/p&gt;

&lt;p&gt;That night taught me something that took two open-source projects to fully solve:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;You need a firewall to stop the bleed. You need a debugger to see why it happened.&lt;/strong&gt; One without the other is half a safety system.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Half one: stop the bleeding
&lt;/h2&gt;

&lt;p&gt;Before that night, I'd tried the standard defenses, in order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Budget alert emails.&lt;/strong&gt; They fire &lt;em&gt;after&lt;/em&gt; the spend happens. At 3 AM the only thing reading them is your inbox.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rate limits.&lt;/strong&gt; They cap &lt;em&gt;speed&lt;/em&gt;, not spend. An agent can trickle its way to the same bill over an hour.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Manual monitoring.&lt;/strong&gt; It fails at exactly the moment you need it most, 3 AM, weekends, holidays, the one night you forget.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;So I built &lt;strong&gt;&lt;a href="https://github.com/kindrat86/agentshield" rel="noopener noreferrer"&gt;AgentShield&lt;/a&gt;&lt;/strong&gt;: a pre-execution spend firewall for AI agents. Every transaction is evaluated against your rules &lt;em&gt;before&lt;/em&gt; the API call goes out. Pure Python stdlib, zero dependencies, &amp;lt;1ms per evaluation, 7 composable rules, transaction limits, daily totals, velocity, merchant allowlists, category blocks, session budgets, cascade cost estimates, plus a kill switch. &lt;code&gt;APPROVED&lt;/code&gt;, &lt;code&gt;BLOCKED&lt;/code&gt;, or &lt;code&gt;FLAGGED&lt;/code&gt;, in priority order, deterministically.&lt;/p&gt;

&lt;p&gt;The 3 AM incident wouldn't have survived contact with a single rule: &lt;code&gt;transaction_limit: max_amount $100&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That solves &lt;em&gt;"what should we block right now?"&lt;/em&gt; It does not solve &lt;em&gt;"why did it happen?"&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Half two: see the loop
&lt;/h2&gt;

&lt;p&gt;Blocking a runaway agent is like unplugging a burning appliance. Safe. Necessary. But you still don't know if it was a bug in your code, a retry loop in a library, a poisoned tool response, or a bad prompt. Without that answer, the agent just runs again tomorrow, and you play firewall whack-a-mole.&lt;/p&gt;

&lt;p&gt;That's the problem &lt;strong&gt;&lt;a href="https://github.com/Jacopos311/Agent-Devtools" rel="noopener noreferrer"&gt;Agent-Devtools&lt;/a&gt;&lt;/strong&gt;, built by Jacopo, solves: a local-first causal debugger for agent runs. It answers "why did my agent behave this way?" with visual replay, behavior diffing, and full visibility into prompts, context, memory, retrieval, and tool calls, the exact execution timeline that led to the bad decision.&lt;/p&gt;

&lt;p&gt;Jacopo found AgentShield through a comment on a LangChain issue about runaway agent loops, saw the same gap I did, and opened an issue: &lt;em&gt;"You stop the bleed, I show the loop. Want to integrate?"&lt;/em&gt; We both said yes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wiring the halves together
&lt;/h2&gt;

&lt;p&gt;The integration is deliberately boring, a shared event schema, two small modules, no hosted services:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;agentshield&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;SpendControlEngine&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;SpendEvaluationEmitter&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;agent_devtools&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;TraceStore&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;agent_devtools.adapters.agentshield&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;make_agentshield_callback&lt;/span&gt;

&lt;span class="n"&gt;store&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;TraceStore&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;cb&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;make_agentshield_callback&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;store&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;engine&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;SpendControlEngine&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;emitter&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;SpendEvaluationEmitter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;engine&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Every evaluation flows into your trace store automatically
&lt;/span&gt;&lt;span class="n"&gt;emitter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;emit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
 &lt;span class="n"&gt;transaction&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;amount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;500.00&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;merchant&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;openai-api&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
 &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;category&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;llm_inference&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
 &lt;span class="n"&gt;rules&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;r1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;transaction_limit&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;priority&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
 &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;params&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;max_amount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;250.00&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;action&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;BLOCK&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
 &lt;span class="n"&gt;trace_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;trace_42&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
 &lt;span class="n"&gt;on_event&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;cb&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's it. Under the hood:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AgentShield emits one &lt;code&gt;agentshield.spend.evaluation&lt;/code&gt; event per evaluation, NDJSON to stdout or a file (for tailing), or an in-process callback for embedding.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;trace_id&lt;/code&gt; joins directly to Agent-Devtools' native &lt;code&gt;run_id&lt;/code&gt;, so spend decisions render &lt;strong&gt;inside the execution timeline&lt;/strong&gt;, the blocked call, the rule that fired, the actual-vs-threshold numbers, and the agent's own reasoning, all in one view.&lt;/li&gt;
&lt;li&gt;Per-rule trace visibility: every rule reports &lt;code&gt;triggered&lt;/code&gt;, &lt;code&gt;passed&lt;/code&gt;, &lt;code&gt;skipped&lt;/code&gt;, or &lt;code&gt;not_reached&lt;/code&gt;. You see near-misses, not just the winning rule.&lt;/li&gt;
&lt;li&gt;Money stays exact: amounts cross the wire as Decimal-safe strings, never floats.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The result is the loop the title promises: &lt;strong&gt;the firewall stops the spend at the moment it happens, and the debugger shows you the replay of why the agent got there.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What the collaboration actually took
&lt;/h2&gt;

&lt;p&gt;The nice thing about this integration is how little ceremony it needed:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Jacopo opened the issue. We agreed on a schema &lt;strong&gt;before&lt;/strong&gt; writing code, the contract took one document.&lt;/li&gt;
&lt;li&gt;AgentShield shipped the emitter (&lt;code&gt;evaluate_with_trace()&lt;/code&gt; + &lt;code&gt;SpendEvaluationEmitter&lt;/code&gt;), additive, so existing &lt;code&gt;evaluate()&lt;/code&gt; behavior didn't change.&lt;/li&gt;
&lt;li&gt;Agent-Devtools shipped the adapter, a parser, a callback, and NDJSON ingestion.&lt;/li&gt;
&lt;li&gt;Then the important part: &lt;strong&gt;independent E2E verification on both sides.&lt;/strong&gt; We ran each other's code against real events, and it surfaced real edge cases: missing &lt;code&gt;trace_id&lt;/code&gt; crashing the host runtime (now falls back to &lt;code&gt;unattributed&lt;/code&gt;), callback events invisible in the run list (runs now auto-create), one malformed NDJSON line dropping the rest of a file (now skip-and-continue), and a Decimal precision bug on our side (now exact).&lt;/li&gt;
&lt;li&gt;Everything merged. Both suites green. 37/37 across the cross-stack harness.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Two repos, one issue thread, under a day from first comment to both sides merged.&lt;/p&gt;

&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;Every agent team eventually has their own 3 AM. The only question is whether it's a $2,800 lesson or a $19/month non-event.&lt;/p&gt;

&lt;p&gt;The pattern worth stealing isn't the code, it's the division of labor: &lt;strong&gt;pre-execution control answers "should this run?"; post-execution visibility answers "why did it run badly?"&lt;/strong&gt; Monitoring tools only tell you what already happened. Budget alerts only tell you what already happened. An agent safety story needs both halves, and now they're one integration.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;AgentShield&lt;/strong&gt; (pre-execution firewall): &lt;a href="https://github.com/kindrat86/agentshield" rel="noopener noreferrer"&gt;github.com/kindrat86/agentshield&lt;/a&gt;, stdlib-only, 7 rules, &amp;lt;1ms, kill switch.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent-Devtools&lt;/strong&gt; (post-execution debugger): &lt;a href="https://github.com/Jacopos311/Agent-Devtools" rel="noopener noreferrer"&gt;github.com/Jacopos311/Agent-Devtools&lt;/a&gt;, local-first visual replay for agent runs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both MIT. Both merged. Star both, wire them together, and sleep through the next 3 AM with confidence.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
      <category>debugging</category>
    </item>
    <item>
      <title>How I Built a Deal-Flow Signal From Public GitHub Data (219 Fundraises Backtested)</title>
      <dc:creator>Maryan K</dc:creator>
      <pubDate>Wed, 12 Aug 2026 21:27:04 +0000</pubDate>
      <link>https://dev.to/maryan_k_bef6cf83fa64e809/how-i-built-a-deal-flow-signal-from-public-github-data-219-fundraises-backtested-467o</link>
      <guid>https://dev.to/maryan_k_bef6cf83fa64e809/how-i-built-a-deal-flow-signal-from-public-github-data-219-fundraises-backtested-467o</guid>
      <description>&lt;p&gt;Eighteen months ago I noticed something odd while stalking a startup's GitHub org before an angel check: their commit activity had tripled about a month before they announced their round. Not after. Before.&lt;/p&gt;

&lt;p&gt;I'm a data person, so I did what data people do. I stopped looking at one company and started looking at all of them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hypothesis
&lt;/h2&gt;

&lt;p&gt;Startups behave differently on GitHub in the weeks before a fundraise. They clean up repos, ship faster, onboard new engineers hired ahead of the announcement, and spin up infrastructure for the growth they're about to buy. All of that is visible in public API data if you know what to measure.&lt;/p&gt;

&lt;p&gt;So I built a tracker. It now watches &lt;strong&gt;4,200+ startup GitHub orgs&lt;/strong&gt; weekly across 20 sectors, computing three signals:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Commit velocity&lt;/strong&gt; — commits across tracked repos in a trailing 14-day window vs. the prior window&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Contributor growth&lt;/strong&gt; — distinct active contributors over 30 days vs. baseline (new hires show up here first)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;New repo creation&lt;/strong&gt; — fresh public repos in 30 days (infrastructure buildout)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;These get combined into a composite score and classified: breakout, acceleration, steady, cooling.&lt;/p&gt;

&lt;h2&gt;
  
  
  Did it actually work?
&lt;/h2&gt;

&lt;p&gt;I backtested the pattern against &lt;strong&gt;219 documented fundraises&lt;/strong&gt;. Companies showing the acceleration pattern raised at roughly &lt;strong&gt;3.4× the base rate&lt;/strong&gt;, with the signal appearing &lt;strong&gt;21–47 days before the public announcement&lt;/strong&gt; in the cases where it fired.&lt;/p&gt;

&lt;p&gt;The full methodology is published as an SSRN preprint (DOI: &lt;a href="https://doi.org/10.2139/ssrn.6606558" rel="noopener noreferrer"&gt;10.2139/ssrn.6606558&lt;/a&gt;) with the backtest dataset on Zenodo (&lt;a href="https://doi.org/10.5281/zenodo.19650920" rel="noopener noreferrer"&gt;10.5281/zenodo.19650920&lt;/a&gt;). I wanted this to be checkable, not another black-box "AI signal."&lt;/p&gt;

&lt;h2&gt;
  
  
  What it caught this quarter
&lt;/h2&gt;

&lt;p&gt;A few live examples from the current Q3 2026 period (all public data, verifiable on GitHub right now):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Fleetbase&lt;/strong&gt; (logistics OS, pre-seed): commit velocity up ~16× over its prior window, 7 active contributors (+73% in 30 days). Classic engineering-hiring-burst shape.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tuva Health&lt;/strong&gt; (healthcare analytics): 48 active contributors, 3 new repos in 30 days. Infrastructure buildout pattern.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Swell&lt;/strong&gt; (e-commerce infra): contributor count doubled in 30 days, 2 new repos. Hiring burst.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reliza&lt;/strong&gt; (DevOps/supply chain, pre-seed): contributors up 5× from a small base, 2 new repos.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Will all of these raise? No. That's the honest part.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest limitations
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;~23% false-positive rate.&lt;/strong&gt; Acceleration sometimes means a big enterprise deal, an open-source push, or a hackathon, not a round.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stealth-first companies are invisible.&lt;/strong&gt; No public repos, no signal.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI-heavy startups are noisy.&lt;/strong&gt; Model releases create commit spikes that look like fundraise prep.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Small-base distortion.&lt;/strong&gt; A 3-contributor org going to 6 is +100% but means little. The composite score penalizes low-base orgs, but it's not perfect.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Anyone who sells you a startup signal without a false-positive number is selling you a story.&lt;/p&gt;

&lt;h2&gt;
  
  
  The stack, briefly
&lt;/h2&gt;

&lt;p&gt;GitHub REST API v3 (&lt;code&gt;search/repositories&lt;/code&gt;, &lt;code&gt;stats/commit_activity&lt;/code&gt;, &lt;code&gt;contributors&lt;/code&gt;), a weekly batch pipeline that processes all 4,200 orgs for under €50/month of compute, and a scoring layer. No scraping, no private data, nothing you couldn't rebuild yourself — the signal computation logic is being open-sourced (MIT).&lt;/p&gt;

&lt;p&gt;There's also a free MCP server, so if you use Claude Desktop or Cursor you can query the live signal data directly from your editor: six tools, no paywall on any of them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this goes
&lt;/h2&gt;

&lt;p&gt;I publish the weekly top movers in a free Sunday email at &lt;a href="https://gitdealflow.com" rel="noopener noreferrer"&gt;gitdealflow.com&lt;/a&gt;. The heavier stuff (full rankings, sector sweeps, dashboards) is paid, which is what funds the compute.&lt;/p&gt;

&lt;p&gt;But the core idea is free and reproducible: &lt;strong&gt;public engineering activity is a leading indicator of private funding events.&lt;/strong&gt; If you're an angel, a scout, or just someone who likes watching startups through their commits, the data is sitting there in the open.&lt;/p&gt;

&lt;p&gt;Happy to answer questions about the methodology, the backtest design, or the pipeline in the comments.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;— The Data Nerd&lt;/em&gt;&lt;/p&gt;

</description>
      <category>github</category>
      <category>datascience</category>
      <category>startup</category>
      <category>opensource</category>
    </item>
    <item>
      <title>I Built a Firewall for AI Agent Spending — Here is What I Learned From 56 Attack Scenarios</title>
      <dc:creator>Maryan K</dc:creator>
      <pubDate>Tue, 11 Aug 2026 21:46:34 +0000</pubDate>
      <link>https://dev.to/maryan_k_bef6cf83fa64e809/i-built-a-firewall-for-ai-agent-spending-here-is-what-i-learned-from-56-attack-scenarios-4jh5</link>
      <guid>https://dev.to/maryan_k_bef6cf83fa64e809/i-built-a-firewall-for-ai-agent-spending-here-is-what-i-learned-from-56-attack-scenarios-4jh5</guid>
      <description>&lt;p&gt;It started with a $2,800 Stripe charge.&lt;/p&gt;

&lt;p&gt;An AI agent we'd deployed to handle customer onboarding hit a flaky webhook endpoint. The standard retry logic kicked in. But instead of retrying once or twice, the agent discovered it could call a &lt;em&gt;different&lt;/em&gt; endpoint to achieve the same result. Then another. Then it started chaining calls — each one triggering a separate Stripe charge because the underlying action wasn't idempotent.&lt;/p&gt;

&lt;p&gt;By the time our billing alert fired four hours later, the damage was done. No single call looked abnormal. The spend curve was smooth and gradual. And our monitoring dashboard showed every single transaction in perfect detail — after the money was already gone.&lt;/p&gt;

&lt;p&gt;That's the day I realized: &lt;strong&gt;observability is not enforcement&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem: Agents Are Autonomous Spenders
&lt;/h2&gt;

&lt;p&gt;AI agents are fundamentally different from traditional applications. A web app makes API calls based on user actions — predictable, bounded, and rate-limited by human interaction speed. An AI agent makes calls autonomously, at machine speed, and can discover new endpoints and tools at runtime.&lt;/p&gt;

&lt;p&gt;This creates a new class of financial risk:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Runaway loops&lt;/strong&gt;: An agent enters a reasoning loop, burning tokens indefinitely&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cascade discovery&lt;/strong&gt;: An agent finds a new API and starts making calls you never anticipated&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool re-execution&lt;/strong&gt;: Retry logic triggers duplicate payments when tools aren't idempotent&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Delegation spirals&lt;/strong&gt;: An agent delegates sub-tasks to itself, each spawning more tool calls&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Infinite reasoning&lt;/strong&gt;: Thinking-mode loops consume token budgets without producing output&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I looked at existing solutions. LangChain had cost tracking. CrewAI had guardrails. AutoGen had governance discussions. But they all shared the same fundamental flaw: &lt;strong&gt;they observed spending instead of preventing it&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A dashboard that tells you "you spent $2,800 today" is useless when the money is already gone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building AgentShield: A Spend Firewall
&lt;/h2&gt;

&lt;p&gt;I needed something that sits between the agent and the tools it calls — a firewall that inspects every transaction &lt;em&gt;before&lt;/em&gt; it executes and blocks anything that violates a budget rule.&lt;/p&gt;

&lt;p&gt;The design was inspired by traditional network firewalls: deterministic, fast, and non-bypassable.&lt;/p&gt;

&lt;h3&gt;
  
  
  The 7 Rule Types
&lt;/h3&gt;

&lt;p&gt;After analyzing our incident and dozens of similar failure modes reported across GitHub issues in CrewAI, AutoGen, LangChain, and Claude Code, I identified seven distinct spending attack patterns. Each maps to a specific rule type:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. transaction_limit&lt;/strong&gt; — Blocks any single transaction above a threshold. Simple but effective. If no individual API call should cost more than $5, this catches the $50 call immediately.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. daily_total&lt;/strong&gt; — Rolling 24-hour cumulative spend cap. The failsafe against any single-day disaster. Our $2,800 incident would have been caught at $50.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. velocity&lt;/strong&gt; — Rate limiting per merchant/tool. Catches burst attacks where an agent fires 100 calls in 30 seconds. This is the rule that catches retry storms.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. merchant_allowlist&lt;/strong&gt; — Only approved endpoints can be called. This is critical for agents that can discover new tools at runtime. If your agent shouldn't be calling unknown APIs, don't let it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. category_block&lt;/strong&gt; — Block entire categories of spending. Useful for compliance ("no cryptocurrency purchases") or cost control ("no premium model calls during batch jobs").&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6. session_budget&lt;/strong&gt; — Per-agent-run budget cap. When an agent session starts, assign it a budget. When it's exhausted, the agent gets a budget-exceeded error and must terminate gracefully.&lt;/p&gt;

&lt;p&gt;This is the rule that catches infinite reasoning loops and delegation spirals. The agent might be stuck reasoning forever, but it can't spend more than its session allows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;7. cascade_cost&lt;/strong&gt; — The most nuanced rule. It detects when total spending is growing super-linearly — the signature of a runaway pattern. If call #1 costs $0.01, call #2 costs $0.01, but by call #50 you're spending $1.00/call, something is wrong.&lt;/p&gt;

&lt;p&gt;This is the rule that would have caught our original incident. The agent's per-call cost didn't change, but the &lt;em&gt;frequency&lt;/em&gt; was accelerating exponentially as it discovered new endpoints.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 56-Scenario Evaluation Gym
&lt;/h2&gt;

&lt;p&gt;Rules are only as good as their test coverage. I built a 56-scenario evaluation harness that simulates every spending attack pattern I could find — from GitHub issues, production incidents, and adversarial testing.&lt;/p&gt;

&lt;p&gt;The scenarios fall into six categories:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Scenarios&lt;/th&gt;
&lt;th&gt;Example&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Single-call attacks&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;Agent attempts $500 single transaction&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Burst/velocity attacks&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;Agent fires 200 calls in 10 seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cumulative spending&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;Agent spends $1/call for 500 calls&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Loop patterns&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;Agent enters infinite reasoning loop&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Discovery attacks&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;Agent finds undocumented endpoint&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-agent cascades&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Agent delegates to sub-agents, each spending&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Each scenario runs the agent against AgentShield with a configured rule set. The test passes if the firewall blocks the transaction &lt;em&gt;before&lt;/em&gt; it reaches the provider, and fails if the spend leaks through.&lt;/p&gt;

&lt;p&gt;Current pass rate: &lt;strong&gt;54/56&lt;/strong&gt; (96.4%). The two remaining edge cases involve multi-agent systems where spending is distributed across agents — we're working on cross-agent budget aggregation for v2.&lt;/p&gt;

&lt;h2&gt;
  
  
  Enforcement vs. Observability
&lt;/h2&gt;

&lt;p&gt;This is the distinction that matters most, and it's the one most teams get wrong.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Observability&lt;/strong&gt; answers: &lt;em&gt;"What happened?"&lt;/em&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Dashboards, alerts, logs&lt;/li&gt;
&lt;li&gt;Cost tracking per provider&lt;/li&gt;
&lt;li&gt;Token usage analytics&lt;/li&gt;
&lt;li&gt;Post-incident analysis&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Enforcement&lt;/strong&gt; answers: &lt;em&gt;"What is allowed?"&lt;/em&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Rule evaluation before execution&lt;/li&gt;
&lt;li&gt;Hard blocks on policy violations&lt;/li&gt;
&lt;li&gt;Budget caps that cannot be exceeded&lt;/li&gt;
&lt;li&gt;Pre-call authorization&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You need both. But enforcement is the one that prevents financial loss.&lt;/p&gt;

&lt;p&gt;Think of it this way: a security camera (observability) helps you catch a thief after they've stolen from you. A locked door (enforcement) stops them from entering in the first place.&lt;/p&gt;

&lt;p&gt;AgentShield is the locked door. It sits in the tool dispatch path and evaluates rules in under 1 millisecond — fast enough that the agent never notices the checkpoint, strict enough that no transaction can bypass it.&lt;/p&gt;

&lt;h2&gt;
  
  
  How It Works in Practice
&lt;/h2&gt;

&lt;p&gt;AgentShield runs as middleware. You configure your rules, and every tool call from your agent passes through the firewall:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;agentshield&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;SpendFirewall&lt;/span&gt;

&lt;span class="n"&gt;firewall&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;SpendFirewall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;rules&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;daily_total&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;limit&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;50.00&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;velocity&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;max_calls&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;window_seconds&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;session_budget&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;budget&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;10.00&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nd"&gt;@firewall.guard&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tool_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;stripe_charge&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;process_payment&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;stripe&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;charges&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;customer&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If a rule is violated, the agent receives a &lt;code&gt;BudgetExceededError&lt;/code&gt; instead of completing the call. The agent can handle this gracefully — log it, notify a human, or terminate the session.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;AgentShield is open source and free to use. The hosted version at &lt;a href="https://agentshield.fly.dev" rel="noopener noreferrer"&gt;agentshield.fly.dev&lt;/a&gt; includes a dashboard for rule management and spend visualization — the observability layer on top of the enforcement engine.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GitHub&lt;/strong&gt;: &lt;a href="https://github.com/kindrat86/agentshield" rel="noopener noreferrer"&gt;kindrat86/agentshield&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Live demo&lt;/strong&gt;: &lt;a href="https://agentshield.fly.dev" rel="noopener noreferrer"&gt;agentshield.fly.dev&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Free tier&lt;/strong&gt;: Covers most side projects and small teams&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you're running AI agents in production, the question isn't &lt;em&gt;whether&lt;/em&gt; you'll have a spending incident — it's &lt;em&gt;when&lt;/em&gt;. A spend firewall doesn't prevent bugs, but it ensures that when something goes wrong, the financial damage is bounded.&lt;/p&gt;

&lt;p&gt;Because the alternative is finding out from a Stripe bill.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;AgentShield is MIT-licensed and runs on Python 3.10+. If you've experienced an agent spending incident, I'd love to hear about it — the 56-scenario eval gym is always accepting new attack patterns.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;🔗 &lt;a href="https://github.com/kindrat86/agentshield" rel="noopener noreferrer"&gt;github.com/kindrat86/agentshield&lt;/a&gt; | &lt;a href="https://agentshield.fly.dev" rel="noopener noreferrer"&gt;agentshield.fly.dev&lt;/a&gt;&lt;/p&gt;

</description>
      <category>aiagents</category>
      <category>costcontrol</category>
      <category>python</category>
      <category>devtools</category>
    </item>
    <item>
      <title>How I Built a 56-Scenario Eval Gym for AI Spend Control</title>
      <dc:creator>Maryan K</dc:creator>
      <pubDate>Tue, 11 Aug 2026 20:50:21 +0000</pubDate>
      <link>https://dev.to/maryan_k_bef6cf83fa64e809/how-i-built-a-56-scenario-eval-gym-for-ai-spend-control-2lpn</link>
      <guid>https://dev.to/maryan_k_bef6cf83fa64e809/how-i-built-a-56-scenario-eval-gym-for-ai-spend-control-2lpn</guid>
      <description>&lt;h1&gt;
  
  
  How I Built a 56-Scenario Eval Gym for AI Spend Control
&lt;/h1&gt;

&lt;p&gt;After my AI agent spent $2,800 in 60 seconds at 3 AM, I knew I needed pre-flight spend control. But how do you test something that's supposed to catch runaway AI spending &lt;em&gt;before&lt;/em&gt; it happens?&lt;/p&gt;

&lt;p&gt;The answer: an eval gym with 56 labeled scenarios across 9 categories.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why 56 Scenarios?
&lt;/h2&gt;

&lt;p&gt;When you're building spend-control logic, you quickly realize that simple rules have subtle edge cases:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A $50 limit with a $50.00 transaction — should it block or approve?&lt;/li&gt;
&lt;li&gt;What about $49.99 vs $50.01?&lt;/li&gt;
&lt;li&gt;Daily total resets at midnight — but what timezone?&lt;/li&gt;
&lt;li&gt;Velocity limits: 5 calls in 5 minutes — is the 6th call blocked or just the 5th?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These edge cases compound across 7 rule types. Here's how the scenarios break down:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Scenarios&lt;/th&gt;
&lt;th&gt;What it tests&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Clean approval&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;Legitimate transactions pass through&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Transaction limit&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;Per-transaction caps ($50, $100, etc.)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Daily total&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;Rolling 24h spending caps&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Velocity&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Rate-of-spending detection&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Merchant allowlist&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;Vendor whitelisting/blocking&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Category blocks&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;Category-level restrictions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Session budget&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Per-session spending ceiling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cascade cost&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Multi-step operation cost analysis&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Edge cases&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Boundary values, nulls, negatives&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The gym runs against a clean engine state every time — no shared state, no ordering dependencies. Every scenario is independent.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Hardest Edge Cases
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Amount exactly at limit&lt;/strong&gt;: If the limit is $50.00 and the transaction is $50.00, should it be approved? We chose &lt;em&gt;approve&lt;/em&gt; — "exceeds" means strictly greater than. This is documented and tested.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Decimal precision&lt;/strong&gt;: All monetary arithmetic uses Python's &lt;code&gt;Decimal&lt;/code&gt; type. Never &lt;code&gt;float&lt;/code&gt;. 0.1 + 0.2 is fine for graphics; it's catastrophic for money.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Priority ordering&lt;/strong&gt;: Rules evaluate in priority order. If rule #1 BLOCKs but rule #2 would APPROVE, the first match wins. The test suite verifies that a BLOCK at priority 1 always beats an APPROVE at priority 5.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Gym is a Universal Benchmark
&lt;/h2&gt;

&lt;p&gt;The eval gym is useful even if you don't use AgentShield. If you're building any kind of spend-control or rate-limiting logic, these 56 scenarios provide a test suite you can run against your own implementation.&lt;/p&gt;

&lt;p&gt;Check the live results: &lt;a href="https://agentshield.fly.dev/eval" rel="noopener noreferrer"&gt;56/56 across 9 categories&lt;/a&gt;. The spec is at &lt;a href="https://agentshield.fly.dev/eval-gym-spec" rel="noopener noreferrer"&gt;agentshield.fly.dev/eval-gym-spec&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;I'm working on adding more scenario categories:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Multi-agent shared budget scenarios&lt;/li&gt;
&lt;li&gt;Timezone-aware daily reset edge cases&lt;/li&gt;
&lt;li&gt;Cascade cost with variable failure rates&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The gym is MIT licensed. Use it. Fork it. Add your own scenarios. If you find a bug in the test suite, open an issue — I'll fix it.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Previously: &lt;a href="https://dev.to/maryan_k_bef6cf83fa64e809/i-built-a-firewall-for-ai-agent-spending-heres-the-architecture-2560"&gt;I built a firewall for AI agent spending — here's the architecture&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>ai</category>
      <category>testing</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
