<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ryosuke Tsuji</title>
    <description>The latest articles on DEV Community by Ryosuke Tsuji (@ryantsuji).</description>
    <link>https://dev.to/ryantsuji</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3843591%2F8b126f91-f561-4e6b-8492-814b18d680ec.jpg</url>
      <title>DEV Community: Ryosuke Tsuji</title>
      <link>https://dev.to/ryantsuji</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ryantsuji"/>
    <language>en</language>
    <item>
      <title>When Shipping Gets Cheap, Deciding Gets Expensive: Guardrails for AI-Era Cloud Costs</title>
      <dc:creator>Ryosuke Tsuji</dc:creator>
      <pubDate>Wed, 05 Aug 2026 23:40:03 +0000</pubDate>
      <link>https://dev.to/ryantsuji/when-shipping-gets-cheap-deciding-gets-expensive-guardrails-for-ai-era-cloud-costs-5g95</link>
      <guid>https://dev.to/ryantsuji/when-shipping-gets-cheap-deciding-gets-expensive-guardrails-for-ai-era-cloud-costs-5g95</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;AI assistance disclosure: This article was drafted with the help of Claude. All technical content, design decisions, code references, and screenshots reflect production systems I designed and operate at airCloset; the prose was revised by me prior to publication.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hi, I'm &lt;a href="https://x.com/ryantsuji" rel="noopener noreferrer"&gt;Ryan&lt;/a&gt;, CTO at airCloset.&lt;/p&gt;

&lt;p&gt;We spent a week going through the cloud bill for our internal AI platform. The results, up front:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Cloud Run instance-based billing down 84% (44% on the single service we tested first)&lt;/li&gt;
&lt;li&gt;Container image storage down 65% after revisiting how many versions we keep&lt;/li&gt;
&lt;li&gt;The thing that mattered most wasn't any single cut. It was the gate at merge time and the daily visibility&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But the savings aren't really what I want to write about. &lt;strong&gt;When implementation gets cheap, the cost of deciding whether to build something stays exactly where it was, and suddenly it's the expensive part.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Slow implementation used to be a decision gate
&lt;/h2&gt;

&lt;p&gt;This isn't a post about AI, exactly. But AI is what makes this particular thing grow.&lt;/p&gt;

&lt;p&gt;Building anything used to cost real time. You'd estimate it, review the design, carve out the engineering time. That process is tedious, and it also had a side effect nobody designed for: it forced you to ask whether the thing was worth building at all. &lt;strong&gt;The slowness itself was the gate.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Make implementation fast and that gate comes off. You can ship an idea the same day you have it. That's genuinely good, and it makes "just build it and find out" the right call far more often than it used to be.&lt;/p&gt;

&lt;p&gt;What also comes off, though, is any reason to hold back. Each individual thing is small enough that nobody objects. And what's left behind isn't engineering hours, it's the running cost of whatever you shipped. That part doesn't show up on the day you build it. It accumulates quietly afterward.&lt;/p&gt;

&lt;h2&gt;
  
  
  The starting point: we couldn't attribute spend to an app
&lt;/h2&gt;

&lt;p&gt;What kicked this off was a handful of occasions where usage-based spend grew more than we expected after a particular change. BigQuery scan volume, and Vertex AI / Gemini calls.&lt;/p&gt;

&lt;p&gt;The frustrating part was that &lt;strong&gt;we couldn't attribute the spend back to a specific app after the fact&lt;/strong&gt;. BigQuery does let you label query jobs, and job history is there. But several of our apps run under the same service account, so knowing which principal ran a query didn't tell us which app was responsible. We could push labels through every call path, but until that's fully rolled out you're in the same position.&lt;/p&gt;

&lt;p&gt;If you can't attribute after the fact, you stop it before it grows. That's where the idea of putting a gate at the entrance came from.&lt;/p&gt;

&lt;h2&gt;
  
  
  Visibility comes before any of the cuts
&lt;/h2&gt;

&lt;p&gt;Here's the conclusion first: &lt;strong&gt;the highest-leverage part of this whole week was the gate at merge time and the daily reporting, not any individual optimization.&lt;/strong&gt; Cuts are one-time. Left alone, the same things pile up again.&lt;/p&gt;

&lt;h3&gt;
  
  
  The gate: changes above a threshold go to a human
&lt;/h3&gt;

&lt;p&gt;We added cost as a review dimension for our AI reviewer. If a PR pushes usage-based spend past a threshold, it gets handed to a human. Three thresholds:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Threshold&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Ongoing daily cost increase&lt;/td&gt;
&lt;td&gt;¥2,000/day (~$13)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One-off execution cost&lt;/td&gt;
&lt;td&gt;¥10,000 (~$65)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Monthly storage growth&lt;/td&gt;
&lt;td&gt;¥1,000/month (~$6)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The design choice I care about here: &lt;strong&gt;the AI doesn't do the arithmetic&lt;/strong&gt;. It reads the change and measures quantities like scan volume, call counts, and token counts. Converting to currency and comparing against the threshold happens in a deterministic script. Hand the money math to a model and you get occasional arithmetic slips, the wrong unit price, and answers that drift between runs. A gate that gives different answers on different days isn't a gate.&lt;/p&gt;

&lt;p&gt;Unit prices come from our actual billing export (spend ÷ usage), not from the published rate card. Hardcoding list prices means carrying three separate sources of drift: USD conversion, exchange rate, and committed-use discounts.&lt;/p&gt;

&lt;p&gt;Storage gets its own threshold because &lt;strong&gt;cumulative cost behaves differently from execution cost&lt;/strong&gt;. Stop a job and its daily cost goes to zero. Data you've written keeps billing until someone deletes it. So for storage we ask a different question: is there an expiration policy, and what does this add per month going forward? That's a &lt;strong&gt;forward projection&lt;/strong&gt;, computed from call volume times payload size, not a measurement of what already happened.&lt;/p&gt;

&lt;h3&gt;
  
  
  Daily reporting: post to Slack twice a day
&lt;/h3&gt;

&lt;p&gt;We post per-app cost to Slack twice a day. A few decisions in there worth mentioning.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Post even on days when nothing crossed a threshold.&lt;/strong&gt; "Nothing to see today" is information, and more importantly, the reader needs to be able to tell the difference between a quiet day and a broken reporter. A monitor that only speaks up when something's wrong can't tell you it's dead.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Compare against the same weekday last week, not yesterday.&lt;/strong&gt; Batch volume varies by day of week, so a day-over-day comparison flags every Monday as a spike, and people stop reading it within a week.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Treat "no data last week" as an increase.&lt;/strong&gt; This is the dangerous case. A batch that just started running is exactly the thing that surprises you. Percent-change metrics tend to drop these on the floor because the math is undefined when the baseline is zero, so we handle that case explicitly.&lt;/p&gt;

&lt;h3&gt;
  
  
  The backstop: put a ceiling on usage-based services
&lt;/h3&gt;

&lt;p&gt;Gates and reporting still leak. There are paths that don't go through a PR at all, like manual runs, external triggers, and usage patterns nobody anticipated.&lt;/p&gt;

&lt;p&gt;So for anything that bills by usage, &lt;strong&gt;set a ceiling that matches how you actually operate&lt;/strong&gt;. BigQuery, Vertex AI, and Gemini all let you cap this with quotas.&lt;/p&gt;

&lt;p&gt;The important part is to &lt;strong&gt;derive the ceiling from measured usage, not from what the platform allows&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Take BigQuery. On-demand pricing has a default cap of &lt;strong&gt;200 TiB per day, per project&lt;/strong&gt; (that's been the default since September 2025; before that it was unlimited). Per-TiB pricing varies by region, but at a few dollars per TiB, &lt;strong&gt;that ceiling works out to somewhere north of a thousand dollars a day&lt;/strong&gt;. Run that for a full month and you're at tens of thousands of dollars.&lt;/p&gt;

&lt;p&gt;Which means the default setting is: spend that much, every month, without anyone being told. That isn't a ceiling.&lt;/p&gt;

&lt;p&gt;We cap BigQuery scan volume at roughly 1.5x our normal usage. "Ten times normal, to be safe" has the same problem as the default. A ceiling only means something if it sits where a runaway actually stops.&lt;/p&gt;

&lt;p&gt;One thing to be clear about: &lt;strong&gt;we can do this because this is internal infrastructure.&lt;/strong&gt; Capping at 1.5x means accepting that things stop when they hit the cap. For an internal platform, you rerun it tomorrow. Do the same thing on a customer-facing path and you've built yourself an outage. &lt;strong&gt;Only put this kind of ceiling on things that are allowed to stop.&lt;/strong&gt; For anything touching end users, set the ceiling much higher and handle it with alerting and autoscaling instead.&lt;/p&gt;

&lt;p&gt;Given that constraint, the side effects are fine. A legitimate job occasionally hits the cap, and when it does, that's a prompt to go find out why something needed 1.5x normal volume. Which is its own kind of detection. &lt;strong&gt;Alerts tell you afterward. A quota stops it while it's happening.&lt;/strong&gt; With usage-based billing that difference matters.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd8zztdjjbrjz79uc6lc6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd8zztdjjbrjz79uc6lc6.png" alt="Three layers that stop cost from piling up" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you try to run any of this as a habit rather than a system, it evaporates the first busy week. &lt;strong&gt;It only works once it's structural.&lt;/strong&gt; Gate at the entrance, daily reporting for the continuous view, quota as the hard stop. Three layers, and only together do they make the invisible visible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Case 1: when the default is the expensive one, speed builds debt
&lt;/h2&gt;

&lt;p&gt;Now the things we actually found. Cloud Run first.&lt;/p&gt;

&lt;p&gt;Cloud Run has two billing models: instance-based, where CPU is always allocated, and request-based, where it's allocated only while handling a request. &lt;strong&gt;97% of our Cloud Run spend was instance-based.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The cause was structural. Looking at where we define Cloud Run services in our repo, &lt;strong&gt;most of them didn't specify &lt;code&gt;cpuIdle&lt;/code&gt; at all&lt;/strong&gt;. There's no shared factory, so every new service silently landed on the expensive side unless someone thought to set it.&lt;/p&gt;

&lt;p&gt;Worth knowing: &lt;strong&gt;deploy from the console or gcloud and the default is request-based&lt;/strong&gt;, the cheap one. Define the same service declaratively in IaC with explicit resource limits and &lt;code&gt;cpuIdle&lt;/code&gt; ends up false, which puts you on instance-based. So &lt;strong&gt;click it together and you get the cheap default; write it as code and you get the expensive one&lt;/strong&gt;. The more disciplined your infrastructure practice, the easier this is to miss.&lt;/p&gt;

&lt;p&gt;Which is the whole point from earlier. &lt;strong&gt;When the default sits on the expensive side, going faster builds debt automatically.&lt;/strong&gt; The per-service difference is small enough that nobody notices.&lt;/p&gt;

&lt;h3&gt;
  
  
  The name reads backwards
&lt;/h3&gt;

&lt;p&gt;There's a trap in the naming. The setting is called &lt;code&gt;cpuIdle&lt;/code&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;th&gt;Behavior&lt;/th&gt;
&lt;th&gt;Billing&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;true&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;CPU allocated only during requests&lt;/td&gt;
&lt;td&gt;Request-based (cheaper)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;false&lt;/code&gt; (IaC default)&lt;/td&gt;
&lt;td&gt;CPU always allocated&lt;/td&gt;
&lt;td&gt;Instance-based (pricier)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Read &lt;code&gt;true&lt;/code&gt; as "keeps running while idle" and you have it exactly backwards. And the failure mode is nasty: &lt;strong&gt;no exception, no error, just the work that happens after the response quietly not happening&lt;/strong&gt;. Flip a service that does fire-and-forget background work to &lt;code&gt;true&lt;/code&gt; and the work disappears without a trace in the logs.&lt;/p&gt;

&lt;p&gt;So we reduced the decision to one question: does this service keep working after it returns a response?&lt;/p&gt;

&lt;h3&gt;
  
  
  The exclusion list is narrower than it looks
&lt;/h3&gt;

&lt;p&gt;When we picked which services to switch, we started with a generous exclusion list of anything that "probably needs CPU all the time." Most of it turned out not to. Two things about Cloud Run's behavior are not obvious.&lt;/p&gt;

&lt;p&gt;First, &lt;strong&gt;CPU is allocated during container startup regardless of this setting&lt;/strong&gt;. A service that loads a large dataset from BigQuery at boot, synchronously, before the web server starts, looks like an obvious "needs constant CPU" case. But that work happens during startup, so switching it is fine.&lt;/p&gt;

&lt;p&gt;Second, &lt;strong&gt;long timeouts are not a reason to exclude something&lt;/strong&gt;. A batch that takes 3,600 seconds of synchronous work is handling a request that entire time, so it has CPU. "It's a heavy job, so it needs constant CPU" doesn't hold.&lt;/p&gt;

&lt;p&gt;What actually needs excluding is &lt;strong&gt;services that keep working after the response goes out&lt;/strong&gt;, the fire-and-forget ones. That's the whole list. Getting there took several rewrites of the exclusion set.&lt;/p&gt;

&lt;p&gt;Settings like this always over-exclude when you decide by intuition. &lt;strong&gt;Exclude based on a condition you can state, not on a service feeling like it needs it.&lt;/strong&gt; Whether you can compress the rule into one line is a decent proxy for whether you've understood it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Results, and stopping the regression
&lt;/h3&gt;

&lt;p&gt;The first service we switched came down &lt;strong&gt;44% on its own&lt;/strong&gt;. Rolling it out brought &lt;strong&gt;instance-based billing down 84%&lt;/strong&gt; overall. Everything is switched except the handful of services doing fire-and-forget work.&lt;/p&gt;

&lt;p&gt;Then we added a CI guard that fails on service definitions with no &lt;code&gt;cpuIdle&lt;/code&gt;. The part I like: &lt;strong&gt;specifying &lt;code&gt;cpuIdle: false&lt;/code&gt; explicitly is also a violation.&lt;/strong&gt; If a service genuinely needs constant CPU, it goes in an allowlist with a written reason. The point is to make someone type out why.&lt;/p&gt;

&lt;p&gt;And with guards like this, we now &lt;strong&gt;write a violation on purpose and confirm the build actually fails&lt;/strong&gt;. A guard that's been written isn't necessarily a guard that runs, and there's nothing to tell you when it isn't. All you're left with is the belief that you're covered, which is worse than knowing you aren't.&lt;/p&gt;

&lt;h2&gt;
  
  
  Case 2: the data can't tell you how recovery actually works
&lt;/h2&gt;

&lt;p&gt;Container image storage was the other large line item. The fix was trivial: &lt;strong&gt;we went from keeping 10 image versions to keeping 2&lt;/strong&gt;. Storage dropped 65%.&lt;/p&gt;

&lt;p&gt;The interesting part is how you get to that number. From the data alone, all you can say is "maybe we need 10." How far back you might roll to isn't something the data records.&lt;/p&gt;

&lt;p&gt;Here's the actual reasoning. &lt;strong&gt;If the image is gone, you restore the code from git and deploy it again.&lt;/strong&gt; So keeping images isn't an archive that lets you return to any point in history. It's a way to get back to the previous version quickly in an emergency. Framed that way, two is enough.&lt;/p&gt;

&lt;p&gt;What's doing the work there isn't data, it's knowing &lt;strong&gt;how recovery actually happens in your organization&lt;/strong&gt;. Whether restoring from git is a safe assumption. How long a deploy takes. Whether anyone has ever needed to go further back than one version in an emergency. &lt;strong&gt;None of that is in your logs or your metrics.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Ask an AI to look into it and you'll usually get "keep more versions, to be safe," because that's what the data supports on its own. What a human brings is knowing what the retention is &lt;em&gt;for&lt;/em&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Case 3: we were running full CI for every review round
&lt;/h2&gt;

&lt;p&gt;The CI change that mattered most was about &lt;strong&gt;when&lt;/strong&gt; things run, not what.&lt;/p&gt;

&lt;p&gt;A reviewer flags something, you push a fix, they look again. &lt;strong&gt;Every one of those pushes ran the full set of jobs&lt;/strong&gt;: build, lint, test, and knip.&lt;/p&gt;

&lt;p&gt;The structure is the same with human reviewers. It's just that AI review makes the rounds more frequent and much faster, so waste that was always there becomes obvious. Measured, we average 3.4 rounds per PR (counting a round as one flag-and-fix cycle), with 6.8 minutes of heavy jobs per round. Which means running the same test suite three or four times before review even converged. &lt;strong&gt;We were paying for a full test run against code that still had open comments on it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So we moved the heavy jobs behind review approval. Every push runs a lightweight check set (1.2 minutes), and when the reviewer approves, a bot triggers the heavy CI.&lt;/p&gt;

&lt;p&gt;Light checks per round, heavy jobs once at the end. &lt;code&gt;3.4 × 6.8 min&lt;/code&gt; becomes &lt;code&gt;3.4 × 1.2 min + 6.8 min&lt;/code&gt;, roughly &lt;strong&gt;53% less&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Skipping and not-running are not the same thing
&lt;/h3&gt;

&lt;p&gt;There's an implementation trap here. Our first attempt used &lt;code&gt;if:&lt;/code&gt; conditions to skip jobs inside the same workflow, and that was &lt;strong&gt;the wrong design&lt;/strong&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;State&lt;/th&gt;
&lt;th&gt;Branch protection treats it as&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Job skipped via &lt;code&gt;if:&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Passing&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Workflow never runs (no check exists)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Not reported, blocks the merge&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Skip a required check and you've effectively disabled branch protection. The green check appears without the tests having run, and the PR is mergeable. That fails open.&lt;/p&gt;

&lt;p&gt;Never start the workflow and no check is created, so the unreported required check blocks the merge. That fails closed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That asymmetry is the whole design.&lt;/strong&gt; The heavy jobs now live in a separate workflow triggered only by &lt;code&gt;workflow_dispatch&lt;/code&gt;. If the dispatch fails, the checks stay unreported and the merge stays blocked.&lt;/p&gt;

&lt;p&gt;The cost optimization and the safety net came out of the same decision. "Spend less" and "stop when something's wrong" tend to share a shape.&lt;/p&gt;

&lt;h3&gt;
  
  
  Also: deciding not to do something, with arithmetic
&lt;/h3&gt;

&lt;p&gt;We also skip the heavy jobs on PRs that only touch documentation. The design question there was whether to split the docs-only detection into its own job.&lt;/p&gt;

&lt;p&gt;A dedicated job is structurally cleaner, but &lt;strong&gt;every PR then pays for one more runner start, about 25 seconds&lt;/strong&gt;. Docs-only PRs are around 4% of recent merges, and they save roughly 300 seconds of heavy jobs each.&lt;/p&gt;

&lt;p&gt;Run the expected value: 96% of PRs paying 25 seconds outweighs 4% of PRs saving 300. So the detection happens inside an existing job and we skipped the dedicated one.&lt;/p&gt;

&lt;p&gt;"Make the structure cleaner" always sounds right. In a cost context, &lt;strong&gt;you can decide against it with arithmetic&lt;/strong&gt;. Go by which version feels more elegant and you'll miss reversals like this one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Case 4: it looks free at first, then the data grows
&lt;/h2&gt;

&lt;p&gt;One more: &lt;strong&gt;BigQuery MERGE&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;When you want to insert data without creating duplicates, MERGE is the obvious choice. Match on a key, update if it's there, insert if it isn't. One statement, and it's idempotent.&lt;/p&gt;

&lt;p&gt;The catch is that &lt;strong&gt;MERGE reads the target every time, no matter how many rows you're inserting&lt;/strong&gt;. Even for a single row, it has to scan the target side to confirm that row isn't already there.&lt;/p&gt;

&lt;p&gt;That's awkward from a cost perspective. &lt;strong&gt;While the table is small, it looks like nothing.&lt;/strong&gt; Nothing is wrong on the day you write it. But as the data accumulates, the scan per run grows with it. Your execution frequency hasn't changed, and the cost climbs anyway.&lt;/p&gt;

&lt;p&gt;And because nobody edited any code, &lt;strong&gt;this never shows up in PR review&lt;/strong&gt;. A gate at merge time asks "how much does this change add," so anything that grows without being changed is out of scope by construction. This is the purest version of what I described at the top: invisible on the day you build it, accumulating quietly afterward.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where the duplicate check should live
&lt;/h3&gt;

&lt;p&gt;So we moved the duplicate check to &lt;strong&gt;Firestore, and BigQuery only gets the insert&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Look the key up in Firestore, and only insert the rows that come back as new. BigQuery becomes append-only, and the scan for reconciliation goes away entirely.&lt;/p&gt;

&lt;p&gt;Same shape as Case 2. Not "how do we optimize the MERGE" but &lt;strong&gt;does BigQuery need to be the thing doing the duplicate check at all&lt;/strong&gt;. BigQuery is excellent at scanning large volumes and aggregating. Asking it to confirm whether one key exists is using it against the grain. Firestore is the opposite: point lookups are its job, aggregation isn't. &lt;strong&gt;Push each to the side it's good at and both end up doing something natural.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Why you need something that notices
&lt;/h3&gt;

&lt;p&gt;The point here isn't "don't use MERGE." It has its place, and we haven't replaced all of ours.&lt;/p&gt;

&lt;p&gt;The point is that &lt;strong&gt;cost grows in two different ways: expensive from day one, and expensive after it grows into it&lt;/strong&gt;. The first kind, review catches. The second kind, review never sees.&lt;/p&gt;

&lt;p&gt;Which is what the daily reporting is for. Comparing against the same weekday last week, "this app's cost keeps climbing and nobody has touched it" eventually shows up on its own. That's the reason a gate alone isn't enough.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the freed-up time is actually for
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;When implementation gets cheap, the gate that slow implementation used to provide comes off with it&lt;/li&gt;
&lt;li&gt;What's left is the running cost of everything you shipped, which is invisible on the day you ship it&lt;/li&gt;
&lt;li&gt;And it grows in two ways: &lt;strong&gt;expensive from day one&lt;/strong&gt;, and &lt;strong&gt;expensive once the data grows into it&lt;/strong&gt;. Review only ever catches the first kind&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;So the time you gained goes into asking whether it should exist&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Habits erode, so make it structural: a gate at merge, daily reporting, quotas on usage-based services, guards in CI&lt;/li&gt;
&lt;li&gt;And accept that some calls can't be automated: how recovery works, how far back you'd roll, whether this thing needs to exist at all&lt;/li&gt;
&lt;li&gt;Push what can be systematized into systems, and spend human attention on what can't. That's what reallocating the time means&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Shipping faster is unambiguously good. But spend all of the gains on shipping more and the weight of what you shipped catches up with you later. We put a week into this cleanup. The next move isn't to schedule another one, it's to build more of the machinery that keeps things from piling up in the first place.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>cloudrun</category>
      <category>devops</category>
      <category>gcp</category>
    </item>
    <item>
      <title>GitHub Actions Getting Expensive? We Cut CI Costs to a Quarter With a One-Line Change</title>
      <dc:creator>Ryosuke Tsuji</dc:creator>
      <pubDate>Wed, 29 Jul 2026 23:52:17 +0000</pubDate>
      <link>https://dev.to/ryantsuji/github-actions-getting-expensive-we-cut-ci-costs-to-a-quarter-with-a-one-line-change-31a4</link>
      <guid>https://dev.to/ryantsuji/github-actions-getting-expensive-we-cut-ci-costs-to-a-quarter-with-a-one-line-change-31a4</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;AI assistance disclosure: This article was drafted with the help of Claude. All technical content, design decisions, code references, and screenshots reflect production systems I designed and operate at airCloset; the prose was revised by me prior to publication.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hi, I'm &lt;a href="https://x.com/ryantsuji" rel="noopener noreferrer"&gt;Ryan&lt;/a&gt;, CTO at airCloset.&lt;/p&gt;

&lt;p&gt;We've migrated our GitHub Actions runners twice: &lt;strong&gt;GitHub-hosted to Blacksmith, then Blacksmith to Namespace&lt;/strong&gt;. The results, up front:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Per-run CI cost is roughly a quarter of what we paid on GitHub-hosted&lt;/li&gt;
&lt;li&gt;The slow tail (p90) is 37% shorter&lt;/li&gt;
&lt;li&gt;Runs that silently never finish went from 32 to zero&lt;/li&gt;
&lt;li&gt;Each migration was a one-line change&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This post is the measured record: how we measured, what we found, and the one migration that failed. If you're deciding whether to move off GitHub-hosted runners, the numbers here should help.&lt;/p&gt;

&lt;h2&gt;
  
  
  With AI in the loop, CI swells quietly
&lt;/h2&gt;

&lt;p&gt;First, why this matters now. Once AI agents become the main driver of development, CI cost grows on two axes at the same time.&lt;/p&gt;

&lt;p&gt;The first axis is &lt;strong&gt;run count&lt;/strong&gt;. Agents open PRs and push fixes faster than humans do, so the number of pushes climbs. In our environment, a single test workflow runs &lt;strong&gt;3,000 to 3,500 times a month&lt;/strong&gt;. Reviews are done by AI too, and every fix push triggers CI again, so this number only grows as development accelerates.&lt;/p&gt;

&lt;p&gt;The second axis is &lt;strong&gt;task count&lt;/strong&gt;. When development gets faster, you start wanting CI to do all the things that never used to be worth the cost. Another lint layer, stricter coverage gates, docs consistency checks. Each run gets longer.&lt;/p&gt;

&lt;p&gt;It's run count multiplied by task count. Leave it unattended and the bill grows quietly, but steadily.&lt;/p&gt;

&lt;h2&gt;
  
  
  First things first: if you fit in the free tier, stay put
&lt;/h2&gt;

&lt;p&gt;Let me state the opposite conclusion before anything else: &lt;strong&gt;if you fit inside the free tier, there is no reason to migrate&lt;/strong&gt;. Public repositories get standard runners for free, and private ones come with 2,000 minutes a month on the Free plan and 3,000 on Team.&lt;/p&gt;

&lt;p&gt;In fact, across airCloset, some repositories are still on GitHub-hosted runners today. Anything that fits inside the 3,000 free minutes stays where it is. There is no reason to move something that runs for free onto something you pay for.&lt;/p&gt;

&lt;p&gt;This post is for teams that blew past the free tier a long time ago and are watching the bill climb every month.&lt;/p&gt;

&lt;h2&gt;
  
  
  The migration path
&lt;/h2&gt;

&lt;p&gt;Here's the road we took.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Period&lt;/th&gt;
&lt;th&gt;Runner&lt;/th&gt;
&lt;th&gt;Trigger&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Until late March&lt;/td&gt;
&lt;td&gt;GitHub-hosted (2-core)&lt;/td&gt;
&lt;td&gt;--&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Late March to June&lt;/td&gt;
&lt;td&gt;Blacksmith (4 vCPU)&lt;/td&gt;
&lt;td&gt;Speed and cost&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Late June to now&lt;/td&gt;
&lt;td&gt;Namespace (4 vCPU/8GB)&lt;/td&gt;
&lt;td&gt;Speed, cost, and the hang problem below&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Along the way we also evaluated &lt;strong&gt;self-hosting on AWS spot instances, and decided against it because every job would pay the instance startup overhead&lt;/strong&gt;. CI is a world where boot latency compounds, so always-pooled runners win.&lt;/p&gt;

&lt;p&gt;Each migration is literally a one-line change to &lt;code&gt;runs-on&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt; jobs:
   test:
&lt;span class="gd"&gt;-    runs-on: ubuntu-latest
&lt;/span&gt;&lt;span class="gi"&gt;+    runs-on: nscloud-ubuntu-24.04-amd64-4x8
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The rest is a one-time GitHub App connection in the provider's console, and the whole edit takes a few minutes. Standard features like &lt;code&gt;actions/cache&lt;/code&gt; keep working as-is. With migration cost this close to zero, the decision comes down to measurements.&lt;/p&gt;

&lt;h2&gt;
  
  
  Migration 1: GitHub-hosted to Blacksmith
&lt;/h2&gt;

&lt;p&gt;The motivation was simple: speed and cost. We compared successful runs of the same test workflow in the two weeks on either side of the migration, a window where &lt;strong&gt;the workflow content was identical&lt;/strong&gt;, using run history from the GitHub API. Note that run counts were far lower back then, which is why the sample sizes look modest.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Median&lt;/th&gt;
&lt;th&gt;p90&lt;/th&gt;
&lt;th&gt;n&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GitHub-hosted (2-core)&lt;/td&gt;
&lt;td&gt;340s&lt;/td&gt;
&lt;td&gt;598s&lt;/td&gt;
&lt;td&gt;237&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Blacksmith (4 vCPU)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;139s (-59%)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;157s (-74%)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;130&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;To be clear about what this is: it is not "Blacksmith is 2.4x faster at the same size." We doubled the machine, from 2-core to 4 vCPU.&lt;/p&gt;

&lt;p&gt;The real point is the unit price. GitHub-hosted 2-core costs $0.006/min, and the 4-core larger runner costs $0.012/min. Blacksmith charges $0.004/min for 2 vCPU and $0.008/min for 4 vCPU. In other words, &lt;strong&gt;for 1.33x what GitHub charges for 2 cores, you rent twice the machine&lt;/strong&gt;. Run time shrank to 41%, so per-run cost worked out to roughly &lt;strong&gt;45% less&lt;/strong&gt; (1.33x price times 0.41x time), while runs got 2.4x faster. Faster and cheaper. "The same money buys twice the machine" is the actual nature of this kind of migration.&lt;/p&gt;

&lt;h2&gt;
  
  
  The snag with Blacksmith: flaky install hangs
&lt;/h2&gt;

&lt;p&gt;Blacksmith earned its keep at that price, but one problem in daily operation was impossible to ignore: &lt;strong&gt;dependency installation (the npm install step) would occasionally hang in silence, and CI would never finish&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;We never pinned down a reproduction; it was flaky. GitHub Actions jobs default to a 6-hour timeout, so left alone, a hung run keeps billing and occupying a runner for 6 hours. For a while, people noticed stuck runs, cancelled them by hand, and re-ran them, and that happened a lot. We then tightened the timeout to 30 minutes, and even after that, &lt;strong&gt;32 runs died at the timeout without anyone noticing, over about two and a half months&lt;/strong&gt;. Manual cancels aren't in that count, so 32 is a floor. In our first three weeks on Namespace: &lt;strong&gt;zero&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The raw number may look small. But when this lands in the middle of an autonomous agent loop, a silent stall becomes a stall of the whole loop. A human notices "huh, CI is stuck," cancels, and re-runs. An agent loop needs detection and retry machinery built separately, and that quietly piled up operational cost.&lt;/p&gt;

&lt;p&gt;One note in fairness: this is what we observed during our usage window, March through June of this year. It may well be improved by now, and I can't guarantee it reproduces elsewhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  Migration 2: Blacksmith to Namespace, including the failure
&lt;/h2&gt;

&lt;p&gt;The motivation was again speed and cost, plus the hang problem above. But this migration &lt;strong&gt;failed once, and we rolled it back&lt;/strong&gt;. The failure seems worth sharing as-is, so here is what happened.&lt;/p&gt;

&lt;h3&gt;
  
  
  We didn't estimate our concurrency
&lt;/h3&gt;

&lt;p&gt;Namespace's plans cap &lt;strong&gt;concurrent capacity, counted in vCPUs&lt;/strong&gt;. The Developer plan (the cheapest, with a free trial) allows 32 vCPUs on Linux. With 4 vCPU/8GB machines, that's &lt;strong&gt;8 machines at a time&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Because AI agents develop in parallel, our CI concurrency is high, and peaks blow well past 8 machines. From day one of the migration, runs piled up in the queue and CI clogged. &lt;strong&gt;We rolled back to Blacksmith after two days&lt;/strong&gt;, upgraded to the Business plan (160 vCPUs, or 40 machines), re-migrated, and it has been stable since.&lt;/p&gt;

&lt;p&gt;The lesson is simple: &lt;strong&gt;measure your repository's peak concurrency before you switch&lt;/strong&gt;. Pull run history from the GitHub API and the peak number of simultaneously running jobs falls out mechanically. Pick a plan on price and speed alone and you'll take the same detour we did.&lt;/p&gt;

&lt;h3&gt;
  
  
  Measured: the median shrank, and the tail shrank more
&lt;/h3&gt;

&lt;p&gt;Amusingly, the rollback handed us a &lt;strong&gt;clean controlled experiment&lt;/strong&gt;. The two rollback days and the five days right after re-migration ran exactly the same workflow content, so the runner is the only variable.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Median&lt;/th&gt;
&lt;th&gt;p90&lt;/th&gt;
&lt;th&gt;n&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Blacksmith (4 vCPU, rollback window)&lt;/td&gt;
&lt;td&gt;378s&lt;/td&gt;
&lt;td&gt;839s&lt;/td&gt;
&lt;td&gt;335&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Namespace (4 vCPU/8GB)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;339s (-10%)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;526s (-37%)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;311&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The median improved by about 10%, but the bigger story is the &lt;strong&gt;tail&lt;/strong&gt;. p90 dropped from 839s to 526s, 37% shorter, and the hangs went to zero. "Occasionally slow, occasionally stuck" simply disappeared, and for agent-loop operation that is worth more than a median improvement, because the duration of a loop iteration is set by its slowest link.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cost math
&lt;/h2&gt;

&lt;p&gt;Published unit prices (Linux x64, as of July 2026):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Runner&lt;/th&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;th&gt;Price/min&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GitHub-hosted standard&lt;/td&gt;
&lt;td&gt;2-core&lt;/td&gt;
&lt;td&gt;$0.006&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GitHub larger runner&lt;/td&gt;
&lt;td&gt;4-core&lt;/td&gt;
&lt;td&gt;$0.012&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Blacksmith&lt;/td&gt;
&lt;td&gt;2 vCPU&lt;/td&gt;
&lt;td&gt;$0.004&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Blacksmith&lt;/td&gt;
&lt;td&gt;4 vCPU&lt;/td&gt;
&lt;td&gt;$0.008&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Namespace&lt;/td&gt;
&lt;td&gt;4 vCPU/8GB&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;$0.004&lt;/strong&gt; (prepaid)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A 4 vCPU Namespace machine costs &lt;strong&gt;less per minute than GitHub's 2-core&lt;/strong&gt;. That's the single biggest lever. Normalize per run with the measured times, and GitHub-hosted to Blacksmith cut about 45%, Blacksmith to Namespace roughly halved it again. &lt;strong&gt;In total, about a quarter of the GitHub-hosted era.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Three caveats.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Namespace has a plan fee&lt;/strong&gt;: $100/month on Team, $250/month on Business, each with the same amount of usage included. And if you fit in the free tier, GitHub Actions costs zero. No unit price beats free. That's why the earlier "stay put" advice comes first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;$0.004/min is the prepaid rate&lt;/strong&gt; (the usage included in your plan). Overage runs at $0.006/min, which is the same as GitHub's 2-core. The assumption is that you pick a plan your usage fits inside.&lt;/li&gt;
&lt;li&gt;Prices change. These are published rates at the time of writing, so check the pricing pages (&lt;a href="https://docs.github.com/en/billing/reference/actions-runner-pricing" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; / &lt;a href="https://www.blacksmith.sh/pricing" rel="noopener noreferrer"&gt;Blacksmith&lt;/a&gt; / &lt;a href="https://namespace.so/pricing" rel="noopener noreferrer"&gt;Namespace&lt;/a&gt;) when you decide.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Which tier are you in?
&lt;/h2&gt;

&lt;p&gt;Three stages, by scale.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;You fit in the free tier&lt;/strong&gt;: stay on GitHub-hosted. Do nothing&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You're past the free tier and the bill is starting to sting&lt;/strong&gt;: Blacksmith or Namespace, either way a one-line &lt;code&gt;runs-on&lt;/code&gt; change buys twice the machine for the same money. Blacksmith has 3,000 free minutes a month and Namespace has a 30-day trial, so trying costs nothing&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI agents are your main developers and CI concurrency is high&lt;/strong&gt;: choose on tail latency and stability, and &lt;strong&gt;estimate peak concurrency before you migrate&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Wrap-up
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;With AI-driven development, CI cost grows as run count times task count&lt;/li&gt;
&lt;li&gt;GitHub-hosted to Blacksmith: 1.33x the unit price for twice the machine, 59% faster runs, about 45% cheaper per run&lt;/li&gt;
&lt;li&gt;Blacksmith to Namespace: median -10%, p90 -37%, hangs from 32 to zero, at half the unit price&lt;/li&gt;
&lt;li&gt;Namespace caps concurrency by plan. &lt;strong&gt;Estimate your peak concurrency before switching&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;If you fit in the free tier, don't migrate. If you're past it, start with one line of &lt;code&gt;runs-on&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If your GitHub Actions bill keeps growing, Namespace is a choice I can recommend with measurements behind it. I hope the numbers help you make your own call.&lt;/p&gt;

</description>
      <category>ci</category>
      <category>devops</category>
      <category>githubactions</category>
      <category>namespace</category>
    </item>
    <item>
      <title>AI-Native Redesign: The Principles Don't Change — Only the Machinery Does</title>
      <dc:creator>Ryosuke Tsuji</dc:creator>
      <pubDate>Tue, 28 Jul 2026 00:13:26 +0000</pubDate>
      <link>https://dev.to/ryantsuji/ai-native-redesign-the-principles-dont-change-only-the-machinery-does-34mk</link>
      <guid>https://dev.to/ryantsuji/ai-native-redesign-the-principles-dont-change-only-the-machinery-does-34mk</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;AI assistance disclosure: This article was drafted with the help of Claude. All technical content, design decisions, code references, and screenshots reflect production systems I designed and operate at airCloset; the prose was revised by me prior to publication.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hi, I'm &lt;a href="https://x.com/ryantsuji" rel="noopener noreferrer"&gt;Ryan&lt;/a&gt;, CTO at airCloset (a fashion-rental subscription service based in Japan).&lt;/p&gt;

&lt;p&gt;"Everything changes with AI" is the prevailing mood. My experience building and then running an internal AI platform (cortex) points the other way. The principles don't change at all. Only the machinery does. This post is about what I've come to treat as principle, what I've concluded should be broken, and the thinking behind that split.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Disclaimer&lt;/strong&gt;: "cortex" in this article is the internal codename for the AI platform built in-house at airCloset. It is unrelated to existing commercial services like Snowflake Cortex or Palo Alto Networks Cortex.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I've written about the individual pieces before: &lt;a href="https://dev.to/ryantsuji/building-one-knowledge-graph-across-46-repositories-with-static-analysis-part-1-egm"&gt;code-graph&lt;/a&gt;, &lt;a href="https://dev.to/ryantsuji/the-heart-of-the-ai-harness-a-knowledge-graph-of-the-ai-by-the-ai-for-the-ai-series-part-2-53bm"&gt;product-graph&lt;/a&gt;, &lt;a href="https://dev.to/ryantsuji/democratizing-internal-data-building-an-mcp-server-that-lets-you-search-991-tables-in-natural-1da5"&gt;db-graph&lt;/a&gt;, &lt;a href="https://dev.to/ryantsuji/we-built-a-custom-graph-rag-to-let-ai-answer-did-that-initiative-actually-work-3oda"&gt;biz-graph&lt;/a&gt;, &lt;a href="https://dev.to/ryantsuji/observability-design-for-the-ai-era-application-infrastructure-ci-llm-each-in-its-own-56eg"&gt;AI-Observability&lt;/a&gt;, the &lt;a href="https://dev.to/ryantsuji/human-on-the-loop-ai-reviewing-ai-prs-at-cortex-769-prsmonth-while-raising-the-quality-bar-4lh5"&gt;auto-review harness&lt;/a&gt;, and &lt;a href="https://dev.to/ryantsuji/fixed-before-anyone-notices-stronger-after-every-fix-self-healing-recurrence-prevention-series-1e86"&gt;Self-Healing&lt;/a&gt;. This post isn't about any of them. It's about the design principle sitting behind all of them, one abstraction level up, more essay than build log.&lt;/p&gt;

&lt;p&gt;The principle, in one sentence: &lt;strong&gt;how do we make accurate information accessible?&lt;/strong&gt; It's an old question. Libraries, legal case books, encyclopedias, search engines — every era has had its own answer using whatever tools that era gave it. Even the technology revolutions people call "paradigm shifts" mostly just changed the &lt;em&gt;means&lt;/em&gt;. The underlying question didn't move.&lt;/p&gt;

&lt;p&gt;Now AI has arrived, and my read (probably not a controversial one) is that its shift is at least on the scale of the internet, possibly larger. As with every previous paradigm shift, the means of answering "how do we make accurate information accessible?" will get redesigned from the ground up. That's what this post is about: &lt;strong&gt;AI-Native Redesign&lt;/strong&gt; — a view where you rebuild the whole design with AI treated as a given, instead of bolting an "AI tool" or a RAG layer onto a design that was optimized for humans doing everything.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Underlying Principle
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Splitting the Principle into Three Nodes
&lt;/h3&gt;

&lt;p&gt;The frame I use: split the principle into three nodes — &lt;strong&gt;creation / maintenance / consumption&lt;/strong&gt;. Someone (or something) creates it. Someone maintains it. Someone consumes it.&lt;/p&gt;

&lt;p&gt;A library: creation = the catalog and classification system, maintenance = adding new titles and updating the shelves, consumption = someone borrowing a book. Legal case books and lawyers: creation = courts writing their rulings, maintenance = new cases getting added to the corpus, consumption = a lawyer looking up a matching case for a client. Engineering docs at a company: creation = writing the spec or README, maintenance = updating it when the code changes, consumption = another engineer reading it while implementing. In every case, each of the three nodes has someone owning it, and the whole thing only works when all three keep turning.&lt;/p&gt;

&lt;p&gt;As long as the three of them are held together by a chain of trust and incentive, the system holds itself up. If one of them drops out, the whole thing slides into a negative spiral and quality decays. I'll come back to the specific mechanism later in this chapter.&lt;/p&gt;

&lt;p&gt;"Information" here changes shape depending on the domain — knowledge in books, legal precedent, internal documentation at a company, runtime behavior of a service, customer trends, the reasoning behind a decision. What they all share is this: &lt;strong&gt;as long as it lives only in one person's head, the moment scale kicks in the whole system stops working&lt;/strong&gt;. It only starts to work at scale once the information is stored in a form you can look up, and reachable when someone needs it.&lt;/p&gt;

&lt;h3&gt;
  
  
  What Has Changed, What Hasn't
&lt;/h3&gt;

&lt;p&gt;The means have changed a lot. Clay tablets, oral tradition, manuscripts, dictionaries and indices, the printing press, legal codes and lawyers, library classification, then in the last few decades: wikis and paper docs, APM and structured logs, searchable knowledge bases, distributed tracing, web-scale search engines, BI dashboards, RAG. Each era tried to optimize the three-node balance with whatever tools it had.&lt;/p&gt;

&lt;p&gt;Wikis fit the 90s because "humans write, humans update, humans search" was the only shape available. APM appeared in the 2000s because storage got cheap enough to hold the telemetry that machines generate. Each generation had its own subject and its own tooling to answer with.&lt;/p&gt;

&lt;p&gt;But the underlying question hasn't moved. Every era in every domain was solving the same problem — "make accurate information accessible" — with that era's tools. And in some sense, each generation was consciously designing the three-node balance. But once humans are in the loop, they cut corners on writing, forget to update, get things wrong, eventually stop maintaining, and the whole thing breaks. Some systems have held up (library systems, legal frameworks, established academic disciplines), but most have fallen into the negative-spiral side because the human limit was the binding constraint.&lt;/p&gt;

&lt;p&gt;What makes AI different is that it dramatically widens the range where creation and maintenance can run without much human labor. Automating creation and maintenance is itself old news — deterministic systems have covered a huge amount of it, so widely that we don't even notice anymore. It's not just engineering infrastructure (APM, CI/CD, log collection, schema validation). Media services run on the same shape: articles get created, updated, and deleted through a system, then distribution pushes them to the app, the website, and the print edition. Bank transactions, e-commerce catalogs, routing data in a map app, the timeline in a social network. Deterministic automation is everywhere, quiet enough that it's easy to forget it's there.&lt;/p&gt;

&lt;p&gt;But deterministic automation has a ceiling. There's a class of work it can't touch: qualitative judgment, articulating design context, updating documentation to follow a code change, adding annotations — anything that needs interpretation and context. AI is the first thing that reaches into that zone. And it can even sit on the consumption side (in the APM example, this means AI running the "look at the dashboards, close the improvement loop" work that humans used to do). Once the human requirement drops structurally, many domains can start running a positive spiral for the first time. But &lt;em&gt;how&lt;/em&gt; to wire this expanded capability into your own organization's loops is still a human design call. That's what this post — &lt;strong&gt;AI-Native Redesign&lt;/strong&gt; — is about.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Where this post lands&lt;/strong&gt;&lt;br&gt;
AI isn't a replacement for deterministic automation. It's a new capability that brings the zones deterministic tools never reached into the automation envelope — but redesigning your organization's loops around that expanded capability is what turns it into a change on the ground.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  A Closer Look at Each Node
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Creation.&lt;/strong&gt; Putting information into a form that can be looked up later. Writing the structure of your code down as docs. Instrumenting the runtime to emit logs and traces. Leaving the reasoning behind a decision as a design doc. Turning a customer trend into a KPI definition. Anything that leaves something behind in a form someone (or something) can come back to.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Maintenance.&lt;/strong&gt; Keeping what you already wrote from drifting away from its source as that source changes. Code changes, the doc should follow. Service topology changes, the metric definitions should follow. Customer trends shift, KPIs should follow. Decisions get overridden, the record should follow. If creation is a one-shot act, maintenance is the standing chore that never ends.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Consumption.&lt;/strong&gt; Actually reaching for what's stored and using it to decide or act. Humans reading. Machines querying. Alerts firing. AI agents pulling context. All of it counts.&lt;/p&gt;

&lt;p&gt;These aren't sequential phases (creation → maintenance → consumption). They're &lt;strong&gt;a single system held together by mutual trust and incentive&lt;/strong&gt; — you can't evaluate any one of them in isolation.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Negative Spiral — Why It Collapses in Most Real Places
&lt;/h3&gt;

&lt;p&gt;Concretely, the interdependence looks like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;If it doesn't get consumed, no one is motivated to create it.&lt;/strong&gt; Nobody keeps writing docs no one reads. No engineer polishes a dashboard no one looks at.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If it doesn't get maintained, it becomes unfit to consume.&lt;/strong&gt; A doc from six months ago gets a "probably stale" tag in someone's head and is skipped. A dashboard whose metric definition drifted since the last product change quietly seeds wrong decisions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If it doesn't get created, there's nothing to maintain in the first place.&lt;/strong&gt; Information that was never stored in a queryable form can't be tracked as it changes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;As long as all three sides believe "my part is worth doing," the loop keeps turning. But if any one of them gets too expensive, the other two lose their incentive too. "Nobody's going to read it, and it'll rot anyway," "I updated it but no one cared," "when I search it's stale or wrong" — three separate excuses that reinforce each other, and the negative spiral picks up speed.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How this collapses in most places&lt;/strong&gt;&lt;br&gt;
If any one of the three nodes gets too expensive, the other two lose their reason to invest. The reason documentation cultures don't stick, monitoring stacks go stale, and knowledge bases quietly hollow out isn't a tool problem — it's structural. The balance between the three nodes has to hold, or the whole thing decays in silence.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's why patching a single node never fixes it. Many people have seen the "our Confluence has 100k pages and no one reads or updates them" version of this. Not a Confluence problem — what happens when at least two of the three nodes (creation and maintenance) stay expensive, and the fix that gets applied is a search feature on the consumption side. The loop never closes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnp1lqb84kpuol4uu56z5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnp1lqb84kpuol4uu56z5.png" alt="Three-node loop: negative spiral vs. positive spiral. A system built on the assumption that humans handle every node slides into a negative spiral where each side reinforces the others' reasons not to invest. Rebuilding with AI as a given — each of the three nodes as an AI × deterministic hybrid, where the split is a design call and deterministic is preferred — turns it into a self-sustaining positive spiral" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The Rest of This Post
&lt;/h3&gt;

&lt;p&gt;Everything from here on works off the same principle: make information accessible, and run it as a three-node loop.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Chapter 2&lt;/strong&gt; lines up concrete cases from around us (internal docs, monitoring, business data, schema management, security) and shows how each one has been carrying its own version of the three-node problem.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chapter 3&lt;/strong&gt; argues why AI is the pivot, from the angle of "the first thing that brings zones deterministic automation couldn't reach into range of automation."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chapter 4&lt;/strong&gt; walks through the AI-native implementations in cortex (code-graph / db-graph / biz-graph / product-graph / Observability) side by side.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chapter 5&lt;/strong&gt; goes into the difference between "adding an AI tool" and AI-Native Redesign, and why keeping pre-AI design in place while adding AI is a losing pattern.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chapter 6&lt;/strong&gt; sketches what widens between organizations that have multi-layered self-sustaining loops and organizations that don't — the evolution-speed gap I think grows non-linearly over time.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each chapter reads independently, but read in order, they form a single chain: unchanging principle → common applications → what's special about AI → concrete implementations → the redesign argument → what's next.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Deterministic Automation Couldn't Reach
&lt;/h2&gt;

&lt;p&gt;I said earlier that deterministic automation has already reached almost everywhere. And I said that deterministic automation still has zones it can't touch. This chapter breaks those remaining zones down.&lt;/p&gt;

&lt;p&gt;If you look at each domain and split it into the three nodes, two categories show up.&lt;/p&gt;

&lt;h3&gt;
  
  
  Type 1: Domains Where One Node Is Still Human-Only
&lt;/h3&gt;

&lt;p&gt;Domains where deterministic automation covers some nodes but not others. Laying them out side by side:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Domain&lt;/th&gt;
&lt;th&gt;Creation&lt;/th&gt;
&lt;th&gt;Maintenance&lt;/th&gt;
&lt;th&gt;Consumption&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Internal docs&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Manual (specs / design docs / runbooks)&lt;/td&gt;
&lt;td&gt;Manual (keeping them current)&lt;/td&gt;
&lt;td&gt;Manual (reading them)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Observability&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Automatic (APM / logs / metrics)&lt;/td&gt;
&lt;td&gt;Automatic (metric def changes still manual)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Manual (dashboard triage, improvement cycles)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Business data (BI reports)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Automatic (SQL / ETL)&lt;/td&gt;
&lt;td&gt;Automatic (schema changes still manual)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Manual (analysis, interpretation)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Schema management (basic CRUD)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Automatic (ORM generation)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Manual (migration decisions)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Automatic (validation)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Security posture&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Automatic (logs, anomaly detection)&lt;/td&gt;
&lt;td&gt;Automatic (rule updates still manual)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Manual (threat assessment, alert triage)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every remaining human dependency is in &lt;strong&gt;the parts that need qualitative judgment or contextual interpretation&lt;/strong&gt;. In most of these, the bottleneck settles on the consumption side — "the information is there, but no one uses it" is the classic symptom. Humans interpreting a room full of dashboards and log streams at once hits a cognitive ceiling, and that ceiling has been the structural bottleneck. Internal docs is a special case where all three nodes are human, which is why it's the most obvious pain point.&lt;/p&gt;

&lt;h3&gt;
  
  
  Type 2: Domains Where Building the Artifact Wasn't Worth the Return
&lt;/h3&gt;

&lt;p&gt;Separate from Type 1, there's a deeper category: &lt;strong&gt;the zones we never tried in the first place&lt;/strong&gt;. The individual data and code are there, but the cost of building the "system that structures the relationships between them" never matched the payoff. This covers two subtypes — relationships that already exist but aren't visible to humans (like connections across code, or across tables) and relationships that were never even defined (like the causal link between a marketing initiative and a KPI). I'll come back to Type 2a and Type 2b in a moment.&lt;/p&gt;

&lt;p&gt;What's interesting is that the underlying work — static analysis, SQL aggregation, graph-building pipelines — is often perfectly writable in deterministic logic. The reason nothing existed here wasn't "we couldn't build it." It was "the ROI didn't clear the bar."&lt;/p&gt;

&lt;p&gt;Break down why no one built it and you get problems on both sides — a double whammy:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cost side.&lt;/strong&gt; Even if deterministic logic can build it, standing up the system to do so (the static analyzer, the extraction pipeline, the graph substrate) is real engineering effort.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Payoff side.&lt;/strong&gt; Even if you build it, humans can't consume it usefully. A human traversing a graph node-by-node isn't a real workflow, and no semantic search layer existed to sit on top.&lt;/li&gt;
&lt;li&gt;When cost is high and payoff is thin, no one signs up.&lt;/li&gt;
&lt;li&gt;So the thing that would have been useful just... didn't exist.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Type 2a: Making existing relationships visible (an analysis problem).&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Relationships that are already there in the code or the data, but that humans can't see or follow. For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Boundary connections across a codebase (which API is called from other repos, and where)&lt;/li&gt;
&lt;li&gt;Semantic relationships between tables (which set of tables represents the same business entity)&lt;/li&gt;
&lt;li&gt;Log correlation across distributed services (the causal chain of logs across services)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are all "the information already exists, but no human can analyze it" failure modes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Type 2b: Designing new relationships (a design problem).&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Relationships that don't exist anywhere yet — you have to design a conceptual model first, then extract against it. For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The causal link between an initiative and a KPI (which campaign moved which number)&lt;/li&gt;
&lt;li&gt;The pairing of a decision to its outcome (how the call in a design doc played out after implementation)&lt;/li&gt;
&lt;li&gt;The structure connecting customer segments to behavior patterns&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are a deeper failure mode: "the information doesn't exist at all — you have to design it before you can extract it." The classic example is relationships that individuals hold in spreadsheets or in their heads, which don't scale up to an organization.&lt;/p&gt;

&lt;p&gt;In my earlier biz-graph post (&lt;a href="https://dev.to/ryantsuji/we-built-a-custom-graph-rag-to-let-ai-answer-did-that-initiative-actually-work-3oda"&gt;Making Initiative Impact Analysis Explorable with Graph RAG + MCP&lt;/a&gt;), I put the difference this way:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;db-graph made existing relationships discoverable. biz-graph designed relationships that didn't exist yet and produced them. The first is an analysis problem, the second is a design problem.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Type 1 vs. Type 2
&lt;/h3&gt;

&lt;p&gt;Even though they're both "zones deterministic automation couldn't reach," they behave differently:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Type 1&lt;/strong&gt;: "this piece is human, so it's slow or stuck" — a bottleneck in an existing workflow.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Type 2&lt;/strong&gt;: "this didn't exist to begin with" — opening a domain that was never charted.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Type 2 has the larger ceiling. Removing a bottleneck makes an existing workflow faster. Charting new territory makes decisions and insights possible that weren't possible before.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdr8yp92o1nx8041vqubn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdr8yp92o1nx8041vqubn.png" alt="Type 1 vs. Type 2. Type 1 removes the human-judgment bottleneck sitting in the middle of an existing flow by widening it with AI. Type 2 takes scattered points (code, tables, initiatives, KPIs) that were previously unconnected and weaves them into a semantic web using an AI × deterministic combination" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The Pattern Underneath Both
&lt;/h3&gt;

&lt;p&gt;What Type 1 and Type 2 share: &lt;strong&gt;the parts that resisted rule-based encoding&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Type 1's leftover work: "is this dashboard anomaly a real incident or a known false positive?" "when and how do we run this migration?" "what should this doc say to its reader?"&lt;/li&gt;
&lt;li&gt;Type 2's leftover work: "which pieces of code represent the same boundary?" "which initiative moved which KPI?" "how do you lay out a conceptual schema like Week × MetricDomain?"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The common thread is &lt;strong&gt;qualitative or contextual judgment, or semantic connection&lt;/strong&gt;. Those can't be written as rules, which is why they've been left standing.&lt;/p&gt;

&lt;p&gt;That framing sets up the next chapter — why AI is the pivot. AI:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;covers the remaining qualitative-judgment nodes from Type 1, and&lt;/li&gt;
&lt;li&gt;shifts both sides of the Type 2 double whammy (it lowers the cost of building deterministic extraction systems, and it becomes the missing consumer that can pick up the resulting artifact semantically).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Not only "the zones deterministic automation couldn't reach," but also "the zones deterministic automation could reach but wasn't worth building" — both come into range of automation for the first time. That's what makes AI new.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why AI Is the Pivot
&lt;/h2&gt;

&lt;p&gt;The previous chapter split "zones deterministic automation couldn't reach" into two types: the leftover qualitative-judgment nodes in Type 1, and the cost-vs.-payoff bind in Type 2. This chapter goes into why AI is the first thing that can address both.&lt;/p&gt;

&lt;h3&gt;
  
  
  What's Genuinely New About AI
&lt;/h3&gt;

&lt;p&gt;Every automation technology before AI stayed inside the region where &lt;strong&gt;you can write deterministic logic&lt;/strong&gt;. Enumerate rules, put in branches, match patterns. It's a very powerful stack, and as I said earlier, it's threaded through nearly everything in modern life.&lt;/p&gt;

&lt;p&gt;But it hit a class of work it couldn't touch — the same class the previous chapter arrived at from a different angle: &lt;strong&gt;qualitative judgment, contextual interpretation, semantic connection&lt;/strong&gt;. A judgment that a human "kind of understands" explodes into an unmanageable condition tree the moment you try to write it as code. What Type 1's leftover human-only nodes and Type 2's un-built structuring systems had in common was that neither survived rule-based encoding.&lt;/p&gt;

&lt;p&gt;What AI brings for the first time is &lt;strong&gt;making that whole class machine-processable&lt;/strong&gt;. An LLM handles judgments that don't survive being turned into code by treating them as statistical patterns. This isn't an extension of the deterministic stack. It's a new capability, orthogonal to it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Three Directions of Change
&lt;/h3&gt;

&lt;p&gt;This new capability shifts systems in three distinct directions. They map onto the Type 1 / Type 2 split from the previous chapter: direction 1 addresses Type 1, and directions 2 and 3 address Type 2.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Direction 1: automate the remaining qualitative-judgment nodes in Type 1.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In Type 1 domains, the human-only node — dashboard interpretation, threat assessment, documentation updates, KPI analysis — is now something AI can carry.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Direction 2: cut the cost of building the structuring system.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The Type 2 side "deterministic logic could build this, but the effort was too much." AI assists at each step of building the system (spec design, code generation, testing, debugging), which drops the barrier to standing up these systems. Extraction pipelines that used to take months now regularly land in days. As a concrete data point, the initial version of cortex's biz-graph (the MCP server that handles initiative × KPI causality) went from implementation to Pulumi deployment in one day.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Direction 3: consume structured data semantically.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This one dissolves Type 2's "even if we build it, no one can use it." A graph a human couldn't traverse by hand, AI walks with a mix of graph traversal and semantic search. A question like "which of last month's marketing initiatives contributed most to new-user acquisition?" gets answered with both structures working together.&lt;/p&gt;

&lt;h3&gt;
  
  
  Deterministic-First
&lt;/h3&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Design principle: deterministic-first&lt;/strong&gt;&lt;br&gt;
If a thing can be written deterministically, write it deterministically. Keep the surface where AI does inference as narrow as necessary — this is what "containing hallucinations" actually means in practice.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Let AI touch parts that deterministic logic could have handled, and you widen the hallucination surface for nothing. The "containing hallucinations" phrase I've used across earlier posts is really this call — where to draw the line between deterministic and AI.&lt;/p&gt;

&lt;p&gt;Leaning deterministic also gets you side benefits:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Idempotency.&lt;/strong&gt; Same input, same output, every time. Critical for testing, auditing, and reproducibility.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost, by orders of magnitude.&lt;/strong&gt; Inference calls are token-metered. Deterministic execution is typically 10× or more cheaper.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Where AI goes is a decision, not a default. "AI can do it" isn't the criterion. Deterministic where deterministic works. AI only where deterministic doesn't. That division of labor sits at the core of AI-Native Redesign.&lt;/p&gt;

&lt;p&gt;Some concrete pairings:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Static analysis, SQL, pipelines → deterministic (Type 2 creation side).&lt;/li&gt;
&lt;li&gt;The structured artifact those produce → AI consumes it semantically (Type 2 consumption side).&lt;/li&gt;
&lt;li&gt;Wherever qualitative judgment is what's actually needed → AI takes it (Type 1).&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Why AI Requires Whole-System Redesign
&lt;/h3&gt;

&lt;p&gt;With earlier automation tools (APM, CI/CD, monitoring, BI), the standard move was to drop them on top of existing operations. They stood alone and didn't disturb the existing workflow, so adding them was enough.&lt;/p&gt;

&lt;p&gt;AI is different. "Add an AI tool to the existing system" won't solve what I described above. The reason is that AI's value doesn't come from any single tool feature — it comes from shifting the balance across all three nodes.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Replace just the human node in Type 1 with AI, and if the downstream consumption workflow doesn't line up, no one uses the AI's output.&lt;/li&gt;
&lt;li&gt;Stand up new Type 2 structure, and if consumption is still shaped around a human doing the judging, the artifact never actually gets used.&lt;/li&gt;
&lt;li&gt;If you don't redesign the boundary between "AI handles this" and "deterministic handles this," ownership gets ambiguous fast.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is why "add an AI tool" isn't enough. The whole system has to be redesigned with AI as a given. That's the &lt;strong&gt;AI-Native Redesign&lt;/strong&gt; thesis, and I go into it in more detail later on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Examples from cortex
&lt;/h2&gt;

&lt;p&gt;This chapter walks through the concrete systems running in cortex and shows how the principles I've laid out play out in each of them. There's a dedicated post for each system that I'll link inline. Here I'm keeping the focus on the same three questions: where's the deterministic layer, where's AI, and how was the division decided.&lt;/p&gt;

&lt;h3&gt;
  
  
  code-graph
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://dev.to/ryantsuji/building-one-knowledge-graph-across-46-repositories-with-static-analysis-part-1-egm"&gt;code-graph&lt;/a&gt; surfaces the code connections across 46 repositories (API boundaries, DB boundaries, event boundaries) into a single knowledge graph. Type 2a — making existing relationships visible.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Creation.&lt;/strong&gt; Deterministic (tree-sitter static analysis). Function, class, and import call relationships get extracted as written.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Maintenance.&lt;/strong&gt; Static analysis re-runs on code changes. Maintenance stays deterministic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Consumption.&lt;/strong&gt; AI over MCP, combining semantic search and graph traversal.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hallucination containment.&lt;/strong&gt; Boundary nodes (API endpoint, DB table, event topic) are materialized explicitly during static analysis, so AI's inference range is scoped to "stop at the boundary."&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  db-graph
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://dev.to/ryantsuji/democratizing-internal-data-building-an-mcp-server-that-lets-you-search-991-tables-in-natural-1da5"&gt;db-graph&lt;/a&gt; covers relationships between database tables and the business context around them (Type 2a). Both the ORM-level JOIN relationships and the business-entity-level semantic relationships get graphed.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Creation (structure).&lt;/strong&gt; Deterministic (static analysis of the ORM, schema extraction) + &lt;strong&gt;human review&lt;/strong&gt; as the guarantee.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Creation (business context).&lt;/strong&gt; AI generation + &lt;strong&gt;human review&lt;/strong&gt; as the guarantee.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Maintenance.&lt;/strong&gt; ORM and schema changes get detected deterministically; drift in the business context is caught in human review.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Consumption.&lt;/strong&gt; AI answers natural-language questions like "which tables are involved in the customer purchase cycle" by combining graph traversal and semantic search.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hallucination containment.&lt;/strong&gt; Structure stays deterministic. Business context is AI-generated but gated by human review as the final check. The reason db-graph goes through human review while cortex-product-graph runs on AI review comes down to blast radius: wrong DB structure or wrong business context feeds directly into organizational decision-making, so the risk of being wrong is much larger than for a code error. The call isn't just "if AI is accurate enough, let AI handle it" — it's paying the cost of human review whenever the downside of being wrong is large.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  biz-graph
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://dev.to/ryantsuji/we-built-a-custom-graph-rag-to-let-ai-answer-did-that-initiative-actually-work-3oda"&gt;biz-graph&lt;/a&gt; covers the causal relationship between initiatives and KPIs (Type 2b — designing new relationships). Unlike db-graph, there's no "JOIN target" sitting between initiatives and KPIs to begin with. The relationship has to be designed by a human first.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Creation.&lt;/strong&gt; AI (extracting structure from initiative slide decks) + deterministic (KPI data extraction, embedding-based similarity edges) + &lt;strong&gt;human schema design&lt;/strong&gt; (someone defines conceptual anchors like Week and MetricDomain).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Maintenance.&lt;/strong&gt; New initiative decks and KPI updates keep flowing in to keep the graph current — slide parsing on the AI side, the value pipeline deterministic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Consumption.&lt;/strong&gt; An AI agent handling "what's the causal relationship between last month's social campaigns and this week's new-user counts?" traverses Initiative → Week → co-occurring KPIs on the graph.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hallucination containment.&lt;/strong&gt; The human-designed schema (conceptual anchors like Week and MetricDomain) is the deterministic frame around AI. AI can only reason inside that frame. KPI extraction, similarity edges, graph construction — all deterministic. AI is confined to the consumption side and to the judgment calls during slide extraction.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  cortex-product-graph
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://dev.to/ryantsuji/the-heart-of-the-ai-harness-a-knowledge-graph-of-the-ai-by-the-ai-for-the-ai-series-part-2-53bm"&gt;cortex-product-graph&lt;/a&gt; is cortex's main knowledge graph, unifying cortex's own code, DB schema, docs, and Pulumi IaC. AI is used heavily in cortex development itself, and this system is a good example of how the AI-Native Redesign principles land in a working setup.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Creation (structure).&lt;/strong&gt; Deterministic (ts-morph extracts @graph-* JSDoc from code and Pulumi IaC, then merges with the documentation and with db-graph).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Creation (code + annotation).&lt;/strong&gt; Developer + AI assist (Claude Code / Codex generate code and the annotations at the same time).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Maintenance.&lt;/strong&gt; cortex's AI review runs per-PR and checks code logic, doc consistency, and annotation drift together, filing REQUEST_CHANGES when they don't line up.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Consumption.&lt;/strong&gt; AI over MCP, combining semantic and structural search.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hallucination containment.&lt;/strong&gt; cortex's AI review looks at code + annotation + docs together on every PR. Even PRs that get merged without a human reviewer go through AI as a review layer. Because the code and its @graph-* annotations sit next to each other in JSDoc (code and intent as an SSoT inside the same file), AI spots the gap between the code change and its intent immediately, which is why AI review stays accurate. This is also a bet on &lt;strong&gt;AI-Readability&lt;/strong&gt; — writing code in a form that's not only readable by humans but also structurally parseable by AI agents.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Observability + Self-Healing
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://dev.to/ryantsuji/observability-design-for-the-ai-era-application-infrastructure-ci-llm-each-in-its-own-56eg"&gt;AI-Observability&lt;/a&gt; handles the four monitoring axes (Application / Infrastructure / CI / LLM), and the loop from there to &lt;a href="https://dev.to/ryantsuji/fixed-before-anyone-notices-stronger-after-every-fix-self-healing-recurrence-prevention-series-1e86"&gt;Self-Healing&lt;/a&gt; is where AI-Native Redesign is at its clearest. Type 1 — removing the consumption-side bottleneck.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Creation + maintenance.&lt;/strong&gt; Deterministic (OpenTelemetry, metrics, logs, traces, deterministic alerts).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Consumption (detection).&lt;/strong&gt; Deterministic alert thresholds fire incidents.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Consumption (judgment and response).&lt;/strong&gt; AI cross-references log and trace context (pulled via Grafana MCP in practice) with the relevant source code / tables / docs (found by traversing cortex-product-graph), produces a root-cause hypothesis, and then a fix PR.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Maintenance loop.&lt;/strong&gt; The generated fix PR is quality-checked by the AI review flow that runs on top of cortex-product-graph before it merges.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hallucination containment.&lt;/strong&gt; AI never touches production directly. AI's output is always a PR — something reviewable. cortex-product-graph + AI review + auto-merge chain acts as the final gate.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is the fully-formed loop of AI-Native Redesign: monitoring → detection → AI inference → PR generation → AI review → merge. The whole loop turns, and cortex ends up in the "fixed before we notice" state (the title of the Self-Healing post).&lt;/p&gt;

&lt;h3&gt;
  
  
  The Pattern Underneath All Five
&lt;/h3&gt;

&lt;p&gt;Building these five systems, some design calls kept recurring by choice.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Creation side stays deterministic by default.&lt;/strong&gt; Anything that reads "the data as written" — static analysis, ORM, OTel, SQL, extraction pipelines — leans deterministic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI is confined to consumption and to the meaning layer on top of structured artifacts.&lt;/strong&gt; Graph traversal, semantic search, annotation generation — the judgment work that doesn't survive being written as rules — is where AI sits.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;There's always a hallucination-containment mechanism.&lt;/strong&gt; Boundary nodes, AI review, human review — the specific mechanism differs, but every system has some form of lid on AI output. AI's free-writing zone is kept narrow, and even inside that zone, its output goes through a review layer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI review quality depends on the context foundation.&lt;/strong&gt; AI can do comprehensive PR review in cortex because cortex-product-graph exists as the structured context foundation. Without it, AI would only see local information from the PR diff, and couldn't judge consistency with the rest of the codebase or the docs. Before "where and how do we use AI," the question that comes first is: "what context can we give AI to reason over?"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The review choice has a risk-profile component.&lt;/strong&gt; Even when AI review would be accurate enough, if the downside of being wrong is large, human review can still be the right call. As I mentioned with db-graph, it's a comparison between the cost of being wrong and the cost of human labor.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Human design judgment sits above the whole thing.&lt;/strong&gt; Schema design (biz-graph), guideline definitions (auto-review), monitoring target selection (Observability) — these are human calls.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The pattern to notice: the same AI-Native Redesign principles land in different shapes depending on the subject and the goal.&lt;/p&gt;

&lt;h2&gt;
  
  
  AI-Native Redesign vs. "Adding an AI Tool"
&lt;/h2&gt;

&lt;p&gt;So far the argument has been: AI is the first thing that reaches into the zones deterministic automation couldn't, and making that actually work needs whole-system design — a context foundation like cortex-product-graph, or a closed loop like Observability + Self-Healing.&lt;/p&gt;

&lt;p&gt;This chapter goes into why the "just add AI to the existing system" approach — the approach that skips the whole-system redesign — falls short.&lt;/p&gt;

&lt;h3&gt;
  
  
  What "Adding an AI Tool" Usually Looks Like
&lt;/h3&gt;

&lt;p&gt;Over the past year I've put a range of AI tools into the organization myself. Here are the shapes I've tried at least once:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;AI summary on dashboards.&lt;/strong&gt; AI reads the dashboard and gives you the takeaway.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI-generated docs.&lt;/strong&gt; Docs get produced from code.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI PR review.&lt;/strong&gt; AI reads the PR and comments.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI search on the internal knowledge base.&lt;/strong&gt; Natural-language queries against internal knowledge.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI-assisted coding.&lt;/strong&gt; Claude Code, Cursor, etc.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each one was useful in isolation. But honestly, the results topped out around 1.x — worth the deployment cost, but nowhere near a paradigm shift.&lt;/p&gt;

&lt;p&gt;That was the point where I had to step back and ask what it would take for the organization to use AI better. What came out of that were the three failure modes below, and the AI-Native Redesign direction that follows from them.&lt;/p&gt;

&lt;h3&gt;
  
  
  Failure Mode 1: The Three-Node Balance Stays Optimized for Humans
&lt;/h3&gt;

&lt;p&gt;Existing systems were shaped, within the capabilities of their era, around "humans handle every node." That assumption is baked into things you don't usually see as design choices:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Dashboards: density and count set for what a human can consume.&lt;/li&gt;
&lt;li&gt;Documentation: structure and granularity set for what a human can read.&lt;/li&gt;
&lt;li&gt;Code: conventions and granularity set for what humans write and review.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Add AI on top and AI has to operate inside "the balance optimized for humans." AI's actual strengths — watching a room full of dashboards at once, traversing docs structurally, verifying code exhaustively — get suppressed because the surrounding two nodes stay shaped for humans.&lt;/p&gt;

&lt;h3&gt;
  
  
  Failure Mode 2: AI Is Asked to Judge Without a Context Foundation
&lt;/h3&gt;

&lt;p&gt;As the earlier examples showed, AI review works comprehensively only when there's a context foundation like cortex-product-graph. Ask AI to judge without one, and it only sees local information, and its actual value doesn't come out.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;PR review AI: seeing only the PR diff, all it can do is comment on coding style.&lt;/li&gt;
&lt;li&gt;Dashboard AI summary: summarizes the numbers on that one dashboard; the relationship to the rest of the system is invisible.&lt;/li&gt;
&lt;li&gt;Doc AI search: keyword match or semantic search, each a local result.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The feeling "AI is shallow" is usually not about the AI. It's about the missing context foundation the AI was supposed to reason over.&lt;/p&gt;

&lt;p&gt;This is adjacent to what's now being called context engineering. I've written about the retrieval-side design in &lt;a href="https://dev.to/ryantsuji/graph-rag-isnt-a-one-shot-anymore-the-case-for-agentic-graph-rag-mcps-1dj5"&gt;an earlier agentic Graph RAG post&lt;/a&gt;, but in the three-node frame, retrieval quality is only the consumption node. The failure here is upstream — the creation side never produced a form that could be pulled as context.&lt;/p&gt;

&lt;h3&gt;
  
  
  Failure Mode 3: The Creation Side Doesn't Shift into a Form AI Can Consume
&lt;/h3&gt;

&lt;p&gt;Of the three directions I laid out earlier, Direction 1 (automate the leftover qualitative-judgment nodes from Type 1) can be reached by adding AI to an existing system. But Direction 2 (cut the cost of building structuring systems) and Direction 3 (consume structured data semantically) both require the creation side to change into a form AI can consume.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;JSDoc @graph-* annotations on code express structure and business intent as an SSoT (Single Source of Truth) → AI can understand structure and intent together.&lt;/li&gt;
&lt;li&gt;Logs emitted as structured events, correlated with traces and metrics from other services → AI can follow causal chains across a distributed system.&lt;/li&gt;
&lt;li&gt;Docs restructured from standalone files into data tied to code, design, and domain → AI can pull them as context.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are creation-side design changes. Adding a feature to the existing system doesn't produce them. "Add AI on the consumption side" alone caps AI's ceiling at whatever the input-side constraints are — a rough format, implicit context, purely local data.&lt;/p&gt;

&lt;h3&gt;
  
  
  What AI-Native Redesign Actually Is
&lt;/h3&gt;

&lt;p&gt;Flip the three failure modes and you get what AI-Native Redesign is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Rebalance all three nodes around "AI + human" as the assumption.&lt;/strong&gt; Which node gets AI and which stays human is redesigned from scratch. Not "add AI to a human-balanced system" — draw a new balance where AI-carried nodes are first-class parts of it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Build the context foundation first.&lt;/strong&gt; The structured context AI can reason over (something like cortex-product-graph) gets built first, and AI review and self-repair go on top of it. The opposite order — "put AI in, then notice context is missing" — is what fails.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Change the creation side.&lt;/strong&gt; Reshaping the creation side into a form AI can consume — annotations, structured events, fine-grained docs — is part of the redesign. Consumption-side additions alone aren't enough.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The five cortex implementations from earlier all did these three. code-graph shaped structure through static analysis into a form AI could reach (creation-side change). cortex-product-graph became the context foundation for judgment itself. Observability + Self-Healing redesigned all three nodes with AI in the mix, from monitoring through to auto-repair.&lt;/p&gt;

&lt;p&gt;This is the underlying reason the evolution-speed gap I sketch in the next chapter widens over time. The gap between "add an AI tool" and "AI-Native Redesign" is bigger than a linear-vs.-exponential ROI gap, from what I've seen.&lt;/p&gt;

&lt;h2&gt;
  
  
  Life After AI-Native Redesign
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What Building cortex Has Actually Felt Like
&lt;/h3&gt;

&lt;p&gt;Some things I couldn't see back when we were just deploying individual AI tools have come into view while building and running cortex.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Fix PRs now ride straight into AI review and auto-merge with no human in the path — 115 of them via Self-Healing alone in the last 30 days (&lt;a href="https://dev.to/ryantsuji/fixed-before-anyone-notices-stronger-after-every-fix-self-healing-recurrence-prevention-series-1e86"&gt;details in the Self-Healing post&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;Drift between docs, annotations, and code gets repaired automatically in places that used to sit uncorrected.&lt;/li&gt;
&lt;li&gt;The cost of chasing "how was this actually built again?" through past decisions has dropped.&lt;/li&gt;
&lt;li&gt;Individual writing speed hasn't changed much. What did change: the quality bar for what actually ships, and the fact that people other than me can now contribute. Monthly merged PRs going from the 10–23 range through March to 518 in April 2026 came from the workflow switch (main push → PR + AI review + auto-merge), not from writing more. The number is really the shape of "the ceiling of 'reviewed manually by me' came off, so this scale and this quality bar can now be sustained by more people than just me" (&lt;a href="https://dev.to/ryantsuji/building-a-real-ai-harness-auto-reviewed-prs-self-healing-ops-and-non-engineer-contributors-3lfa"&gt;data in the harness intro post&lt;/a&gt;).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This isn't "add an AI tool, get 1.x." It's what happens when multiple self-sustaining loops start turning across layers. Qualitatively a different kind of change from a one-off productivity bump.&lt;/p&gt;

&lt;h3&gt;
  
  
  What Multi-Layered Self-Sustaining Loops Actually Look Like
&lt;/h3&gt;

&lt;p&gt;Each of the cortex systems has its own self-sustaining loop.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;code-graph.&lt;/strong&gt; Every code change updates the graph, AI reviews using the updated graph.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;cortex-product-graph.&lt;/strong&gt; Every PR keeps annotations and code aligned, and AI review accuracy tightens with each pass.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Observability + Self-Healing.&lt;/strong&gt; Monitoring detects an incident, AI produces a fix PR, AI review checks it, and it merges.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;biz-graph.&lt;/strong&gt; Initiative-to-KPI relationships get extracted continuously and stay in a form usable for decision-making.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These loops run independently, but they're connected through cortex-product-graph as a shared context foundation. The output of one loop becomes the input to another — that shape of connection.&lt;/p&gt;

&lt;p&gt;Once this kind of multi-layer loop starts running inside the organization, it changes how time gets spent at a fundamental level. Not "AI helps" — closer to "a substantial share of daily work completes inside the loops."&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4xw3rverjvbxe91e5j72.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4xw3rverjvbxe91e5j72.png" alt="Multi-layered self-sustaining loops — code-graph, db-graph, biz-graph, and Observability + Self-Healing each turn as satellite loops around cortex-product-graph as the shared context foundation. In Self-Healing, both AI inference (Grafana MCP logs + cortex-product-graph) and AI review (cortex-product-graph as context) reference the shared foundation, and the loop closes without a human in the path" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Where This Structure Could Rot
&lt;/h3&gt;

&lt;p&gt;If the three-node symmetry is the axis, then AI-Native systems' own negative spiral is a question that has to be asked too. The circular dependency I presented as a virtuous cycle (AI review maintains cortex-product-graph, cortex-product-graph supports AI review) turns into a self-amplifying error loop the moment errors get into the foundation. AI reasoning confidently and consistently wrong on top of a contaminated context foundation is a real, symmetric failure mode of this design, not a hypothetical.&lt;/p&gt;

&lt;p&gt;The defenses split three ways. The reason db-graph puts human review at the final gate is entry-side containment — narrowing the flow of contamination into the foundation wherever the downside of being wrong is large. Materializing boundary nodes explicitly through static analysis is inference-range containment — narrowing where AI is allowed to reason. And keeping cortex-product-graph in a form that can always be regenerated from deterministic extraction + annotations + docs is recoverability — a way back once contamination is detected. All three are held together by "the foundation is never something AI alone writes into." The signal that this failure mode is starting to show is drift in AI review's own accuracy metrics (REQUEST_CHANGES rate, false positive / negative rate) that no one can explain — which is the symmetric-side extension of what I meant in the &lt;a href="https://dev.to/ryantsuji/observability-design-for-the-ai-era-application-infrastructure-ci-llm-each-in-its-own-56eg"&gt;AI-Observability post&lt;/a&gt; when I argued LLMs should be the fourth monitoring axis.&lt;/p&gt;

&lt;h3&gt;
  
  
  How I Read the Evolution-Speed Gap
&lt;/h3&gt;

&lt;p&gt;The rest is my read. I don't know how much of this generalizes to other organizations.&lt;/p&gt;

&lt;p&gt;One objection I want to head off: "this sounds like something you need a dedicated platform team for." airCloset doesn't have one. I built cortex solo, alongside my CTO duties, and now people from the business side, not just engineering, are shipping on top of that harness. The fact that this was reachable at all is itself Direction 2 (AI cutting the cost of building structuring systems) in action.&lt;/p&gt;

&lt;p&gt;Between organizations running AI as individual tools and organizations that have assembled multiple self-sustaining loops across layers, my sense is the evolution-speed gap widens over time.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;In the first, AI is a useful tool, but decision-and-implementation speed itself doesn't change much.&lt;/li&gt;
&lt;li&gt;In the second, the full cycle — decision → implementation → detection → repair — gets an order of magnitude faster.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Compound that gap over time and how much you can get done in the same period starts to diverge — gradually or sharply. "Exponentially" would be overstating it, but at least the way I see the gap widening, "linear" doesn't describe it either.&lt;/p&gt;

&lt;h3&gt;
  
  
  What This Post Was Trying to Say
&lt;/h3&gt;

&lt;p&gt;Take the timeless question — "how do we make accurate information accessible?" — and redesign for it with AI as a given capability. That's the idea at the center of AI-Native Redesign.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AI isn't a replacement for deterministic automation. It's a new capability that brings zones deterministic automation couldn't reach into the automation envelope for the first time.&lt;/li&gt;
&lt;li&gt;Adding AI in isolation doesn't work. The positive spiral only kicks in when all three nodes are rebuilt together.&lt;/li&gt;
&lt;li&gt;Doing that requires a context foundation AI can reason over (something like cortex-product-graph), built ahead of the AI layer.&lt;/li&gt;
&lt;li&gt;The whole thing is a stack of self-sustaining loops, and the more of them turn together, the more the organization's evolution speed changes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One last thing. I scoped this post to information access, but the creation / maintenance / consumption loop shows up far more widely than that. Cultural transmission, the survival of a business, the growth of an academic field, a living language itself — all of them turn on the same structure, where creation dries up without consumption and quality decays without maintenance. Dead languages, lost traditions, failed companies, the internal wiki no one reads — push far enough and they hollow out through the same mechanism. Which raises a question I can't answer here: how far past information systems does "AI structurally lowers the cost of creation and maintenance" actually reach? I'll leave that one open.&lt;/p&gt;

&lt;p&gt;The cortex build is one instance of trying this out. Different organizations and different subjects will land somewhere else, but the underlying question — "how do we make information accessible?" — should be the same. If this post is useful as material for laying an AI-Native answer over other contexts, that's what I was hoping for.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>knowledgegraph</category>
      <category>observability</category>
    </item>
    <item>
      <title>Observability Design for the AI Era — Reconciling PII Protection With AI Searchability, and Driving Self-Healing</title>
      <dc:creator>Ryosuke Tsuji</dc:creator>
      <pubDate>Mon, 13 Jul 2026 23:50:57 +0000</pubDate>
      <link>https://dev.to/ryantsuji/observability-design-for-the-ai-era-reconciling-pii-protection-with-ai-searchability-and-driving-2f0l</link>
      <guid>https://dev.to/ryantsuji/observability-design-for-the-ai-era-reconciling-pii-protection-with-ai-searchability-and-driving-2f0l</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;AI assistance disclosure: This article was drafted with the help of Claude. All technical content, design decisions, code references, and screenshots reflect production systems I designed and operate at airCloset; the prose was revised by me prior to publication.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hi, I'm &lt;a href="https://x.com/ryantsuji" rel="noopener noreferrer"&gt;Ryan&lt;/a&gt;, CTO at airCloset.&lt;/p&gt;

&lt;p&gt;In &lt;a href="https://dev.to/ryantsuji/observability-design-for-the-ai-era-application-infrastructure-ci-llm-each-in-its-own-56eg"&gt;Part 1&lt;/a&gt;, I walked through the four monitoring axes (application / infrastructure / CI / LLM) and the deliberately different shape each one ends up in. That's the &lt;strong&gt;write-side&lt;/strong&gt; of the observability stack, more or less wrapped up.&lt;/p&gt;

&lt;p&gt;But shaping the write side isn't the end of the story. The moment &lt;strong&gt;production data flows through the stack&lt;/strong&gt;, you have to block the path PII can take to slip in — and that's true with or without AI. It's the kind of classic observability problem where, if you cut corners, you walk straight into a leak incident.&lt;/p&gt;

&lt;p&gt;Historically, the set of people who could read logs mostly overlapped with the set who could read the DB. For engineers with DB access, logs weren't an &lt;em&gt;additional&lt;/em&gt; path to personal data — which put log-side defenses in a position where hardening them didn't meaningfully move the overall defense line for most organizations.&lt;/p&gt;

&lt;p&gt;AI breaks that premise. Non-engineers pulling logs over MCP don't have DB access. Logs became, for the first time, &lt;strong&gt;a path where someone without DB access can reach personal data&lt;/strong&gt;. On top of that, log content now flows into AI's input, which introduces new exposure surfaces: transmission to the model, and re-surfacing in the model's output. Log PII protection has shifted from "hygiene worth doing" to &lt;strong&gt;"required as a trust-boundary redesign."&lt;/strong&gt; That's the premise this post starts from.&lt;/p&gt;

&lt;p&gt;And on top of that, if &lt;strong&gt;the observability stack isn't queryable by AI&lt;/strong&gt;, the whole "AI-consumable observability" goal from Part 1 falls apart.&lt;/p&gt;

&lt;p&gt;Part 2 is about how I reconciled these two — &lt;strong&gt;protecting PII while keeping searchability for AI&lt;/strong&gt; — and how that combination ends up driving &lt;strong&gt;Self-Healing from CI failure to PR proposal&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Observability Stack Is a Natural Path for PII
&lt;/h2&gt;

&lt;p&gt;App emits a log → it lands in Loki → AI queries it through MCP. Stand up this naive flow and you get:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Customer email addresses and phone numbers in error logs&lt;/li&gt;
&lt;li&gt;Order response payloads riding inside traces&lt;/li&gt;
&lt;li&gt;DB query logs that emit full table rows&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Plain-text PII pooling in the observability stack means &lt;strong&gt;AI can search it directly&lt;/strong&gt;. This isn't really an AI problem, it's an observability problem: the stack itself becomes a PII conduit. At the same time, if you scrub PII completely, you lose &lt;strong&gt;"I want to investigate Customer A's support ticket"&lt;/strong&gt; as a query, which is a normal support workflow.&lt;/p&gt;

&lt;p&gt;cortex (the internal AI platform) had to reconcile both. The key principle was: &lt;strong&gt;don't make "block the PII path" and "search by PII" mutually exclusive&lt;/strong&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;: "cortex" here refers to airCloset's internal AI platform codename. Unrelated to Snowflake Cortex, Palo Alto Networks Cortex, etc.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Multi-Layer PII Design — Six Layers
&lt;/h2&gt;

&lt;p&gt;cortex's PII handling is six layers, each with a different role:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;th&gt;Mechanism&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Write: BQ Policy Tag&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Column-level access control&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;pii_high&lt;/code&gt; / &lt;code&gt;pii_medium&lt;/code&gt; / &lt;code&gt;pii_low&lt;/code&gt; three-tier taxonomy. Without fine-grained reader on the column, SELECT errors out with &lt;code&gt;Access Denied&lt;/code&gt; (pure CLS (Column-Level Security) — no dynamic masking)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Write: ETL DLP&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Strip plain-text PII from derived tables&lt;/td&gt;
&lt;td&gt;Cloud DLP redacts during transforms (customer support data, etc.). Placeholders like &lt;code&gt;[EMAIL_ADDRESS]&lt;/code&gt; / &lt;code&gt;[PHONE_NUMBER]&lt;/code&gt; preserve the structure&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Write: log hashing&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Plain text never reaches Loki&lt;/td&gt;
&lt;td&gt;App-side hash via &lt;code&gt;hashEmail&lt;/code&gt; (HMAC-SHA256 → 12-char prefix; key lives outside the observability stack) before log emit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Search: same function on both sides&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Look up a specific customer's logs without ever touching plain text&lt;/td&gt;
&lt;td&gt;Query-side runs the same &lt;code&gt;hashEmail&lt;/code&gt; before sending to Loki&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Output: MCP masking&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Mask when AI consumes&lt;/td&gt;
&lt;td&gt;Column-name detection masks the local part (e.g. &lt;code&gt;r***@air-closet.com&lt;/code&gt;), keeping &lt;code&gt;@domain&lt;/code&gt; so first-response triage can still tell which domain the account belonged to&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Identity separation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Internal staff email is handled in a separate track from customer PII&lt;/td&gt;
&lt;td&gt;HMAC-signed by Edge Router as auth attribution; not part of the masking pipeline&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The fourth row — &lt;strong&gt;search with the same function on both sides&lt;/strong&gt; — is where the security / usability tradeoff gets really tight.&lt;/p&gt;

&lt;p&gt;I'll use email as the running example, but the six layers guard more than email. PII spans &lt;strong&gt;names (including phonetic readings), phone numbers, addresses, postal codes, dates of birth, card and bank details, external-service IDs&lt;/strong&gt;, and more. The anonymization technique varies by the nature of the field — same-function hashing to preserve correlation (email, phone), partial masking (names, addresses), full redaction (card numbers, tokens) — and that call is made per field. What stays constant is &lt;strong&gt;the structure: which of the six layers guards it, and how&lt;/strong&gt;. That's the reusable part of the design.&lt;/p&gt;

&lt;p&gt;And this anonymization isn't confined to observability logs (Loki) either. An MCP tool that queries a service DB, for instance, pulls customer names, addresses, and phone numbers into its result set, so the same PII anonymization rules run before anything is handed back to the AI. The consistent rule is &lt;strong&gt;"anonymize PII on every data path that reaches the AI,"&lt;/strong&gt; applied across data-source types, not just one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hash on Both the Write and Search Sides
&lt;/h2&gt;

&lt;p&gt;Naively "remove PII from logs" and you can no longer answer "let me look up Customer A's logs." But if you &lt;strong&gt;hash at write time and store that hash in the log&lt;/strong&gt;, the search side can run &lt;strong&gt;the same hash function over the input&lt;/strong&gt; and find the matching record. Plain-text email never touches either end.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo0u6gk65c0y5wa7fmi91.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo0u6gk65c0y5wa7fmi91.png" alt="Hash on both write and search to keep plain-text PII out of the observability stack while preserving search" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Concretely:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Write side:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Application code&lt;/span&gt;
&lt;span class="nx"&gt;logger&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Subscription updated&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;user&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;hashEmail&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;user&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;email&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="c1"&gt;// → '7a3f9c2e0b1d' (HMAC-SHA256 12-char prefix)&lt;/span&gt;
  &lt;span class="na"&gt;plan&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;monthly&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="c1"&gt;// → Only the hashEmail result ends up in Loki&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Search side (when you want to pull a specific customer's logs):&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here's the awkward part. "Pull up Customer A's logs" — the naive way to build it hands the raw email to the AI, which then passes it to an MCP tool to search. But that means &lt;strong&gt;handing plain-text PII to the AI (the model, and the vendor behind it)&lt;/strong&gt;. Guard the inside of Loki with hashes all you want; it leaks at the search input, one step earlier.&lt;/p&gt;

&lt;p&gt;So in cortex the search tool takes a &lt;strong&gt;non-PII ID, resolves it to an email inside the MCP server, hashes it there, and returns only the hash&lt;/strong&gt;. The email exists only inside the MCP server and never reaches the model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// MCP tool resolve_email_hash (runs server-side)&lt;/span&gt;
&lt;span class="c1"&gt;// Input is an ID (non-PII). The email is never returned to the caller = the AI.&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;email&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;resolveEmailById&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// resolved from the DB, server-side&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;hash&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;hashEmail&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;email&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;secret&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;        &lt;span class="c1"&gt;// same function, same key as the write side&lt;/span&gt;
&lt;span class="c1"&gt;// → the AI gets back only the hash, never the email&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The AI takes that &lt;code&gt;hash&lt;/code&gt; and searches Loki via Grafana MCP as &lt;code&gt;{service_name="subscription"} |~ "${hash}"&lt;/code&gt;. Both the write side and the search side run &lt;strong&gt;the same &lt;code&gt;hashEmail&lt;/code&gt; with the same key&lt;/strong&gt;, so logs from the same customer collapse to the same hash. Meanwhile:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Plain-text email never enters Loki&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The query string Loki sees doesn't contain plain-text email either&lt;/strong&gt; (only the hashed value reaches it)&lt;/li&gt;
&lt;li&gt;And &lt;strong&gt;the AI (the model) never receives plain-text email either&lt;/strong&gt;. All it touches is a non-PII ID and hashes that already live in Loki. The plain-text email never leaves the trust boundary of the MCP server.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enumeration resistance comes from keeping the HMAC key outside the stack&lt;/strong&gt;. Email is a low-entropy, enumerable input space, so &lt;strong&gt;a bare one-way hash (plain SHA-256, etc.) is breakable&lt;/strong&gt;. The hash function is public, so once logs leak, an attacker just hashes a list of likely emails on their own machine and matches against the leaked values, no key required. &lt;strong&gt;HMAC folds a secret key into the hash computation itself&lt;/strong&gt;, so an attacker who doesn't have the key can't even turn a candidate email into "the same shape as the leaked hash." They never get onto the brute-force field. Keep the key only at the write side and the search tool, never in Loki itself, and you get "a log leak alone doesn't expose the plaintext unless the key leaks too", one more condition an attacker has to satisfy&lt;/li&gt;
&lt;li&gt;Truncating to a 12-char prefix (48 bits) means collisions are possible in theory, but negligible at customer-base scale. By the birthday problem, the 50% collision point sits around 20M records (≈ 2^24.5), and below that the expected collision count stays tiny. More to the point, a collision &lt;strong&gt;wouldn't leak plaintext anyway&lt;/strong&gt;: this hash is a correlation key for identifying a customer's logs, not a security boundary, so the worst case is "another customer's logs occasionally land on the same hash", a degradation of correlation accuracy, not a disclosure&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This reuses the property "same input → same hash" of hash functions in the form "&lt;strong&gt;the same function on both sides makes search work&lt;/strong&gt;." The security / debug usability tradeoff compresses cleanly.&lt;/p&gt;

&lt;p&gt;And of course, this is all just the &lt;strong&gt;app log layer&lt;/strong&gt;. The BQ side is protected by Policy Tag-based column-level access control as its own layer (rows 1–2 of the table above). The whole thing is multi-layered.&lt;/p&gt;

&lt;p&gt;What makes the "take an ID, resolve and hash inside" shape work is that &lt;strong&gt;plain-text email never crosses the trust boundary of the MCP server&lt;/strong&gt;. The easy implementation (hand the AI a raw email, let the tool search) leaks the plaintext to the model at the search input, no matter how well you guard the inside of Loki. You could argue "the vendor's terms say it won't leave," but that's a dependency on terms, and it's weak under audit. Take an ID and hash inside, and you keep plaintext away from the model &lt;strong&gt;structurally&lt;/strong&gt;, not contractually. When I said up top that PII protection has become "a trust-boundary redesign," this is the kind of design call I meant.&lt;/p&gt;

&lt;p&gt;An aside: when I was working this out, I asked an AI for help, and it suggested building an admin screen where a human manually turns emails into hashes. That's one way to keep PII away from the model, sure, but it &lt;strong&gt;doesn't fit autonomous operation&lt;/strong&gt; — a human has to step in before any investigation can start. cortex is built to run all the way through to "fixed before anyone notices" self-healing, so a solution that inserts a human isn't on the table. "Take an ID, hash inside the MCP server" came out of that constraint. What counts as an acceptable solution was, in the end, a design judgment on my side.&lt;/p&gt;

&lt;h2&gt;
  
  
  Integration Surface — "Humans = Web, AI = MCP" on the Same Backend
&lt;/h2&gt;

&lt;p&gt;Three backends (Prometheus / BigQuery / Loki) now carry the observable data, and PII is handled. The next question is &lt;strong&gt;who queries them, and how&lt;/strong&gt;. The common trap is to build "human dashboard aggregations" and "AI data feeds" separately. The moment you do:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Two implementations chasing the same question&lt;/li&gt;
&lt;li&gt;Numbers drift between them&lt;/li&gt;
&lt;li&gt;It becomes unclear which is canonical&lt;/li&gt;
&lt;li&gt;Aggregations for AI and for humans update on different schedules&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;cortex's choice: &lt;strong&gt;share one observability backend; only the consumer-facing interface differs.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo42slptcxj968ziffw6k.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo42slptcxj968ziffw6k.png" alt="Same observability backend (Prometheus / BQ / Loki) — humans through the web dashboard, AI through MCP" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Human side: AI Operations Portal
&lt;/h3&gt;

&lt;p&gt;There's an internal portal (codenamed PI Lab) that aggregates dashboards by monitoring target:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Claude Code usage&lt;/strong&gt; (the cc-usage screen from Part 1)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MCP tool usage&lt;/strong&gt; (by server / tool / user / team)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Infrastructure cost&lt;/strong&gt; (Gemini / GCP / AWS / GitHub on one screen)&lt;/li&gt;
&lt;li&gt;Alert state, deploy history, etc.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Here's what the MCP usage dashboard actually looks like:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frod3isyf56srk50gatuf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frod3isyf56srk50gatuf.png" alt="MCP tool usage dashboard — call count per server / tool plus average execution time" width="800" height="444"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Over the past 30 days, &lt;code&gt;service-product-graph&lt;/code&gt; had 37,946 calls (with 7,106 errors), &lt;code&gt;gws&lt;/code&gt; had 19,350, &lt;code&gt;db-graph&lt;/code&gt; had 17,297 — and that's just the top. &lt;strong&gt;Which MCP is used how much, where the failures are showing up&lt;/strong&gt; — all visible at a daily glance. (The "high error rate" some servers seem to have is partly typed errors counted in — expected rejections like "permission denied" — so the interpretation needs care.) The "annotation graph MCP, ~50,000 calls / 73 users" figure from the previous series came from this same view.&lt;/p&gt;

&lt;p&gt;These pages on the React side pull from BQ / Prometheus / Loki through an internal API. The aggregation logic lives at the API layer.&lt;/p&gt;

&lt;h3&gt;
  
  
  AI side: MCP
&lt;/h3&gt;

&lt;p&gt;When AI agents need the same data, they go through purpose-specific MCPs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Grafana MCP&lt;/strong&gt; — LogQL / PromQL queries against Loki / Mimir / Prometheus / Tempo. Natural-language questions like "What time window had the most errors on Service X last week?" are the agent's job to translate into LogQL / PromQL before they go over MCP&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;BQ MCP&lt;/strong&gt; (via cortex-product-graph) — SQL queries against &lt;code&gt;claude_usage.claude_usage&lt;/code&gt; / &lt;code&gt;cortex.mcp_tool_calls&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The design pivot: &lt;strong&gt;the human dashboard and the AI MCP share the same backend.&lt;/strong&gt; No separate "AI aggregation table" and "human aggregation table." Build the observability backend once, then provide &lt;strong&gt;a consumer-specific interface layer&lt;/strong&gt; (web dashboard / MCP) on top.&lt;/p&gt;

&lt;p&gt;In DDD terms, MCP and the web dashboard are both just &lt;strong&gt;presentation layers&lt;/strong&gt; — different I/O channels into the same domain (the observability backend). Treating MCP as "something special" leads to duplicate implementations; treating it as one presentation layer form keeps the design clean.&lt;/p&gt;

&lt;p&gt;That's exactly why "the observability stack is visible to AI" actually holds. Build the backend, but without &lt;strong&gt;an AI-facing presentation layer (= MCP)&lt;/strong&gt;, AI can't query it. MCP is the piece that makes "hand it to AI" actually work.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Driver of Self-Healing
&lt;/h2&gt;

&lt;p&gt;The layer that keeps the observability stack from being "just a screen to look at" is Self-Healing. I covered the full picture in &lt;a href="https://dev.to/ryantsuji/fixed-before-anyone-notices-stronger-after-every-fix-self-healing-recurrence-prevention-series-1e86"&gt;AI Harness Series Part 4&lt;/a&gt;, so I'll skip the details here, but from the observability side, the start and end of the chain are clear:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdxyljbedbscqmpzbb99e.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdxyljbedbscqmpzbb99e.png" alt="Self-Healing chain from CI failure / production alert to PR proposal" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The flow:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Detect&lt;/strong&gt; — Production alert / CI failure fires a Loki LogQL alert&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deliver&lt;/strong&gt; — POST to event-relay (the internal webhook hub)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Launch&lt;/strong&gt; — auto-review bot starts up (= an agent backed by Claude Code)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gather context&lt;/strong&gt; — The bot pulls full logs via &lt;strong&gt;Grafana MCP&lt;/strong&gt;, traces related PR / commit / code via &lt;strong&gt;Product Graph MCP&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Propose&lt;/strong&gt; — File a fix PR&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verify&lt;/strong&gt; — If CI passes, the bot auto-merges; if not, another bot reviews&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;So the starting point of Self-Healing is &lt;strong&gt;whether the observability stack can hand "what broke" to AI in the right shape&lt;/strong&gt;. If errors aren't recognized / stacktraces aren't preserved / related code (PR / commit / graph) isn't reachable — any of those missing and the chain stops cold. (The specific failure modes are in the next section.) Put another way:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The quality of observability is the ceiling for AI autonomous operation.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's the central claim of Part 2. Reframe the observability stack as &lt;strong&gt;"input that drives AI,"&lt;/strong&gt; not "monitoring infrastructure," and the priorities of your design decisions shift accordingly.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Still Open — Defining "What Counts as an Error" and the Stacktrace Design
&lt;/h2&gt;

&lt;p&gt;The biggest remaining issue, honest version.&lt;/p&gt;

&lt;p&gt;You can polish the observability stack to a mirror finish, but if the design of &lt;strong&gt;what counts as an error&lt;/strong&gt; and &lt;strong&gt;whether the stacktrace survives&lt;/strong&gt; falls apart, all of it is wasted. I touched on this earlier in &lt;a href="https://dev.to/ryantsuji/the-heart-of-the-ai-harness-a-knowledge-graph-of-the-ai-by-the-ai-for-the-ai-series-part-2-53bm"&gt;AI Harness Series Part 2&lt;/a&gt; in the context of cortex's internal knowledge graph, and it shows up on the observability side too.&lt;/p&gt;

&lt;p&gt;Concretely, here are the failure modes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;try ~ catch&lt;/code&gt; swallows the error without logging → nothing reaches the observability stack&lt;/li&gt;
&lt;li&gt;catch &lt;em&gt;does&lt;/em&gt; log, but at &lt;code&gt;console.log&lt;/code&gt;-equivalent info level → not recognized as an error&lt;/li&gt;
&lt;li&gt;Error gets emitted, but only &lt;code&gt;error.message&lt;/code&gt; is written; stacktrace is dropped → AI can't trace back to the original code&lt;/li&gt;
&lt;li&gt;An async error goes unhandled and the process falls over&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are all &lt;strong&gt;problems at the code that creates the observability entry point&lt;/strong&gt;, not at the observability stack itself. No matter how polished the stack is, if the faucet at the entry point is broken, nothing flows out.&lt;/p&gt;

&lt;p&gt;What's in place today is three layers, none of them complete:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;lint (static)&lt;/strong&gt; — The &lt;code&gt;no-silent-catch&lt;/code&gt; rule blocks empty catches and &lt;code&gt;.catch(() =&amp;gt; null)&lt;/code&gt;-style swallows. But once there's &lt;em&gt;any&lt;/em&gt; function call inside the catch, lint is satisfied — so patterns like "demote to &lt;code&gt;logger.info(err.message)&lt;/code&gt;" or "log only &lt;code&gt;error.message&lt;/code&gt; and drop the stacktrace" slip through statically&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Guideline document&lt;/strong&gt; — Rules like "use &lt;code&gt;serializeError(error)&lt;/code&gt; to store stacktrace as a structured field" and "dropping &lt;code&gt;stack&lt;/code&gt; via &lt;code&gt;logger.error(err.message)&lt;/code&gt; is a Major violation" are written down in the internal guidelines. But static checking can't enforce these; they rely on human / AI review&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI auto-review&lt;/strong&gt; — The PR auto-review bot does look at test coverage including "are error cases being tested," but it has no observability-specific checklist, so it can't systematically catch stacktrace design quality&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In other words: &lt;strong&gt;"There's a guideline, lint catches some, AI review catches some, but it's not airtight"&lt;/strong&gt; is the honest description. The real gap is that &lt;strong&gt;at the moment new code is being written, there isn't a harness that proactively suggests / completes "this should be treated as an error, this should keep its stacktrace."&lt;/strong&gt; Auto-review picks things up at PR time, but a proactive harness for the observability entry-point design itself isn't built yet.&lt;/p&gt;

&lt;p&gt;"Observability stack: done. Observability target design: still on humans." That's the honest picture. Closing that gap with a harness is the next step.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing — Static Edition + Dynamic Edition Are Lined Up; Merging Them Is the Next Series
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://dev.to/ryantsuji/building-one-knowledge-graph-across-46-repositories-with-static-analysis-part-1-egm"&gt;The code-graph series&lt;/a&gt; was about reshaping a static analysis graph so AI could query it — &lt;strong&gt;handing the structure of code as fact&lt;/strong&gt;. This two-part series was about &lt;strong&gt;handing what's happening in production right now, also as fact&lt;/strong&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Shape&lt;/th&gt;
&lt;th&gt;What's Handed Over&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Static edition (code-graph + db-graph + annotation graph)&lt;/td&gt;
&lt;td&gt;3-graph parallel + SAME_ENTITY&lt;/td&gt;
&lt;td&gt;Code and meaning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dynamic edition (Part 1 + this post)&lt;/td&gt;
&lt;td&gt;Prometheus / BQ / Loki + MCP&lt;/td&gt;
&lt;td&gt;Production behavior and cost&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The honest part: these two &lt;strong&gt;still sit side by side&lt;/strong&gt;, not joined. For cortex's stated principle of "&lt;strong&gt;don't let AI infer — hand it facts&lt;/strong&gt;" to truly reach completion, the next step is to &lt;strong&gt;pour dynamic data into the static graph and merge them&lt;/strong&gt;. This is the exact same gap I flagged as the "absence of dynamic analysis" open issue at the end of code-graph Part 2: putting "how often is this edge actually used in production?" on the static graph's nodes. That's when "hand it as fact" reaches its final form.&lt;/p&gt;

&lt;p&gt;Layer Self-Healing on top of static + dynamic and you get "AI autonomously operates," which works today. But &lt;strong&gt;merging the two editions into one graph is still ahead — that's the next series.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And one more time, observability target design (what counts as an error, whether stacktrace survives) is what really sets the ceiling. Harness-ifying that is the next homework item.&lt;/p&gt;

&lt;p&gt;Thanks for reading this far.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>observability</category>
      <category>typescript</category>
    </item>
    <item>
      <title>Observability Design for the AI Era — Application / Infrastructure / CI / LLM, Each in Its Own Shape</title>
      <dc:creator>Ryosuke Tsuji</dc:creator>
      <pubDate>Mon, 06 Jul 2026 23:44:23 +0000</pubDate>
      <link>https://dev.to/ryantsuji/observability-design-for-the-ai-era-application-infrastructure-ci-llm-each-in-its-own-56eg</link>
      <guid>https://dev.to/ryantsuji/observability-design-for-the-ai-era-application-infrastructure-ci-llm-each-in-its-own-56eg</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;AI assistance disclosure: This article was drafted with the help of Claude. All technical content, design decisions, code references, and screenshots reflect production systems I designed and operate at airCloset; the prose was revised by me prior to publication.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hi, I'm &lt;a href="https://x.com/ryantsuji" rel="noopener noreferrer"&gt;Ryan&lt;/a&gt;, CTO at airCloset.&lt;/p&gt;

&lt;p&gt;In the previous series, &lt;a href="https://dev.to/ryantsuji/making-the-context-across-46-repositories-semantically-searchable-for-ai-part-2-51d9"&gt;code-graph deep dive (Part 2)&lt;/a&gt;, I wrote about making a 46-repo codebase semantically searchable for AI. The final issue I left open in that piece was &lt;strong&gt;the absence of dynamic analysis&lt;/strong&gt;:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What lives on the graph is the fact that "this edge exists statically." How often that edge actually gets used in production isn't recorded.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A graph that gives you static facts is one thing. Telling AI &lt;strong&gt;what's actually happening in production right now&lt;/strong&gt; is a separate problem. So the same shaping discipline I applied to the static graph needs to apply to the observability stack too.&lt;/p&gt;

&lt;p&gt;This post is the first half of that story. I split it into two: Part 1 (this post) covers &lt;strong&gt;how I shape four different monitoring surfaces&lt;/strong&gt; (application / infrastructure / CI / LLM). &lt;a href="https://dev.to/ryantsuji/observability-design-for-the-ai-era-reconciling-pii-protection-with-ai-searchability-and-driving-2f0l"&gt;Part 2&lt;/a&gt; covers PII handling, the integration surface, and Self-Healing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Does "Observable to AI" Even Mean?
&lt;/h2&gt;

&lt;p&gt;The biggest lesson from the code-graph series was: &lt;strong&gt;the data has to be shaped before AI can consume it&lt;/strong&gt;. Throwing 46 repositories of source at a model blows past the context window and invites hallucination. So we shaped it — static analysis into a graph, boundary nodes given meaning, SAME_ENTITY joins between graphs — and only then handed it over.&lt;/p&gt;

&lt;p&gt;The observability stack has the exact same problem. Throw raw production logs at AI and you get:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Sheer log volume that drowns the context window&lt;/li&gt;
&lt;li&gt;No way for the model to tell errors from noise&lt;/li&gt;
&lt;li&gt;Metrics, logs, and traces that don't link to each other&lt;/li&gt;
&lt;li&gt;Questions like "what are we spending right now" that raw logs don't answer at all&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In other words, &lt;strong&gt;logs have to be reshaped before AI can use them.&lt;/strong&gt; Same problem, different domain.&lt;/p&gt;

&lt;p&gt;The catch is that the &lt;em&gt;right&lt;/em&gt; shape depends on &lt;strong&gt;what you want AI to answer&lt;/strong&gt;. At cortex (the internal AI platform), I split the monitoring surface into four axes and let each one settle into its own form:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;: "cortex" here refers to airCloset's internal AI platform codename. Unrelated to Snowflake Cortex, Palo Alto Networks Cortex, etc.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn2bfma8hhmtirm6lc5zq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn2bfma8hhmtirm6lc5zq.png" alt="Four monitoring axes, each shaped to the question's nature, then handed to AI" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Monitoring target&lt;/th&gt;
&lt;th&gt;What you want AI to answer&lt;/th&gt;
&lt;th&gt;Shape&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Application&lt;/td&gt;
&lt;td&gt;"What's happening in production right now?" (exploration)&lt;/td&gt;
&lt;td&gt;log + trace&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Infrastructure&lt;/td&gt;
&lt;td&gt;"Do we have enough resources? Anything down?" (time series)&lt;/td&gt;
&lt;td&gt;metric&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CI&lt;/td&gt;
&lt;td&gt;"What broke? Since when?" (alert + history)&lt;/td&gt;
&lt;td&gt;log + alert&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LLM&lt;/td&gt;
&lt;td&gt;"How much are we spending? Who's using how much?" (real-time + structured aggregation)&lt;/td&gt;
&lt;td&gt;metric + structured records&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;"Just push everything through OTel and dump it all in Loki" is an option. But the moment you do, you're asking one backend to answer wildly different kinds of questions — real-time "what's spending right now" alongside "monthly cost broken down by team via SQL" — and one of them is going to suffer. Splitting by purpose is the choice I made.&lt;/p&gt;

&lt;p&gt;Let me walk through each of the four axes. Application and infrastructure are the foundation, so I'll keep those brief. CI and LLM are where the AI-era design judgments actually surface, so I'll dig into those.&lt;/p&gt;

&lt;h2&gt;
  
  
  Application — OTel + Loki + Tempo, the Standard Stack
&lt;/h2&gt;

&lt;p&gt;The foundation is unremarkable. Every cortex application is instrumented with &lt;a href="https://opentelemetry.io/" rel="noopener noreferrer"&gt;OpenTelemetry&lt;/a&gt;, with traces going to Tempo, logs to Loki, and metrics to Mimir — the standard Grafana Cloud setup.&lt;/p&gt;

&lt;p&gt;There's no special trick here. What matters is the discipline: &lt;strong&gt;every app emits logs and traces in the same shape&lt;/strong&gt;. That uniformity is what lets AI later run something like &lt;code&gt;{service_name="&amp;lt;service&amp;gt;"} |~ "error"&lt;/code&gt; through MCP and investigate across services.&lt;/p&gt;

&lt;p&gt;I covered the actual instrumentation in &lt;a href="https://dev.to/ryantsuji/fixed-before-anyone-notices-stronger-after-every-fix-self-healing-recurrence-prevention-series-1e86"&gt;AI Harness Series Part 4 (Self-Healing)&lt;/a&gt;, so I'll leave the details there. The point worth repeating is: &lt;strong&gt;a standard OTel stack, properly laid down, is the precondition for everything AI-driven that comes later&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Infrastructure — Cloud Run / BigQuery / Pub/Sub Metrics, All Into Mimir
&lt;/h2&gt;

&lt;p&gt;cortex runs on GCP and stitches together Cloud Run, Cloud Run Jobs, BigQuery, Pub/Sub, Cloud Tasks, and the usual suspects. Each GCP resource's metrics (CPU, memory, execution count, latency, queue dwell time, etc.) flow through Cloud Monitoring into Mimir.&lt;/p&gt;

&lt;p&gt;Nothing special here either — just standard GCP metrics, all gathered into one Mimir instance. But that "one place" property pays off later: AI can answer "which service used the most CPU last week?" or "is there a worker with a clogged queue?" naturally, because everything is queryable from a single store. MCP picks it up from there.&lt;/p&gt;

&lt;p&gt;That's it for the foundation. Standard observability stacks are well-documented elsewhere; go read Grafana's and OpenTelemetry's docs if you want the details.&lt;/p&gt;

&lt;p&gt;The interesting AI-era design judgments are in the next two axes — CI and LLM.&lt;/p&gt;

&lt;h2&gt;
  
  
  CI — Ship Logs to Loki via Post-Hoc Pull, Not Webhook Push
&lt;/h2&gt;

&lt;p&gt;cortex runs CI on GitHub Actions, and I ship every CI log into Grafana Loki.&lt;/p&gt;

&lt;p&gt;"Why? GitHub Actions has a perfectly good UI for that" is a reasonable question. The reasons are concrete:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Having AI hit the GitHub Actions API on every investigation is slow and auth-heavy. Ingesting into Loki once means AI can &lt;strong&gt;query it ad-hoc&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;One Loki instance holds CI logs and application logs together, so you can &lt;strong&gt;cross-query&lt;/strong&gt; them&lt;/li&gt;
&lt;li&gt;LogQL alerts turn CI failure into a structured signal&lt;/li&gt;
&lt;li&gt;AI can ask "any tests that have been broken since last week?" in natural language&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But the shipping mechanism is unusual. The choice cortex made:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Don't push logs from inside the CI run. After the run finishes, pull them from the GitHub API.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1gf6svjrpf0vosl8e7xe.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1gf6svjrpf0vosl8e7xe.png" alt="Shipping CI logs via post-hoc pull instead of webhook push" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Concretely:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;When the Test job ends, a &lt;code&gt;workflow_run&lt;/code&gt; event fires&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;separate workflow&lt;/strong&gt; dedicated to log shipping triggers&lt;/li&gt;
&lt;li&gt;That workflow pulls logs from the GitHub API (&lt;code&gt;/repos/.../actions/jobs/.../logs&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;Ships them to Grafana Cloud as structured JSON (job / status / ref / pr / commit / output, etc.) via OTLP &lt;code&gt;/v1/logs&lt;/code&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Filter on &lt;code&gt;{service_name="ci", ref="main", status="failure"}&lt;/code&gt; and you get just the main-branch CI failures, cleanly.&lt;/p&gt;

&lt;p&gt;Why pull instead of push:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;CI execution and observability decouple.&lt;/strong&gt; If shipping fails, the test run is unaffected. You can also retry / replay shipping independently&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No path for PR code to touch the API key.&lt;/strong&gt; The shipping workflow runs in the default-branch context and uses base-repo secrets, not whatever a fork PR brought. The test workflow itself never touches the Grafana API key — that's a structural guarantee, not a "we trust it won't leak"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Shipping failure becomes observable.&lt;/strong&gt; If shipping lives inside CI, a shipping bug means the observability stack goes silent — and you don't notice. Split them, and the shipping workflow's success / failure is itself something you can alert on&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The moment a main-branch failure shows up, a LogQL alert fires and Slack gets pinged. That's the trigger for Self-Healing, which I cover in &lt;a href="https://dev.to/ryantsuji/observability-design-for-the-ai-era-reconciling-pii-protection-with-ai-searchability-and-driving-2f0l"&gt;Part 2&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  LLM — Gemini and Claude Code, Two Different Shapes
&lt;/h2&gt;

&lt;p&gt;The last axis is LLM observability. cortex uses both Gemini API and Claude Code (Anthropic's official CLI) heavily, and &lt;strong&gt;since both cost money, I want visibility into how they're used&lt;/strong&gt; (though the billing models differ — Gemini is pay-per-use, Claude Code is a subscription, and that difference matters later). The reason I shape them differently isn't really about "what kind of question" — it's about &lt;strong&gt;where you can instrument — the instrumentation locus&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Gemini&lt;/strong&gt; — I own the calling code, so I can wrap every call with a common helper and emit metrics inline. Prometheus is the natural fit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Code&lt;/strong&gt; — It's an external CLI; I can't wrap its calls from the inside. Usage shows up as records after the fact. A structured store (BigQuery) is the natural fit.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The "real-time vs SQL aggregation" framing of the question is a consequence of where you can instrument, not the cause. With that clarified, here's how each one plays out.&lt;/p&gt;

&lt;h3&gt;
  
  
  Gemini — Prometheus, Cost Visible in Real Time via Client-Side Estimation
&lt;/h3&gt;

&lt;p&gt;cortex uses Gemini everywhere: db-graph table description generation, code-graph field type inference, general context generation. What I want to see is &lt;strong&gt;what's expensive right now, with no lag&lt;/strong&gt;. If a runaway prompt or batch job kicks off, I don't want to wait until tomorrow's billing report.&lt;/p&gt;

&lt;p&gt;So every Gemini call goes through a common wrapper (&lt;code&gt;traceGeminiCall&lt;/code&gt;) that emits four metrics per call:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;gemini.tokens.total&lt;/code&gt; — cumulative tokens (labels: &lt;code&gt;model&lt;/code&gt; / &lt;code&gt;service&lt;/code&gt; / &lt;code&gt;type=prompt|completion&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;gemini.requests.total&lt;/code&gt; — request count (labels: &lt;code&gt;model&lt;/code&gt; / &lt;code&gt;service&lt;/code&gt; / &lt;code&gt;status&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;gemini.request.duration&lt;/code&gt; — latency histogram&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;gemini.cost.usd&lt;/code&gt; — estimated cost (labels: &lt;code&gt;model&lt;/code&gt; / &lt;code&gt;service&lt;/code&gt;)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The design choice that splits opinions is: &lt;strong&gt;who computes the cost?&lt;/strong&gt; Two options:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A. Pull from Google Cloud Billing API after the fact&lt;/strong&gt; — accurate, but billing lags by hours to a day, and &lt;strong&gt;there's no per-task cost granularity&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;B. Compute client-side from token counts × a price table&lt;/strong&gt; — instant, with &lt;strong&gt;per-task granularity attached by you&lt;/strong&gt;, but the price table needs upkeep&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I picked B. The price table lives in a constant called &lt;code&gt;GEMINI_PRICING&lt;/code&gt; and gets manually bumped whenever Google moves prices. Just &lt;code&gt;gemini-3-flash&lt;/code&gt; / &lt;code&gt;gemini-3-pro&lt;/code&gt; with input/output unit prices each. Nothing fancy.&lt;/p&gt;

&lt;p&gt;The real reason for B is &lt;strong&gt;real-time visibility&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Billing lags by hours to a day.&lt;/strong&gt; A runaway prompt or batch bleeds cost all night before tomorrow's billing surfaces it. Computing client-side, tokens times price right after the call, lets you see "what's expensive right now" at the &lt;code&gt;service&lt;/code&gt; level (&lt;code&gt;code-graph&lt;/code&gt; / &lt;code&gt;gcs-transformer&lt;/code&gt; / &lt;code&gt;db-dictionary&lt;/code&gt; and so on, app/pipeline-grained) within minutes. That's a speed billing can never match.&lt;/li&gt;
&lt;li&gt;Price table maintenance is light (Google doesn't change prices often), so the upkeep cost is trivial.&lt;/li&gt;
&lt;li&gt;Cloud Billing API authentication, fetching, normalization, fan-out is its own pipeline of weight you'd have to maintain.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then I emit &lt;code&gt;gemini_cost_usd_USD_total&lt;/code&gt; as a cumulative Prometheus counter (the doubled &lt;code&gt;usd_USD&lt;/code&gt; comes from OTel meter name &lt;code&gt;gemini.cost.usd&lt;/code&gt; combined with the unit &lt;code&gt;USD&lt;/code&gt; during Prometheus exporter conversion) and PromQL can answer "how much did we spend in the last hour" directly: &lt;code&gt;sum(increase(gemini_cost_usd_USD_total[1h]))&lt;/code&gt;. Alert fires at $1/hour, info severity, into Slack. In practice this is less an aggregation surface I query after the fact and more a tripwire: the threshold-crossing Slack alert is how a runaway gets caught.&lt;/p&gt;

&lt;p&gt;One line worth drawing here: the &lt;code&gt;gemini.cost.usd&lt;/code&gt; counter carries exactly two labels, &lt;code&gt;model&lt;/code&gt; and &lt;code&gt;service&lt;/code&gt;, and &lt;code&gt;service&lt;/code&gt; is &lt;strong&gt;coarse&lt;/strong&gt; (a bounded set of app/pipeline names). Try to push call-site-level identity onto the label, "what did that one prompt cost," and the label combinations blow up across many repos and inference types until the time-series DB can't absorb them. So the Prometheus side stays a tripwire: coarse &lt;code&gt;service&lt;/code&gt; granularity, immediate alerting, nothing finer. The per-prompt attribution question, "which prompt burned the most this week," isn't a time-series question at all, it's a &lt;strong&gt;SQL&lt;/strong&gt; one. That wants the token records in BigQuery with as much call-site context as you care to attach, which is the same reason Claude Code goes to BQ below. "I can instrument this call" and "this should live as a time series" are separate claims, and the fine-grained aggregation is where Gemini and Claude Code converge back onto the same backend.&lt;/p&gt;

&lt;p&gt;Prometheus is what you want when the question is "right now."&lt;/p&gt;

&lt;h3&gt;
  
  
  Claude Code — Send to BigQuery, Built for SQL Aggregation
&lt;/h3&gt;

&lt;p&gt;Every developer at the company uses Claude Code. But the economics differ from Gemini: it's a subscription, so token usage doesn't translate straight into a dollar figure. What I'm after here is less the cost itself and more the &lt;strong&gt;usage picture&lt;/strong&gt; — &lt;strong&gt;who's using how much, how many tokens per repo, how well the cache is landing&lt;/strong&gt; — so I can turn it into better usage.&lt;/p&gt;

&lt;p&gt;The question that split opinion: "Should Claude Code usage go to Loki too?"&lt;/p&gt;

&lt;p&gt;The answer: &lt;strong&gt;No, into BigQuery.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Why? Because Claude Code usage is, fundamentally, a &lt;strong&gt;structured ledger&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;email&lt;/code&gt; — the user&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;repository&lt;/code&gt; — which repo it was used in&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;timestamp&lt;/code&gt; — when&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;input_tokens&lt;/code&gt; / &lt;code&gt;output_tokens&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;cache_creation_input_tokens&lt;/code&gt; / &lt;code&gt;cache_read_input_tokens&lt;/code&gt; — prompt-cache effectiveness included&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And the questions you want to ask look like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"Last week, what's the cumulative spend for Team A members?"&lt;/li&gt;
&lt;li&gt;"How much did edits on Repo X cost over the past month?"&lt;/li&gt;
&lt;li&gt;"What's the prompt-cache hit ratio difference between teams?"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All of these are &lt;strong&gt;SQL aggregation questions&lt;/strong&gt;. LogQL aggregation and joins on Loki are painful. BigQuery, with a DAY partition and email as the primary key, just writes naturally.&lt;/p&gt;

&lt;p&gt;So the Claude Code → BigQuery pipeline runs in four stages:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Emit&lt;/strong&gt; — A bundled analyzer in Claude Code POSTs &lt;code&gt;UsageInput&lt;/code&gt; (token info only, no email) to an internal endpoint&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Auth proxy&lt;/strong&gt; — A Cloudflare Edge Router worker validates &lt;code&gt;CORTEX_API_KEY&lt;/code&gt; and stamps the user's email onto the request as &lt;code&gt;X-Cortex-User-Email&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ingest&lt;/strong&gt; — A Cloud Run API dedupes and publishes to Pub/Sub&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Persist&lt;/strong&gt; — A Cloud Run worker pulls from Pub/Sub, validates the schema, and streaming-inserts to BigQuery&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Two structural points worth calling out:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Identity authority lives at the Edge Router.&lt;/strong&gt; User identity is resolved exactly once, there. The emit side (Claude Code) never holds the email. This shuts down whole classes of client-side id-spoofing and social-engineering paths structurally&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pub/Sub gives async decoupling.&lt;/strong&gt; Ingest and worker are separate, so backpressure on the worker doesn't affect ingest response times. On failure, Pub/Sub DLQ retries up to five times&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What sits in BigQuery is visible day-by-day through the internal portal I'll cover in Part 2. Here's what it actually looks like:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwcvlbckus0t4aqx7s98s.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwcvlbckus0t4aqx7s98s.png" alt="Claude Code usage dashboard — 78.0B tokens over the past 30 days, 96% of which is cache read" width="799" height="445"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The numbers are interesting enough to mention briefly: in the last 30 days, &lt;strong&gt;78.0B tokens / 384K messages / 47 users / 79 repositories&lt;/strong&gt;. The one to focus on is &lt;strong&gt;Cache Read Input at 75.1B (96% of total)&lt;/strong&gt; — prompt-cache is dramatically effective. On a subscription this doesn't show up as a dollar figure, but cache read tokens carry roughly 1/10 the effective input rate, so if you were paying per-token API pricing for the same usage, this works out to roughly &lt;strong&gt;7× more efficient at the blended input level&lt;/strong&gt; versus the cache-less counterfactual. Being able to see usage efficiency as a concrete number like this is the point of the visualization; "aggregation-shaped backend matched to the question" is the design choice that makes this kind of metric &lt;strong&gt;fall out of SQL naturally and show up daily&lt;/strong&gt;. Doing the same thing in LogQL would be a battle.&lt;/p&gt;

&lt;p&gt;As a side note: &lt;strong&gt;MCP tool-call logs&lt;/strong&gt; end up in BigQuery too (&lt;code&gt;cortex.mcp_tool_calls&lt;/code&gt;), but via a simpler path — each MCP server just writes records directly, no OTel in the loop. The "annotation graph MCP used ~50,000 times by ~73 people" figure from the previous series came from this exact table.&lt;/p&gt;

&lt;p&gt;The core point of this layer is: &lt;strong&gt;don't dogmatically force everything through OTel — match the tool to the qualitative nature of the aggregation.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  To Be Continued
&lt;/h2&gt;

&lt;p&gt;That's the four axes (application / infrastructure / CI / LLM) and the design judgments behind each. The &lt;strong&gt;write-side&lt;/strong&gt; of the observability stack is wrapped up.&lt;/p&gt;

&lt;p&gt;But shaping the write side isn't the whole story. The moment production data flows through the stack, &lt;strong&gt;PII&lt;/strong&gt; becomes a constraint you have to design around. And the data has to actually be &lt;strong&gt;consumable by AI&lt;/strong&gt; through MCP, with a thoughtful integration surface for both humans (web dashboards) and AI (MCP). Connect all of that, and &lt;strong&gt;the real driver of Self-Healing&lt;/strong&gt; comes into focus from the observability side. That's the Part 2 story.&lt;/p&gt;

&lt;p&gt;Thanks for reading. Part 2, "&lt;a href="https://dev.to/ryantsuji/observability-design-for-the-ai-era-reconciling-pii-protection-with-ai-searchability-and-driving-2f0l"&gt;Observability Design for the AI Era — Reconciling PII Protection With AI Searchability, and Driving Self-Healing&lt;/a&gt;," is out now. Read on.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>observability</category>
      <category>typescript</category>
    </item>
    <item>
      <title>Making the Context Across 46 Repositories Semantically Searchable for AI</title>
      <dc:creator>Ryosuke Tsuji</dc:creator>
      <pubDate>Mon, 29 Jun 2026 23:50:05 +0000</pubDate>
      <link>https://dev.to/ryantsuji/making-the-context-across-46-repositories-semantically-searchable-for-ai-part-2-51d9</link>
      <guid>https://dev.to/ryantsuji/making-the-context-across-46-repositories-semantically-searchable-for-ai-part-2-51d9</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;AI assistance disclosure: This article was drafted with the help of Claude. All technical content, design decisions, code references, and screenshots reflect production systems I designed and operate at airCloset; the prose was revised by me prior to publication.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hi, I'm &lt;a href="https://x.com/ryantsuji" rel="noopener noreferrer"&gt;Ryan&lt;/a&gt;, CTO at airCloset.&lt;/p&gt;

&lt;p&gt;In &lt;a href="https://dev.to/ryantsuji/building-one-knowledge-graph-across-46-repositories-with-static-analysis-part-1-egm"&gt;Part 1&lt;/a&gt;, I wrote about unifying 46 repositories of production code into a single knowledge graph via static analysis. The graph itself got built, but I closed the post with &lt;strong&gt;four open issues&lt;/strong&gt;: no semantic search, node explosion, having to open the file to actually know what a function does, and the cost of writing a new parser every time a new boundary pattern showed up.&lt;/p&gt;

&lt;p&gt;This Part 2 is about &lt;strong&gt;how I solved the first one — the entry-point problem (no semantic search).&lt;/strong&gt; The other three are left exactly as Part 1 described them — I'll come back to them at the end, together with the new issues that surfaced once the entry-point problem was out of the way.&lt;/p&gt;

&lt;p&gt;The reason to start with the entry-point problem is simple: if the graph exists but the only way to reach it is grep, the model ends up inferring anyway. The whole point — &lt;strong&gt;"give the model verified facts, not inference"&lt;/strong&gt; — falls apart. So the entry-point problem had to be solved before the others.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Hint Was in db-graph
&lt;/h2&gt;

&lt;p&gt;Months earlier, I'd already solved the same structural problem in a different domain — the &lt;a href="https://dev.to/ryantsuji/democratizing-internal-data-building-an-mcp-server-that-lets-you-search-991-tables-in-natural-1da5"&gt;db-graph project&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Internally, we had a large number of DB tables spread across many services, and &lt;strong&gt;no single person had the full picture&lt;/strong&gt;. Different people knew different pieces well, but the whole map didn't fit in anyone's head. So I built db-graph: extract schemas statically from ORM definitions, generate per-table descriptions with Gemini, embed them as 768-dimensional vectors in the graph, and make the whole thing semantically searchable in natural language.&lt;/p&gt;

&lt;p&gt;At the time of that article it covered 991 tables. Today it spans &lt;strong&gt;21 schemas / 1,133 tables / 10,815 columns&lt;/strong&gt;, and finding data in natural language without knowing table names is just how people work now.&lt;/p&gt;

&lt;p&gt;The pattern that proved out there:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Static-analysis graph + AI-generated context = natural-language semantic search works.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Bringing the Same Pattern to code-graph
&lt;/h2&gt;

&lt;p&gt;If it worked for db-graph, it should work for code-graph. The moment that thought landed, I noticed something:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;code-graph already contains "DB table nodes" as boundary nodes&lt;/strong&gt; — they're one of the boundary node types I covered in Part 1.&lt;/p&gt;

&lt;p&gt;So if I just &lt;strong&gt;join&lt;/strong&gt; code-graph and db-graph, code-graph automatically inherits db-graph's semantic context. Without writing a single annotation, the existing assets alone make the graph meaningfully richer.&lt;/p&gt;

&lt;p&gt;That's where the idea of "joining graphs" first came up — not treating each graph as its own island, but designing the joins between them.&lt;/p&gt;

&lt;h2&gt;
  
  
  But API / Event / Page Still Need Meaning — and Annotating Every Function Is Off the Table
&lt;/h2&gt;

&lt;p&gt;Joining db-graph took care of DB context. But the remaining boundaries (API / Event) and the graph's entry-point type (Page) still need meaning attached. Static analysis alone can't pull intent out of those, so context has to come from somewhere else.&lt;/p&gt;

&lt;p&gt;The choice was clear: &lt;strong&gt;write the intent directly into the code via annotations&lt;/strong&gt; (the same approach used by cortex's internal knowledge graph, which I covered in &lt;a href="https://dev.to/ryantsuji/the-heart-of-the-ai-harness-a-knowledge-graph-of-the-ai-by-the-ai-for-the-ai-series-part-2-53bm"&gt;AI Harness Series, Part 2&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The catch: you can't annotate all the functions across 46 repos. There must be tens of thousands of them. Asking established teams running an existing production codebase to retroactively annotate everything is just not realistic.&lt;/p&gt;

&lt;p&gt;But here's the second realization:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What matters is just the boundary nodes.&lt;/strong&gt; So if I only annotate around the boundaries, that's enough.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;When an AI agent asks "what breaks if I change this code" or "what other repos call this API," what it needs isn't a per-function logic explanation. It needs &lt;strong&gt;boundary intent&lt;/strong&gt; — what is this screen for, what does this API return, what milestone in the business does this Event mark.&lt;/p&gt;

&lt;p&gt;= &lt;strong&gt;Minimum annotations, maximum meaning.&lt;/strong&gt; That became the heart of the design.&lt;/p&gt;

&lt;h2&gt;
  
  
  Designing the annotation graph
&lt;/h2&gt;

&lt;p&gt;Putting it together (internally we call this annotation graph &lt;strong&gt;service-product-graph&lt;/strong&gt;, or SPG):&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frg6cw54333sci83uz1g1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frg6cw54333sci83uz1g1.png" alt="Three graphs joined as peers form a knowledge graph that carries meaning" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Three graphs sit &lt;strong&gt;as peers, joined by SAME_ENTITY edges&lt;/strong&gt;. There's no hierarchy — &lt;strong&gt;you can start from any graph and reach the others&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;code-graph (structure)&lt;/strong&gt; — functions / classes / boundary nodes from static analysis (46 repos)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;db-graph (DB context)&lt;/strong&gt; — 1,133 tables, semantically described&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;annotation graph (intent)&lt;/strong&gt; — &lt;code&gt;@graph-*&lt;/code&gt; tags written only around boundaries&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The entry point for AI agents is a single &lt;strong&gt;MCP server&lt;/strong&gt; that traverses all three graphs. AI agents never hit db-graph directly — the annotation graph's MCP server proxies db-graph calls on their behalf.&lt;/p&gt;

&lt;p&gt;The annotation graph has 7 node types: Page / Section / Dialog / Field / Action / Api / Task. The early version was screen-focused and called &lt;code&gt;screen-graph&lt;/code&gt;, but once it grew to cover backend Api / Task, it was renamed to service-product-graph.&lt;/p&gt;

&lt;h2&gt;
  
  
  An Annotation Example
&lt;/h2&gt;

&lt;p&gt;Here's what an annotation looks like (fictional, but close in shape to the real ones):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="cm"&gt;/**
 * @graph-page /home
 * @graph-business Main screen. Members can see what they're currently renting, buy items, and initiate returns.
 * @graph-label Home Screen
 * @graph-has-section banners, wearing-items, wearing-return, delivery-status
 * @graph-has-dialog buying-modal, return-modal
 * @graph-navigates-to /return-procedure, /checkout, /my-karte
 * @graph-calls GET /api/v1/wearing
 * @graph-reads admin_delivery_orders, admin_rental_items
 * @graph-flow styling-loop
 * @graph-status monthly-member
 */&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things matter here:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;@graph-business&lt;/code&gt;&lt;/strong&gt; carries the intent text (in our actual codebase it's written in Japanese). This is exactly what gets vectorized — it's the substance of semantic search.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;@graph-flow&lt;/code&gt; / &lt;code&gt;@graph-status&lt;/code&gt;&lt;/strong&gt; carry where this sits in the member lifecycle (free signup → monthly subscription → styling loop → cancellation, etc.) and which member segment it's for. They add a second dimension of meaning: "this screen shows up inside the styling loop for monthly members."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There's also &lt;code&gt;@graph-case&lt;/code&gt; (the conditional pattern tag that test cases derive from), but that's for another time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running Annotations Without Interfering With the Day-to-Day Dev Workflow
&lt;/h2&gt;

&lt;p&gt;This is where it gets practical.&lt;/p&gt;

&lt;p&gt;Once I committed to building annotation graph, here were the constraints:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Engineers run normal product dev with human code review&lt;/li&gt;
&lt;li&gt;AI review isn't wired up on every repo yet — cortex's fully automated review (covered in &lt;a href="https://dev.to/ryantsuji/ai-isnt-something-to-trust-its-something-to-design-series-final-30aa"&gt;AI Harness Series, Part 6&lt;/a&gt;) only works inside the cortex monorepo&lt;/li&gt;
&lt;li&gt;Asking humans to review annotations on top of their normal review load is a non-starter&lt;/li&gt;
&lt;li&gt;Even a split like "humans review the code, the AI reviews the annotations" inside the same PR mixes two review streams together and just confuses everyone&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In other words: &lt;strong&gt;don't mix humans and AI inside the same PR&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The solution was to physically separate annotations onto their own branch.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff5sbie95vgkzk6wz5pdn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff5sbie95vgkzk6wz5pdn.png" alt="Separate the AI-managed annotation branch from the human-managed main branch" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Leave main untouched; engineers' normal flow stays exactly as it was&lt;/li&gt;
&lt;li&gt;Stand up a separate &lt;strong&gt;annotation branch&lt;/strong&gt; that's the AI's exclusive territory&lt;/li&gt;
&lt;li&gt;When main changes, a webhook fires&lt;/li&gt;
&lt;li&gt;The annotation branch handles &lt;strong&gt;generating&lt;/strong&gt; the diff annotations &lt;strong&gt;and reviewing&lt;/strong&gt; them — the AI does both, end-to-end&lt;/li&gt;
&lt;li&gt;From the engineer's side, they only touch main and don't even need to know annotations exist&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is the "every line of code passes through an AI gate" ideal from &lt;a href="https://dev.to/ryantsuji/ai-isnt-something-to-trust-its-something-to-design-series-final-30aa"&gt;AI Harness Series, Part 6&lt;/a&gt;, adapted to the constraints of an existing organization. cortex (the internal AI platform) is a monorepo I assemble from scratch, so "every commit passes the AI gate" actually holds there. For the 46-repo production system, that precondition doesn't hold. So instead of giving up on the ideal, I split it: &lt;strong&gt;engineers' workflow on one branch, AI's annotation workflow on another, both running in parallel&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Protecting Cross-Graph Consistency With an SLO
&lt;/h3&gt;

&lt;p&gt;Just running the annotation pipeline doesn't guarantee the &lt;strong&gt;quality of the joins&lt;/strong&gt; between the three graphs (code-graph / db-graph / annotation graph). So there's a set of SLOs that automatically check the consistency across the entire graph.&lt;/p&gt;

&lt;p&gt;The main rules:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;API chain connectivity&lt;/strong&gt; — at least &lt;strong&gt;95%&lt;/strong&gt; of &lt;code&gt;HANDLES_API&lt;/code&gt; handlers must have downstream function calls (= no handlers that receive an API and then do nothing)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DB access completeness&lt;/strong&gt; — at least &lt;strong&gt;80%&lt;/strong&gt; of DB read/write edges must be &lt;strong&gt;joined to db-graph column nodes&lt;/strong&gt; (= code-graph's DB boundaries are connected to db-graph's meaning)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Event field resolution&lt;/strong&gt; — at least &lt;strong&gt;70%&lt;/strong&gt; of Event edges must carry field-level information&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No ambiguous edges&lt;/strong&gt; — name-resolution-ambiguous edges must be &lt;strong&gt;0&lt;/strong&gt; (severity: error)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are really just &lt;strong&gt;a naive question — "shouldn't the boundaries connect to each other?" — turned into an SLO.&lt;/strong&gt; If anything drops below threshold, an alert fires, and the trustworthiness of the whole graph gets defended every day.&lt;/p&gt;

&lt;p&gt;The daily boundary-analysis cron from Part 1 (5% connection-rate drop = alert) was code-graph-only. This is a &lt;strong&gt;cross-graph SLO&lt;/strong&gt; — it guards the joins between graphs themselves. Add a parser to one repo, write a new annotation, change a schema — whatever happens, by the next morning a quality drop in any join becomes visible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Joining the Static Graph and the Annotation Graph via SAME_ENTITY Bridges
&lt;/h2&gt;

&lt;p&gt;I've been writing "join" casually, but the actual joining wasn't that straightforward.&lt;/p&gt;

&lt;p&gt;Static-analysis API / Page / Task nodes and annotation graph API / Page / Task nodes are created as &lt;strong&gt;separate nodes&lt;/strong&gt;. They mean the same thing, but their names / paths / identifiers don't match by themselves — there's nothing automatic about lining them up.&lt;/p&gt;

&lt;p&gt;To connect them, we generate a separate edge type called SAME_ENTITY. There are three bridges:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;API bridge&lt;/strong&gt; — API path normalization with a 4-stage fallback

&lt;ol&gt;
&lt;li&gt;Per-repo prefix conversion (e.g., normalize console-side &lt;code&gt;/console/api/&lt;/code&gt; to &lt;code&gt;/api/&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;Version stripping (&lt;code&gt;/v1.x/&lt;/code&gt; → &lt;code&gt;/&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;Parameter normalization (unify &lt;code&gt;/:id&lt;/code&gt;, &lt;code&gt;/{id}&lt;/code&gt; to &lt;code&gt;/:dynamic&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;Exact match → tolerate trailing &lt;code&gt;?&lt;/code&gt; → strip trailing &lt;code&gt;:dynamic?&lt;/code&gt; → finally fall back to a dynamic-dispatch boundary &lt;code&gt;:dynamic&lt;/code&gt;, loosening progressively&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Page bridge&lt;/strong&gt; — 6 strategies applied in priority order (URL direct match, component path match, itemId match, PascalCase normalization match, parent-directory linking, strip dynamic segments and match parent URL)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Task bridge&lt;/strong&gt; — 8 per-repo patterns&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There was also one operational footgun. The first implementation used &lt;code&gt;INSERT NOT EXISTS&lt;/code&gt; to avoid duplicates. But BigQuery's streaming-buffer visibility lag let duplicates slip in — in one repo the edges doubled from 106 to 214 overnight. We fixed it by rewriting to &lt;code&gt;MERGE INTO&lt;/code&gt; to make the operation idempotent.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Result: Entering the Graph from "the subscription-fee calculation"
&lt;/h2&gt;

&lt;p&gt;With all of this in place, the entry-point problem from the end of Part 1 was finally solved:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"the subscription-fee calculation for members seems off"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Throw this natural-language query at annotation graph and vector search returns the related nodes (Page / Api / Function / DB table) &lt;strong&gt;as facts&lt;/strong&gt;. From there, SAME_ENTITY takes you over to code-graph functions, including callers and callees in other repos. From the DB boundaries in code-graph, you can cross into db-graph and pull the relevant columns.&lt;/p&gt;

&lt;p&gt;The entry point can be anywhere — "what calls this table?" starts from db-graph, "what's the blast radius of this function?" starts from code-graph, both walk the same connected network. From a single natural-language query, or from a specific node, &lt;strong&gt;you can now traverse all three graphs and get every relevant piece of code plus every relevant DB schema&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The Part 1 lament — &lt;strong&gt;"the graph is there but the entry point is missing"&lt;/strong&gt; — could finally be put to bed.&lt;/p&gt;

&lt;h3&gt;
  
  
  Real Usage Numbers
&lt;/h3&gt;

&lt;p&gt;From 2026-04-16 (first production deployment) to the time of writing — about 2.5 months — the annotation graph's MCP server has handled &lt;strong&gt;~50,000 calls from ~73 users&lt;/strong&gt;. The breakdown:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Engineers (PI Division + QA + relevant engineering teams)&lt;/strong&gt; — ~47,000 calls / 51 users&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Non-engineers (stylists / customer support / mall operations / executives / administration)&lt;/strong&gt; — ~2,800 calls / 21 users&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The interesting line is the second one. "Search the codebase in natural language" is usually an engineer's tool — but once the entry-point problem was solved, &lt;strong&gt;people outside engineering&lt;/strong&gt; started using it too, asking things like "how does this feature actually work?" or "what's in this DB?" in their own words.&lt;/p&gt;

&lt;p&gt;This is adjacent to the "non-engineers writing specs with AI" trend I covered in &lt;a href="https://dev.to/ryantsuji/the-author-doesnt-have-to-be-an-engineer-how-the-harness-holds-quality-series-part-5-12e4"&gt;AI Harness Series, Part 5&lt;/a&gt; — &lt;strong&gt;a graph that can be queried by meaning starts to matter org-wide.&lt;/strong&gt; Call volume is overwhelmingly dominated by engineers, of course. The interesting thing is the &lt;strong&gt;range of job roles starting to pick it up&lt;/strong&gt;. That's the real impact of solving the entry-point problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  MCP as the Single Front Door
&lt;/h2&gt;

&lt;p&gt;The MCP server is the cross-graph entry point. It exposes six tools — service search / service detail / API detail / data-flow tracing / impact-radius tracing / business-rule full-text search — and that's the only entry point AI agents ever touch.&lt;/p&gt;

&lt;p&gt;One design choice worth calling out: &lt;strong&gt;AI agents never talk to db-graph directly.&lt;/strong&gt; The annotation graph's MCP proxies db-graph calls. From the agent's side, the mental model stays simple: "ask one MCP and get everything back."&lt;/p&gt;

&lt;p&gt;That makes the full chain — "Screen → API → Code → DB → Column" — traversable in a single MCP tool call.&lt;/p&gt;

&lt;h2&gt;
  
  
  April–May Timeline of Trial and Error
&lt;/h2&gt;

&lt;p&gt;Same approach as Part 1 (pulling commits from Jan–Mar). For Part 2, the key commits are from April–May.&lt;/p&gt;

&lt;h3&gt;
  
  
  April: Expansion and the First Bridges
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;2026-04-14&lt;/strong&gt; ─ &lt;code&gt;refactor(graph): rename screen-graph to service-product-graph&lt;/code&gt; — declaration that the scope expands from screen-only to whole-service&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-04-15&lt;/strong&gt; ─ &lt;code&gt;feat(graph): add Api and Task node types to service-product-graph parser&lt;/code&gt; — Api / Task node types added&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-04-15&lt;/strong&gt; ─ &lt;code&gt;feat(mcp): add cross-graph tools to service-product-graph MCP&lt;/code&gt; — &lt;strong&gt;cross-graph tools land&lt;/strong&gt; (the single front door across all three graphs)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-04-15&lt;/strong&gt; ─ &lt;code&gt;feat(graph): add SAME_ENTITY bridge edges between service-product-graph and code-graph&lt;/code&gt; — &lt;strong&gt;first bridges&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-04-18&lt;/strong&gt; ─ &lt;code&gt;feat(graph): resolve Redis keys to code-graph boundary nodes&lt;/code&gt; — boundary resolution through Redis&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-04-19&lt;/strong&gt; ─ &lt;code&gt;feat(service-product-graph): add EventBridge EMITS_TO support + SAME_ENTITY bridge&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-04-20&lt;/strong&gt; ─ &lt;code&gt;feat(code-graph, service-product-graph): improve SAME_ENTITY boundary bridge coverage&lt;/code&gt; — 4-stage fallback locked in&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-04-21&lt;/strong&gt; ─ &lt;code&gt;feat(auto-review): SPG annotation auto-maintenance pipeline&lt;/code&gt; — &lt;strong&gt;AI auto-maintenance pipeline&lt;/strong&gt; (= what Part 1 hinted at with "humans alone can't, but AI can")&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-04-22&lt;/strong&gt; ─ &lt;code&gt;feat(service-product-graph): add Task SAME_ENTITY bridge to code-graph&lt;/code&gt; — all three bridges in place&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  May: Stabilizing and Expanding
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;2026-05-01&lt;/strong&gt; ─ Annotation generation moves from local execution to a Cloud Run Job; operation stabilizes&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-05-05&lt;/strong&gt; ─ &lt;code&gt;feat(spg): add mall repos to SPG indexing&lt;/code&gt; — mall repos indexed&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-05-06&lt;/strong&gt; ─ &lt;code&gt;feat(spg): add Go-aware parser&lt;/code&gt; — &lt;strong&gt;Go support&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-05-06 to 08&lt;/strong&gt; ─ Page bridge strategies expanded to six, connection rate hits 100%&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  What This Timeline Says
&lt;/h3&gt;

&lt;p&gt;April 15 was the day "expansion + cross-graph tools + bridges" landed in close succession. Over the next week, "Redis / EventBridge / Task bridges / annotation auto-maintenance" stacked up week over week.&lt;/p&gt;

&lt;p&gt;In particular, &lt;strong&gt;the annotation auto-maintenance pipeline on April 21&lt;/strong&gt; is where the "humans alone can't do this, but AI can" promise from Part 1 got cashed in. From that point on, annotation shifted from "humans grind through writing them" to "design the whole operation assuming AI writes them."&lt;/p&gt;

&lt;h2&gt;
  
  
  What Still Isn't Solved
&lt;/h2&gt;

&lt;p&gt;Solving the entry-point problem didn't make everything clean. A few issues remain.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Maintaining Annotation Coverage
&lt;/h3&gt;

&lt;p&gt;The frontend side is annotated heavily. Backend / Go / batch are still thin. &lt;strong&gt;Some nodes will always be missing annotations&lt;/strong&gt; — that's structural, and you can't drive it to zero. It's an ongoing operational issue.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Bridge Mis-Joins Aren't Fully Eliminated Structurally
&lt;/h3&gt;

&lt;p&gt;The Page bridge in particular has cases where multiple annotation Pages map to the same boundary — that's structural and unavoidable. Adding more strategies got coverage to 100%, but &lt;strong&gt;guaranteeing "every join is correct" 100% is hard&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. No Dynamic Analysis
&lt;/h3&gt;

&lt;p&gt;The graph only carries the fact that "this edge exists statically." How often that edge actually &lt;strong&gt;gets used in production&lt;/strong&gt; isn't recorded. Piping production execution counts back into the static graph and surfacing dead-code edges as a separate signal — that's still untouched.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Onboarding Cost When a New Repo Joins Production
&lt;/h3&gt;

&lt;p&gt;Every time a new repo enters production, the bridge normalization rules and per-repo patterns need adjusting. This is the annotation-graph-side version of Part 1's fourth issue (the cost of adding a new parser for every new boundary pattern).&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing: Not "Thrown Away," but "Evolved"
&lt;/h2&gt;

&lt;p&gt;In Part 1's closing note, I touched on the fact that the cortex side (the internal AI platform) bailed out of the code-graph approach &lt;strong&gt;early&lt;/strong&gt; and bet on an annotation-based knowledge graph instead. The bail-out was fast enough that calling it "thrown away" wouldn't be wrong — but looking back across this whole series, the more accurate word is &lt;strong&gt;"evolved."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;What it evolved into, in the end, is &lt;strong&gt;three graphs joined as peers&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;code-graph (structure)&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;db-graph (DB context)&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;annotation graph (boundary intent)&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Joined by SAME_ENTITY, served to the agent through MCP. The thing static analysis alone couldn't deliver — querying by meaning — became workable by reusing the db-graph success pattern and adding minimal annotations only at the boundaries.&lt;/p&gt;

&lt;p&gt;And one more framing: paired with the &lt;a href="https://dev.to/ryantsuji/ai-isnt-something-to-trust-its-something-to-design-series-final-30aa"&gt;AI Harness Series, Parts 1–6&lt;/a&gt;, this series sits as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;AI Harness series&lt;/strong&gt; — how to live with AI when you're assembling the system from scratch yourself&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;code-graph-deep-dive series (Part 1 + Part 2)&lt;/strong&gt; — how to live with AI inside an existing organization's running production system&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;= the same philosophy (design without trusting AI), implemented under two different sets of constraints.&lt;/p&gt;

&lt;p&gt;Thanks for reading this far.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>knowledgegraph</category>
      <category>staticanalysis</category>
      <category>typescript</category>
    </item>
    <item>
      <title>Got the Top 7 Badge — honestly thrilled 🙌</title>
      <dc:creator>Ryosuke Tsuji</dc:creator>
      <pubDate>Wed, 24 Jun 2026 00:21:42 +0000</pubDate>
      <link>https://dev.to/ryantsuji/got-the-top-7-badge-honestly-thrilled-535</link>
      <guid>https://dev.to/ryantsuji/got-the-top-7-badge-honestly-thrilled-535</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/devteam/top-7-featured-dev-posts-of-the-week-55d8" class="crayons-story__hidden-navigation-link"&gt;Top 7 Featured DEV Posts of the Week&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
      &lt;a href="https://dev.to/devteam/top-7-featured-dev-posts-of-the-week-55d8" class="crayons-article__context-note crayons-article__context-note__feed"&gt;&lt;p&gt;Cyberpunk cat RPGs and robot personalities&lt;/p&gt;

&lt;/a&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;
          &lt;a class="crayons-logo crayons-logo--l" href="/devteam"&gt;
            &lt;img alt="The DEV Team logo" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Forganization%2Fprofile_image%2F1%2Fd908a186-5651-4a5a-9f76-15200bc6801f.jpg" class="crayons-logo__image" width="800" height="800"&gt;
          &lt;/a&gt;

          &lt;a href="/jess" class="crayons-avatar  crayons-avatar--s absolute -right-2 -bottom-2 border-solid border-2 border-base-inverted  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F264%2Fb75f6edf-df7b-406e-a56b-43facafb352c.jpg" alt="jess profile" class="crayons-avatar__image" width="400" height="400"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/jess" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Jess Lee
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Jess Lee
                &lt;a href="/++"&gt;&lt;img alt="Subscriber" class="subscription-icon" src="https://assets.dev.to/assets/subscription-icon-805dfa7ac7dd660f07ed8d654877270825b07a92a03841aa99a1093bd00431b2.png" width="166" height="102"&gt;&lt;/a&gt;
              
              &lt;div id="story-author-preview-content-3971915" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/jess" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F264%2Fb75f6edf-df7b-406e-a56b-43facafb352c.jpg" class="crayons-avatar__image" alt="" width="400" height="400"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Jess Lee&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

            &lt;span&gt;
              &lt;span class="crayons-story__tertiary fw-normal"&gt; for &lt;/span&gt;&lt;a href="/devteam" class="crayons-story__secondary fw-medium"&gt;The DEV Team&lt;/a&gt;
            &lt;/span&gt;
          &lt;/div&gt;
          &lt;a href="https://dev.to/devteam/top-7-featured-dev-posts-of-the-week-55d8" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Jun 23&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/devteam/top-7-featured-dev-posts-of-the-week-55d8" id="article-link-3971915"&gt;
          Top 7 Featured DEV Posts of the Week
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag crayons-tag--filled  " href="/t/discuss"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;discuss&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/top7"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;top7&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/devteam/top-7-featured-dev-posts-of-the-week-55d8" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/raised-hands-74b2099fd66a39f2d7eed9305ee0f4553df0eb7b4f11b01b6b1b499973048fe5.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/exploding-head-daceb38d627e6ae9b730f36a1e390fca556a4289d5a41abb2c35068ad3e2c4b5.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;34&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/devteam/top-7-featured-dev-posts-of-the-week-55d8#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              7&lt;span class="hidden s:inline"&gt;&amp;nbsp;comments&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            2 min read
          &lt;/small&gt;
            
              &lt;span class="bm-initial crayons-icon c-btn__icon"&gt;
                

              &lt;/span&gt;
              &lt;span class="bm-success crayons-icon c-btn__icon"&gt;
                

              &lt;/span&gt;
            
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
      <category>community</category>
      <category>devjournal</category>
      <category>watercooler</category>
      <category>writing</category>
    </item>
    <item>
      <title>Building One Knowledge Graph Across 46 Repositories With Static Analysis</title>
      <dc:creator>Ryosuke Tsuji</dc:creator>
      <pubDate>Mon, 22 Jun 2026 23:54:01 +0000</pubDate>
      <link>https://dev.to/ryantsuji/building-one-knowledge-graph-across-46-repositories-with-static-analysis-part-1-egm</link>
      <guid>https://dev.to/ryantsuji/building-one-knowledge-graph-across-46-repositories-with-static-analysis-part-1-egm</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;AI assistance disclosure: This article was drafted with the help of Claude. All technical content, design decisions, code references, and screenshots reflect production systems I designed and operate at airCloset; the prose was revised by me prior to publication.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hi, I'm &lt;a href="https://x.com/ryantsuji" rel="noopener noreferrer"&gt;Ryan&lt;/a&gt;, CTO at airCloset.&lt;/p&gt;

&lt;p&gt;This post is about unifying a production codebase spanning &lt;strong&gt;46 repositories&lt;/strong&gt; across multiple services into one knowledge graph, using static analysis.&lt;/p&gt;

&lt;p&gt;Internally we call it &lt;strong&gt;code-graph&lt;/strong&gt;, and I built it between January and March of this year.&lt;/p&gt;

&lt;p&gt;Three things I want to write down:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why "just letting AI read the code" isn't enough, and why I had to chase down the connections that cross repository boundaries&lt;/li&gt;
&lt;li&gt;How I extracted boundaries across 46 repos and a zoo of frameworks (jQuery / AngularJS / Express / NestJS / TypeORM / Redux Axios ...)&lt;/li&gt;
&lt;li&gt;What 3 months of trial and error solved, and what it didn't&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is Part 1, covering the construction of code-graph itself, the painful parts, and the issues that remained. &lt;a href="https://dev.to/ryantsuji/making-the-context-across-46-repositories-semantically-searchable-for-ai-part-2-51d9"&gt;Part 2&lt;/a&gt; is about &lt;strong&gt;service-product-graph (SPG)&lt;/strong&gt; — a layer I built on top of code-graph to compensate for what static analysis couldn't do alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Was This For?
&lt;/h2&gt;

&lt;p&gt;A long-running production codebase usually looks something like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Multiple services and multiple teams touching it&lt;/li&gt;
&lt;li&gt;Each era's framework still alive and mixed in&lt;/li&gt;
&lt;li&gt;Dependencies via API, DB, and Event are &lt;strong&gt;tangled&lt;/strong&gt; — not clean 1:1 front-to-back relationships:

&lt;ul&gt;
&lt;li&gt;The same API gets called from multiple repositories (= n:1 callers)&lt;/li&gt;
&lt;li&gt;The same DB table is written to and read from across multiple repositories (= n:n)&lt;/li&gt;
&lt;li&gt;For Events, just looking at the emit side doesn't tell you how completely the subscribe side is covered — it's practically untraceable&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The starting point was wanting to ask AI: "show me the blast radius," "tell me what breaks if I change this," — across this entire codebase.&lt;/p&gt;

&lt;p&gt;The naive answer is: &lt;strong&gt;"just hand all 46 repositories worth of code to AI and let it analyze."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;But that doesn't work, for two reasons:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Context window&lt;/strong&gt;: 46 repositories × years of accumulated code is just not a size you can hand to an AI in one shot&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hallucination&lt;/strong&gt;: even if you could, "read everything and extract the relationships" is an inference task. It misses things, it makes mistakes. That's not usable for impact analysis on a production system&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the first idea I landed on was: &lt;strong&gt;build a knowledge graph externally, via static analysis&lt;/strong&gt;. That's the starting point of code-graph.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scale: 46 Repositories
&lt;/h2&gt;

&lt;p&gt;The target splits into two graphs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;air-closet graph (37 repos)&lt;/strong&gt;: a graph that spans multiple services like airCloset, Men's, WMS, and more&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;mall graph (9 repos)&lt;/strong&gt;: airCloset Mall and related&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So &lt;strong&gt;46 repositories&lt;/strong&gt; in total.&lt;/p&gt;

&lt;p&gt;The thing to notice is that this isn't "one service with 37 repos." It's a &lt;strong&gt;collection of multiple services&lt;/strong&gt; that adds up to that scale. Making the dependencies that cross service boundaries visible as cross-repo edges is exactly what the &lt;strong&gt;boundary nodes&lt;/strong&gt; discussion below is about.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Boundary Nodes Matter
&lt;/h2&gt;

&lt;p&gt;This is the heart of the article.&lt;/p&gt;

&lt;p&gt;When you want AI to understand code, getting it to "read what's in front of it, plus what's next to it" is honestly not hard. grep, open the file, hand it to the model — that works fine.&lt;/p&gt;

&lt;p&gt;For a small codebase, that's enough. But at scale, you hit the context window and hallucination problems mentioned above. I suspect most readers can relate.&lt;/p&gt;

&lt;p&gt;One way to improve this is to &lt;strong&gt;statically analyze the codebase, convert it into a knowledge graph, and serve it to AI through MCP&lt;/strong&gt;. That's the approach.&lt;/p&gt;

&lt;p&gt;The first step was static analysis with &lt;strong&gt;tree-sitter&lt;/strong&gt; (an OSS library that parses source code into syntax trees — it supports a lot of languages and is what VS Code and similar editors use for syntax highlighting; I genuinely recommend it if you want to build something in this space). It's a great tool, but on its own it doesn't solve everything.&lt;/p&gt;

&lt;p&gt;What it doesn't solve is &lt;strong&gt;tracing relationships that cross boundaries — APIs, databases, and so on&lt;/strong&gt;. tree-sitter can extract the relationships between variables, functions, and other in-language constructs. But it can't extract those boundaries.&lt;/p&gt;

&lt;p&gt;The thing that humans and AI alike get stuck on, in practice, is exactly that — code connections that cross boundaries:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The same API is being called from another repo you weren't looking at&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;The frontend in repo A and the nightly batch in repo C might both hit &lt;code&gt;/api/v1/users/me&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Looking at just one of the repos, AI has no way of knowing&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The same DB table is being read or written by some batch process you don't know about&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;When you're modifying service-side code, some batch in a different location might be reading and writing the same table&lt;/li&gt;
&lt;li&gt;Misjudge the blast radius and you get data inconsistency&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The subscribers for this event might not be fully accounted for&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;With distributed pub/sub, looking only at the emit side doesn't let you cover the subscribe side&lt;/li&gt;
&lt;li&gt;Something runs somewhere you don't know about&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In short: getting AI to understand &lt;strong&gt;the code on the other side of a boundary&lt;/strong&gt;, without hallucinating. That's the goal.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr9ptjgc8b4o04tahj6qr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr9ptjgc8b4o04tahj6qr.png" alt="Boundary nodes bridge code across repository walls" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you have boundary nodes, AI can answer "this API is also called from repo X" &lt;strong&gt;as a fact&lt;/strong&gt;. Instead of asking AI to infer, you hand it &lt;strong&gt;a fact that's already been resolved&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Yes, there is inference during the extraction phase — TypeScript Compiler and Gemini both contribute. But the results are persisted as confirmed values in the graph, and a daily boundary-analysis cron (covered below) lets us notice drift the next morning. By the time AI consumes the graph, only verified facts flow to it.&lt;/p&gt;

&lt;p&gt;AI has a tendency to answer "with whatever it can see" rather than saying "I don't know." That's where silent hallucinations creep in — wrong answers that neither AI nor the human catches. Boundary nodes are what physically prevents that. They give AI a verified place to stand.&lt;/p&gt;

&lt;h2&gt;
  
  
  Construction: tree-sitter Base, With TypeScript Compiler and Gemini Where Needed
&lt;/h2&gt;

&lt;p&gt;Normal code structure (function calls, class inheritance, imports) is &lt;strong&gt;relatively straightforward&lt;/strong&gt; to extract with tree-sitter. Walk the AST, turn functions / methods / classes / fields into nodes, connect references with edges. Just grind through it.&lt;/p&gt;

&lt;p&gt;The catch is that while tree-sitter is great at building syntax trees, it's &lt;strong&gt;weak on type information and scope resolution&lt;/strong&gt;. To accurately follow a field access chain like &lt;code&gt;user.preferences.theme&lt;/code&gt;, you need to resolve what type the variable &lt;code&gt;user&lt;/code&gt; is and where it's defined. tree-sitter alone can't reach that.&lt;/p&gt;

&lt;p&gt;So for field-access resolution we use &lt;strong&gt;TypeScript Compiler API&lt;/strong&gt; and &lt;strong&gt;Gemini&lt;/strong&gt; in combination. tree-sitter extracts the structure → TypeScript Compiler resolves variables and types → for the dynamic cases that even that can't reach, Gemini infers. Three stages with distinct responsibilities, which is how we push field-access accuracy up.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkfhv76j1mrukazipy94i.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkfhv76j1mrukazipy94i.png" alt="3-stage field access resolution: tree-sitter, TypeScript Compiler, Gemini" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We define 21 edge types:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;CALLS&lt;/code&gt; (function call) / &lt;code&gt;EXTENDS&lt;/code&gt; (inheritance) / &lt;code&gt;IMPLEMENTS&lt;/code&gt; (interface implementation), etc. — the basic structure tree-sitter can give us&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;CALLS_API&lt;/code&gt; (caller) / &lt;code&gt;HANDLES_API&lt;/code&gt; (handler) — API boundary&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;EMITS_TO&lt;/code&gt; (emitter) / &lt;code&gt;SUBSCRIBES_TO&lt;/code&gt; (subscriber) — Event boundary&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;WRITES_TO&lt;/code&gt; / &lt;code&gt;READS_FROM&lt;/code&gt; — DB boundary&lt;/li&gt;
&lt;li&gt;and more&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The real battle starts when you try to extract the boundary edges (&lt;code&gt;CALLS_API&lt;/code&gt; / &lt;code&gt;HANDLES_API&lt;/code&gt; / &lt;code&gt;EMITS_TO&lt;/code&gt; / &lt;code&gt;SUBSCRIBES_TO&lt;/code&gt; / &lt;code&gt;WRITES_TO&lt;/code&gt; / &lt;code&gt;READS_FROM&lt;/code&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  Extracting and Joining Boundary Nodes: 3 Months of Trial and Error (Jan–Mar)
&lt;/h2&gt;

&lt;p&gt;Unlike normal code, boundaries (API endpoints, DB tables, Event topics) are &lt;strong&gt;written in wildly different ways&lt;/strong&gt; depending on the framework, language, technical area, library, repository, and the person who wrote it.&lt;/p&gt;

&lt;p&gt;Take "define an API endpoint": is it Express? NestJS with a &lt;code&gt;@Get()&lt;/code&gt; decorator? A Fastify route? Each one produces a completely different AST shape. And the same repo can contain multiple patterns simultaneously.&lt;/p&gt;

&lt;p&gt;And it's not just extraction that's hard. &lt;strong&gt;Joining the extracted boundaries on the graph&lt;/strong&gt; is its own headache. For the same API path or DB table name, you get:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Casing variation: camelCase / snake_case / PascalCase&lt;/li&gt;
&lt;li&gt;Trailing-slash variation (&lt;code&gt;/users/me&lt;/code&gt; vs. &lt;code&gt;users/me&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;The boundary name itself is a variable (&lt;code&gt;${baseUrl}/users/me&lt;/code&gt;)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;…all mixed together. Normalizing all of that and correctly joining caller to handler, emitter to subscriber, writer to reader was genuinely the painful part.&lt;/p&gt;

&lt;p&gt;And this had to happen across all 46 repositories × the framework zoo.&lt;/p&gt;

&lt;p&gt;Looking back at the actual git history from that period, you see new parsers and detectors being added almost every week, noise filters going in, and concept renames landing. Here are the main commits from January through March, in order (the commit prefix starts as &lt;code&gt;graph-rag&lt;/code&gt; — the stack was originally named after the "knowledge graph + RAG for LLM consumption" framing — and is renamed to &lt;code&gt;code-graph&lt;/code&gt; on February 15; a few late-February commits still carry a short-lived &lt;code&gt;graph&lt;/code&gt; prefix from that transition):&lt;/p&gt;

&lt;h3&gt;
  
  
  January: Starting Out, and Realizing tree-sitter Alone Isn't Enough
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;2026-01-15&lt;/strong&gt; ─ &lt;code&gt;feat(graph-rag): add TypeScript parser with tree-sitter&lt;/code&gt; — the starting commit&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-01-15&lt;/strong&gt; ─ &lt;code&gt;feat(graph-rag): add graph builder with BigQuery storage&lt;/code&gt; — graph data is written to BigQuery&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-01-19&lt;/strong&gt; ─ &lt;code&gt;feat(graph-rag): add TypeScript Compiler-based variable resolution for field extraction&lt;/code&gt; — realized that &lt;strong&gt;tree-sitter alone couldn't resolve variable types&lt;/strong&gt;, brought in the TypeScript Compiler API alongside it&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  February: Framework Diversity, Fighting Noise
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;2026-02-02&lt;/strong&gt; ─ &lt;code&gt;feat(graph-rag): add frontend parser for jQuery/Vanilla JS codebase&lt;/code&gt; — &lt;strong&gt;jQuery / Vanilla JS&lt;/strong&gt; frontend code&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-02-03&lt;/strong&gt; ─ &lt;code&gt;feat(graph-rag): add AngularJS Page detection for frontend BFS&lt;/code&gt; — &lt;strong&gt;AngularJS&lt;/strong&gt; page detection (older framework, still very much running)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-02-15&lt;/strong&gt; ─ &lt;code&gt;refactor(code-graph): consolidate 18 MCP tools into 5 with deep subgraph traversal&lt;/code&gt; — the toolset had ballooned to 18, consolidated to 5 (also the moment the stack was unified under the name &lt;code&gt;code-graph&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-02-18&lt;/strong&gt; ─ &lt;code&gt;fix(code-graph): reduce graph noise by filtering Type nodes, external lib CALLS, and Storybook files&lt;/code&gt; — noise reduction: filter out Type nodes, external library CALLS, Storybook files&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-02-19&lt;/strong&gt; ─ &lt;code&gt;fix(code-graph): extract path aliases from tsconfig paths in addition to make-symlink&lt;/code&gt; + &lt;code&gt;fix(code-graph): resolve @alias path imports for CommonJS symlink patterns&lt;/code&gt; — &lt;strong&gt;the path-alias pain&lt;/strong&gt;: tsconfig paths, make-symlink, and on top of that the CommonJS symlink pattern — three different mechanisms to support&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-02-19&lt;/strong&gt; ─ &lt;code&gt;feat(code-graph): add stop_at=boundary option to trace_connections&lt;/code&gt; — option to stop traversal at boundary nodes (explicit traversal scoping / node-explosion mitigation)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-02-21&lt;/strong&gt; ─ &lt;code&gt;feat(graph): add typeORM JOIN detection, NestJS decorator parsing, Fetcher API detection&lt;/code&gt; — &lt;strong&gt;TypeORM JOINs / NestJS decorators / Fetcher API&lt;/strong&gt; support&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-02-21&lt;/strong&gt; ─ &lt;code&gt;fix(graph): pass fullFileCode to Redux Axios variable resolver for scope-based extraction&lt;/code&gt; — &lt;strong&gt;Redux Axios&lt;/strong&gt; variable resolver fix&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  March: Concept Cleanup and Precision
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;2026-03-08&lt;/strong&gt; ─ &lt;code&gt;refactor(code-graph): rename __external__ to __boundary__&lt;/code&gt; — &lt;strong&gt;concept cleanup&lt;/strong&gt;: standardize on "boundary node" rather than "external resource"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-03-16&lt;/strong&gt; ─ &lt;code&gt;refactor: remove db-dictionary from code-graph stack&lt;/code&gt; — split the DB schema dictionary (the layer that lets you look up table / column definitions) off into its own graph to evolve independently&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-03-24&lt;/strong&gt; ─ &lt;code&gt;fix(code-graph): infer table names from dynamic variable names&lt;/code&gt; — table-name inference from dynamic variable names&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-03-24&lt;/strong&gt; ─ &lt;code&gt;feat(code-graph): add orphan boundary node cleanup script&lt;/code&gt; — cleanup script for orphan boundary nodes&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  What This Timeline Tells You
&lt;/h3&gt;

&lt;p&gt;Every single week there's a new framework or pattern being handled. The work of "extracting boundary nodes" is, fundamentally, &lt;strong&gt;adding parsers for each new way people write the boundary&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Just listing the frameworks / mechanisms that showed up:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;tree-sitter (TypeScript / JavaScript / Go / Dart (Flutter))&lt;/li&gt;
&lt;li&gt;TypeScript Compiler (variable resolution)&lt;/li&gt;
&lt;li&gt;jQuery / Vanilla JS&lt;/li&gt;
&lt;li&gt;AngularJS&lt;/li&gt;
&lt;li&gt;Express / Koa / Fastify&lt;/li&gt;
&lt;li&gt;NestJS (decorator parsing)&lt;/li&gt;
&lt;li&gt;TypeORM (DB JOIN detection)&lt;/li&gt;
&lt;li&gt;Fetcher API&lt;/li&gt;
&lt;li&gt;Redux Axios (variable resolver)&lt;/li&gt;
&lt;li&gt;3 different path-alias schemes (tsconfig paths / make-symlink / CommonJS symlink)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This isn't a "TypeScript / JavaScript / Go / Dart static analysis" story you can wrap up in one sentence. The air-closet codebase is a collection of long-running production systems where every era's framework still coexists. We had to pick up, from the AST, the era-specific meaning of "here's an API endpoint," "here's a DB call," "here's an Event subscription."&lt;/p&gt;

&lt;h3&gt;
  
  
  Why I Was So Particular About Accuracy
&lt;/h3&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;90% is completely unusable.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Take "list every piece of code that calls this API." If you recall only 90% of the callers, then 10% of the relevant code is invisible to AI. When you're using code-graph for blast-radius investigation, &lt;strong&gt;that invisible 10% is what causes the incident&lt;/strong&gt;. That's single-hop recall.&lt;/p&gt;

&lt;p&gt;And it gets worse the further you walk. For multi-hop graph traversal, every hop multiplies in: at 0.9 per hop you get 0.81 at 2 hops, 0.729 at 3, ~0.59 at 5, ~0.35 at 10 — after just a handful of hops you're at less than half. Push it to 0.99 and you get 0.98 at 2 hops, 0.95 at 5, ~0.90 at 10. &lt;strong&gt;Whether the system is usable in practice is decided by that single-digit difference between 90% and 99%&lt;/strong&gt; — and it bites you on both axes: single-hop recall when you're enumerating, multi-hop confidence when you're traversing.&lt;/p&gt;

&lt;p&gt;So every time a new boundary pattern showed up, we'd add a new custom parser, &lt;strong&gt;aiming to keep the boundary connection rate above 99%&lt;/strong&gt;. We can't measure extraction recall directly — there's no ground-truth "every boundary that should exist" denominator — so the indicator we actually measure daily is "what fraction of callers / handlers are correctly connected on the graph" = the connection rate. The next section is about how that's monitored.&lt;/p&gt;

&lt;h2&gt;
  
  
  Boundary Analysis Is Running Today
&lt;/h2&gt;

&lt;p&gt;The code-graph we built is still running daily.&lt;/p&gt;

&lt;p&gt;Concretely, a &lt;strong&gt;boundary-analysis cron&lt;/strong&gt; runs at JST 7:00 every morning. What it does:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;API boundaries&lt;/strong&gt;: match &lt;code&gt;CALLS_API&lt;/code&gt; (caller) with &lt;code&gt;HANDLES_API&lt;/code&gt; (handler), and aggregate cross-repo connection rates&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Event boundaries&lt;/strong&gt;: match &lt;code&gt;EMITS_TO&lt;/code&gt; (emit) with &lt;code&gt;SUBSCRIBES_TO&lt;/code&gt; (subscribe)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DB boundaries&lt;/strong&gt;: aggregate cases where &lt;code&gt;WRITES_TO&lt;/code&gt; and &lt;code&gt;READS_FROM&lt;/code&gt; from &lt;strong&gt;different repositories&lt;/strong&gt; touch the same table (= implicit cross-repo DB dependency)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The day-over-day numbers get compared, and if the connection rate drops by more than 5%, we get a Grafana alert.&lt;/p&gt;

&lt;p&gt;This whole thing only makes sense &lt;strong&gt;because we have boundary nodes to compare against&lt;/strong&gt;. We're monitoring the &lt;strong&gt;quality of the extracted boundaries themselves&lt;/strong&gt; on a daily cadence. The kind of drift the connection rate catches by the next morning: "a parser fell behind a new pattern and a class of boundaries went invisible," "the repository layout changed and path aliases stopped resolving." There are failure modes the connection rate alone can't see — a caller-side parser regression that drops callers entirely will leave the surviving handlers still looking "connected" to whatever callers remain, and the missing ones slip out silently. That's a separate axis we cover with day-over-day absolute node counts per repo / pattern.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Still Doesn't Work
&lt;/h2&gt;

&lt;p&gt;Even after all that, a handful of issues remain that I can't solve at the root.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. No Semantic Search (an Entry-Point Problem)
&lt;/h3&gt;

&lt;p&gt;The search MCP tool only does LIKE-based substring matching.&lt;/p&gt;

&lt;p&gt;If you're in the middle of development and want to follow connections starting from a function you're already looking at, that's fine — you can pull it up by function name or filename directly.&lt;/p&gt;

&lt;p&gt;The problem shows up when you're investigating a production bug or a customer support ticket. You have no idea what filenames or function names are involved at the start. When the input is "the subscription-fee calculation for members seems off," and you want to walk to the related code from there, no natural-language query into the graph means &lt;strong&gt;you can't find the entry point in the first place&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The intent was: "instead of grepping the whole codebase, navigate relevance via graph RAG." What we ended up with is a structure where you have to grep at the entry point and infer your way in.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Node Explosion
&lt;/h3&gt;

&lt;p&gt;If you naively turn the AST into a graph, every builtin function, anonymous function, and internal utility becomes a node. The &lt;code&gt;map&lt;/code&gt; call you don't care about, the internal helper you don't care about — they're all nodes.&lt;/p&gt;

&lt;p&gt;Trigger a traversal starting from one node, and within a few hops you're dragging in helpers, types, and primitives until the node count explodes. There's no axis built into the graph structure for "filter by relevance."&lt;/p&gt;

&lt;p&gt;We work around it with explicit controls like stopping traversal at boundary nodes, but that's a workaround, not a root fix.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. To Know What a Function Actually Does, You Still Have to Read the File
&lt;/h3&gt;

&lt;p&gt;The graph tells you "something is here," "this calls out to another repo." But what the function actually &lt;strong&gt;does&lt;/strong&gt; still requires opening the file.&lt;/p&gt;

&lt;p&gt;That makes the graph slow on its own. The codebase-investigation tool we built later uses the graph to narrow down candidate files and then hands those to Git Server MCP to actually read — but the underlying graph-only resolution limit doesn't go away.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Operational Cost of Adding Parsers for Every New Boundary Pattern
&lt;/h3&gt;

&lt;p&gt;Every time a new framework or library enters the codebase, we have to learn "how do they write boundaries in this thing" and add a new parser.&lt;/p&gt;

&lt;p&gt;The parser directory already has 10+ custom detectors / extractors. There's no sign of the maintenance and extension cost going down — &lt;strong&gt;every time a new tech stack enters the codebase, the same work repeats&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Side Note: A Different Call Elsewhere — cortex
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;: "cortex" in this section is the internal codename for an AI platform I've been building in-house at airCloset. Unrelated to existing commercial products like Snowflake Cortex or Palo Alto Networks Cortex.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Setting code-graph aside for a moment: I also have a separate project — &lt;strong&gt;cortex&lt;/strong&gt; — where I'm building an in-house AI platform from scratch (currently a single monorepo with 100+ apps).&lt;/p&gt;

&lt;p&gt;On that project I did initially try the same approach as code-graph, but bailed out early and went with an &lt;strong&gt;annotation-based knowledge graph&lt;/strong&gt; instead:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It's a monorepo I'm assembling myself, so I can realistically annotate everything at once&lt;/li&gt;
&lt;li&gt;Use JSDoc tags to write intent directly into the code, and build the graph from that&lt;/li&gt;
&lt;li&gt;Vectorize that intent and store it on the node, so semantic search works&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The decision to "write intent into the code and graph it" — and the trial and error that led to it — I covered in detail in a separate series. If interested: &lt;a href="https://dev.to/ryantsuji/the-heart-of-the-ai-harness-a-knowledge-graph-of-the-ai-by-the-ai-for-the-ai-series-part-2-53bm"&gt;AI Harness Series, Part 2 (The Knowledge Graph at the Heart of cortex)&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Annotation-Based Won't Work for Production Systems
&lt;/h2&gt;

&lt;p&gt;And no, you can't take the same approach for the production-side codebase that code-graph deals with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Annotating all 46 repos at once isn't realistic&lt;/li&gt;
&lt;li&gt;Long-running production systems, touched by multiple teams, with mixed frameworks&lt;/li&gt;
&lt;li&gt;The precondition "put annotations into the code" doesn't hold&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the choice was: &lt;strong&gt;keep code-graph (static analysis) as the base, and evolve by layering on additional graph layers&lt;/strong&gt; to compensate.&lt;/p&gt;

&lt;p&gt;How we're trying to solve the issues above, I cover in &lt;a href="https://dev.to/ryantsuji/making-the-context-across-46-repositories-semantically-searchable-for-ai-part-2-51d9"&gt;Part 2&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  To Be Continued
&lt;/h2&gt;

&lt;p&gt;That's it for Part 1. &lt;a href="https://dev.to/ryantsuji/making-the-context-across-46-repositories-semantically-searchable-for-ai-part-2-51d9"&gt;Part 2&lt;/a&gt; covers how we try to get past the issues above.&lt;/p&gt;

&lt;p&gt;The real story is less "thrown away" and more "&lt;strong&gt;evolved&lt;/strong&gt;."&lt;/p&gt;

&lt;p&gt;Thanks for reading this far.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>knowledgegraph</category>
      <category>staticanalysis</category>
      <category>typescript</category>
    </item>
    <item>
      <title>AI Isn't Something to Trust — It's Something to Design</title>
      <dc:creator>Ryosuke Tsuji</dc:creator>
      <pubDate>Tue, 16 Jun 2026 00:02:03 +0000</pubDate>
      <link>https://dev.to/ryantsuji/ai-isnt-something-to-trust-its-something-to-design-series-final-30aa</link>
      <guid>https://dev.to/ryantsuji/ai-isnt-something-to-trust-its-something-to-design-series-final-30aa</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;AI assistance disclosure: This article was drafted with the help of Claude. All technical content, design decisions, code references, and screenshots reflect production systems I designed and operate at airCloset; the prose was revised by me prior to publication.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hi, I'm &lt;a href="https://x.com/ryantsuji" rel="noopener noreferrer"&gt;Ryan&lt;/a&gt;, CTO at airCloset.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Disclaimer&lt;/strong&gt;: "cortex" in this article is the internal codename for an AI platform built in-house at airCloset. It is unrelated to existing commercial services like Snowflake Cortex or Palo Alto Networks Cortex.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Across the five posts of this series I've worked through how cortex's harness is put together, one piece at a time: the overall picture, the knowledge graph, Auto Review, Self-Healing + Recurrence Prevention, and non-engineer PRs. Having walked through all of them, I want to step one level down for the wrap-up. &lt;strong&gt;Why am I building this thing in the first place?&lt;/strong&gt; That's what this post is about.&lt;/p&gt;

&lt;p&gt;The five posts might look independent, but the root is one thing, and the series doesn't close cleanly without that one thing being put into words. Together with the philosophy, I want to look back at the failures that don't show up when you only write about what worked — what I threw away, where I tripped — as a reference point for anyone trying something similar.&lt;/p&gt;

&lt;h2&gt;
  
  
  Series Index
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;#&lt;/th&gt;
&lt;th&gt;Theme&lt;/th&gt;
&lt;th&gt;Key scene&lt;/th&gt;
&lt;th&gt;Article&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Series intro: cortex's harness&lt;/td&gt;
&lt;td&gt;PRs auto-merge / incidents self-heal before you notice&lt;/td&gt;
&lt;td&gt;&lt;a href="https://dev.to/ryantsuji/building-a-real-ai-harness-auto-reviewed-prs-self-healing-ops-and-non-engineer-contributors-3lfa"&gt;ai-harness-intro&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Product Graph (cpg)&lt;/td&gt;
&lt;td&gt;Code, docs, DB, infra unified into one graph&lt;/td&gt;
&lt;td&gt;&lt;a href="https://dev.to/ryantsuji/the-heart-of-the-ai-harness-a-knowledge-graph-of-the-ai-by-the-ai-for-the-ai-series-part-2-53bm"&gt;cortex-product-graph&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;AI PR review&lt;/td&gt;
&lt;td&gt;webhook → AI review → auto-fix → squash merge&lt;/td&gt;
&lt;td&gt;&lt;a href="https://dev.to/ryantsuji/human-on-the-loop-ai-reviewing-ai-prs-at-cortex-769-prsmonth-while-raising-the-quality-bar-4lh5"&gt;cortex-auto-review&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Self-Healing + observability + auto-added guardrails&lt;/td&gt;
&lt;td&gt;Alert → AI investigates → fix PR + new lint/type gate → auto redeploy&lt;/td&gt;
&lt;td&gt;&lt;a href="https://dev.to/ryantsuji/fixed-before-anyone-notices-stronger-after-every-fix-self-healing-recurrence-prevention-series-1e86"&gt;cortex-self-healing&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Democratizing the maintenance phase&lt;/td&gt;
&lt;td&gt;Domain experts open PRs to production; the harness owns the quality gate&lt;/td&gt;
&lt;td&gt;&lt;a href="https://dev.to/ryantsuji/the-author-doesnt-have-to-be-an-engineer-how-the-harness-holds-quality-series-part-5-12e4"&gt;cortex-non-engineer-prs&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Series wrap-up&lt;/td&gt;
&lt;td&gt;The underlying philosophy plus a retrospective on the failures and lessons&lt;/td&gt;
&lt;td&gt;This post ← you are here&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Origin — What I Was Thinking About in 2025
&lt;/h2&gt;

&lt;p&gt;When I started building cortex, there was one question I wanted to answer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How do I get AI to understand the system accurately?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If AI could understand the system accurately, then PR review, bug investigation, and fixes could all be delegated, and even non-engineers could open up their own development. Conversely, as long as I was stuck on "understand it accurately," everything downstream was sitting on unstable ground. So I spent a lot of time on &lt;strong&gt;the prerequisite layer&lt;/strong&gt; before any of the individual mechanisms.&lt;/p&gt;

&lt;p&gt;The two obvious approaches both hit walls.&lt;/p&gt;

&lt;h3&gt;
  
  
  Wall 1: The Context Window Limit
&lt;/h3&gt;

&lt;p&gt;The first reflex is "just give it all the information it might need." Stuff the codebase, docs, DB schema, infra definitions all into the prompt, and AI gets the whole picture.&lt;/p&gt;

&lt;p&gt;That fails on size. Codebase + docs + schemas + infra at our company doesn't come close to fitting into any realistic context window.&lt;/p&gt;

&lt;p&gt;"Surely context windows will keep growing, and this'll work eventually?" — the more I thought about it, the less of a future I saw in that direction.&lt;/p&gt;

&lt;p&gt;Even with a model whose context window is very large like Gemini, behavior gets unstable when you push it close to the limit. Middle information gets dropped, irrelevant tokens skew the conclusion sideways. This isn't a model-selection problem; it's a structural attention problem. The more unrelated tokens you mix in, the more the attention ratio toward relevant tokens drops mechanically. This is the documented &lt;strong&gt;"lost in the middle"&lt;/strong&gt; phenomenon (information placed at the start and end of long inputs gets used; &lt;strong&gt;information placed in the middle is effectively ignored&lt;/strong&gt;), and &lt;strong&gt;stuff the context window full and you routinely end up in a state where the information you thought you handed over isn't actually visible to the model&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;"Lost in the middle" itself may get mitigated as long-context models improve, so I treat it as empirical supporting evidence rather than the core argument. The real wall is the &lt;strong&gt;recursive&lt;/strong&gt; one beneath it: even if "size" is solved, you immediately need &lt;strong&gt;a higher-level context to judge which tokens are necessary and which aren't&lt;/strong&gt;. That problem is recursive and &lt;strong&gt;can't be resolved by context window size, in principle&lt;/strong&gt;. Information has to be structured, or AI doesn't make correct judgments. That's true of humans too — but humans are a notch better off, because &lt;strong&gt;LLMs don't notice they don't know, and they answer with confidence anyway&lt;/strong&gt;. Silently wrong is worse than visibly stuck.&lt;/p&gt;

&lt;p&gt;The keep-growing-context-windows path didn't have a real resolution in sight.&lt;/p&gt;

&lt;h3&gt;
  
  
  Wall 2: Don't Lean on Learning Either
&lt;/h3&gt;

&lt;p&gt;The other obvious move is to make AI itself learn. Fine-tune per organization, teach it our codebase, our docs, our business. I considered it. Currently not doing it.&lt;/p&gt;

&lt;p&gt;Two reasons. One: getting learning into actual production was still research-phase (in 2025 then; still in 2026 as I write this) and the road to real deployment is still long. The other is thornier: &lt;strong&gt;even if you could learn it, "forgetting" is extremely hard&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A business system has to reflect "the current truth." When the design changes, the DB schema changes, the business rules change, &lt;strong&gt;you want to actively erase old knowledge&lt;/strong&gt;. But "delete just this piece of what's baked into the LLM weights" is unsolved at the research level — there's even a field name for it, &lt;strong&gt;machine unlearning&lt;/strong&gt;, which tells you how hard it is. And on top of that, teaching the model new things also &lt;strong&gt;destroys unrelated existing knowledge&lt;/strong&gt; (called &lt;strong&gt;destructive interference&lt;/strong&gt; / catastrophic forgetting). Lean on learning and both hit at once: the cost of keeping things consistent explodes.&lt;/p&gt;

&lt;p&gt;Rather than treating "doesn't learn" as a downside, I came around to: &lt;strong&gt;because it doesn't learn, swapping out the external knowledge is enough to reflect the current state, and the consistency story is much simpler&lt;/strong&gt;. That was the call at the time.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Way Out — GraphRAG + MCP
&lt;/h3&gt;

&lt;p&gt;With no future in the context-window direction or the learning direction, I came across the &lt;strong&gt;GraphRAG&lt;/strong&gt; concept.&lt;/p&gt;

&lt;p&gt;GraphRAG itself is widely discussed elsewhere; for me, what it meant was the framing: "&lt;strong&gt;supply only the context that's needed, at the moment it's needed&lt;/strong&gt;." Combined with &lt;strong&gt;MCP&lt;/strong&gt; (Anthropic's protocol for connecting LLMs to external tools), AI can go fetch what it needs on its own.&lt;/p&gt;

&lt;p&gt;What was decisive was that this structure lets AI traverse the graph agentically. Rather than "read everything and find related parts by inference," AI &lt;strong&gt;gets to the node it needs and pulls the fact out&lt;/strong&gt;. Which leads to:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Instead of making AI infer, supply facts as context.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That one sentence became the core of cortex's entire design philosophy.&lt;/p&gt;

&lt;p&gt;The first thing I built was a static-analysis-based &lt;strong&gt;code-graph&lt;/strong&gt;, which I then threw away after trial and error, and arrived at the annotation-based &lt;strong&gt;product-graph (cpg)&lt;/strong&gt; — details in the trial-and-error section.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F416pzfhy6z28i4wa2g72.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F416pzfhy6z28i4wa2g72.png" alt="2025 origin. Neither growing the context window nor relying on learning had a future; GraphRAG + MCP became the way through." width="800" height="490"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  I Don't Trust AI to Begin With
&lt;/h2&gt;

&lt;p&gt;The origin section in one line:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;I don't trust AI to fill in the blanks for me.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;"Don't trust" here is not the same as "have no faith in." This isn't doubting Claude / GPT / Gemini's generation quality. What I mean is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;It doesn't know context it wasn't handed.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;It doesn't, on its own and without being told, produce the ideal state.&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The first one is a truth no amount of model progress will change. Architecturally, LLMs can't know things that weren't in the training data and aren't in this session's context. "Surely smarter models will pick up on it" — I don't think that future is coming. Smarter is a real direction; smart alone doesn't compensate for not knowing.&lt;/p&gt;

&lt;p&gt;The second one is about responsibility, and humans owning it. AI can't decide on its own what "ideal" means. When it tries, it lands on a generic best-practice answer slightly off from the actual situation. Ideal depends on the business, the organization, the moment in time — none of which is visible to AI unless a human verbalizes it and hands it over.&lt;/p&gt;

&lt;p&gt;So that conviction is &lt;strong&gt;not underestimating AI's capability; it's a design decision to not let AI auto-complete the prerequisites&lt;/strong&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Mastering AI is not about giving it freedom — it's about confining its output to a predictable range.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The mechanism for confining it is the harness this series has been describing.&lt;/p&gt;

&lt;h2&gt;
  
  
  So I Build Harnesses to Hold AI to Determinism
&lt;/h2&gt;

&lt;p&gt;Reading each post through the lens of "don't make AI infer; lean on determinism" surfaces that the five of them are all the same conviction showing up in different layers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Part 2 — Knowledge Graph&lt;/strong&gt;: Instead of making AI search the codebase, this mechanism tilts toward making the codebase legible. With &lt;code&gt;@graph-*&lt;/code&gt; annotations, code / docs / DB / infra are unified into one graph, so AI doesn't have to grep + infer to find related parts. This is the direct implementation of "supply facts as context" from the origin section. → &lt;a href="https://dev.to/ryantsuji/the-heart-of-the-ai-harness-a-knowledge-graph-of-the-ai-by-the-ai-for-the-ai-series-part-2-53bm"&gt;cortex-product-graph&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Part 3 — Auto Review Dimensions&lt;/strong&gt;: Nine review dimensions (responsibility / severity / type SSoT / etc.) are fixed in advance. When AI does the review, what to check isn't something it gets to infer. "Looking at the PR as a whole" gives AI too much room for inference, so dimensions are split and &lt;strong&gt;each is judged as its own question&lt;/strong&gt;. &lt;strong&gt;Dimensions = locked by the harness, evaluation = AI's job.&lt;/strong&gt; → &lt;a href="https://dev.to/ryantsuji/human-on-the-loop-ai-reviewing-ai-prs-at-cortex-769-prsmonth-while-raising-the-quality-bar-4lh5"&gt;cortex-auto-review&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Part 4 — Self-Healing + Recurrence Prevention&lt;/strong&gt;: Alert → investigation → fix PR → redeploy. The flow itself is fixed. AI doesn't get to think through "how should we respond to incidents" each time. And Recurrence Prevention — adding lint / CI gates so the same trap can't be stepped on twice — is &lt;strong&gt;mechanical refusal at the gate, not trust-AI-not-to-do-it-again&lt;/strong&gt;. Or put differently: I don't expect AI never to repeat a mistake. → &lt;a href="https://dev.to/ryantsuji/fixed-before-anyone-notices-stronger-after-every-fix-self-healing-recurrence-prevention-series-1e86"&gt;cortex-self-healing&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Part 5 — Non-Engineer PRs&lt;/strong&gt;: If the harness weren't holding quality, business-side folks opening PRs directly to production wouldn't survive a single day. Conversely, with the three mechanisms above stacked up (context locked, dimensions locked, traps locked out mechanically), the person closest to the requirements can ship the change directly. The translation layer and the engineering priority queue disappeared as a downstream consequence of the determinism push. → &lt;a href="https://dev.to/ryantsuji/the-author-doesnt-have-to-be-an-engineer-how-the-harness-holds-quality-series-part-5-12e4"&gt;cortex-non-engineer-prs&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So what's covered across the five posts is "don't make AI infer; lean on determinism" implemented at different layers. The root is one conviction.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "Don't Make AI Infer, Lean on Determinism" Actually Means
&lt;/h2&gt;

&lt;p&gt;Let me sharpen this phrase that's come up a few times.&lt;/p&gt;

&lt;p&gt;"Lean on determinism" does &lt;strong&gt;not&lt;/strong&gt; mean "give AI zero room to infer." Code generation, judging review findings, hypothesizing root causes from error logs — these are domains where AI not inferring is the same as no work getting done.&lt;/p&gt;

&lt;p&gt;Where I want to lean on determinism is in domains where &lt;strong&gt;variance isn't allowed&lt;/strong&gt;. Specifically:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Which part of the codebase to look at&lt;/strong&gt; — don't have AI guess by analogy; pull it deterministically from the knowledge graph&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Which review dimensions to apply&lt;/strong&gt; — don't let AI pick "the important-looking dimensions"; lock the dimension list in advance&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;How to respond to incidents&lt;/strong&gt; — don't make AI think through the workflow each time; fix the alert → fix PR path&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not stepping on the same trap twice&lt;/strong&gt; — don't ask AI to "try to be careful"; let lint / CI mechanically refuse it&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What implements this line — where inference is allowed vs. where it isn't — is the harness. To borrow the metaphor from Part 5, the harness lays down &lt;strong&gt;rails you can't fall off&lt;/strong&gt;. On top of the rails, AI runs free (inference works as inference); but it can't fall off the rails sideways.&lt;/p&gt;

&lt;p&gt;Put differently, this is equivalent to the framing in &lt;a href="https://dev.to/ryantsuji/the-heart-of-the-ai-harness-a-knowledge-graph-of-the-ai-by-the-ai-for-the-ai-series-part-2-53bm"&gt;Part 2 (cortex-product-graph)&lt;/a&gt;: "where hallucination gets confined." Saying "no inference allowed" isn't quite right — the harness isn't a thing that makes hallucination go to zero. It's a thing that confines hallucination to places where hallucination is OK (i.e., the inference-allowed zone). The structure and facts about the codebase are pulled deterministically, so &lt;strong&gt;the retrieval process itself has no opening for hallucination&lt;/strong&gt;; hallucinations on the judgment side get filtered downstream by tests / lint / dimension-by-dimension reviews. The places where hallucination is allowed and the places where it isn't are &lt;strong&gt;physically split by the harness&lt;/strong&gt;. That's the continuation of the Part 2 framing.&lt;/p&gt;

&lt;p&gt;Step back one more level and what the harness is really doing is &lt;strong&gt;shifting when inference happens&lt;/strong&gt;. The annotations and descriptions on the graph were also written by AI originally — there is inference baked into them. But that inference is &lt;strong&gt;write-time&lt;/strong&gt; — happens once, reviewed, then frozen — not &lt;strong&gt;read-time&lt;/strong&gt; (happens every query, &lt;strong&gt;unverified at the point of use&lt;/strong&gt;). The graph is &lt;strong&gt;frozen, reviewed inference&lt;/strong&gt;, which is exactly why the read side can treat it as fact. "Leaning on determinism" can be rephrased as &lt;strong&gt;not letting unverified inference run on every query&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fiffzn4u10c1qk6t6dgze.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fiffzn4u10c1qk6t6dgze.png" alt="Inference-allowed zone (top, green) and inference-forbidden zone (bottom, orange). The harness implements this boundary — i.e., decides where hallucination gets confined." width="800" height="490"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is also the underlying basis for &lt;a href="https://dev.to/ryantsuji/building-a-real-ai-harness-auto-reviewed-prs-self-healing-ops-and-non-engineer-contributors-3lfa"&gt;Part 1 (Series Intro)&lt;/a&gt;'s claim "models commoditize; harnesses differentiate." Model-side quality is converging across Claude / GPT / Gemini, but the harness is &lt;strong&gt;codebase-specific and business-specific&lt;/strong&gt;, so this is where org-level differentiation actually comes from.&lt;/p&gt;

&lt;p&gt;Worth flagging: &lt;strong&gt;the position of this boundary moves with model capability&lt;/strong&gt;. As agentic search and reasoning get stronger, today's "must be deterministic" zone might be tomorrow's "inference is good enough" zone — and in fact cortex itself depends on AI's ability to traverse the graph agentically. But &lt;strong&gt;the boundary itself never disappears&lt;/strong&gt;. How information is structured, where the line gets drawn between fact and inference — &lt;strong&gt;whether you hold that line explicitly as a design decision&lt;/strong&gt; is what differentiates organizations, across every model generation.&lt;/p&gt;

&lt;p&gt;A note: this framing isn't confined to cortex's harness. The same stance shapes &lt;a href="https://dev.to/ryantsuji/democratizing-internal-data-building-an-mcp-server-that-lets-you-search-991-tables-in-natural-1da5"&gt;db-graph MCP&lt;/a&gt;, the natural-language interface over internal DB schemas, and &lt;a href="https://dev.to/ryantsuji/bridging-i-want-to-build-and-i-want-to-publish-safely-for-non-engineers-sandbox-mcp-392a"&gt;Sandbox MCP&lt;/a&gt;, which lets non-engineers safely publish AI-built apps. &lt;strong&gt;It's the through-line in any platform we build that's based on AI doing meaningful work.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;One level more abstract: &lt;strong&gt;the individual features aren't where the value is.&lt;/strong&gt; The value sits in the conviction itself. cortex / db-graph / Sandbox MCP are all that one conviction translated into our own use cases.&lt;/p&gt;

&lt;p&gt;The way I think about "design":&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Design is translating an abstract principle into a concrete implementation that fits your own use cases.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It's not drawing class diagrams, and it's not laying out architecture diagrams — it's the translation work of &lt;strong&gt;"how does this principle take shape under our business / codebase / constraints?"&lt;/strong&gt; That's where each organization's distinctiveness lives, and that's the value that can't be copied.&lt;/p&gt;

&lt;p&gt;Said the other way: another organization copying cortex's surface doesn't reproduce the substance. What gets asked of every org is &lt;strong&gt;how it translates this principle into its own use cases&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trial and Error That Got Me Here
&lt;/h2&gt;

&lt;p&gt;Everything I've written about above is the form that ended up working. Getting to that form involved &lt;strong&gt;a lot of throwing away&lt;/strong&gt;. Two representative examples worth keeping on record, plus one shorter one.&lt;/p&gt;

&lt;h3&gt;
  
  
  I Spent Two Months on Static-Analysis code-graph, Then Threw It Out
&lt;/h3&gt;

&lt;p&gt;The first thing I built was static-analysis-based &lt;strong&gt;code-graph&lt;/strong&gt;: extracting AST data — imports, call graphs, type dependencies — and putting that into a graph DB. At a glance, the obvious implementation of "make AI understand the codebase."&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Why two months?&lt;/strong&gt; code-graph wasn't just cortex; it spanned our consumer-facing services and internal-system repositories too — &lt;strong&gt;over 40 repos in total&lt;/strong&gt; (cortex being one of them). The mechanically-extractable AST data (imports / call graphs / type dependencies) was usable as-is via tree-sitter, but each repo had its own API endpoints / DB schema / event definitions / Pub/Sub topology, and &lt;strong&gt;extracting those boundary nodes (where an app meets the outside) goes beyond mechanical AST analysis and had to be implemented per-repo-type&lt;/strong&gt; — that's where the time went.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So I spent two months out of the first three on this, and got something that worked end-to-end.&lt;/p&gt;

&lt;p&gt;And then I threw it away.&lt;/p&gt;

&lt;p&gt;Why: static analysis is great at capturing &lt;strong&gt;structure&lt;/strong&gt;, but it can't traverse on &lt;strong&gt;intent or business context&lt;/strong&gt;. Concretely, three things broke:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No semantic entry point for search&lt;/strong&gt; — if I want to query the codebase with "show me the function calculating member subscription billing," I can't get there unless I already know the function name or file. A graph built only from static analysis has no semantic-tag entry pointing to "what is this code for?"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The graph contains only code&lt;/strong&gt; — internal helpers / utilities / types / arguments all become nodes, so traversal from any function &lt;strong&gt;blows up within a few hops&lt;/strong&gt;, dragging in helpers and primitives. There's no axis to filter on semantic relatedness&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What I actually wanted was code + DB schema + docs + infra on one graph&lt;/strong&gt; — given a function, I want to pull, in one query, the DB tables it touches, the docs where the design lives, and the linked business requirement. A code-only graph just can't do that&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;→ I switched to the annotation-based approach (&lt;code&gt;@graph-*&lt;/code&gt; JSDoc tags write the business intent into the code, and that gets unified with DB schema / docs / infra into one graph). Searchable semantically, and when you traverse, only related stuff comes back. That's the current &lt;strong&gt;product-graph (cpg)&lt;/strong&gt;. &lt;strong&gt;Don't drag sunk cost forward and you'll get to the final form&lt;/strong&gt; — discarding two months of investment instead of trying to recoup it was the foundation for everything that came after.&lt;/p&gt;

&lt;h3&gt;
  
  
  Setting Coverage 90% as a Solo Target Broke the Implementations
&lt;/h3&gt;

&lt;p&gt;Test coverage is still gated at 90%+ (as covered in &lt;a href="https://dev.to/ryantsuji/human-on-the-loop-ai-reviewing-ai-prs-at-cortex-769-prsmonth-while-raising-the-quality-bar-4lh5"&gt;Part 3&lt;/a&gt;). That part hasn't changed. But there was a period when &lt;strong&gt;Coverage was treated as a standalone target&lt;/strong&gt;, and during that period the implementation visibly got worse.&lt;/p&gt;

&lt;p&gt;Specifically:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Heavy default-value use that hides branches&lt;/strong&gt;: &lt;code&gt;function(input = {})&lt;/code&gt; style writes the missing-input branch out of the test path. Coverage goes up, protection against unexpected input is gone&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Catch-and-swallow over throw&lt;/strong&gt;: try / catch returning &lt;code&gt;null&lt;/code&gt;. Don't throw → no need to test "doesn't throw," and Coverage is satisfied. Invalid state silently propagates&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Early returns that flatten too much&lt;/strong&gt;: dump complex conditions through an "early return" escape. Tests pass; what should have been validation just isn't there anymore&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Result: &lt;strong&gt;Coverage 90%, quality lower than before&lt;/strong&gt;. When you look at Coverage alone, the shortest path to "satisfy it" is &lt;strong&gt;a weaker implementation that passes the tests&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Two lessons:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Set a number as a target, and the number becomes the goal&lt;/strong&gt;. Coverage is a "minimum floor" — not "a goal to hit"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't evaluate any metric alone&lt;/strong&gt;. Coverage has to be evaluated alongside responsibility separation / exception design / boundary value coverage / etc.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then, as the follow-up: I added linting that &lt;strong&gt;mechanically closes off the routes that let you weaken implementations to satisfy Coverage&lt;/strong&gt;. Two specific examples:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;no-silent-catch&lt;/code&gt;&lt;/strong&gt;: AST-level ban on empty catch and silent-handler patterns like &lt;code&gt;.catch(() =&amp;gt; null)&lt;/code&gt;. Catch bodies have to have a &lt;strong&gt;function call (logger included) / re-throw / new / await&lt;/strong&gt; — otherwise it's an error. Catches the "weaken throws to satisfy Coverage but lose observability in production" pattern structurally. The violation message routes you to &lt;code&gt;@cortex/otel/logger&lt;/code&gt; for structured logging, so the chain through Cloud Run OTel → Loki / Grafana stays intact&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;vitest-strong-matchers&lt;/code&gt;&lt;/strong&gt;: bans weak matchers like &lt;code&gt;toBeTruthy&lt;/code&gt; / &lt;code&gt;toBeDefined&lt;/code&gt; / &lt;code&gt;toContain&lt;/code&gt; / &lt;code&gt;toBe(true|false)&lt;/code&gt; / &lt;code&gt;expect.any&lt;/code&gt; / &lt;code&gt;expect.objectContaining&lt;/code&gt;. Catches "any assertion that passes" patterns at the AST level, and points you instead toward &lt;code&gt;toStrictEqual&lt;/code&gt; / &lt;code&gt;toMatchInlineSnapshot&lt;/code&gt; that pin down the full output. This is one notch above Coverage — a &lt;strong&gt;test quality&lt;/strong&gt; concern — but it lines up because the same reflection applies: &lt;strong&gt;don't let a number become the goal&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On top of that, cortex's &lt;a href="https://github.com/air-closet/cortex/blob/main/docs/guidelines/testing.md" rel="noopener noreferrer"&gt;testing guideline&lt;/a&gt; opens with "&lt;strong&gt;Coverage is not the goal, just a supporting indicator&lt;/strong&gt;," and threshold-lowering / &lt;code&gt;istanbul ignore&lt;/code&gt; workarounds get bounced as Critical in Auto Review. So even when Coverage is satisfied, "this is intentionally deleting a branch" / "this is swallowing the exception" comes back as a Major finding.&lt;/p&gt;

&lt;p&gt;From the lesson "a single metric warps implementation," we descended through &lt;strong&gt;guideline that states the principle → lint that mechanically rejects → Auto Review that evaluates as a dimension&lt;/strong&gt; before Coverage 90% finally functioned as the "minimum floor" it should have been all along. This too sits in the lineage of the &lt;strong&gt;Recurrence Prevention&lt;/strong&gt; mechanism from Part 4 (so the same trap can't be stepped on twice).&lt;/p&gt;

&lt;h3&gt;
  
  
  Parallel Sub-Agent → Sequential Evaluation
&lt;/h3&gt;

&lt;p&gt;Third: an internal-structure call about Auto Review. &lt;strong&gt;Distribute the 9 dimensions to parallel sub-agents and evaluate concurrently&lt;/strong&gt; — the plausible-looking design ("parallel = faster, parallel should also hold quality") I tried first and ended up throwing out.&lt;/p&gt;

&lt;p&gt;What actually happened: &lt;strong&gt;time, cost, and accuracy all got worse&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Time got worse&lt;/strong&gt;: each sub-agent has its own startup, its own context load, its own result aggregation overhead. "9-way parallel = 9x faster" didn't hold; there were even cases where sequential evaluation in one session ended up faster&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost got worse&lt;/strong&gt;: each sub-agent loads PR diff + guidelines + related code independently — common context loads ran 9 times. Token consumption measured at &lt;strong&gt;just under 4x — not the naive 9x&lt;/strong&gt; (the context other than diff is shared across many dimensions, which is what kept it from blowing up to a full 9x)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Accuracy didn't hold&lt;/strong&gt;: parallel sub-agents don't see each other's verdicts, so the same problem comes back as "APPROVE" from one and "REQUEST_CHANGES" from another. Duplicate findings show up too. Without a "what kind of PR is this as a whole?" pass to anchor on, dimensional findings drift toward local optima and the overall picture gets worse&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Switching to sequential evaluation: same session goes through 9 dimensions in sequence, so context loads once, and each dimension's call has the previous dimension's verdict in front of it. &lt;strong&gt;All three — time, cost, accuracy — improve simultaneously.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Of course, sequential evaluation introduces &lt;strong&gt;order dependence between dimensions&lt;/strong&gt; — earlier verdicts can shape later ones. That's a real trade-off, and I accepted it knowingly. &lt;strong&gt;Inter-dimension consistency at the cost of some order sensitivity&lt;/strong&gt; is more useful as a 9-dimension review than fully independent dimensions that contradict each other.&lt;/p&gt;

&lt;p&gt;The takeaway: the distributed-systems intuition that "&lt;strong&gt;parallel = faster, parallel = quality holds&lt;/strong&gt;" &lt;strong&gt;breaks its own assumptions in an AI harness&lt;/strong&gt;. Unlike parallelizing across CPU cores on your machine, with AI &lt;strong&gt;the context isn't shared memory; it's per-process state&lt;/strong&gt;. Sequential evaluation in one session ends up better on speed, token efficiency, and inter-dimension consistency at the same time — a structural property that's easy to miss at design time.&lt;/p&gt;

&lt;h3&gt;
  
  
  What This Section Is Really Saying
&lt;/h3&gt;

&lt;p&gt;The form I've described across the series is &lt;strong&gt;the result of a lot of trial and error&lt;/strong&gt;. Not starting with the right answer and laying it out from there. The decisions of throwing things away with sunk costs included, the trap of letting a metric I chose turn into the goal, the distributed setup that looked natural but worked against me — those are the things I walked through before landing at the current shape.&lt;/p&gt;

&lt;p&gt;Not easy. I don't pretend it was. But &lt;strong&gt;if you do walk through it, real results follow&lt;/strong&gt; — that's the honest read on it now.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing the Series
&lt;/h2&gt;

&lt;p&gt;What I most wanted to communicate across these six posts comes down to one thing:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;AI coding is not about "how to use AI" — it's about designing the environment AI runs in.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Or, put another way:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;AI isn't something to trust. It's something to design.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Assuming a large codebase&lt;/strong&gt;: prompt engineering / model selection / tool selection — each matters individually, but polishing them alone doesn't get you to auto-merging PRs, auto-healing incidents, or non-engineer development. Getting there requires building &lt;strong&gt;a codebase / business flow / observability / repair cycle where AI doesn't need to infer&lt;/strong&gt;. That's not an individual AI skill — that's &lt;strong&gt;an environment-design problem&lt;/strong&gt; (conversely, for a small project of a few dozen files, today's AI models work fine standalone. &lt;strong&gt;Harnesses become essential when scale exceeds what one person can hold in their head.&lt;/strong&gt;).&lt;/p&gt;

&lt;p&gt;And the conviction at the root of environment design is, repeating myself, "&lt;strong&gt;I don't trust AI to fill in the blanks for me&lt;/strong&gt;" — looking the reality in the face that context that wasn't handed over isn't known, and the ideal state doesn't happen without being told. Once you accept that premise, &lt;strong&gt;what to build clarifies naturally&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Looking back, four decisions ended up being the ones that mattered:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Locked the conviction first&lt;/strong&gt;: putting words to the root ("AI isn't something to trust") gave priority order to every mechanism. If I'd started from technique, I don't think I'd have made it to the current form&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Invested with throwing-out as the default&lt;/strong&gt;: like I did with code-graph at the two-month mark, I went into things with "throwing this out is OK." Drag sunk cost forward and you can't move forward&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Refused standalone numerical targets&lt;/strong&gt;: the moment a metric like Coverage 90% becomes the goal on its own, implementations warp. Designed the system so it gets evaluated alongside other dimensions&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Designed for "no inference," not around AI's capability&lt;/strong&gt;: I prioritized building structure where AI doesn't have to infer, instead of relying on what AI can do. That's what made the system stable end-to-end, I think&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If even one of these is useful to someone starting on something similar, that would be great.&lt;/p&gt;




&lt;h2&gt;
  
  
  Afterword — Where Engineering Careers Are Heading
&lt;/h2&gt;

&lt;p&gt;Slipping off the wrap-up topic — this is something I've been turning over recently, written here in a "loosely held thought" tone, so feel free to skim.&lt;/p&gt;

&lt;p&gt;As harnesses mature, I think &lt;strong&gt;engineering work splits along two directions&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;One direction is &lt;strong&gt;value creation from problem identification and business design&lt;/strong&gt;. In the world of Part 5 — where non-engineer PRs work — "writing code" stops being scarce, and the actual scarce thing becomes &lt;strong&gt;the ability to define what to build&lt;/strong&gt;. The person closest to the requirements (a PMO, a business manager, a domain-deep engineer) ends up driving Claude Code through to the merged PR themselves. This direction looks less like "engineer" and more like a &lt;strong&gt;business designer who moves between domain and implementation&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The other direction is &lt;strong&gt;building the foundation that lets all of that happen safely and quickly&lt;/strong&gt;. Non-engineers can open PRs to the production repo only because the harness underneath holds quality — knowledge graph, Auto Review, Self-Healing, Recurrence Prevention, lint, CI, tests, observability stack, all interlocked. Designing / maintaining / evolving that gets &lt;em&gt;harder&lt;/em&gt;, not easier. As the &lt;strong&gt;house-builder side, rail-layer side&lt;/strong&gt;, this demands deep infra understanding / security instinct / observability design / a feel for AI's architectural quirks.&lt;/p&gt;

&lt;p&gt;I'm building cortex, so I'm spending more time on the latter; building "a foundation where the business can run its own changes" is genuinely fun for me. &lt;strong&gt;That said, I'm not the type who fully commits to one side&lt;/strong&gt; — I move between listening to business questions and assembling the foundation, and the satisfaction from each is its own kind. This isn't a "which is more important?" question — the harness exists precisely so the former is possible, and the former being alive is what gives the latter meaning. They're mutually dependent.&lt;/p&gt;

&lt;p&gt;Maybe the era of polishing &lt;strong&gt;just&lt;/strong&gt; "coding ability" is shifting slightly. Where to put your value — or whether to move between both directions — becomes a question more engineers will need to choose into intentionally.&lt;/p&gt;




&lt;p&gt;Six posts in, thanks for sticking with me to the end.&lt;/p&gt;




&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;#&lt;/th&gt;
&lt;th&gt;Theme&lt;/th&gt;
&lt;th&gt;Key scene&lt;/th&gt;
&lt;th&gt;Article&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Series intro: cortex's harness&lt;/td&gt;
&lt;td&gt;PRs auto-merge / incidents self-heal before you notice&lt;/td&gt;
&lt;td&gt;&lt;a href="https://dev.to/ryantsuji/building-a-real-ai-harness-auto-reviewed-prs-self-healing-ops-and-non-engineer-contributors-3lfa"&gt;ai-harness-intro&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Product Graph (cpg)&lt;/td&gt;
&lt;td&gt;Code, docs, DB, infra unified into one graph&lt;/td&gt;
&lt;td&gt;&lt;a href="https://dev.to/ryantsuji/the-heart-of-the-ai-harness-a-knowledge-graph-of-the-ai-by-the-ai-for-the-ai-series-part-2-53bm"&gt;cortex-product-graph&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;AI PR review&lt;/td&gt;
&lt;td&gt;webhook → AI review → auto-fix → squash merge&lt;/td&gt;
&lt;td&gt;&lt;a href="https://dev.to/ryantsuji/human-on-the-loop-ai-reviewing-ai-prs-at-cortex-769-prsmonth-while-raising-the-quality-bar-4lh5"&gt;cortex-auto-review&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Self-Healing + observability + auto-added guardrails&lt;/td&gt;
&lt;td&gt;Alert → AI investigates → fix PR + new lint/type gate → auto redeploy&lt;/td&gt;
&lt;td&gt;&lt;a href="https://dev.to/ryantsuji/fixed-before-anyone-notices-stronger-after-every-fix-self-healing-recurrence-prevention-series-1e86"&gt;cortex-self-healing&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Democratizing the maintenance phase&lt;/td&gt;
&lt;td&gt;Domain experts open PRs to production; the harness owns the quality gate&lt;/td&gt;
&lt;td&gt;&lt;a href="https://dev.to/ryantsuji/the-author-doesnt-have-to-be-an-engineer-how-the-harness-holds-quality-series-part-5-12e4"&gt;cortex-non-engineer-prs&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Series Final&lt;/td&gt;
&lt;td&gt;The underlying philosophy plus a retrospective on the failures and lessons&lt;/td&gt;
&lt;td&gt;This post&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>devops</category>
      <category>engineering</category>
    </item>
    <item>
      <title>Part 5 ("The Author Doesn't Have to Be an Engineer") has been generating sharp comments.
Worth a read for the thread alone.</title>
      <dc:creator>Ryosuke Tsuji</dc:creator>
      <pubDate>Thu, 11 Jun 2026 01:06:48 +0000</pubDate>
      <link>https://dev.to/ryantsuji/part-5-the-author-doesnt-have-to-be-an-engineer-has-been-generating-sharp-comments-worth-a-1pao</link>
      <guid>https://dev.to/ryantsuji/part-5-the-author-doesnt-have-to-be-an-engineer-has-been-generating-sharp-comments-worth-a-1pao</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/ryantsuji/the-author-doesnt-have-to-be-an-engineer-how-the-harness-holds-quality-series-part-5-12e4" class="crayons-story__hidden-navigation-link"&gt;The Author Doesn't Have to Be an Engineer: How the Harness Holds Quality&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
      &lt;a href="https://dev.to/ryantsuji/the-author-doesnt-have-to-be-an-engineer-how-the-harness-holds-quality-series-part-5-12e4" class="crayons-article__context-note crayons-article__context-note__feed"&gt;&lt;p&gt;Self-healing guardrails for business-side PRs&lt;/p&gt;

&lt;/a&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/ryantsuji" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3843591%2F8b126f91-f561-4e6b-8492-814b18d680ec.jpg" alt="ryantsuji profile" class="crayons-avatar__image"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/ryantsuji" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Ryosuke Tsuji
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Ryosuke Tsuji
                
              
              &lt;div id="story-author-preview-content-3849367" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/ryantsuji" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3843591%2F8b126f91-f561-4e6b-8492-814b18d680ec.jpg" class="crayons-avatar__image" alt=""&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Ryosuke Tsuji&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/ryantsuji/the-author-doesnt-have-to-be-an-engineer-how-the-harness-holds-quality-series-part-5-12e4" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Jun 8&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/ryantsuji/the-author-doesnt-have-to-be-an-engineer-how-the-harness-holds-quality-series-part-5-12e4" id="article-link-3849367"&gt;
          The Author Doesn't Have to Be an Engineer: How the Harness Holds Quality
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/devops"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;devops&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/engineering"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;engineering&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/github"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;github&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/ryantsuji/the-author-doesnt-have-to-be-an-engineer-how-the-harness-holds-quality-series-part-5-12e4" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/exploding-head-daceb38d627e6ae9b730f36a1e390fca556a4289d5a41abb2c35068ad3e2c4b5.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/multi-unicorn-b44d6f8c23cdd00964192bedc38af3e82463978aa611b4365bd33a0f1f4f3e97.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;19&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/ryantsuji/the-author-doesnt-have-to-be-an-engineer-how-the-harness-holds-quality-series-part-5-12e4#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              32&lt;span class="hidden s:inline"&gt;&amp;nbsp;comments&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            16 min read
          &lt;/small&gt;
            
              &lt;span class="bm-initial crayons-icon c-btn__icon"&gt;
                

              &lt;/span&gt;
              &lt;span class="bm-success crayons-icon c-btn__icon"&gt;
                

              &lt;/span&gt;
            
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
      <category>career</category>
      <category>discuss</category>
      <category>softwareengineering</category>
      <category>testing</category>
    </item>
    <item>
      <title>The Author Doesn't Have to Be an Engineer: How the Harness Holds Quality</title>
      <dc:creator>Ryosuke Tsuji</dc:creator>
      <pubDate>Mon, 08 Jun 2026 23:32:30 +0000</pubDate>
      <link>https://dev.to/ryantsuji/the-author-doesnt-have-to-be-an-engineer-how-the-harness-holds-quality-series-part-5-12e4</link>
      <guid>https://dev.to/ryantsuji/the-author-doesnt-have-to-be-an-engineer-how-the-harness-holds-quality-series-part-5-12e4</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;AI assistance disclosure: This article was drafted with the help of Claude. All technical content, design decisions, code references, and screenshots reflect production systems I designed and operate at airCloset; the prose was revised by me prior to publication.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hi, I'm &lt;a href="https://x.com/ryantsuji" rel="noopener noreferrer"&gt;Ryan&lt;/a&gt;, CTO at airCloset.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Disclaimer&lt;/strong&gt;: "cortex" in this article is the internal codename for an AI platform built in-house at airCloset. It is unrelated to existing commercial services like Snowflake Cortex or Palo Alto Networks Cortex.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;In &lt;a href="https://dev.to/ryantsuji/building-a-real-ai-harness-auto-reviewed-prs-self-healing-ops-and-non-engineer-contributors-3lfa"&gt;Part 1 (Series Intro)&lt;/a&gt;, I wrote about how cortex's harness has matured to the point where &lt;strong&gt;non-engineers (business-side managers, PMOs, and the like) can open PRs to the production repository&lt;/strong&gt;. The harness here is the runtime foundation for AI in production -- the combination of the knowledge graph, Auto Review, Self-Healing, and Recurrence Prevention covered across Parts 1 through 4.&lt;/p&gt;

&lt;p&gt;Part 5 is what comes next: &lt;strong&gt;that harness has now reached the layer of who actually writes the code&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;"Surely an engineer is still checking afterward, right?" -- I expect a lot of readers will land here with that question. So this post leads with &lt;strong&gt;one concrete example&lt;/strong&gt; before anything else.&lt;/p&gt;

&lt;p&gt;Part 5 covers:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;What kinds of PRs are actually shipping&lt;/strong&gt; -- two recent ones in detail&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What works and what doesn't&lt;/strong&gt; -- the boundary between adding on top of an existing stack and standing up new infrastructure&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Why this holds for non-engineers&lt;/strong&gt; -- how the four mechanisms from Parts 1-4 carry it&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What's next -- into toC services&lt;/strong&gt; -- the direction of travel for consumer-facing scale&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The deeper toC implementation story will live in a separate post; here you'll get the framing and the direction.&lt;/p&gt;

&lt;h2&gt;
  
  
  Series
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;#&lt;/th&gt;
&lt;th&gt;Theme&lt;/th&gt;
&lt;th&gt;Key scene&lt;/th&gt;
&lt;th&gt;Article&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Series intro: cortex harness&lt;/td&gt;
&lt;td&gt;PRs merging unattended / incidents fixed before anyone notices&lt;/td&gt;
&lt;td&gt;&lt;a href="https://dev.to/ryantsuji/building-a-real-ai-harness-auto-reviewed-prs-self-healing-ops-and-non-engineer-contributors-3lfa"&gt;ai-harness-intro&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Product Graph (cpg)&lt;/td&gt;
&lt;td&gt;Code / docs / DB / infra unified into one graph&lt;/td&gt;
&lt;td&gt;&lt;a href="https://dev.to/ryantsuji/the-heart-of-the-ai-harness-a-knowledge-graph-of-the-ai-by-the-ai-for-the-ai-series-part-2-53bm"&gt;cortex-product-graph&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Auto PR review&lt;/td&gt;
&lt;td&gt;webhook -&amp;gt; AI review -&amp;gt; auto-fix -&amp;gt; squash merge&lt;/td&gt;
&lt;td&gt;&lt;a href="https://dev.to/ryantsuji/human-on-the-loop-ai-reviewing-ai-prs-at-cortex-769-prsmonth-while-raising-the-quality-bar-4lh5"&gt;cortex-auto-review&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Self-Healing + observability + auto-added guardrails&lt;/td&gt;
&lt;td&gt;Alert -&amp;gt; AI investigates -&amp;gt; fix PR + new lint/type gate -&amp;gt; auto redeploy + same pattern auto-rejected from then on&lt;/td&gt;
&lt;td&gt;&lt;a href="https://dev.to/ryantsuji/fixed-before-anyone-notices-stronger-after-every-fix-self-healing-recurrence-prevention-series-1e86"&gt;cortex-self-healing&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Democratizing the maintenance phase&lt;/td&gt;
&lt;td&gt;Domain experts open PRs to production; the harness owns the quality gate&lt;/td&gt;
&lt;td&gt;This article ← you are here&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Series Final&lt;/td&gt;
&lt;td&gt;The underlying philosophy plus a retrospective on the failures and lessons&lt;/td&gt;
&lt;td&gt;&lt;a href="https://dev.to/ryantsuji/ai-isnt-something-to-trust-its-something-to-design-series-final-30aa"&gt;cortex-philosophy&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Start with one scene
&lt;/h2&gt;

&lt;p&gt;A &lt;strong&gt;+1,742 line / 41 file&lt;/strong&gt; PR lands on the internal dashboard web app. Title: "PL dashboard ver.2". The change opens up project visibility to managers and team leads across multiple business units, scoping what each person sees to their own division or team. It adds an SSoT in the shared types package, new routes on the API server with SQL involving &lt;code&gt;INNER JOIN&lt;/code&gt; and &lt;code&gt;LEFT JOIN&lt;/code&gt;, new pages and view-state on the web app, and a personal-settings surface -- the whole stack of things you'd expect for a real feature.&lt;/p&gt;

&lt;p&gt;The point is, this isn't a typo fix or a string swap. Entities, repositories, API routes, screens, filters, personal settings -- every layer you'd normally touch for a feature got touched. &lt;strong&gt;A few days of work for an experienced engineer, in scale terms.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The review-fix cycle ran like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;PR open&lt;/strong&gt; (+1,742 / 41 files)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;auto-review pass 1&lt;/strong&gt;: Major finding (a permission-scope fall-through -- data from other divisions leaking into the view that shouldn't be there) plus a handful of Minor items&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;author bot push&lt;/strong&gt;: closes the scope fall-through, addresses the Minor items&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;auto-review pass 2&lt;/strong&gt;: Nit items remaining, plus a lint catch (&lt;code&gt;no-empty-function&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;author bot push&lt;/strong&gt;: lint clean&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;auto-review pass 3&lt;/strong&gt;: still some COMMENTED nits, not yet APPROVE&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;author bot push (iteration 2)&lt;/strong&gt;: hardens loading skeleton, reverts an unnecessary JSDoc tweak&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;auto-review pass 4: APPROVED&lt;/strong&gt; → CI green + APPROVE both met → auto-merge → production&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;From PR open to merge: &lt;strong&gt;four review-fix rounds, three author-bot pushes, zero human reviewers in the loop.&lt;/strong&gt; The reviews come from the auto-review bot, the fixes come from an author bot (an automated review-response agent that the PR author has running on their machine), the final APPROVE is submitted by the AI, and an auto-merge script picks it up the instant CI is green. Production lands with &lt;strong&gt;56/56 shared type checks (SSoT), 2,284/2,284 API tests, 1,113/1,113 web specs, and 0 lint errors.&lt;/strong&gt; (cortex splits the lint job between &lt;a href="https://oxc.rs/docs/guide/usage/linter" rel="noopener noreferrer"&gt;oxlint&lt;/a&gt; for general checks and a custom eslint plugin for the &lt;code&gt;@graph-*&lt;/code&gt; rules.)&lt;/p&gt;

&lt;p&gt;The second review pass is worth noting. "Scope fall-through" is a somewhat technical finding -- a hole in the permission filter meant data from divisions other than your own could leak into the view. This is an internal dashboard, so it's not an external-leak incident, but &lt;strong&gt;"only see what's relevant to you" is the whole point of a dashboard like this&lt;/strong&gt; -- losing it doesn't just risk an information slip, it drowns the user in noise that they shouldn't be filtering through in the first place. That's the kind of issue that's easy to merge by mistake and painful to notice in production. &lt;strong&gt;The fact that auto-review caught it on pass one and bounced it back for the author side to fix is what makes this whole flow viable for non-engineers.&lt;/strong&gt; Without that loop, a PR of this size from a non-engineer would be a bad bet.&lt;/p&gt;

&lt;p&gt;And: &lt;strong&gt;the author of this PR is not an engineer&lt;/strong&gt;. A business-side teammate handed a feature description to Claude Code, leaned on the knowledge graph (covered in &lt;a href="https://dev.to/ryantsuji/the-heart-of-the-ai-harness-a-knowledge-graph-of-the-ai-by-the-ai-for-the-ai-series-part-2-53bm"&gt;Part 2&lt;/a&gt;) to pull in the relevant existing code, and the +1,742 line PR is what came back. The four review-fix rounds above are what happened next.&lt;/p&gt;

&lt;p&gt;That setup lines up directly with the central claim of this post: &lt;strong&gt;the person who knows the business requirements best, instead of organizing them and handing them to an engineer, runs them through Claude Code to production themselves.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Quick clarification on "write." When I say "write" in this article, I don't mean &lt;strong&gt;typing line by line in an editor&lt;/strong&gt;. I mean &lt;strong&gt;handing the business requirements to Claude Code, judging the resulting diffs and AI review comments with domain knowledge, and seeing it through to a production merge&lt;/strong&gt; -- the whole arc. Most of the actual diff is written by Claude Code; review feedback is handled by the author bot. What the human does is three things: put what they want into words, make the judgment calls along the way ("does this fit, is this off"), and sign off when it's ready to merge. None of that is implementation work in the technical sense. That's what "write" means here.&lt;/p&gt;

&lt;p&gt;There's still a learning curve, of course -- the prompts you give Claude Code, where to point it for context. But &lt;strong&gt;none of that is learning to program.&lt;/strong&gt; What you need is the ability to articulate what you want clearly, not syntax or framework knowledge.&lt;/p&gt;

&lt;p&gt;The harness covers quality, so even at +1,742 lines / 41 files, this works.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Out of scope&lt;/strong&gt;: this post does &lt;em&gt;not&lt;/em&gt; cover the path where non-engineers freely ship apps to a sandbox environment instead of opening PRs against the production repo. That's a different mechanism, covered in an earlier post: &lt;a href="https://dev.to/ryantsuji/bridging-i-want-to-build-and-i-want-to-publish-safely-for-non-engineers-sandbox-mcp-392a"&gt;Bridging "I Want to Build" and "I Want to Publish Safely" for Non-Engineers with a Custom Sandbox MCP&lt;/a&gt;. This post is specifically about &lt;strong&gt;opening PRs against the production repo&lt;/strong&gt; -- the front door that's traditionally been engineer-only.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  When you need a change, you can make it yourself
&lt;/h2&gt;

&lt;p&gt;The point of the previous section is this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;When you need a change, you make it yourself, without flagging an engineer.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;When that holds, work like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"I want a new metric on the dashboard"&lt;/li&gt;
&lt;li&gt;"The aggregation filter doesn't match how the business actually operates"&lt;/li&gt;
&lt;li&gt;"I want a small business-support feature embedded in the production app"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;stops queueing behind whatever an engineer is in the middle of. The fix lands when the need lands.&lt;/p&gt;

&lt;p&gt;Think about the old flow. Someone on the business side notices a small thing that needs to change. They write the requirements up. They open a ticket or a Slack thread for an engineer. The engineer is in the middle of something else, so it queues. When they finally get to it, the interpretation drifts from what the business actually meant, there's a back-and-forth, a review pass, and only then does it ship. Even a small change takes days to a week in wall-clock time.&lt;/p&gt;

&lt;p&gt;That's the cost of a &lt;strong&gt;translation layer between business understanding and code&lt;/strong&gt;, and it gets worse the busier the engineer is. The business's improvement cycle ends up paced by engineering's backlog.&lt;/p&gt;

&lt;p&gt;When the person who knows the requirements writes the change themselves, that translation layer and that queue both disappear.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F9qelj5u7og19uc7gogaw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F9qelj5u7og19uc7gogaw.png" alt="Business request to production -- the translation layer and queue go away, taking the cycle from days to hours" width="800" height="410"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here are two recent examples of that working.&lt;/p&gt;

&lt;h3&gt;
  
  
  Two non-engineer PRs that recently shipped
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;PR&lt;/th&gt;
&lt;th&gt;Kind&lt;/th&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;th&gt;What changed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;PR 1&lt;/td&gt;
&lt;td&gt;Deep bug fix&lt;/td&gt;
&lt;td&gt;+348 -177 / 7 files&lt;/td&gt;
&lt;td&gt;The dashboard's actuals number was unfairly exceeding the target. Root cause: the "which teams to aggregate" definition was asymmetric between target side and actuals side. Fix lifts the shared "teams to include" list into its own file and points both sides at it. &lt;strong&gt;Tests added too.&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PR 2&lt;/td&gt;
&lt;td&gt;Feature build on top of existing stack&lt;/td&gt;
&lt;td&gt;+1,742 -227 / 41 files&lt;/td&gt;
&lt;td&gt;The PL dashboard v2 from the opening scene. &lt;strong&gt;Entities, repositories, API, UI -- all touched&lt;/strong&gt;, but the web app itself (the stack) was already standing; this rides on top.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Different shapes, but both are non-engineer PRs that made it all the way to merge.&lt;/p&gt;

&lt;h4&gt;
  
  
  PR 1 -- a deep root-cause fix
&lt;/h4&gt;

&lt;p&gt;This started from "the number looks wrong" on the business side, and the PR went all the way down to a data-integrity issue. The surface symptom: "the actuals number on the dashboard exceeds the monthly target, with the achievement reading 101% even though the team knows that's not real." The lazy fix would be a fudge factor or a clamp on the display. That's not what happened.&lt;/p&gt;

&lt;p&gt;The author dug into the aggregation queries and pinned the real cause: &lt;strong&gt;the actuals side and the target side were reading from different tables, and the definition of "which teams count" wasn't symmetric between them.&lt;/strong&gt; Teams that don't carry a target value (designers, PMOs, and so on) didn't show up on the target side but were getting counted on the actuals side, so the numerator was inflated against the denominator.&lt;/p&gt;

&lt;p&gt;The fix is structural, not cosmetic. A single file defines "the teams in scope for this aggregation" as a shared list, and both sides reference it. &lt;strong&gt;No future drift between target-side and actuals-side definitions&lt;/strong&gt; -- it's locked in by the shared constant.&lt;/p&gt;

&lt;p&gt;The handling of "what data falls out of an aggregation" and "are the target and actuals sides really symmetric" is the kind of thing engineers miss too. &lt;strong&gt;A non-engineer working through it down to the structural level and fixing it there&lt;/strong&gt; is what stands out about this PR.&lt;/p&gt;

&lt;h4&gt;
  
  
  PR 2 -- a big feature build on top of an existing stack
&lt;/h4&gt;

&lt;p&gt;This is the PR the opening scene walked through. +1,742 / 41 files spanning entity, repository, API, and UI -- &lt;strong&gt;a scale of change that's well past what people usually mean when they say "modification."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;What lets a non-engineer ship this size of change is that &lt;strong&gt;the web app itself (the stack) is already standing.&lt;/strong&gt; Nobody's standing up a new app, no new Cloud Run service needed defining, no new dependency packages, no new directory structure. The change adds a route, a page, and a repository entry inside the existing structure that's already there. It rides on what's been built.&lt;/p&gt;

&lt;p&gt;This is the "on top of an existing stack" range. That's where the boundary is, and the next section spells it out.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note on terminology&lt;/strong&gt;: "modification" in this article is broader than "small tweaks to existing logic." It includes adding new entities, new endpoints, and new pages on top of an existing stack. The line I'm drawing is between &lt;strong&gt;building on top of a stack&lt;/strong&gt; vs. &lt;strong&gt;standing the stack up in the first place.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What works, what doesn't
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The principle: standing up a stack is hard, building on top of one isn't
&lt;/h3&gt;

&lt;p&gt;The cleanest dividing line for non-engineer development isn't "modification vs. new development." It's &lt;strong&gt;"on top of an existing stack" vs. "stand up a new stack."&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Standing up a new stack&lt;/strong&gt; (work that starts from infrastructure: a new web app from scratch, a new Cloud Run service defined from a Dockerfile, a brand-new BigQuery pipeline) → engineering work&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Adding to an existing stack&lt;/strong&gt; (a new page in an app that already exists, a new endpoint on an existing API, a new data source on an existing pipeline) → non-engineers can do this&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All three of the example PRs above sit on the second side. The stack itself was already built (by me, for the most part), so they get to work inside it. "Stand up a new app from scratch" or "define infrastructure (Dockerfile / IaC) from zero" are still engineer territory.&lt;/p&gt;

&lt;p&gt;Put another way: &lt;strong&gt;renovations and new rooms inside an existing house are open to anyone. Building the house itself is engineering.&lt;/strong&gt; Get the structure wrong -- the load-bearing parts, the wiring, the plumbing -- and the cost of recovery is high. That's the part of stack design where there's still too much "if this is wrong, everything downstream breaks" risk to hand to AI.&lt;/p&gt;

&lt;h3&gt;
  
  
  What's left for engineers: laying the rails -- the stack and the harness itself
&lt;/h3&gt;

&lt;p&gt;The flip side: &lt;strong&gt;the rail-laying work&lt;/strong&gt; -- standing up a stack, and &lt;strong&gt;extending the harness itself&lt;/strong&gt; -- is what non-engineers don't touch yet. Both require a different kind of knowledge:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Infrastructure&lt;/strong&gt;: containers, IaC, the operational characteristics of cloud services. Cloud Run resource ceilings, cold starts, Pub/Sub at-least-once semantics, BigQuery partition / cluster design, how Pulumi stacks split. Get this wrong and a thing that compiles can still fall over in production&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Authentication for external integrations&lt;/strong&gt;: OAuth, webhooks, how you handle API keys and where they sit in Secret Manager. One small slip leaks credentials into the repo or lets a webhook fire something you didn't intend&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security fundamentals&lt;/strong&gt;: what to never expose, where to sanitize, where the privilege boundary cuts. SQL injection, XSS, SSRF, broken authorization -- "it works" isn't enough here&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Harness design and extension&lt;/strong&gt;: adding a new Auto Review dimension, changing Self-Healing logic, writing a new lint rule (e.g. in &lt;code&gt;eslint-plugin-graph&lt;/code&gt;), structuring guidelines. &lt;strong&gt;Decisions that require understanding how the whole flywheel hangs together&lt;/strong&gt; -- the most meta layer&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last bullet -- harness extension -- has an important implication: &lt;strong&gt;for non-engineers to keep being able to ship to production, someone has to keep the harness evolving.&lt;/strong&gt; Recurrence Prevention (Part 4) is the automatic loop that adds lint / CI guards / guidelines per trap. But the architecture of the harness itself -- the structure of dimensions, the calibration of judgment, the design of the Self-Healing flow, the shape of the knowledge graph -- those are a meta layer that still requires engineering judgment.&lt;/p&gt;

&lt;p&gt;Concrete case: the current nine Auto Review dimensions (&lt;code&gt;[Graph]&lt;/code&gt; / &lt;code&gt;[Architecture]&lt;/code&gt; / &lt;code&gt;[Security]&lt;/code&gt; / &lt;code&gt;[Test]&lt;/code&gt; / &lt;code&gt;[Doc]&lt;/code&gt; / &lt;code&gt;[Impact]&lt;/code&gt; / &lt;code&gt;[Observability]&lt;/code&gt; / &lt;code&gt;[AI-Antipattern]&lt;/code&gt; / &lt;code&gt;[Recurrence]&lt;/code&gt;) were designed by observing past incidents and fix patterns. When a tenth dimension becomes necessary -- say, a "breaking change check on dependency upgrades" axis -- decisions about responsibility splits with existing dimensions and where to set thresholds are made by looking at the whole structure. That's the kind of engineering work that stays.&lt;/p&gt;

&lt;p&gt;The harness provides "rails you can't derail from." Laying those rails -- and laying the foundation those rails sit on -- is a different job, and it's still on engineering. &lt;strong&gt;Engineers lay the rails; anyone can run on them.&lt;/strong&gt; That's the boundary today.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fn0pvl864tgh1o4vpzlru.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fn0pvl864tgh1o4vpzlru.png" alt="Three layers -- the upper layer is the non-engineer surface; the lower two (harness and stack) are engineering work" width="800" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this works for non-engineers
&lt;/h2&gt;

&lt;p&gt;This is a short recap, because everything that makes it work was already covered in Parts 1 through 4. &lt;strong&gt;Four mechanisms reinforcing each other&lt;/strong&gt; -- that's what lets non-engineers operate safely on top of an existing stack.&lt;/p&gt;

&lt;h3&gt;
  
  
  ① The knowledge graph pulls relevant code from "what you want to do"
&lt;/h3&gt;

&lt;p&gt;cortex-product-graph from &lt;a href="https://dev.to/ryantsuji/the-heart-of-the-ai-harness-a-knowledge-graph-of-the-ai-by-the-ai-for-the-ai-series-part-2-53bm"&gt;Part 2&lt;/a&gt; -- the unified graph fusing code, docs, DB schema, and infrastructure into one knowledge base (implementation name: cpg) -- carries this layer.&lt;/p&gt;

&lt;p&gt;Non-engineers don't need to know function names or repo structure. A natural-language question like "I want to add a metric column to the dashboard" goes to Claude Code, which hits the knowledge graph with a semantic search and gets back the relevant nodes -- the screen, the API, the DB, the docs -- in one or two hops. &lt;strong&gt;You can get started without knowing the technical vocabulary.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For PR 2: the author told Claude Code "I want PL dashboard v2 with division/team scoping for non-PI-Div PMOs and team leads," and the knowledge graph pulled the existing &lt;code&gt;/projects&lt;/code&gt; route, &lt;code&gt;project-repository.ts&lt;/code&gt;, &lt;code&gt;FilterHeaders.tsx&lt;/code&gt;, and &lt;code&gt;ProjectTable.tsx&lt;/code&gt; as the relevant nodes. The author never needed to know what file to edit. &lt;strong&gt;That's how the translation layer drops out.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  ② Auto Review enforces quality at the gate
&lt;/h3&gt;

&lt;p&gt;The 9-dimension automated review from &lt;a href="https://dev.to/ryantsuji/human-on-the-loop-ai-reviewing-ai-prs-at-cortex-769-prsmonth-while-raising-the-quality-bar-4lh5"&gt;Part 3&lt;/a&gt; is the next layer. &lt;code&gt;[Graph]&lt;/code&gt; / &lt;code&gt;[Architecture]&lt;/code&gt; / &lt;code&gt;[Security]&lt;/code&gt; / &lt;code&gt;[Test]&lt;/code&gt; / &lt;code&gt;[Doc]&lt;/code&gt; / &lt;code&gt;[Impact]&lt;/code&gt; / &lt;code&gt;[Observability]&lt;/code&gt; / &lt;code&gt;[AI-Antipattern]&lt;/code&gt; / &lt;code&gt;[Recurrence]&lt;/code&gt; -- the AI returns REQUEST_CHANGES on what's missing and loops with the author bot until APPROVE -- the four-round example from the opening scene is exactly this in motion.&lt;/p&gt;

&lt;p&gt;The point is this: &lt;strong&gt;the first PR doesn't have to be perfect.&lt;/strong&gt; The author doesn't need to ship a completed, security-hole-free version on the first try. Push the initial PR and the rest gets sorted by the auto-review and the author bot bouncing off each other. The reason &lt;strong&gt;the author bot doesn't spin off into a loop of confused fixes&lt;/strong&gt; is that the knowledge graph holds the full codebase context: changes are made with structural awareness of what they touch, so misreadings of the review feedback don't compound.&lt;/p&gt;

&lt;h3&gt;
  
  
  ③ Self-Healing catches what slips through to production
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://dev.to/ryantsuji/fixed-before-anyone-notices-stronger-after-every-fix-self-healing-recurrence-prevention-series-1e86"&gt;Part 4&lt;/a&gt; covered Self-Healing. If something does break in production, the AI starts from the alert, investigates root cause, opens a fix PR, and gets it auto-redeployed -- the entire loop runs without humans. &lt;strong&gt;Incidents triggered by a non-engineer's change recover on their own, hands-off.&lt;/strong&gt; That's what makes the bar to opening a PR feel survivable.&lt;/p&gt;

&lt;p&gt;This isn't "non-engineers are safe because nothing can go wrong." It's "even if something goes wrong, the harness has it covered." The system is designed to &lt;strong&gt;minimize damage&lt;/strong&gt;, not eliminate failure. The three-layer construction (Observation → Repair → Strengthening) from Part 4 is what makes that net real.&lt;/p&gt;

&lt;h3&gt;
  
  
  ④ Recurrence Prevention keeps the trap count from growing
&lt;/h3&gt;

&lt;p&gt;The Recurrence Prevention loop from the back half of Part 4. &lt;strong&gt;Every trap that gets stepped on gets nailed down in the same PR&lt;/strong&gt;, so the next attempt at the same pattern gets caught. The form depends: mechanizable traps become lint or CI guards; less-mechanizable ones become entries in the guideline docs (&lt;code&gt;docs/gotchas&lt;/code&gt;, severity docs) that the AI reviewer reads. Either way, the catch happens before merge. Non-engineers contribute to this loop too -- when they hit a trap, the doc entry that prevents the next person from hitting it can come from them.&lt;/p&gt;

&lt;p&gt;As this compounds, &lt;strong&gt;the rails get denser.&lt;/strong&gt; Where there was once a loose "don't go that way" guideline, every incident adds another small rail saying "or this way, or this way, or this way," and the lane that's safe to walk gets clearer. The denser the rails, the safer non-engineers are in the lane.&lt;/p&gt;

&lt;p&gt;→ The four pieces aren't independent components. &lt;strong&gt;Each one's output feeds the next one's input.&lt;/strong&gt; This is the Guides + Sensors flywheel from Part 1 in action. I won't re-explain the details since they're in the prior posts, but &lt;strong&gt;non-engineers shipping to production is the result of all four wheels turning together.&lt;/strong&gt; Take any one out and the level of upfront knowledge required to write to production jumps, and the whole thing collapses.&lt;/p&gt;

&lt;h2&gt;
  
  
  Next -- carrying the pattern to consumer-facing services
&lt;/h2&gt;

&lt;p&gt;cortex is an internal AI platform, so the system as it stands can't be lifted into a toC production service as-is. &lt;strong&gt;The biggest issue is the difference in quality bar.&lt;/strong&gt; For toC, "detect after user impact → Self-Healing fix" is too late. The requirement becomes: incidents don't happen, and when something is about to ship, there's review and testing on top of human sign-off.&lt;/p&gt;

&lt;p&gt;That said, the &lt;strong&gt;shape&lt;/strong&gt; of the harness -- a knowledge graph for context, 9-dimension AI review, an author bot responding to feedback -- carries over directly. The thing that changes is &lt;strong&gt;the final step&lt;/strong&gt;: cortex's auto-merge becomes "&lt;strong&gt;AI does the prep, a human signs off&lt;/strong&gt;." Not by giving up the AI's range, but by having the AI handle the heavy lifting (test writing, environment setup, test runs, the 9-dimension review) and leaving only the final APPROVE on a human. "If the human sign-off stays, engineer time doesn't really decrease, does it?" -- but historically engineers were spending the bulk of their time on the implementation, the test writing, the environment setup, the self-review, the back-and-forth on review. Sign-off itself is the smallest piece of that pie. With AI doing the prep work, what an engineer spends time on shifts from implementation labor to &lt;strong&gt;quality judgment&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A caveat on the knowledge graph: &lt;strong&gt;it only earns its keep at large codebase scale.&lt;/strong&gt; If the codebase fits in one AI context window, a cross-repo graph is unnecessary. The reason cortex (100+ apps) and the toC side (40+ repos) need one is because the scale forces it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F59dzskllpk4kp8jirits.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F59dzskllpk4kp8jirits.png" alt="cortex's shape carried into toC services -- internal knowledge graph → service-side knowledge graph / auto-merge → AI-prep + human sign-off / autonomous Self-Healing → human final call -- three things shift, the rest holds" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The concrete plan is real (extending the knowledge graph across the toC side's 40+ repositories, designing the AI-prep flow, etc.), and the full version goes in a separate post.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrap-up
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The person who knows the business requirements best, instead of writing them up for an engineer, runs them through to production directly.&lt;/strong&gt; Quality is held by the harness, so what's required from the writer is domain knowledge and the ability to direct an AI well. Business asks stop queuing behind engineering, and the cycle speeds up&lt;/li&gt;
&lt;li&gt;The four mechanisms from Parts 1-4 (knowledge graph / Auto Review / Self-Healing / Recurrence Prevention) form a reinforcing flywheel. &lt;strong&gt;The first PR doesn't have to be perfect, and what does break is repaired automatically.&lt;/strong&gt; That's the design&lt;/li&gt;
&lt;li&gt;The boundary: &lt;strong&gt;engineers lay the rails, anyone can run on them.&lt;/strong&gt; Standing up the stack (infrastructure, authentication, security) and extending the harness itself (new lint rules, new review dimensions, Self-Healing flow design) stay on engineering&lt;/li&gt;
&lt;li&gt;Carrying this to consumer-facing toC services, &lt;strong&gt;the knowledge graph (a 40+ repo cross-repo graph on the service side) covers the context layer, but the quality bar shifts, so auto-merge becomes "AI prep + human sign-off."&lt;/strong&gt; Details in a separate post&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;In &lt;strong&gt;Part 6&lt;/strong&gt; I'll wrap the series with the philosophy at the foundation -- &lt;strong&gt;why this design, what got given up, what got kept&lt;/strong&gt;. The series so far has been about "the parts that are working"; Part 6 puts the failures and the dead ends on the table too, including the gap between the philosophy and the actual implementation. A retrospective for myself, and -- I hope -- a reference for anyone heading down a similar road.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>engineering</category>
      <category>github</category>
    </item>
    <item>
      <title>Fixed Before Anyone Notices, Stronger After Every Fix: Self-Healing + Recurrence Prevention</title>
      <dc:creator>Ryosuke Tsuji</dc:creator>
      <pubDate>Mon, 01 Jun 2026 23:57:25 +0000</pubDate>
      <link>https://dev.to/ryantsuji/fixed-before-anyone-notices-stronger-after-every-fix-self-healing-recurrence-prevention-series-1e86</link>
      <guid>https://dev.to/ryantsuji/fixed-before-anyone-notices-stronger-after-every-fix-self-healing-recurrence-prevention-series-1e86</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;AI assistance disclosure: This article was drafted with the help of Claude. All technical content, design decisions, code references, and screenshots reflect production systems I designed and operate at airCloset; the prose was revised by me prior to publication.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hi, I'm &lt;a href="https://x.com/ryantsuji" rel="noopener noreferrer"&gt;Ryan&lt;/a&gt;, CTO at airCloset.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Disclaimer&lt;/strong&gt;: "cortex" in this article is the internal codename for an AI platform built in-house at airCloset. It is unrelated to existing commercial services like Snowflake Cortex or Palo Alto Networks Cortex.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;In &lt;a href="https://dev.to/ryantsuji/human-on-the-loop-ai-reviewing-ai-prs-at-cortex-769-prsmonth-while-raising-the-quality-bar-4lh5"&gt;Part 3&lt;/a&gt; I covered &lt;strong&gt;AI reviewing AI PRs&lt;/strong&gt; -- the auto-review pipeline that defends quality &lt;strong&gt;at the PR stage&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This post is the other side: &lt;strong&gt;defending quality in production&lt;/strong&gt;, via &lt;strong&gt;Self-Healing&lt;/strong&gt;. A production alert fires, an AI investigates it, opens a fix PR, the PR goes through the same auto-review pipeline from Part 3, gets auto-merged and auto-redeployed. And the same fix PR is &lt;strong&gt;required to add a new Guide -- whether that's a lint rule, CI guard, type constraint, or guideline update&lt;/strong&gt; -- so the same anti-pattern gets auto-rejected from then on. The guardrails grow every time.&lt;/p&gt;

&lt;p&gt;"Incidents get fixed automatically" is catchy on its own, but on its own it's probably not enough in the long run. You have to &lt;strong&gt;close the recurrence class while you fix the incident&lt;/strong&gt; -- self-healing &lt;strong&gt;plus&lt;/strong&gt; self-strengthening -- before the quality gates start to compound over time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with last month's numbers
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;115 Self-Healing PRs merged in the past 30 days.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Effectively all of them merged and deployed without human involvement.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Humans only step in when the AI judges "this is not something code can fix."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's the current state of "incident response" at cortex.&lt;/p&gt;

&lt;p&gt;Don't read "115 = 115 user-impacting incidents" though. Roughly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;About half (54) are Deploy Failed-style alerts&lt;/strong&gt; -- CI / Pulumi deploy step caught a failure, the AI absorbed it &lt;strong&gt;before it shipped to production&lt;/strong&gt;. Recently the &lt;code&gt;[Recurrence]&lt;/code&gt; loop (covered later) has been piling up countermeasures here, so this bucket is trending down anecdotally&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The remaining 61 are production-runtime alerts&lt;/strong&gt; (Service Error Log Detected / Pipeline Failure / Generator Failure etc.) -- the service is running in production, but an error-log threshold or consecutive-failure threshold tripped. The AI absorbed them before they propagated to user impact&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So it's less "incident response" than "&lt;strong&gt;production anomalies that monitoring caught, fixed 115 times by AI before anyone woke up&lt;/strong&gt;." The number of incidents humans actually have to acknowledge is in the low single digits per month.&lt;/p&gt;

&lt;p&gt;There's also a clear pattern of &lt;strong&gt;the same service firing repeatedly&lt;/strong&gt; (one ETL-ish service alone accounts for 25 of the 61) -- which is exactly what the &lt;code&gt;[Recurrence]&lt;/code&gt; loop covered later is supposed to &lt;strong&gt;eliminate by turning into lint or type gates&lt;/strong&gt;. That's the back half of this post.&lt;/p&gt;

&lt;p&gt;One more honest note: &lt;strong&gt;the recent month's number is slightly inflated&lt;/strong&gt;. The codebase had a fair number of "silent catch" patterns -- catch blocks that swallow exceptions without logging anything. We added the &lt;code&gt;no-silent-catch&lt;/code&gt; lint rule and &lt;strong&gt;swept the existing silent catches in batches&lt;/strong&gt;, which exposed previously hidden production errors as alerts. So part of the spike is "monitoring caught up to reality." Once the &lt;code&gt;[Recurrence]&lt;/code&gt; loop converts these into lint over time, the number should converge. &lt;strong&gt;"Things we couldn't see, we can see now" is a quality improvement&lt;/strong&gt; -- what we're seeing is the catch-up phase.&lt;/p&gt;

&lt;p&gt;One more thing worth saying: doing this by hand is utterly unsustainable. Running 115 manual cycles of "ack alert -&amp;gt; read logs -&amp;gt; context switch -&amp;gt; understand the code -&amp;gt; fix -&amp;gt; open PR -&amp;gt; review -&amp;gt; deploy" would bankrupt any team's engineering bandwidth. &lt;strong&gt;The system absorbs them without anyone noticing, and converts the fix into a new Guide (lint / CI guard / type constraint / guideline) at the same time&lt;/strong&gt; -- that's the actual subject of this post.&lt;/p&gt;

&lt;p&gt;The moment an alert fires, the AI starts an investigation, traces Loki / Product Graph / git blame to root cause, opens a fix PR, runs it through the auto-review from &lt;a href="https://dev.to/ryantsuji/human-on-the-loop-ai-reviewing-ai-prs-at-cortex-769-prsmonth-while-raising-the-quality-bar-4lh5"&gt;Part 3&lt;/a&gt;, APPROVE -&amp;gt; auto-merge -&amp;gt; auto-redeploy. One full loop.&lt;/p&gt;

&lt;h2&gt;
  
  
  Series
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;#&lt;/th&gt;
&lt;th&gt;Theme&lt;/th&gt;
&lt;th&gt;Key scene&lt;/th&gt;
&lt;th&gt;Article&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Series intro: cortex harness&lt;/td&gt;
&lt;td&gt;PRs merging unattended / incidents fixed before anyone notices&lt;/td&gt;
&lt;td&gt;&lt;a href="https://dev.to/ryantsuji/building-a-real-ai-harness-auto-reviewed-prs-self-healing-ops-and-non-engineer-contributors-3lfa"&gt;ai-harness-intro&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Product Graph (cpg)&lt;/td&gt;
&lt;td&gt;Code / docs / DB / infra unified into one graph&lt;/td&gt;
&lt;td&gt;&lt;a href="https://dev.to/ryantsuji/the-heart-of-the-ai-harness-a-knowledge-graph-of-the-ai-by-the-ai-for-the-ai-series-part-2-53bm"&gt;cortex-product-graph&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Auto PR review&lt;/td&gt;
&lt;td&gt;webhook -&amp;gt; AI review -&amp;gt; auto-fix -&amp;gt; squash merge&lt;/td&gt;
&lt;td&gt;&lt;a href="https://dev.to/ryantsuji/human-on-the-loop-ai-reviewing-ai-prs-at-cortex-769-prsmonth-while-raising-the-quality-bar-4lh5"&gt;cortex-auto-review&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Self-Healing + observability + auto-added guardrails&lt;/td&gt;
&lt;td&gt;Alert -&amp;gt; AI investigates -&amp;gt; fix PR + new lint/type gate -&amp;gt; auto redeploy + same pattern auto-rejected from then on&lt;/td&gt;
&lt;td&gt;This article ← you are here&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Democratizing the maintenance phase&lt;/td&gt;
&lt;td&gt;Domain experts open PRs to production; the harness owns the quality gate&lt;/td&gt;
&lt;td&gt;&lt;a href="https://dev.to/ryantsuji/the-author-doesnt-have-to-be-an-engineer-how-the-harness-holds-quality-series-part-5-12e4"&gt;cortex-non-engineer-prs&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Series Final&lt;/td&gt;
&lt;td&gt;The underlying philosophy plus a retrospective on the failures and lessons&lt;/td&gt;
&lt;td&gt;&lt;a href="https://dev.to/ryantsuji/ai-isnt-something-to-trust-its-something-to-design-series-final-30aa"&gt;cortex-philosophy&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Big picture -- the three layers: Observation, Repair, Strengthening
&lt;/h2&gt;

&lt;p&gt;For Self-Healing to work, you need an &lt;strong&gt;Observation layer&lt;/strong&gt; in front and a &lt;strong&gt;Strengthening layer&lt;/strong&gt; (recurrence prevention) behind it. Self-Healing itself is the middle &lt;strong&gt;Repair layer&lt;/strong&gt;. The "self-healing + self-strengthening" loop only spins up when all three are in place.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Prerequisites&lt;/strong&gt;: The three layers only stand up on top of two prior pieces: &lt;strong&gt;cpg&lt;/strong&gt; (the unified code / docs / DB / infra knowledge graph from &lt;a href="https://dev.to/ryantsuji/the-heart-of-the-ai-harness-a-knowledge-graph-of-the-ai-by-the-ai-for-the-ai-series-part-2-53bm"&gt;Part 2&lt;/a&gt;) and the &lt;strong&gt;Observability stack&lt;/strong&gt; covered in this post.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No Observability&lt;/strong&gt; -&amp;gt; the observation layer is empty, nothing gets detected -&amp;gt; the repair layer never even fires&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No cpg&lt;/strong&gt; -&amp;gt; the AI cannot see "where else does this trap exist" -&amp;gt; the repair layer does symptom-level patching at best, and the strengthening layer's horizontal expansion stops working&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Put differently: &lt;strong&gt;trying to copy this setup without those two will just multiply incidents&lt;/strong&gt;. An AI that blindly looks at error logs and rewrites production code is just speeding up the rate at which &lt;code&gt;gh pr create&lt;/code&gt; ships accidents. cpg and Observability are the &lt;strong&gt;minimum bar&lt;/strong&gt; for being able to delegate auto-repair to AI.&lt;/p&gt;

&lt;p&gt;Note also that cortex is a &lt;strong&gt;several-hundred-thousand-line codebase&lt;/strong&gt;, and at that scale loading the whole codebase as AI context is &lt;strong&gt;impossible for the AI as well&lt;/strong&gt; (let alone for a human). Tell the AI to trace impact with just grep and file reads, and it'll run out of context window before it finds anything. cpg is what lets it ask "which other code does this function's change ripple into" and get the answer in one hop. Small repos may not need this. Past a certain scale, cpg is not optional, it's &lt;strong&gt;required&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;In Fowler's Guides / Sensors terms from Part 1, cpg and Observability are &lt;strong&gt;the substrate that supports both Guides (pre-execution controls like lint) and Sensors (post-execution gates like auto-review and Self-Healing)&lt;/strong&gt;. Observability feeds Sensors via firing alerts; cpg feeds the Guides side by supplying the auto-review with impact-scoping context. &lt;strong&gt;Neither belongs on one side only&lt;/strong&gt; -- they're foundational to both, and Self-Healing and auto-review only function on top of this substrate. That's the structural claim this post is built around.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F9oqc8k3k9f6ljwddm50b.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F9oqc8k3k9f6ljwddm50b.png" alt="Three layers -- Observation -&gt; Repair -&gt; Strengthening loop"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;th&gt;Key components&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Observation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Real-time detection of production anomalies&lt;/td&gt;
&lt;td&gt;OTel SDK / Loki / Mimir / Tempo / Faro / Grafana / Pino logs with trace_id&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Repair&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;AI receives the alert, investigates root cause, opens a fix PR, auto-review, auto-merge, auto-redeploy&lt;/td&gt;
&lt;td&gt;Event Relay -&amp;gt; SSE -&amp;gt; &lt;code&gt;self-healing&lt;/code&gt; mode script -&amp;gt; claude -p (worktree) -&amp;gt; gh pr create&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Strengthening&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The fix PR is required to add a new Guide (lint / CI guard / type constraint / guideline). The same anti-pattern can't reach production again&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;@cortex/eslint-plugin-graph&lt;/code&gt; (26 rules), &lt;code&gt;scripts/check-*.ts&lt;/code&gt; (13 guards), &lt;a href="https://github.com/air-closet/cortex-review-guidelines/blob/main/en/guidelines/recurrence-prevention.md" rel="noopener noreferrer"&gt;&lt;code&gt;recurrence-prevention.md&lt;/code&gt;&lt;/a&gt;, the &lt;code&gt;[Recurrence]&lt;/code&gt; lens of auto-review&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I'll walk through them in order.&lt;/p&gt;

&lt;h2&gt;
  
  
  Observation -- where do the alerts come from?
&lt;/h2&gt;

&lt;p&gt;cortex's production observability is built on &lt;strong&gt;Grafana Cloud + OpenTelemetry&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;OTel SDK&lt;/strong&gt; (the shared &lt;code&gt;@cortex/otel&lt;/code&gt; package) -- every service calls &lt;code&gt;initOtel({ serviceName })&lt;/code&gt; at its entry point. Trace / metric / log all go out via OTLP to Grafana Cloud&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Loki&lt;/strong&gt; (logs) -- Pino structured logs get &lt;code&gt;trace_id&lt;/code&gt; automatically. trace and log are cross-referenced&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mimir&lt;/strong&gt; (metrics) -- Cloud Run / pipeline / Gemini API token usage, etc.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tempo&lt;/strong&gt; (traces) -- distributed tracing&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Faro&lt;/strong&gt; (frontend) -- captures browser JS errors / performance / network failures&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Grafana&lt;/strong&gt; -- dashboards + Alert Rules + Notification Policy&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We also have &lt;strong&gt;a strict definition of log levels, anchored on business impact&lt;/strong&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Level&lt;/th&gt;
&lt;th&gt;Definition&lt;/th&gt;
&lt;th&gt;Examples&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;warn&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Business-foreseeable, &lt;strong&gt;does not need immediate action&lt;/strong&gt; (retryable / self-recovers).&lt;/td&gt;
&lt;td&gt;Search query returned 0 results, optional field unset, short retry due to rate limit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;error&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Data recovery / re-run will definitely be needed afterward&lt;/strong&gt;. Impact expected to be under 20%.&lt;/td&gt;
&lt;td&gt;"User record that should exist isn't there," BigQuery insert failure, per-record enrichment failure&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;fatal&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The feature as a whole &lt;strong&gt;fails for 20%+ of requests&lt;/strong&gt;. Service-continuity broken, fatal config missing, full upstream outage.&lt;/td&gt;
&lt;td&gt;OTel init failure, required secret missing at startup, full input data source outage for a pipeline&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The key point is to &lt;strong&gt;not pick the level mechanically based on the exception class name&lt;/strong&gt; like &lt;code&gt;NotFoundError&lt;/code&gt;. Same "record not found" situation: "this record must exist and doesn't" is &lt;code&gt;error&lt;/code&gt; / &lt;code&gt;fatal&lt;/code&gt;; "user search returned 0 hits" is &lt;code&gt;warn&lt;/code&gt;. &lt;strong&gt;The level is decided by business impact&lt;/strong&gt; -- "does this require data recovery later," "is the whole feature down" -- not by the type. Without this discipline you simultaneously get monitoring fatigue and missed critical incidents. Self-Healing reacts mainly to &lt;code&gt;error&lt;/code&gt;-threshold trips; &lt;code&gt;fatal&lt;/code&gt; is the human-escalation side.&lt;/p&gt;

&lt;p&gt;Alert Rules are &lt;strong&gt;managed declaratively in Pulumi&lt;/strong&gt;, grouped by service into categories like &lt;code&gt;BOT / Pipeline / Transformer / Generator / Gemini / CI / Deploy / Service Catch-All&lt;/code&gt;. When we add a new service, one line in infra code spins up the dashboards and alerts automatically.&lt;/p&gt;

&lt;p&gt;This is "the infrastructure that lets &lt;strong&gt;the AI see the same things humans see&lt;/strong&gt;." Self-Healing picks up alerts coming off this stack.&lt;/p&gt;

&lt;h3&gt;
  
  
  What Observability can't catch, Self-Healing can't fix either
&lt;/h3&gt;

&lt;p&gt;Honest disclaimer: Self-Healing can only react to &lt;strong&gt;what the observation layer can detect as an anomaly&lt;/strong&gt;. "Observability is everything" is literally true here.&lt;/p&gt;

&lt;p&gt;What the current stack catches is roughly &lt;strong&gt;logic-level errors&lt;/strong&gt; -- exceptions, error logs, deploy failures, external-API call failures, threshold-based metric anomalies.&lt;/p&gt;

&lt;p&gt;What it doesn't catch:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;UI errors&lt;/strong&gt; -- the logic ran, no error logs, but the screen &lt;strong&gt;shows something different from intent / shows the wrong value&lt;/strong&gt;. Faro catches client-side JS exceptions and network failures, but "the logic ran and the output is just wrong" never fires an alert&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Silent data corruption&lt;/strong&gt; -- aggregated values slowly drift, bad values get into a table. Unless it crosses a threshold or schema check, nothing detects it&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Perceived UX degradation&lt;/strong&gt; -- requests feel slow, the UX feels off. Only catchable once SLO / latency thresholds trip&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So Self-Healing is "&lt;strong&gt;AI replacing the human in the loop for incidents the observation layer can catch&lt;/strong&gt;." &lt;strong&gt;The coverage of the observation layer itself is the prerequisite.&lt;/strong&gt; Holes in observation stay as blind spots that neither auto-review nor Self-Healing reaches.&lt;/p&gt;

&lt;p&gt;This isn't really a limitation of Self-Healing -- it's the &lt;strong&gt;importance of growing the observation stack&lt;/strong&gt;, which cortex keeps investing in continuously. (From &lt;a href="https://dev.to/ryantsuji/building-a-real-ai-harness-auto-reviewed-prs-self-healing-ops-and-non-engineer-contributors-3lfa"&gt;Part 1&lt;/a&gt;, Observability is one of the "supporting foundations" beneath the flywheel.)&lt;/p&gt;

&lt;h2&gt;
  
  
  Repair -- the Self-Healing flow
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;MODE=self-healing&lt;/code&gt; runs the same &lt;code&gt;webhook-server&lt;/code&gt; script as the auto-review setup from &lt;a href="https://dev.to/ryantsuji/human-on-the-loop-ai-reviewing-ai-prs-at-cortex-769-prsmonth-while-raising-the-quality-bar-4lh5"&gt;Part 3&lt;/a&gt;, but listening for Grafana firing alerts.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fjoa2w71p3guljzdg6sou.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fjoa2w71p3guljzdg6sou.png" alt="Self-Healing full flow -- median 30 min to 1 hr from firing alert to production recovery"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The textual flow looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[Grafana Alert Rule firing]
   ↓ POST /webhook/grafana
[Event Relay (in-house)] -- persisted in Firestore
   ↓ SSE push (event: grafana-alert)
[self-healing mode script]
   ↓ throttle check (same fingerprint skipped for 4h)
   ↓ 👀 reaction in Slack to signal "I'm on it"
   ↓ git worktree add -b hotfix/auto-alert-{service}-{ts} origin/main
   ↓ run claude -p inside the worktree
     - search related code via Product Graph MCP
     - pull error logs from Loki via Grafana MCP
     - identify root cause and fix
     - update tests as needed
     - conventional commit
   ↓ git push + gh pr create
[fix PR]
   ↓ auto-review (the Part 3 pipeline)
   ↓ APPROVE -&amp;gt; auto-merge -&amp;gt; auto-redeploy
[recovered]
   ↓ ✅ in the Slack thread
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  What happens when the AI judges "this is not fixable in code"
&lt;/h3&gt;

&lt;p&gt;Not every alert is fixable by code. The implementation has a rule: "if you judge it unfixable, exit without changing anything." In that case Slack gets a notification of the form "&lt;strong&gt;This alert cannot be addressed in code. Investigation: ...&lt;/strong&gt;" -- &lt;strong&gt;including what the AI investigated&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Worth clarifying on the numbers side: the headline &lt;strong&gt;115&lt;/strong&gt; is "Self-Healing runs that reached PR-created -&amp;gt; merged -&amp;gt; deployed." This "unfixable, exit clean" case is a &lt;strong&gt;separate bucket&lt;/strong&gt;, happening several times a month (external transient outages, infra / config issues that aren't code, cases too complex for the AI to judge confidently). &lt;strong&gt;The "humans show up" bucket is this separate one&lt;/strong&gt; -- it isn't a "some of the 115 failed" failure rate.&lt;/p&gt;

&lt;p&gt;Here's what a real Slack message looks like (a case where a transformer service's GitHub PAT had expired):&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;ℹ️ This alert cannot be addressed in code. Investigation:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Investigation summary&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Checked the error logs in Loki for the past hour. The single error source for the target service right now is &lt;strong&gt;expiration / revocation of the GitHub PAT (Personal Access Token)&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I found&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The BQ snapshot pipeline is healthy: 7 consecutive &lt;code&gt;[BQ] Snapshot rows inserted&lt;/code&gt; successes. The earlier autodetect type-drift problem was resolved in an earlier PR that put &lt;code&gt;SAFE_CAST(... AS INT64)&lt;/code&gt; in place.&lt;/li&gt;
&lt;li&gt;The single ongoing error (multiple occurrences in the past hour): &lt;code&gt;GithubUnauthorizedError: GitHub API /commits/master unauthorized (401): PAT likely expired/revoked&lt;/code&gt; -- the &lt;code&gt;/check-drift&lt;/code&gt; handler returns 401 every time Cloud Scheduler invokes it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Root cause&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The GitHub PAT held in Secret Manager for this service has expired or been revoked. The code already classifies the 401 with a dedicated &lt;code&gt;GithubUnauthorizedError&lt;/code&gt; and is already sending the "PAT needs rotation" notification to Slack. &lt;strong&gt;The detection / notification design on the code side is working as intended.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Resolution: exit without changes&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is an external-credential operational issue (PAT expiry), and what needs to happen is &lt;strong&gt;a human rotates the PAT in Secret Manager&lt;/strong&gt;. Not the kind of incident code can resolve. Therefore: no code change, no commit, no PR.&lt;/p&gt;

&lt;p&gt;Operator actions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Add a new version of the target secret in GCP Console / Secret Manager with a valid PAT (&lt;code&gt;repo&lt;/code&gt; scope, read access to the target repository)&lt;/li&gt;
&lt;li&gt;No Cloud Run revision redeploy needed (&lt;code&gt;secretKeyRef version:latest&lt;/code&gt; is referenced)&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;p&gt;At this level of detail, "what was investigated, why code can't fix it, what the human should do" all come out in one Slack message. Open the thread and the operator can act immediately. The productivity gap vs. "alerts just forwarded blindly" is significant.&lt;/p&gt;

&lt;h3&gt;
  
  
  Deduplication
&lt;/h3&gt;

&lt;p&gt;A throttle ensures the same &lt;code&gt;fingerprint&lt;/code&gt; (Grafana's unique alert identifier) is &lt;strong&gt;not re-processed for 4 hours&lt;/strong&gt;. Without this, alerts that fire again before the fix PR has merged would spawn another worktree, another fix PR, and so on -- an easy infinite loop.&lt;/p&gt;

&lt;p&gt;We also &lt;strong&gt;permanently skip&lt;/strong&gt; any &lt;code&gt;alertname&lt;/code&gt; containing &lt;code&gt;credential&lt;/code&gt;. Credential incidents carry leakage risk if the AI touches them, so they're explicitly escalated to humans.&lt;/p&gt;

&lt;h3&gt;
  
  
  Self-Healing and Part 3 auto-review -- "the fixer AI" and "the reviewer AI" are independent
&lt;/h3&gt;

&lt;p&gt;This is the most consequential design choice of the agent setup, so calling it out explicitly.&lt;/p&gt;

&lt;p&gt;PRs opened by Self-Healing are &lt;strong&gt;not special PRs, just fix PRs&lt;/strong&gt;. They go through the Part 3 auto-review pipeline &lt;strong&gt;under exactly the same conditions&lt;/strong&gt; -- the 9 lenses (Graph / Architecture / Security / Test / Doc / Impact / Observability / AI-Antipattern / Recurrence) get checked in order. Critical / Major findings -&amp;gt; &lt;code&gt;REQUEST_CHANGES&lt;/code&gt;; Nit-only / no findings + CI green -&amp;gt; &lt;code&gt;APPROVE&lt;/code&gt; -&amp;gt; auto-merge.&lt;/p&gt;

&lt;p&gt;The important bit: &lt;strong&gt;this is not a monolithic "AI fixing AI" loop&lt;/strong&gt;. The fixer-side AI and the reviewer-side AI are fully independent:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Different process, different session&lt;/strong&gt;: the self-healing-mode AI and the reviewer-mode AI are launched as separate &lt;code&gt;claude -p&lt;/code&gt; processes. They do not share context&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Different input sources&lt;/strong&gt;: the fixer builds the problem from Grafana alert + Loki + cpg. The reviewer judges from the PR diff + cpg + review guidelines&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Different objectives&lt;/strong&gt;: the fixer is optimizing for "stop the incident." The reviewer is judging "does this violate the 9 lenses or the severity contract?" A deliberate separation of concerns where the two roles' incentives are intentionally misaligned&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;As a result, &lt;strong&gt;PRs the fixer dashed off get blocked by the reviewer&lt;/strong&gt; (REQUEST_CHANGES -&amp;gt; back to the fixer). The AI does not approve its own output. "Just-make-it-work" fixes don't get through.&lt;/p&gt;

&lt;p&gt;This is the often-debated &lt;strong&gt;review-independence&lt;/strong&gt; problem in LLM-agent operation, solved here in the obvious way: split the work across separate agents.&lt;/p&gt;

&lt;h3&gt;
  
  
  A concrete example: meet subscription's 409 ALREADY_EXISTS
&lt;/h3&gt;

&lt;p&gt;Take the alert from the Google Meet recording auto-fetch service I covered in &lt;a href="https://dev.to/ryantsuji/how-we-built-an-automated-meeting-intelligence-system-with-google-meet-slack-and-rag-42ln"&gt;the Meeting Intelligence post&lt;/a&gt;. On 2026-05-21, Self-Healing opened a fix PR titled &lt;code&gt;fix(meet-xxx): auto-fix for Service Error Log Detected&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The trigger error from Loki:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Workspace Events API request failed: 409 Conflict
"Subscription associated with the resource already exists."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;How the AI investigated:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Pinned the error in Loki&lt;/strong&gt; -- ran &lt;code&gt;{service_name="meet-xxx"} | json | level=~"ERROR|error|Error"&lt;/code&gt; via Grafana MCP, picked up the &lt;code&gt;Failed to renew Meet subscription&lt;/code&gt; stack trace&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Traced the call path in Product Graph&lt;/strong&gt; -- identified &lt;code&gt;renewSubscriptions&lt;/code&gt; -&amp;gt; &lt;code&gt;createMeetSubscription&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cross-referenced past PRs&lt;/strong&gt; -- the "opposite-direction inconsistency" (name in Firestore but missing from Google = 404) had already been self-healed in another PR with &lt;code&gt;patchMeetSubscriptionTtl&lt;/code&gt; -&amp;gt; null fallback. &lt;strong&gt;The current direction (still on Google's side but missing from Firestore = 409) was the gap&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verdict&lt;/strong&gt;: "the same pattern may exist elsewhere" -- a [Recurrence] decision matrix "&lt;strong&gt;horizontal expansion required&lt;/strong&gt;" case&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Instead of a quick patch, &lt;strong&gt;it implemented the same-direction self-healing symmetrically to the opposite-direction fallback that was already there&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Made &lt;code&gt;createMeetSubscription&lt;/code&gt; idempotent&lt;/li&gt;
&lt;li&gt;If POST returns 409, extract the existing Subscription name from the response and call &lt;code&gt;patchMeetSubscriptionTtl&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;The caller writes the return value back into Firestore, so the next renewal converges to the normal PATCH path (&lt;strong&gt;self-healing&lt;/strong&gt;)&lt;/li&gt;
&lt;li&gt;Per the existing &lt;code&gt;graph/no-silent-catch&lt;/code&gt; lint, JSON.parse failures are also &lt;code&gt;logger.warn&lt;/code&gt; + &lt;code&gt;serializeError&lt;/code&gt; for structured logging&lt;/li&gt;
&lt;li&gt;Three tests added&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is what "Self-Healing pushing all the way to root cause and rolling the fix out horizontally" looks like in practice. &lt;strong&gt;"Close the recurrence class, don't just suppress the symptom"&lt;/strong&gt; (the spirit of &lt;a href="https://github.com/air-closet/cortex-review-guidelines/blob/main/en/guidelines/recurrence-prevention.md" rel="noopener noreferrer"&gt;&lt;code&gt;recurrence-prevention.md&lt;/code&gt;&lt;/a&gt;) executed autonomously by the AI.&lt;/p&gt;

&lt;h2&gt;
  
  
  Strengthening -- Guides (lint + guidelines) grow automatically
&lt;/h2&gt;

&lt;p&gt;This is the layer that &lt;strong&gt;keeps Self-Healing from being just "auto-repair."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In Fowler's Guides / Sensors terms from &lt;a href="https://dev.to/ryantsuji/building-a-real-ai-harness-auto-reviewed-prs-self-healing-ops-and-non-engineer-contributors-3lfa"&gt;Part 1&lt;/a&gt;, the Strengthening layer is &lt;strong&gt;the place where Guides grow&lt;/strong&gt; -- i.e. the pre-execution controls that prevent AI from deviating in the first place. cortex's Guides come in two flavors:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Machine-read Guides&lt;/strong&gt;: lint / type / CI guard / coverage thresholds / Prettier -- enforced at commit / CI time&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Human-and-AI-read Guides&lt;/strong&gt;: guidelines like &lt;a href="https://github.com/air-closet/cortex-review-guidelines/blob/main/en/guidelines/recurrence-prevention.md" rel="noopener noreferrer"&gt;&lt;code&gt;recurrence-prevention.md&lt;/code&gt;&lt;/a&gt;, &lt;a href="https://github.com/air-closet/cortex-review-guidelines/blob/main/en/guidelines/severity.md" rel="noopener noreferrer"&gt;&lt;code&gt;severity.md&lt;/code&gt;&lt;/a&gt;, &lt;a href="https://github.com/air-closet/cortex-review-guidelines/blob/main/en/guidelines/ai-antipattern.md" rel="noopener noreferrer"&gt;&lt;code&gt;ai-antipattern.md&lt;/code&gt;&lt;/a&gt;, etc. -- used as decision criteria by auto-review&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The 9 lenses, severity contract, and no-downgrade rules from &lt;a href="https://dev.to/ryantsuji/human-on-the-loop-ai-reviewing-ai-prs-at-cortex-769-prsmonth-while-raising-the-quality-bar-4lh5"&gt;Part 3&lt;/a&gt; are the latter; the auto-added lints in Part 4 are the former. &lt;strong&gt;Together they form the Guides surface&lt;/strong&gt;. Lints are "formalized guidelines," guidelines are "lints that haven't been formalized yet."&lt;/p&gt;

&lt;p&gt;The Sensors side -- Self-Healing and auto-review -- &lt;strong&gt;grow these Guides every time they run&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Self-Healing's root-cause investigation finds "the same pattern exists elsewhere" -&amp;gt; demands horizontal expansion + a new lint (= new Guide)&lt;/li&gt;
&lt;li&gt;Auto-review's &lt;code&gt;[Recurrence]&lt;/code&gt; lens blocks PRs that fix without adding lint&lt;/li&gt;
&lt;li&gt;Both depend on &lt;a href="https://dev.to/ryantsuji/the-heart-of-the-ai-harness-a-knowledge-graph-of-the-ai-by-the-ai-for-the-ai-series-part-2-53bm"&gt;cpg&lt;/a&gt; to see impact scope across the codebase&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;cpg is what lets the AI ask "where else does this trap exist." Self-Healing and auto-review (= the Sensors side) &lt;strong&gt;share cpg as a substrate, and each run thickens Guides by one notch&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F5m28vtcko417bg44sr3e.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F5m28vtcko417bg44sr3e.png" alt="cpg as the shared substrate; Self-Healing and auto-review (Sensors) grow Guides"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  What happens every time Self-Healing runs (the recurrence-prevention-first flow)
&lt;/h3&gt;

&lt;p&gt;Every fix PR Self-Healing opens is checked for &lt;code&gt;[Recurrence]&lt;/code&gt; by auto-review. The decision matrix:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Situation&lt;/th&gt;
&lt;th&gt;Required action&lt;/th&gt;
&lt;th&gt;Form&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Same trap stepped on 2+ times&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Lint required&lt;/strong&gt; (custom ESLint rule / type constraint / CI guard)&lt;/td&gt;
&lt;td&gt;Machine (new Guide)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pattern may exist elsewhere&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Horizontal expansion required&lt;/strong&gt; (cpg traversal for similar nodes, fix all of them in this PR)&lt;/td&gt;
&lt;td&gt;Investigation + fix&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cannot be machine-checked but worth formalizing&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Add to an existing guideline&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Guideline entry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One-off, no value in formalization&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Nothing&lt;/strong&gt; (bug fix only)&lt;/td&gt;
&lt;td&gt;--&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;When the "stepped on 2+ times" situation applies, &lt;strong&gt;the fix PR can't merge without a new lint included&lt;/strong&gt;. So every Self-Healing run produces:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Horizontal expansion via cpg&lt;/strong&gt; -- not just the immediate fix target, every similar node enumerated&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A new Guide added in the same PR&lt;/strong&gt; -- ESLint custom rule / type constraint / CI guard / guideline entry, one of the four&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;All existing violations cleared in the same PR&lt;/strong&gt; -- no &lt;code&gt;warn&lt;/code&gt;-as-deferral, &lt;code&gt;error&lt;/code&gt; on first introduction&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Auto-review -&amp;gt; auto-merge -&amp;gt; auto-redeploy&lt;/strong&gt; -- the regular Part 3 pipeline&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Going forward, writing the same pattern gets mechanically rejected by CI / lint&lt;/strong&gt; -- the recurrence class is structurally closed&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F3imbzr9glo3hdw7wqwps.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F3imbzr9glo3hdw7wqwps.png" alt="5 steps every Self-Healing run produces -- recurrence-prevention-first flow"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;"Add the guard while you fix the bug" runs as a self-sustaining loop driven by Self-Healing.&lt;/p&gt;

&lt;h3&gt;
  
  
  "We'll do it later" and "introduce as &lt;code&gt;warn&lt;/code&gt;" are banned
&lt;/h3&gt;

&lt;p&gt;A couple of important contract clauses from the guidelines:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"Plan to lint later," "lint when we refactor," "another PR will handle this" -- &lt;strong&gt;all banned&lt;/strong&gt;. If it can be addressed in this PR, it must be&lt;/li&gt;
&lt;li&gt;"Existing violations remain, so introduce as &lt;code&gt;warn&lt;/code&gt; and promote to &lt;code&gt;error&lt;/code&gt; later" -- &lt;strong&gt;not accepted&lt;/strong&gt;. This is deferral in disguise. The responsibility for the &lt;code&gt;warn&lt;/code&gt;-&amp;gt;&lt;code&gt;error&lt;/code&gt; promotion goes nowhere and the rule rots&lt;/li&gt;
&lt;li&gt;If you add a lint rule, &lt;strong&gt;fix all existing violations in the same PR and ship at &lt;code&gt;error&lt;/code&gt;&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These extend the &lt;strong&gt;no-downgrade rules&lt;/strong&gt; from &lt;a href="https://dev.to/ryantsuji/human-on-the-loop-ai-reviewing-ai-prs-at-cortex-769-prsmonth-while-raising-the-quality-bar-4lh5"&gt;Part 3&lt;/a&gt; -- preempting the typical escape hatches.&lt;/p&gt;

&lt;h3&gt;
  
  
  The "step on it, mechanize it" lineage
&lt;/h3&gt;

&lt;p&gt;Custom Guides currently piled up in cortex:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;graph/no-silent-catch&lt;/code&gt;&lt;/strong&gt; (ESLint) -- the source of the "inflated number" mentioned in the intro. Bans catch blocks that swallow exceptions&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stacktrace-preservation guideline&lt;/strong&gt; (codified as a Major violation in &lt;a href="https://github.com/air-closet/cortex-review-guidelines/blob/main/en/guidelines/observability.md" rel="noopener noreferrer"&gt;&lt;code&gt;observability.md&lt;/code&gt;&lt;/a&gt;, caught by auto-review) -- forbids &lt;code&gt;logger.error(err.message)&lt;/code&gt; style logs that drop the stack and keep only the message string. Forces the &lt;code&gt;err&lt;/code&gt; field to hold &lt;code&gt;serializeError(error)&lt;/code&gt; so &lt;code&gt;name&lt;/code&gt; / &lt;code&gt;message&lt;/code&gt; / &lt;code&gt;stack&lt;/code&gt; are preserved as structured fields. &lt;strong&gt;Observability is everything&lt;/strong&gt; here, so logs that drop stack info are treated as inherently broken&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;cortex-quality/require-fetch-timeout&lt;/code&gt;&lt;/strong&gt; (oxlint -- a Rust-implemented JS/TS lint that runs ESLint-compatible rule sets, dozens of times faster than ESLint due to the Rust impl. cortex uses oxlint for the standard ruleset and ESLint for custom rules that need AST-level work) -- mandates &lt;code&gt;signal: AbortSignal.timeout(...)&lt;/code&gt; on external &lt;code&gt;fetch&lt;/code&gt; calls. Born from a case where a no-timeout &lt;code&gt;fetch&lt;/code&gt; hung indefinitely and triggered a Cloud Tasks redelivery storm&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;graph/no-bq-string-timestamp-param&lt;/code&gt;&lt;/strong&gt; (ESLint) -- from a case where passing TIMESTAMP as a string to a BigQuery query parameter NULLed the value out through a serializer bug and silently failed every INSERT&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;graph/require-firestore-ignore-undefined&lt;/code&gt;&lt;/strong&gt; (ESLint) -- forces &lt;code&gt;ignoreUndefinedProperties: true&lt;/code&gt; on &lt;code&gt;new Firestore()&lt;/code&gt;. From a case where a single NULL row caused a 100% failure rate in a sync batch&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;check-otel-env-injection&lt;/code&gt;&lt;/strong&gt; (CI guard) -- the recurrence prevention for the Cloud Run OTel env injection case below&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TypeScript type tightening&lt;/strong&gt; (type level) -- tighter function signatures, branded types for ID disambiguation, exhaustive discriminated unions, etc. Patterns that can't be lint-caught but are catchable at the type level get closed from the type side&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These aren't textbook-learnable rules -- they're "&lt;strong&gt;stepped on once, then mechanized&lt;/strong&gt;." The number of traps the organization has stepped on translates directly into the number of Guides piled up (across ESLint / oxlint / CI guard / types).&lt;/p&gt;

&lt;h3&gt;
  
  
  How does the AI write a lint rule without breaking it?
&lt;/h3&gt;

&lt;p&gt;Three structural things keep this sane:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Existing rules are the template&lt;/strong&gt;: the custom-rule directory already holds 26 custom rules, each as &lt;code&gt;.ts&lt;/code&gt; + &lt;code&gt;.test.ts&lt;/code&gt; pairs. New rules follow the same shape, so the AI never has to write the AST-walking boilerplate from scratch&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tests first&lt;/strong&gt;: violation / pass fixtures go into &lt;code&gt;.test.ts&lt;/code&gt; first, implementation fills in TDD-style. Coverage threshold (90% statements + branches) is gated by the &lt;a href="https://dev.to/ryantsuji/human-on-the-loop-ai-reviewing-ai-prs-at-cortex-769-prsmonth-while-raising-the-quality-bar-4lh5"&gt;Part 3&lt;/a&gt; auto-review, so a lint without tests cannot merge&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;lint / type / CI guard sit in the same "mechanize" bucket&lt;/strong&gt;: the decision matrix in &lt;a href="https://github.com/air-closet/cortex-review-guidelines/blob/main/en/guidelines/recurrence-prevention.md" rel="noopener noreferrer"&gt;&lt;code&gt;recurrence-prevention.md&lt;/code&gt;&lt;/a&gt; groups lint / type constraint / CI guard together as the "lint-required" row, and leaves the choice within that bucket (write it as a lint? express it at the type level? add a separate CI guard?) to the AI based on how much AST work is involved and whether runtime semantics matter. Traps that need AST inspection but actually hinge on runtime behavior usually end up as a type constraint (branded type / discriminated union / signature tightening) rather than a custom lint&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So "AI writes a lint rule" is supported by &lt;strong&gt;existing rule corpus + the test harness + the mechanize-bucket selection criteria&lt;/strong&gt; -- three together. The path where the AI hand-rolls raw ESLint API and bricks something is structurally closed.&lt;/p&gt;

&lt;h3&gt;
  
  
  A concrete example: Cloud Run OTel env injection -&amp;gt; promoted to CI guard
&lt;/h3&gt;

&lt;p&gt;Multiple services hit this trap: when a Cloud Run Service / Job is defined in Pulumi, forgetting to inject &lt;code&gt;OTEL_EXPORTER_OTLP_ENDPOINT&lt;/code&gt; and &lt;code&gt;GRAFANA_CLOUD_API_KEY&lt;/code&gt; via &lt;code&gt;secretKeyRef&lt;/code&gt; causes OTel init to be skipped in production, no trace/log reaches Grafana, and incidents become silently invisible.&lt;/p&gt;

&lt;p&gt;The normal response would be "we'll be more careful next time." At cortex:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Incident surfaces -&amp;gt; Self-Healing opens a fix PR (adds the env injection to the affected service)&lt;/li&gt;
&lt;li&gt;Auto-review's &lt;code&gt;[Recurrence]&lt;/code&gt; decides "same trap stepped on -&amp;gt; lint required"&lt;/li&gt;
&lt;li&gt;The same PR adds &lt;code&gt;scripts/check-otel-env-injection.ts&lt;/code&gt; (CI guard) -- mechanically asserts OTel env injection across all Cloud Run resource definitions under &lt;code&gt;infra/&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;All other existing services get their env injection added in the same PR&lt;/li&gt;
&lt;li&gt;Merge -&amp;gt; deploy -&amp;gt; any future write of the same kind gets rejected by CI&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That's what "the guardrails grow every time Self-Healing runs" looks like in practice. The trap is "stepped on -&amp;gt; mechanically checked from then on."&lt;/p&gt;

&lt;h3&gt;
  
  
  Where Guides stand right now (in numbers)
&lt;/h3&gt;

&lt;p&gt;Snapshot of cortex's Guide inventory:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Count&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Custom ESLint rules&lt;/strong&gt; (&lt;code&gt;@cortex/eslint-plugin-graph&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;26&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;no-silent-catch&lt;/code&gt; / &lt;code&gt;require-firestore-ignore-undefined&lt;/code&gt; / &lt;code&gt;no-bq-string-timestamp-param&lt;/code&gt; etc.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;CI guards&lt;/strong&gt; (&lt;code&gt;scripts/check-*.ts&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;check-otel-env-injection&lt;/code&gt; / &lt;code&gt;check-cloudscheduler-oidctoken-audience&lt;/code&gt; etc.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Standard oxlint rules&lt;/strong&gt; (set to &lt;code&gt;error&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;183&lt;/td&gt;
&lt;td&gt;Base config ships everything at error&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;TypeScript strict gates&lt;/strong&gt; (baseline)&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;strict&lt;/code&gt; / &lt;code&gt;noImplicitAny&lt;/code&gt; / &lt;code&gt;strictNullChecks&lt;/code&gt; / &lt;code&gt;noUncheckedIndexedAccess&lt;/code&gt; etc.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;TypeScript type tightening&lt;/strong&gt; (per-recurrence)&lt;/td&gt;
&lt;td&gt;grows over time&lt;/td&gt;
&lt;td&gt;branded type / discriminated union / function-signature tightening etc. Patterns that can't be lint-caught but can be type-caught are closed from the type side&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Test coverage thresholds&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;statements + branches 90%&lt;/td&gt;
&lt;td&gt;Uniform across all packages&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Prettier&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1 config&lt;/td&gt;
&lt;td&gt;Format auto-fix&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Guidelines&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;the entire review-guidelines repo&lt;/td&gt;
&lt;td&gt;Used as the decision basis by auto-review&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The first two categories plus the type-tightening row -- &lt;strong&gt;Custom ESLint, CI guard, type tightening&lt;/strong&gt; -- are the part that &lt;strong&gt;compounds over time&lt;/strong&gt; through the &lt;code&gt;[Recurrence]&lt;/code&gt; lens every time Self-Healing or auto-review runs. &lt;strong&gt;The guardrails grow with time.&lt;/strong&gt; That's the substance of the Strengthening layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The whole loop, from the top
&lt;/h2&gt;

&lt;p&gt;When you compose the three layers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[production anomaly] -&amp;gt; Observation layer (OTel/Loki/Grafana) -&amp;gt; Alert firing
                                              ↓
                                       Event Relay -&amp;gt; SSE
                                              ↓
[Self-Healing mode script]
   - claude -p in worktree
   - root cause via cpg + Loki + git blame
   - commit fix
   - (if applicable) add new lint / type gate too
   - gh pr create
                                              ↓
[Auto-review (Part 3)] -- 9 lenses in order, especially [Recurrence] forces
                         recurrence-prevention action (lint / horizontal expansion / guideline entry)
                                              ↓
                          APPROVE + CI green
                                              ↓
[auto-merge -&amp;gt; Turborepo build -&amp;gt; Pulumi parallel deploy]
                                              ↓
[production recovered + same anti-pattern mechanically rejected from now on]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The loop &lt;strong&gt;completes without human intervention&lt;/strong&gt;. Not just repair, but the quality gates that grow with every repair -- that's the "auto-recovery + auto-strengthening" substance at cortex.&lt;/p&gt;

&lt;p&gt;That said, as the front of the article spelled out, &lt;strong&gt;the loop is only viable because cpg and Observability exist&lt;/strong&gt;. cpg makes horizontal expansion possible; Observability turns production anomalies into structured data. With those two in place at the foundation, AI can stand on the side that does Repair and Strengthening. &lt;strong&gt;Self-Healing is not a standalone mechanism. It's a Sensor riding on top of cortex's Guides (cpg + Observability + lint + guidelines).&lt;/strong&gt; That's the single most important framing in this post.&lt;/p&gt;

&lt;h2&gt;
  
  
  Self-Healing by the numbers
&lt;/h2&gt;

&lt;p&gt;Breaking the headline down further.&lt;/p&gt;

&lt;h3&gt;
  
  
  Main firing categories
&lt;/h3&gt;

&lt;p&gt;What kicked off Self-Healing in the past 30 days (with the mapping back to the front-of-post 2 buckets):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Bucket&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Service Error Log Detected&lt;/strong&gt; (most frequent)&lt;/td&gt;
&lt;td&gt;Production-runtime (61 side)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Pipeline Failure&lt;/strong&gt; -- data pipeline failing a configured number of times in a row&lt;/td&gt;
&lt;td&gt;Production-runtime (61 side)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Generator Failure&lt;/strong&gt; -- AI generation jobs (embedding / annotation etc.) failing&lt;/td&gt;
&lt;td&gt;Production-runtime (61 side)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Deploy Failed&lt;/strong&gt; -- deploy step failures (Pulumi up / Cloud Run revision failed)&lt;/td&gt;
&lt;td&gt;Deploy step (54 side)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Alert-firing to production-recovery time
&lt;/h3&gt;

&lt;p&gt;Median &lt;strong&gt;30 minutes to 1 hour&lt;/strong&gt;. Roughly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Alert firing -&amp;gt; AI investigation start: under 1 minute (Event Relay + SSE)&lt;/li&gt;
&lt;li&gt;AI investigation + fix + PR open: 3-8 minutes&lt;/li&gt;
&lt;li&gt;Auto-review (including the &lt;a href="https://dev.to/ryantsuji/human-on-the-loop-ai-reviewing-ai-prs-at-cortex-769-prsmonth-while-raising-the-quality-bar-4lh5"&gt;Part 3&lt;/a&gt; &lt;strong&gt;10.8 review-fix iterations on average&lt;/strong&gt;): 20-45 minutes&lt;/li&gt;
&lt;li&gt;Auto-merge + deploy: 3-10 minutes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Many of these finish before anyone wakes up (alert fires early morning -&amp;gt; by the time people come in, there's just a ✅ in Slack).&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed / Bridge to Part 5
&lt;/h2&gt;

&lt;p&gt;We've now covered &lt;strong&gt;the cortex picture across Parts 1-4&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://dev.to/ryantsuji/building-a-real-ai-harness-auto-reviewed-prs-self-healing-ops-and-non-engineer-contributors-3lfa"&gt;Part 1&lt;/a&gt;: the cortex big picture and harness-engineering framing&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/ryantsuji/the-heart-of-the-ai-harness-a-knowledge-graph-of-the-ai-by-the-ai-for-the-ai-series-part-2-53bm"&gt;Part 2&lt;/a&gt;: Product Graph (cpg) -- the AI's "brain"&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/ryantsuji/human-on-the-loop-ai-reviewing-ai-prs-at-cortex-769-prsmonth-while-raising-the-quality-bar-4lh5"&gt;Part 3&lt;/a&gt;: auto-review -- defending quality at the PR stage&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Part 4 (this post): Self-Healing + Observability + auto-added guardrails -- defending quality in production while growing the quality gates themselves&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The engineering role has shifted, over the last half-year, from "&lt;strong&gt;write&lt;/strong&gt;, &lt;strong&gt;review&lt;/strong&gt;, &lt;strong&gt;fix&lt;/strong&gt;, &lt;strong&gt;merge&lt;/strong&gt;, &lt;strong&gt;deploy&lt;/strong&gt;, &lt;strong&gt;incident-respond&lt;/strong&gt;" -- all of that -- toward &lt;strong&gt;looking at the whole system from above and tuning it&lt;/strong&gt;. &lt;code&gt;human-on-the-loop&lt;/code&gt;, working at the Policy layer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Part 5&lt;/strong&gt; covers the harness reaching the "who writes the code" layer. The center of it is &lt;strong&gt;domain experts (business-side managers, PMOs — non-engineers) opening PRs to production&lt;/strong&gt;, with a concrete walk-through of a +1,742 line / 41 file feature PR that landed with zero human reviewers in the loop. What guarantees the quality is the harness stack built across this series — "whoever writes, the harness owns the quality gate" is the Part 5 framing.&lt;/p&gt;

&lt;p&gt;The toC service expansion gets a brief mention at the end for direction, but the full implementation discussion lives in a separate post.&lt;/p&gt;

&lt;p&gt;The actual series wrap-up is &lt;strong&gt;Part 6&lt;/strong&gt;. The center of it is &lt;strong&gt;the underlying philosophy&lt;/strong&gt; -- why I picked this design, what I gave up, what I kept. Alongside that, since the series so far has been mostly "what's working," I want to look back at the failures and dead ends behind that surface, and the gap between the philosophy and the implementation. A retrospective for myself, and -- hopefully -- a reference for anyone starting down a similar path.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>github</category>
      <category>observability</category>
    </item>
  </channel>
</rss>
