<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Marc</title>
    <description>The latest articles on DEV Community by Marc (@marc_kumiko).</description>
    <link>https://dev.to/marc_kumiko</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4001120%2F74d319bd-196a-45cb-843d-d70b2e6a54c5.png</url>
      <title>DEV Community: Marc</title>
      <link>https://dev.to/marc_kumiko</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/marc_kumiko"/>
    <language>en</language>
    <item>
      <title>A green boot does not prove your key manager is being used</title>
      <dc:creator>Marc</dc:creator>
      <pubDate>Sat, 03 Oct 2026 06:15:00 +0000</pubDate>
      <link>https://dev.to/marc_kumiko/a-green-boot-does-not-prove-your-key-manager-is-being-used-3g7b</link>
      <guid>https://dev.to/marc_kumiko/a-green-boot-does-not-prove-your-key-manager-is-being-used-3g7b</guid>
      <description>&lt;p&gt;Five production apps, three encryption keys each, all stored as base64-encoded plaintext in Kubernetes Secrets. The plan was to move the values into Scaleway's Key Manager so Kubernetes would store only ciphertext and the pod would unwrap it at boot. Straightforward work, one afternoon, maybe two.&lt;/p&gt;

&lt;p&gt;It took three days. Every failure mode had the same shape: something booted fine and was quietly wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  The slot that resolves twice
&lt;/h2&gt;

&lt;p&gt;Our framework has an environment schema for each app. A field can be marked as a key-manager slot:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;KUMIKO_SECRETS_MASTER_KEY_V1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;base64Key32&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;AES-256 master key (KEK) for tenant-secrets encryption.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;meta&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;kumiko&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;kms&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;}),&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At boot, the runtime walks those slots, finds the &lt;code&gt;_CIPHERTEXT&lt;/code&gt; twin for each one, calls the key manager, and hands the app an env with the plaintext filled in. One resolution pass, everything resolved, done.&lt;/p&gt;

&lt;p&gt;Two of our five apps also build their own crypto provider before the framework's boot path runs. They need it to construct a context object that the framework receives later. So the same slot gets resolved twice, from two different env objects, in two different places. Our boot log made this visible only because both passes log with a prefix:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight properties"&gt;&lt;code&gt;&lt;span class="err"&gt;[publicstatus]&lt;/span&gt; &lt;span class="err"&gt;KUMIKO_SECRETS_MASTER_KEY_V1&lt;/span&gt; &lt;span class="py"&gt;source&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;key-manager keyId=441a92a0 region=fr-par&lt;/span&gt;
&lt;span class="err"&gt;[runProdApp]&lt;/span&gt;   &lt;span class="err"&gt;KUMIKO_SECRETS_MASTER_KEY_V1&lt;/span&gt; &lt;span class="py"&gt;source&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;key-manager keyId=441a92a0 region=fr-par&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I had been reading only the second line for a week. Before the fix, the first pass was the one that would have crashed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Plaintext next to ciphertext wins, silently
&lt;/h2&gt;

&lt;p&gt;The first version of the Pulumi code set both values, on the theory that having the plaintext around was a safe fallback during the migration. It is the opposite of safe. The resolver checks for a plaintext value first, and if it finds one it never calls the key manager at all. You get a pod that boots, serves traffic, passes every health check, and has never once talked to the KMS you just spent two days wiring up.&lt;/p&gt;

&lt;p&gt;So the deployment code is either/or, with a comment that says why:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Either/or, never both: a plaintext slot left beside its ciphertext&lt;/span&gt;
&lt;span class="c1"&gt;// wins silently, so the pod would boot green while still bypassing the&lt;/span&gt;
&lt;span class="c1"&gt;// Key Manager.&lt;/span&gt;
&lt;span class="p"&gt;...(&lt;/span&gt;&lt;span class="nx"&gt;secretsMasterKeyCiphertext&lt;/span&gt;
  &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;KUMIKO_SECRETS_MASTER_KEY_V1_CIPHERTEXT&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;secretsMasterKeyCiphertext&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;KUMIKO_SECRETS_MASTER_KEY_V1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;masterKey&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;base64&lt;/span&gt; &lt;span class="p"&gt;}),&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you are doing this migration, make the two states mutually exclusive in code before you touch any environment. A fallback that can win silently is worse than no fallback.&lt;/p&gt;

&lt;h2&gt;
  
  
  The schema ate the key before the resolver saw it
&lt;/h2&gt;

&lt;p&gt;The second pass resolves against the &lt;em&gt;parsed&lt;/em&gt; env, and Zod's default object parsing drops unknown keys. The &lt;code&gt;_CIPHERTEXT&lt;/code&gt; twin is generated by the framework for the composed schema, but each app also hand builds a typed schema on top. If the twin is not declared there, the parsed env simply does not contain it, and the resolver gets an env where the ciphertext is missing and the plaintext is still absent. It throws, the pod crashloops, and the error message talks about a missing key rather than a stripped one.&lt;/p&gt;

&lt;p&gt;The fix is one line per app, and it needs an explanation, because the next person will read it as redundant with the framework's generated twin and delete it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;KUMIKO_SECRETS_MASTER_KEY_V1_CIPHERTEXT&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;string&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;optional&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Key-Manager ciphertext, wrapped by the same key as the KEK. Must stay in this schema: bin/main.ts hands the PARSED env to resolvePlatformKeks, which never sees an undeclared key.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I actually deleted it myself, mid migration, after convincing myself it was dead code. It was not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Zod 4 throws on extend for refined schemas
&lt;/h2&gt;

&lt;p&gt;When a slot arrives as ciphertext only, the plaintext field has to be relaxed to optional before parsing, otherwise the required check rejects exactly the deployment the feature exists to enable. The obvious implementation is &lt;code&gt;schema.extend({ [name]: field.optional() })&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That works for four of our apps. The fifth one has an object-wide &lt;code&gt;superRefine&lt;/code&gt; for a cross-field rule, and Zod 4 throws on &lt;code&gt;.extend()&lt;/code&gt; for any schema carrying a refinement. It failed at boot, in production, on a Friday. The version that works uses &lt;code&gt;safeExtend&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nb"&gt;Object&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;keys&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;relaxed&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nx"&gt;schema&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;schema&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;safeExtend&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;relaxed&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A framework helper that manipulates a user-supplied schema has to be tested against a refined schema. That is the case that behaves differently, and it is never the case in your fixtures.&lt;/p&gt;

&lt;h2&gt;
  
  
  The app that did not mount the feature but needed the key
&lt;/h2&gt;

&lt;p&gt;One app never mounts the tenant secrets feature, so its schema had no reason to know about the master key slot. It does use multifactor authentication, and the MFA code envelope encrypts TOTP secrets with the same provider, pulling the key straight out of &lt;code&gt;process.env&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Name-based greps did not find it. The provider resolves its keys by regex over the environment:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;match&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sr"&gt;/^KUMIKO_SECRETS_MASTER_KEY_V&lt;/span&gt;&lt;span class="se"&gt;(\d&lt;/span&gt;&lt;span class="sr"&gt;+&lt;/span&gt;&lt;span class="se"&gt;)&lt;/span&gt;&lt;span class="sr"&gt;$/&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exec&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So no consumer ever types the key name, and a grep for the name finds nothing in the app that most needed the change. What found it was grepping for the provider's constructor instead. If you are auditing who consumes a secret, grep for the thing that reads it, not for the name of the thing being read.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verifying that the key did not change
&lt;/h2&gt;

&lt;p&gt;The migration is safe only if the wrapped value decrypts to the exact same key bytes. Existing rows in the database were encrypted with that key, and a key that is merely valid will decrypt nothing.&lt;/p&gt;

&lt;p&gt;The check is narrow. Encrypt the live value, then decrypt the fresh ciphertext and compare the API's raw &lt;code&gt;plaintext&lt;/code&gt; field with the source string, using the same canonical encoding. Do not compare your own encode of the decode; that proves only that your codec is symmetric. We had already been burned by exactly that once, which is why the wrap script aborts unless the raw fields match:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$roundtrip&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$plaintext&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
  &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"!! &lt;/span&gt;&lt;span class="nv"&gt;$app&lt;/span&gt;&lt;span class="s2"&gt;: decrypt of the fresh ciphertext does NOT match the source value"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&amp;amp;2
  &lt;span class="nb"&gt;exit &lt;/span&gt;1
&lt;span class="k"&gt;fi&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  A monitor that reported green on an unchecked commit
&lt;/h2&gt;

&lt;p&gt;Small one, but it nearly cost me a bad merge. I had a poll loop watching four pull requests, filtering the GitHub check rollup for anything failing and anything pending. No failures and nothing pending reads as green.&lt;/p&gt;

&lt;p&gt;Right after a force push, the rollup is momentarily empty, so for about thirty seconds a branch with zero checks run against it satisfies both conditions. My monitor announced green on a commit that CI had not looked at yet. The rewrite uses the exit code of &lt;code&gt;gh pr checks&lt;/code&gt; instead, where 0 is passed and 8 is pending, so an empty rollup cannot masquerade as success.&lt;/p&gt;

&lt;p&gt;Whenever you derive a positive state from the absence of negatives, ask what an empty input looks like.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually went wrong on rollout day
&lt;/h2&gt;

&lt;p&gt;Two other things went wrong, both unrelated to keys.&lt;/p&gt;

&lt;p&gt;Four apps rolled at once and hit &lt;code&gt;remaining connection slots are reserved for roles with the SUPERUSER attribute&lt;/code&gt;, because the old pods still held their Postgres connections while the new ones started. One app restarted and came back. Worth staging your rollouts if your connection ceiling is anywhere near your replica count times your pool size.&lt;/p&gt;

&lt;p&gt;The rollback path was also different from what I assumed. The Pulumi code reads the ciphertext with &lt;code&gt;config.require(...)&lt;/code&gt;, so removing the config value does not revert anything. It makes the next &lt;code&gt;pulumi up&lt;/code&gt; throw before it can do any work. The real rollback is to deploy the previous revision of the infrastructure code. That is a thing to work out before you need it rather than during.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it landed
&lt;/h2&gt;

&lt;p&gt;All five apps now log &lt;code&gt;source=key-manager&lt;/code&gt; for all three key slots. The two apps with the extra boot path log it in both resolution passes, and the plaintext slot is gone from every Kubernetes secret. The first app went out alone with a targeted apply, the other four followed once its boot log confirmed the path.&lt;/p&gt;

&lt;p&gt;The part I would repeat is the log line that names its own pass. Almost every wrong turn above showed up in that log before it showed up anywhere else, and one showed up only there.&lt;/p&gt;

</description>
      <category>security</category>
      <category>devops</category>
      <category>kubernetes</category>
      <category>typescript</category>
    </item>
    <item>
      <title>We moved OCR out of the app so egress could be zero</title>
      <dc:creator>Marc</dc:creator>
      <pubDate>Wed, 30 Sep 2026 08:15:00 +0000</pubDate>
      <link>https://dev.to/marc_kumiko/we-moved-ocr-out-of-the-app-so-egress-could-be-zero-17m6</link>
      <guid>https://dev.to/marc_kumiko/we-moved-ocr-out-of-the-app-so-egress-could-be-zero-17m6</guid>
      <description>&lt;p&gt;Our document ingest path used to run OCR inside the app pods. The library wanted glibc. The app images were Alpine. Missing language packs got pulled from github.com at parse time. That last part is the one that kept me up: a PDF upload could trigger an outbound fetch from a pod that otherwise had no business talking to the public internet.&lt;/p&gt;

&lt;p&gt;So we moved parsing into its own service, baked the traineddata into the image, and closed egress entirely. Including DNS.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the app was the wrong place
&lt;/h2&gt;

&lt;p&gt;Three separate failures stacked in one process:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The native liteparse binary didn't load on Alpine. You discover that in the target environment, not in a glibc laptop build.&lt;/li&gt;
&lt;li&gt;Tessdata that wasn't already on disk downloaded at runtime. OCR "just worked" until the network policy, the rate limit, or the broken mirror made it not work.&lt;/li&gt;
&lt;li&gt;The parse library accepted path-shaped options (&lt;code&gt;tessdataPath&lt;/code&gt;, &lt;code&gt;imageOutputDir&lt;/code&gt;, &lt;code&gt;ocrServerUrl&lt;/code&gt;). A caller that could set those could point the parser at another filesystem or another host.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;None of those are model problems. They're packaging and trust-boundary problems that show up the first time you treat document upload as untrusted input, which it is.&lt;/p&gt;

&lt;h2&gt;
  
  
  One HTTP service, three caller knobs
&lt;/h2&gt;

&lt;p&gt;The service is a small Bun HTTP app: multipart upload in, pages of text out. Callers may pass only:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strictObject&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;ocrEnabled&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;boolean&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;optional&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
  &lt;span class="na"&gt;ocrLanguage&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;string&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;regex&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/^&lt;/span&gt;&lt;span class="se"&gt;[&lt;/span&gt;&lt;span class="sr"&gt;a-z&lt;/span&gt;&lt;span class="se"&gt;]{3}(\+[&lt;/span&gt;&lt;span class="sr"&gt;a-z&lt;/span&gt;&lt;span class="se"&gt;]{3})&lt;/span&gt;&lt;span class="sr"&gt;*$/&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;optional&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
  &lt;span class="na"&gt;maxPages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;number&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;positive&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;optional&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;strictObject&lt;/code&gt; rejects every other field by name. &lt;code&gt;tessdataPath&lt;/code&gt; is not an option; it comes from &lt;code&gt;LITEPARSE_TESSDATA_PATH&lt;/code&gt; and nowhere else. Languages are checked against the &lt;code&gt;.traineddata&lt;/code&gt; files present at startup. Ask for &lt;code&gt;fra&lt;/code&gt; when only &lt;code&gt;eng&lt;/code&gt; and &lt;code&gt;deu&lt;/code&gt; are baked in and you get &lt;code&gt;400&lt;/code&gt;, not a download.&lt;/p&gt;

&lt;p&gt;The image pins &lt;code&gt;deu&lt;/code&gt; and &lt;code&gt;eng&lt;/code&gt; from &lt;code&gt;tesseract-ocr/tessdata_best&lt;/code&gt; at a fixed commit SHA with sha256 verification. Build time is when we fetch. Runtime is when we refuse to.&lt;/p&gt;

&lt;p&gt;Auth is a shared bearer token, sixteen characters minimum. Health checks stay unauthenticated so Kubernetes can probe without holding the secret. There is no &lt;code&gt;/metrics&lt;/code&gt; scrape path, because nothing scrapes it yet and an open metrics port would be another thing to reason about.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closed egress is the point
&lt;/h2&gt;

&lt;p&gt;In the cluster the service sits in its own namespace. The NetworkPolicy sets &lt;code&gt;egress: []&lt;/code&gt; — no outbound traffic, DNS included. Ingress is an allowlist of caller namespaces from stack config. Empty list means deny-all until the first consumer is declared.&lt;/p&gt;

&lt;p&gt;That shape only works because the container never needs the network. Puppeteer, sitting next to it in our stack, cannot do the same: it renders user HTML with external subresources. OCR can, once the models live in the image.&lt;/p&gt;

&lt;p&gt;App pods call &lt;code&gt;http://liteparse-service…:3000&lt;/code&gt; over the cluster network with the shared token. They never embed the native binary. Alpine stays Alpine. The parse blast radius is one dedicated Deployment with a memory tier and no way out.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we'd do again
&lt;/h2&gt;

&lt;p&gt;Pin traineddata in the image and fail closed on unknown languages. Whitelist request options so a multipart field cannot redirect filesystem or network paths. Put the NetworkPolicy next to the Deployment in the same PR, with ingress empty by default so a forgotten consumer config doesn't open the service to the whole cluster.&lt;/p&gt;

&lt;p&gt;If your OCR still downloads language packs on first use, you don't have an offline parser. You have a deferred &lt;code&gt;curl&lt;/code&gt; with a PDF as the trigger.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>kubernetes</category>
      <category>security</category>
    </item>
    <item>
      <title>Undocumented handlers don't exist to the agent</title>
      <dc:creator>Marc</dc:creator>
      <pubDate>Sun, 27 Sep 2026 06:10:00 +0000</pubDate>
      <link>https://dev.to/marc_kumiko/undocumented-handlers-dont-exist-to-the-agent-e71</link>
      <guid>https://dev.to/marc_kumiko/undocumented-handlers-dont-exist-to-the-agent-e71</guid>
      <description>&lt;p&gt;We wired an in-app AI agent into a product that already had a full write/query surface. Auth was fine. The agent had permission to call writes. Half the handlers it needed simply never showed up in its tool list.&lt;/p&gt;

&lt;p&gt;The agent said it couldn't find them. The handlers existed. What was missing was a description string.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the agent can see
&lt;/h2&gt;

&lt;p&gt;Our agent doesn't get raw TypeScript. It gets a manifest built from every feature the app mounts: handlers, screens, entities. Each entry needs a short description, written at the place the handler is defined. Without that string, the entry is dropped. The agent never learns the name exists.&lt;/p&gt;

&lt;p&gt;That surprised a few people, including me the first time. We treated descriptions as documentation polish. For the agent they are the gate. An endpoint without an OpenAPI description might still be callable if you know the path. An agent-facing handler without a description isn't callable at all, because the model never sees a tool to choose.&lt;/p&gt;

&lt;p&gt;Twenty-one handlers across our AI pipeline and prompt store features were in that state. The code ran for humans in the UI. The agent walked past them.&lt;/p&gt;

&lt;p&gt;You also can't patch this later at the composition site. The description has to live next to the handler definition, so every app that mounts the feature inherits it. We learned that the hard way: people tried to "add docs" in the app and the gap tests kept failing, because the packaged feature still shipped blank.&lt;/p&gt;

&lt;h2&gt;
  
  
  Always allow is a permission, not a mood
&lt;/h2&gt;

&lt;p&gt;Visibility is half of it. The other half is what happens after the agent picks a tool.&lt;/p&gt;

&lt;p&gt;Writes default to risk &lt;code&gt;mid&lt;/code&gt;. Queries default to &lt;code&gt;low&lt;/code&gt;. The UI can offer "always allow" for those. Two of our writes are different:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;delete-golden&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;description&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Permanently deletes a golden fixture, so it can no longer be used for dry-runs. This is a hard delete, not an appended revision, and cannot be undone.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nl"&gt;risk&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;high&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="c1"&gt;// ...&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;edit&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;description&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Saves new content for a prompt template by appending a revision that becomes the active prompt immediately… The owning AI feature uses this content verbatim as its system prompt…&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nl"&gt;risk&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;high&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="c1"&gt;// ...&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;delete-golden&lt;/code&gt; is a hard delete. Prompt &lt;code&gt;edit&lt;/code&gt; is free text that becomes another feature's system prompt on the next run. Both are marked &lt;code&gt;high&lt;/code&gt;. The permission loop refuses &lt;code&gt;always&lt;/code&gt; for those handlers outright. You can still approve a single call. You cannot teach the agent that this shape of write is permanently fine.&lt;/p&gt;

&lt;p&gt;Most of the neighbouring writes append a restorable revision (&lt;code&gt;set-policy&lt;/code&gt;, &lt;code&gt;rollback&lt;/code&gt;, prompt &lt;code&gt;revert&lt;/code&gt;). Those stay at the default. Restorable side effects and irreversible ones share a form in the code; they don't share a permission ceiling.&lt;/p&gt;

&lt;p&gt;Phone apps already do this split. Camera access can be "always". Wipe storage can't. We needed the same ceiling for an agent that can press buttons faster than a human reviews them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make the gap fail a test
&lt;/h2&gt;

&lt;p&gt;Once you decide "no description means invisible", silence becomes the failure mode. The agent looks underpowered. Nobody opens a ticket that says "missing string on handler 14".&lt;/p&gt;

&lt;p&gt;So each feature that should be agent-visible gets a gap test:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;createAiPipelineFeature([]) exposes 0 doc gaps&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;gaps&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;findAgentDocGaps&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="nf"&gt;createAiPipelineFeature&lt;/span&gt;&lt;span class="p"&gt;([])]);&lt;/span&gt;
  &lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;gaps&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toEqual&lt;/span&gt;&lt;span class="p"&gt;([]);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;delete-golden is high risk, set-policy is mid risk&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;feature&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;createAiPipelineFeature&lt;/span&gt;&lt;span class="p"&gt;([]);&lt;/span&gt;
  &lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;resolveAgentExposure&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;feature&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;writeHandlers&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;delete-golden&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;write&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;risk&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toBe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;high&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;resolveAgentExposure&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;feature&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;writeHandlers&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;set-policy&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;write&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;risk&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toBe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;mid&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first test is the smoke alarm for missing descriptions. The second pins the two risk decisions we actually care about, so a refactor can't quietly demote a hard delete back to "always allow"-eligible.&lt;/p&gt;

&lt;p&gt;We also run the same gap lint from a CLI against a whole app config. Shipping a new consumer feature without descriptions turns red before anyone asks the agent to do a demo.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we tell ourselves now
&lt;/h2&gt;

&lt;p&gt;If a human can click it and an agent should be able to call it, write the description where the handler is defined. One sentence that says what changes and what stays reversible is enough; our pipeline descriptions are longer because the agent needs the follow-up query names (&lt;code&gt;revisions&lt;/code&gt; before &lt;code&gt;activate&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;If the write can't be undone, or it rewrites another feature's brain, mark it &lt;code&gt;high&lt;/code&gt;. Don't rely on the operator to remember which "always allow" clicks were a bad idea.&lt;/p&gt;

&lt;p&gt;If you only remember one check: ask the agent to list what it can do, then compare that list to your write handlers. The missing names are usually missing strings, not missing models.&lt;/p&gt;

&lt;p&gt;I wrote earlier about &lt;a href="https://dev.to/marc_kumiko/the-model-obeys-your-schema-not-your-description-1cml"&gt;schemas beating prose for tool outputs&lt;/a&gt;. This is the other side of the same habit. There the description was advice and the schema was the contract. Here the description &lt;em&gt;is&lt;/em&gt; the contract for whether the tool exists at all, and risk is the contract for whether "always" is even legal.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>typescript</category>
      <category>agents</category>
    </item>
    <item>
      <title>A green boot does not prove your key manager is being used</title>
      <dc:creator>Marc</dc:creator>
      <pubDate>Fri, 25 Sep 2026 08:11:00 +0000</pubDate>
      <link>https://dev.to/marc_kumiko/a-green-boot-does-not-prove-your-key-manager-is-being-used-119h</link>
      <guid>https://dev.to/marc_kumiko/a-green-boot-does-not-prove-your-key-manager-is-being-used-119h</guid>
      <description>&lt;p&gt;Five production apps, three encryption keys each, all of them sitting in a Kubernetes secret as base64 plaintext. The plan was to move them into Scaleway's Key Manager so the cluster only ever holds ciphertext and the pod unwraps it at boot. Straightforward work, one afternoon, maybe two.&lt;/p&gt;

&lt;p&gt;It took three days, and every single failure mode was the same shape: something booted fine and was quietly wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  The slot that resolves twice
&lt;/h2&gt;

&lt;p&gt;Our framework has a env schema per app. A field can be marked as a key manager slot:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;KUMIKO_SECRETS_MASTER_KEY_V1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;base64Key32&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;AES-256 master key (KEK) for tenant-secrets encryption.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;meta&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;kumiko&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;kms&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;}),&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At boot, the runtime walks those slots, finds the &lt;code&gt;_CIPHERTEXT&lt;/code&gt; twin for each one, calls the key manager, and hands the app an env with the plaintext filled in. One round trip, everything resolved, done.&lt;/p&gt;

&lt;p&gt;Except two of our five apps also build their own crypto provider before the framework's boot path runs, because they need it to construct a context object the framework then receives. So the same slot gets resolved twice, from two different env objects, in two different places. Our boot log made this visible only because both passes log with a prefix:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[publicstatus] KUMIKO_SECRETS_MASTER_KEY_V1 source=key-manager keyId=441a92a0 region=fr-par
[runProdApp]   KUMIKO_SECRETS_MASTER_KEY_V1 source=key-manager keyId=441a92a0 region=fr-par
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I had been reading only the second line for a week. The first one is the one that would have crashed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Plaintext next to ciphertext wins, silently
&lt;/h2&gt;

&lt;p&gt;The first version of the Pulumi code set both values, on the theory that having the plaintext around was a safe fallback during the migration. It is the opposite of safe. The resolver checks for a plaintext value first, and if it finds one it never calls the key manager at all. You get a pod that boots, serves traffic, passes every health check, and has never once talked to the KMS you just spent two days wiring up.&lt;/p&gt;

&lt;p&gt;So the deployment code is either/or, with a comment that says why:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Either/or, never both: a plaintext slot left beside its ciphertext&lt;/span&gt;
&lt;span class="c1"&gt;// wins silently, so the pod would boot green while still bypassing the&lt;/span&gt;
&lt;span class="c1"&gt;// Key Manager.&lt;/span&gt;
&lt;span class="p"&gt;...(&lt;/span&gt;&lt;span class="nx"&gt;secretsMasterKeyCiphertext&lt;/span&gt;
  &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;KUMIKO_SECRETS_MASTER_KEY_V1_CIPHERTEXT&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;secretsMasterKeyCiphertext&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;KUMIKO_SECRETS_MASTER_KEY_V1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;masterKey&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;base64&lt;/span&gt; &lt;span class="p"&gt;}),&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you are doing this migration, make the two states mutually exclusive in code before you touch any environment. A fallback that can win without telling you is worse than no fallback.&lt;/p&gt;

&lt;h2&gt;
  
  
  The schema ate the key before the resolver saw it
&lt;/h2&gt;

&lt;p&gt;The second pass resolves against the &lt;em&gt;parsed&lt;/em&gt; env, and Zod drops unknown keys. The &lt;code&gt;_CIPHERTEXT&lt;/code&gt; twin is generated by the framework for the composed schema, but each app also hand builds a typed schema on top. If the twin is not declared there, the parsed env simply does not contain it, and the resolver gets an env where the ciphertext does not exist and the plaintext is empty. It throws, the pod crashloops, and the error message talks about a missing key rather than a stripped one.&lt;/p&gt;

&lt;p&gt;The fix is one line per app, and it needs a comment, because the next person will read it as redundant with the framework's generated twin and delete it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;KUMIKO_SECRETS_MASTER_KEY_V1_CIPHERTEXT&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;string&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;optional&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Key-Manager ciphertext, wrapped by the same key as the KEK. Must stay in this schema: bin/main.ts hands the PARSED env to resolvePlatformKeks, which never sees an undeclared key.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I actually deleted it myself, mid migration, after convincing myself it was dead code. It was not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Zod 4 throws on extend for refined schemas
&lt;/h2&gt;

&lt;p&gt;When a slot arrives as ciphertext only, the plaintext field has to be relaxed to optional before parsing, otherwise the required check rejects exactly the deployment the feature exists to enable. The obvious implementation is &lt;code&gt;schema.extend({ [name]: field.optional() })&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That works for four of our apps. The fifth one has an object wide &lt;code&gt;superRefine&lt;/code&gt; for a cross field rule, and Zod 4 throws on &lt;code&gt;.extend()&lt;/code&gt; for any schema carrying a refinement. It failed at boot, in production, on a Friday. The version that works uses &lt;code&gt;safeExtend&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nb"&gt;Object&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;keys&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;relaxed&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nx"&gt;schema&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;schema&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;safeExtend&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;relaxed&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A framework helper that manipulates a user supplied schema has to be tested against a refined schema. That is the case that behaves differently, and it is never the case in your fixtures.&lt;/p&gt;

&lt;h2&gt;
  
  
  The app that did not mount the feature but needed the key
&lt;/h2&gt;

&lt;p&gt;One app never mounts the tenant secrets feature, so its schema had no reason to know about the master key slot. It does use multi factor auth, and the MFA code envelope encrypts TOTP secrets with the same provider, pulling the key straight out of &lt;code&gt;process.env&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Name based greps did not find it. The provider resolves its keys by regex over the environment:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;match&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sr"&gt;/^KUMIKO_SECRETS_MASTER_KEY_V&lt;/span&gt;&lt;span class="se"&gt;(\d&lt;/span&gt;&lt;span class="sr"&gt;+&lt;/span&gt;&lt;span class="se"&gt;)&lt;/span&gt;&lt;span class="sr"&gt;$/&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exec&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So no consumer ever types the key name, and a grep for the name finds nothing in the one app that most needed the change. What found it was grepping for the provider's constructor instead. If you are auditing who consumes a secret, grep for the thing that reads it, not for the name of the thing being read.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verifying that the key did not change
&lt;/h2&gt;

&lt;p&gt;The whole migration is only safe if the wrapped value decrypts back to the exact same bytes. Existing rows in the database were encrypted with that key, and a key that is merely valid will decrypt nothing.&lt;/p&gt;

&lt;p&gt;The check that proves this is narrow. Encrypt the live value, then decrypt the fresh ciphertext and compare the raw &lt;code&gt;plaintext&lt;/code&gt; field from the API response against the source string. Do not compare your own encode of the decode, which passes at any nesting depth and proves only that your codec is symmetric. We had already been burned by exactly that once, which is why the wrap script aborts unless the raw fields match:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$roundtrip&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$plaintext&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
  &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"!! &lt;/span&gt;&lt;span class="nv"&gt;$app&lt;/span&gt;&lt;span class="s2"&gt;: decrypt of the fresh ciphertext does NOT match the source value"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&amp;amp;2
  &lt;span class="nb"&gt;exit &lt;/span&gt;1
&lt;span class="k"&gt;fi&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  A monitor that reported green on an unchecked commit
&lt;/h2&gt;

&lt;p&gt;Small one, but it nearly cost me a bad merge. I had a poll loop watching four pull requests, filtering the GitHub check rollup for anything failing and anything pending. No failures and nothing pending reads as green.&lt;/p&gt;

&lt;p&gt;Right after a force push the rollup is momentarily empty, so for about thirty seconds a branch with zero checks run against it satisfies both conditions. My monitor announced green on a commit that CI had not looked at yet. The rewrite uses the exit code of &lt;code&gt;gh pr checks&lt;/code&gt; instead, where 0 is passed and 8 is pending, so an empty rollup cannot masquerade as success.&lt;/p&gt;

&lt;p&gt;Whenever you derive a positive state from the absence of negatives, ask what an empty input looks like.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually went wrong on rollout day
&lt;/h2&gt;

&lt;p&gt;Two things, neither of them about keys.&lt;/p&gt;

&lt;p&gt;Four apps rolled at once and hit &lt;code&gt;remaining connection slots are reserved for roles with the SUPERUSER attribute&lt;/code&gt;, because the old pods still held their Postgres connections while the new ones started. One app restarted and came back. Worth staging your rollouts if your connection ceiling is anywhere near your replica count times your pool size.&lt;/p&gt;

&lt;p&gt;And the rollback path was not what I assumed. The Pulumi code reads the ciphertext with &lt;code&gt;config.require(...)&lt;/code&gt;, so removing the config value does not revert anything, it makes the next &lt;code&gt;pulumi up&lt;/code&gt; throw before it can do any work. The real rollback is deploying the previous revision of the infrastructure code. That is a thing to work out before you need it rather than during.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it landed
&lt;/h2&gt;

&lt;p&gt;All five apps now log &lt;code&gt;source=key-manager&lt;/code&gt; for all three key slots, in both resolution passes, and the plaintext slot is gone from every Kubernetes secret. The first app went out alone with a targeted apply, the other four followed once its boot log confirmed the path.&lt;/p&gt;

&lt;p&gt;The part I would repeat is the log line that names its own pass. Almost every wrong turn above was visible in that log before it was visible anywhere else, and one of them was only visible there.&lt;/p&gt;

</description>
      <category>security</category>
      <category>devops</category>
      <category>kubernetes</category>
      <category>typescript</category>
    </item>
    <item>
      <title>Your AI eval is green because it never called the model</title>
      <dc:creator>Marc</dc:creator>
      <pubDate>Tue, 22 Sep 2026 06:57:00 +0000</pubDate>
      <link>https://dev.to/marc_kumiko/your-ai-eval-is-green-because-it-never-called-the-model-13k8</link>
      <guid>https://dev.to/marc_kumiko/your-ai-eval-is-green-because-it-never-called-the-model-13k8</guid>
      <description>&lt;p&gt;We run an eval suite on every PR. Seven fixtures, four metrics, deterministic, no API key, no network, costs nothing. It was green for weeks.&lt;/p&gt;

&lt;p&gt;Then a code review found that two of the fixtures contained responses the live tool schema would have rejected outright. One had &lt;code&gt;access&lt;/code&gt; nested inside a &lt;code&gt;definition&lt;/code&gt; wrapper when the real pattern shape has it at the top level. Another declared a field as &lt;code&gt;"json"&lt;/code&gt;, which isn't a valid field type in our framework at all; the real one is &lt;code&gt;"jsonb"&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;These are the canned responses the mock provider hands back. They were never checked against the schema the live model has to satisfy, because nothing in a mock run ever talks to a schema validator. So the suite was green on outputs that could not physically occur in production.&lt;/p&gt;

&lt;p&gt;Every "test your LLM in CI" setup makes this trade, and it's worth stating plainly: a mocked eval tests your pipeline, never your model. Once you accept that, the mock gets more useful, because you stop asking it for something it can't give.&lt;/p&gt;

&lt;h2&gt;
  
  
  One runner, two providers
&lt;/h2&gt;

&lt;p&gt;The entire trick is that the runner doesn't know which mode it's in. It takes a provider as a dependency and applies the same metric pipeline either way:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;RunEvalOptions&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="na"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;LLMProvider&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="na"&gt;fixtures&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="nx"&gt;EvalFixture&lt;/span&gt;&lt;span class="p"&gt;[];&lt;/span&gt;
  &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="nx"&gt;metrics&lt;/span&gt;&lt;span class="p"&gt;?:&lt;/span&gt; &lt;span class="nb"&gt;Readonly&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nb"&gt;Record&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;MetricFn&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="na"&gt;mode&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;mock&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;live&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="nx"&gt;generatedAt&lt;/span&gt;&lt;span class="p"&gt;?:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;mode&lt;/code&gt; is metadata. It lands in the report so you can tell a mock baseline from a live run later, and it changes no behavior. Mode selection lives in the script entry point, which is the only place that knows whether &lt;code&gt;--live&lt;/code&gt; was passed.&lt;/p&gt;

&lt;p&gt;For mock runs, the fixtures script themselves:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;runMockEval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;options&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;RunMockEvalOptions&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;EvalReport&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;provider&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;createMockProvider&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="c1"&gt;// Script every fixture's mockResponse in fixture-order. Runner&lt;/span&gt;
  &lt;span class="c1"&gt;// calls provider.chat() once per fixture, also in fixture-order,&lt;/span&gt;
  &lt;span class="c1"&gt;// so FIFO matches.&lt;/span&gt;
  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;fixture&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;options&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;fixtures&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;script&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;fixture&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;mockResponse&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;runEval&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt;&lt;span class="nx"&gt;options&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;mode&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;mock&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each fixture carries its own &lt;code&gt;mockResponse&lt;/code&gt; next to its expected outcome, so there's no separate mock-fixture directory to keep in sync with the real one. The provider is a FIFO queue and the runner is sequential, so the ordering holds without any matching logic.&lt;/p&gt;

&lt;p&gt;Sequential is deliberate, by the way. Anthropic rate-limits, the eval set is seven fixtures, and parallelism would buy a few seconds in exchange for 429s in live mode.&lt;/p&gt;

&lt;h2&gt;
  
  
  Baseline diffs in review
&lt;/h2&gt;

&lt;p&gt;A test suite that says "still green" after a prompt change is nearly worthless for this kind of work. Prompt and schema changes rarely break things outright, they shift them, and what you want to see is which fixture moved and by how much.&lt;/p&gt;

&lt;p&gt;So the mock report goes into a checked-in JSON baseline, and a test compares a fresh mock run against it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;live&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;summary&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;totalFixtures&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toBe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;baseline&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;summary&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;totalFixtures&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;live&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;summary&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;passing&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toBe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;baseline&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;summary&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;passing&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;live&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;summary&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;failing&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toBe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;baseline&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;summary&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;failing&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;live&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;summary&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;meanScore&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toBeCloseTo&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;baseline&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;summary&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;meanScore&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Plus per-fixture: same id set, same pass state, same aggregate score, same per-metric pass state.&lt;/p&gt;

&lt;p&gt;Edit a fixture, a metric, the system prompt, or a tool definition, and this test goes red until you regenerate the baseline. The regenerated baseline is a file in your diff, so the reviewer sees &lt;code&gt;parse-error&lt;/code&gt; flipping from true to false on one fixture, in the same PR that changed the tool schema. Quality changes become visible artifacts of review instead of something you notice three weeks later in production.&lt;/p&gt;

&lt;p&gt;It also means "regenerate the baseline" is a normal, expected step, which is the part people resist. If regenerating feels like cheating, the baseline is doing its job: you have to look at the diff and decide the change was intended.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two ways we got this wrong
&lt;/h2&gt;

&lt;p&gt;The first was comparing the reason strings. Every metric result carries a human-readable message alongside its pass state and score. Comparing those in the drift test is tempting and wrong. The text drifts whenever someone improves an error message, and then the drift test fails for a reason that has nothing to do with model quality. After the third spurious failure people stop reading the diff, which kills the only thing the test is for. We compare structural data only: pass state, score, per-metric pass state. If a message needs to be load-bearing, it gets a dedicated metric test.&lt;/p&gt;

&lt;p&gt;The second was letting the script pick the baseline filename. Live runs and mock runs write to different files, for the obvious reason that a live run must never clobber the deterministic baseline. That selection originally lived in a side-effectful script body, untested, where a swapped ternary would quietly destroy the mock baseline on the next live run. It's now one function with three tests, the third of which exists purely to state the invariant:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;baselineFileName&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;isLive&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;boolean&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;isLive&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;l2-eval-live-baseline.json&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;l2-eval-baseline.json&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;the two modes never resolve to the same file&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;baselineFileName&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nx"&gt;not&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toBe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;baselineFileName&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A test that asserts two constants differ looks silly right up until someone refactors the ternary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pass gates, score tracks
&lt;/h2&gt;

&lt;p&gt;Every metric returns both a boolean and a 0..1 score, and the split matters more than it looks.&lt;/p&gt;

&lt;p&gt;Pass is what gates CI. Score is what shows you a metric moving before it flips. Coverage of generated pattern kinds is a good example: four of five expected kinds found is 0.8, not zero, and not a pass.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;score&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;found&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nx"&gt;expected&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;pass&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;missing&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The parse metric does the same, degrading instead of going binary. Clean parse is 1.0, N parse errors is &lt;code&gt;max(0, 1 - N / 10)&lt;/code&gt;, and a parser that throws is 0. So a change that takes a fixture from seven parse errors to two shows up as real movement even though it failed both times. Binary metrics hide exactly the progress you need to see while you're iterating on a prompt.&lt;/p&gt;

&lt;p&gt;That metric parses the emitted source with the framework's real parser against an in-memory &lt;code&gt;SourceFile&lt;/code&gt;. It's slower than a regex and it catches the case that matters: the model emits TypeScript that looks fine and our parser rejects, because it used factory style instead of the canonical form, or dropped the schema version header.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the mock can't do
&lt;/h2&gt;

&lt;p&gt;It cannot tell you the model got worse.&lt;/p&gt;

&lt;p&gt;Everything above validates prompt assembly, tool definitions, metric logic, and report aggregation. All of that is code, all of it breaks in ordinary code ways, and testing it for free on every PR is worth doing. But the canned responses are ground truth by fiat. If the model starts emitting different shapes tomorrow, every one of these tests still passes.&lt;/p&gt;

&lt;p&gt;That's what the live mode is for, and it's manual on purpose. A live run costs about $0.18 for the full set, which is nothing; the reason it isn't in CI is flakiness and rate limits, not money. It writes to its own baseline file and you fire it when you've changed something that could plausibly move model behavior. In &lt;a href="https://dev.to/marc_kumiko/the-model-obeys-your-schema-not-your-description-1cml"&gt;the schema-tightening post&lt;/a&gt; the before-and-after numbers all came from live runs of this harness, for about $0.18 a round.&lt;/p&gt;

&lt;p&gt;One detail that made live mode usable: a provider exception doesn't kill the run. The failing fixture gets a synthetic &lt;code&gt;provider-error&lt;/code&gt; metric and the run continues, so a 429 on fixture three costs you one fixture instead of the whole report.&lt;/p&gt;

&lt;p&gt;And the fixtures themselves need review like any other test data. Ours drifted because nobody was checking canned responses against the schema constraint that live mode enforces. That failure mode is intrinsic to mocking an LLM. You can only catch it by reading the fixtures, or by running live often enough that the gap shows up.&lt;/p&gt;

&lt;h2&gt;
  
  
  The short version
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;A mocked eval tests your pipeline. It cannot test your model, and it will happily stay green on responses the real API would reject.&lt;/li&gt;
&lt;li&gt;Keep the runner mode-agnostic and pass the provider in. Both modes then share one metric pipeline, which is the thing you actually want identical.&lt;/li&gt;
&lt;li&gt;Put the value in the checked-in baseline diff, so quality changes show up in review as a file someone has to read.&lt;/li&gt;
&lt;li&gt;Compare structure, never human-readable messages, or the drift test becomes noise and people stop reading it.&lt;/li&gt;
&lt;li&gt;Report pass and score separately: one gates, the other shows movement before it flips.&lt;/li&gt;
&lt;li&gt;Review your fixtures against the constraints live mode enforces. Nothing else will.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>typescript</category>
      <category>testing</category>
    </item>
    <item>
      <title>The model obeys your schema, not your description</title>
      <dc:creator>Marc</dc:creator>
      <pubDate>Fri, 18 Sep 2026 22:40:08 +0000</pubDate>
      <link>https://dev.to/marc_kumiko/the-model-obeys-your-schema-not-your-description-1cml</link>
      <guid>https://dev.to/marc_kumiko/the-model-obeys-your-schema-not-your-description-1cml</guid>
      <description>&lt;p&gt;Two models. Same prompt, same tool description, same request. One of them returned this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"kind"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"entity"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"entityName"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"todo"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"definition"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"fields"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"text"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The other returned this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"kind"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"entity"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"todo"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"fields"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"text"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second one is wrong, and our downstream patcher rejected it with a 422 that told nobody anything useful. What took me a while to work out was why the first model got it right, because the answer turned out to have nothing to do with being smarter.&lt;/p&gt;

&lt;p&gt;This was May 2026, Opus 4.7 and Sonnet 4.6 at the time. The story generalizes to whatever pair of models you're holding today.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;We have a tool called &lt;code&gt;apply_patches&lt;/code&gt;. An LLM reads a user request plus a source file and emits a list of structural change operations: add this entity, replace that handler, remove that metric. Each operation carries a &lt;code&gt;pattern&lt;/code&gt;, the canonical object form of the thing being changed.&lt;/p&gt;

&lt;p&gt;The tool schema for that &lt;code&gt;pattern&lt;/code&gt; parameter was, in effect:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"object"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"properties"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"kind"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"string"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"required"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"kind"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;kind&lt;/code&gt; is a string and everything else is whatever. The actual shape lived in the tool description, a paragraph of prose with examples, the way most people write tool definitions.&lt;/p&gt;

&lt;p&gt;The big model complied anyway. The smaller one didn't. Both had read the same description.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the big model complied
&lt;/h2&gt;

&lt;p&gt;It had seen the shape before.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;entityName&lt;/code&gt; and &lt;code&gt;definition.fields&lt;/code&gt; are our field names, from our framework. To a model with training exposure to that shape, "emit an entity pattern" retrieves a memory. To a model without it, "emit an entity pattern" is a guess from the description text, and if you're guessing what an entity looks like, &lt;code&gt;{ name, fields }&lt;/code&gt; is a better guess than the truth. It's what everyone else's API would call those things.&lt;/p&gt;

&lt;p&gt;So this is an exposure gap rather than a capability gap, which matters because you can't fix an exposure gap by paying for a bigger model. It will show up for any model on any shape that isn't in its training data, which is to say on your proprietary shapes, indefinitely. The less your schema looks like the rest of the internet, the harder your description has to work. And descriptions are not what the model is validated against.&lt;/p&gt;

&lt;p&gt;The tool schema is. So we moved the contract into it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tight on the common kinds, loose on the tail
&lt;/h2&gt;

&lt;p&gt;We have around twenty pattern kinds. Nine of them account for roughly 85% of everything the model emits. The other dozen (&lt;code&gt;relation&lt;/code&gt;, &lt;code&gt;workspace&lt;/code&gt;, &lt;code&gt;secret&lt;/code&gt;, &lt;code&gt;claimKey&lt;/code&gt;, &lt;code&gt;systemScope&lt;/code&gt; and friends) show up rarely.&lt;/p&gt;

&lt;p&gt;Writing strict schemas for all twenty would have been a week of work and a permanent maintenance tax, so we didn't. Each common kind became a discriminated &lt;code&gt;oneOf&lt;/code&gt; branch with a real &lt;code&gt;required&lt;/code&gt; list:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;title&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;EntityPattern&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;properties&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nl"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;const&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;entity&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="nx"&gt;entityName&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nl"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;string&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="nx"&gt;definition&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nl"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;object&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="nx"&gt;required&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;kind&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;entityName&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;definition&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The long tail got one fallback branch that requires nothing but &lt;code&gt;kind&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;title&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;OtherPattern&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;properties&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nl"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;string&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;not&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="na"&gt;enum&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
          &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;entity&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;requires&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;toggleable&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;nav&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;writeHandler&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
          &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;queryHandler&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;hook&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;notification&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;metric&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;],&lt;/span&gt;
      &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="nx"&gt;required&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;kind&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Rare kinds still go through unvalidated at the schema layer, and the runtime patcher catches them. That split, tight on the discriminator values you see constantly and permissive on the ones you don't, is the part worth stealing. It costs an afternoon instead of a week and it targets the failures you actually get.&lt;/p&gt;

&lt;p&gt;We did the same for the natural keys that &lt;code&gt;replace&lt;/code&gt; and &lt;code&gt;remove&lt;/code&gt; operations use, and pinned the per-operation requirements with &lt;code&gt;allOf&lt;/code&gt; plus &lt;code&gt;if&lt;/code&gt;/&lt;code&gt;then&lt;/code&gt;, so the model can't hand us a &lt;code&gt;replace&lt;/code&gt; with nothing to replace:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nx"&gt;allOf&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;if&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;properties&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;op&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;const&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;replace&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="na"&gt;then&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;required&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;id&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;pattern&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;if&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;properties&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;op&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;const&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;add&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;     &lt;span class="na"&gt;then&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;required&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;pattern&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;if&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;properties&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;op&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;const&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;remove&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;  &lt;span class="na"&gt;then&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;required&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;id&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both Anthropic and OpenAI honor &lt;code&gt;oneOf&lt;/code&gt; and &lt;code&gt;allOf&lt;/code&gt;/&lt;code&gt;if&lt;/code&gt;/&lt;code&gt;then&lt;/code&gt; in tool input schemas. Most people skip them because the flat version works well enough on whichever model they tested with.&lt;/p&gt;

&lt;h2&gt;
  
  
  Did it work?
&lt;/h2&gt;

&lt;p&gt;The evidence is thinner than I'd like, and it points the right way.&lt;/p&gt;

&lt;p&gt;We ran three fixtures live, twice, for about $0.18 total. Before the change: two passed, one failed. The failure was the rename-entity case, emitting &lt;code&gt;{ kind: "entity", name, fields }&lt;/code&gt;. After: three passed, and the fixture that had been failing emitted &lt;code&gt;{ kind: "entity", entityName: "todo", definition: { fields: ... } }&lt;/code&gt;, byte for byte the shape the big model had been producing all along.&lt;/p&gt;

&lt;p&gt;Three fixtures is not a benchmark. What convinced me was which failure disappeared and what replaced it. The smaller model stopped inventing field names and started producing the canonical shape, on the exact case that had been failing.&lt;/p&gt;

&lt;p&gt;The real payoff came later. Once the smaller model could reliably emit the structured shape, it became viable as the default. That's usually the whole business case for schema work: a tight schema is what makes the cheap model good enough.&lt;/p&gt;

&lt;h2&gt;
  
  
  Footgun 1: oneOf is strict XOR
&lt;/h2&gt;

&lt;p&gt;Two weeks later, a code review caught something the tests hadn't.&lt;/p&gt;

&lt;p&gt;For &lt;code&gt;replace&lt;/code&gt; and &lt;code&gt;remove&lt;/code&gt; operations we have a parallel set of variants describing just the natural key. One of them is a singleton fallback: a &lt;code&gt;kind&lt;/code&gt; and nothing else. And &lt;code&gt;{ kind: "entity" }&lt;/code&gt; matched both the entity branch and the fallback branch.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;oneOf&lt;/code&gt; means exactly one. Two matches is a violation. Anthropic tolerated it; OpenAI's strict mode rejects the schema outright. Same schema, one provider silently fine, the other refusing to run.&lt;/p&gt;

&lt;p&gt;The fix is the &lt;code&gt;not.enum&lt;/code&gt; you saw above, where the fallback explicitly excludes every kind that has its own branch. Worth internalizing if you're building discriminated unions in JSON Schema: a fallback branch is not automatically disjoint from the specific ones. You have to make it disjoint by hand, and a test has to hold it that way, because the day someone adds a tenth specific branch and forgets the exclusion list is the day one of your providers starts 400ing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Footgun 2: field order is a token budget problem
&lt;/h2&gt;

&lt;p&gt;Different bug, same week. A &lt;code&gt;generate_feature&lt;/code&gt; call came back with &lt;code&gt;featureName&lt;/code&gt;, &lt;code&gt;packageDescription&lt;/code&gt;, &lt;code&gt;rationale&lt;/code&gt;, and no &lt;code&gt;source&lt;/code&gt;. Source being the entire point of the call.&lt;/p&gt;

&lt;p&gt;The model had emitted &lt;code&gt;rationale&lt;/code&gt; first, written about 700 tokens of thoughtful design commentary, and hit &lt;code&gt;maxTokens: 4000&lt;/code&gt; before it got to the field that mattered. Nothing errored. We just got a well-argued explanation of a file that didn't exist.&lt;/p&gt;

&lt;p&gt;Three fixes, in descending order of how much I trust them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;source.minLength: 100&lt;/code&gt; and &lt;code&gt;rationale.maxLength: 600&lt;/code&gt;, plus a blunt "BRIEF, 2-3 sentences" in the description. These are real constraints and they're what actually holds.&lt;/li&gt;
&lt;li&gt;Raised &lt;code&gt;maxTokens&lt;/code&gt; to 8000 for this one tool. The bounded tools stayed at 4000, since an unbounded budget everywhere just makes the truncation rarer and weirder.&lt;/li&gt;
&lt;li&gt;Declared &lt;code&gt;source&lt;/code&gt; first in the schema. Anthropic has been observed to emit arguments in declaration order. Observed, not documented. It costs nothing, it might help, and if a provider update changes it tomorrow nothing fails loudly. I left a comment in the source saying exactly that, and you should treat it the same way: a hint, not a contract.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Scores on the affected fixture went from 0.50 to 0.85, on both models, which is the tell that this was never a model-quality issue. The remaining 0.15 is genuine content quality: the model doesn't reach for one of our helpers when it should. That's a prompt and few-shot problem, and no amount of schema work will fix it. Knowing which of your failures are schema-shaped saves you from tightening things that were never loose.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part nobody solves for you
&lt;/h2&gt;

&lt;p&gt;These schemas are a hand-written mirror of our framework's real pattern types. Two definitions of the same shape, in two files, with nothing but discipline between them. When someone adds a required field to the real type, the schema doesn't know.&lt;/p&gt;

&lt;p&gt;We pin what we can in a contract test: the number of variants, the required-field list per kind, the &lt;code&gt;not.enum&lt;/code&gt; exclusion list. Drift fails a test instead of confusing a model in production six weeks later. That's a smoke alarm rather than a solution. If you generate your tool schemas from your actual types, you're ahead of us. If you're hand-writing them like we are, at least pin them.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs
&lt;/h2&gt;

&lt;p&gt;The obvious objection: you just made your prompt bigger, on every single call.&lt;/p&gt;

&lt;p&gt;About 3KB bigger, in our case. With prompt caching that's one cache write at 1.25× input rate, roughly $0.0002 for the schema chunk, and every subsequent call in the cache window reads it at 10%. The schema sits in the stable prefix, which is exactly where caching is designed to put it.&lt;/p&gt;

&lt;p&gt;Structured output used to carry a real tradeoff between schema size and bill. With caching it mostly doesn't. If token cost is what's keeping your tool definitions vague, go measure it. The number is probably smaller than the cost of one confusing 422 in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  The short version
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The description is advice. The schema is the contract. Models validate against one of those.&lt;/li&gt;
&lt;li&gt;A model complying with your undocumented shape may just be recognizing it from training. That's not a capability you can rely on, and your next model isn't guaranteed to have it.&lt;/li&gt;
&lt;li&gt;Tight schemas on the handful of kinds that dominate your traffic, one loose fallback for the tail. Don't schema the whole universe.&lt;/li&gt;
&lt;li&gt;Make your fallback branch explicitly disjoint, or &lt;code&gt;oneOf&lt;/code&gt; will bite you on the strictest provider you support.&lt;/li&gt;
&lt;li&gt;Length constraints hold. Field order is a hint. Know which is which.&lt;/li&gt;
&lt;li&gt;The reason to do any of this is usually that it makes the cheaper model good enough.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>typescript</category>
      <category>json</category>
    </item>
    <item>
      <title>You asked how I know what the model did. Here are the answers, with numbers.</title>
      <dc:creator>Marc</dc:creator>
      <pubDate>Thu, 06 Aug 2026 15:44:56 +0000</pubDate>
      <link>https://dev.to/marc_kumiko/you-asked-how-i-know-what-the-model-did-here-are-the-answers-with-numbers-5576</link>
      <guid>https://dev.to/marc_kumiko/you-asked-how-i-know-what-the-model-did-here-are-the-answers-with-numbers-5576</guid>
      <description>&lt;p&gt;I wrote about &lt;a href="https://dev.to/marc_kumiko/we-cut-our-ai-pipeline-costs-25-without-losing-accuracy-and-the-fix-wasnt-a-cheaper-model-4l5n"&gt;our AI pipeline costs&lt;/a&gt; a while back. The comments were better than the post.&lt;/p&gt;

&lt;p&gt;Valentin Monteiro made the point that cache alerts should be keyed per prompt version rather than on a global ratio, because a global number moves for boring reasons. Tae Kim pointed at how prompt caching fails silently, where a breakpoint on anything per-request gives you a miss on every call that logs exactly like a hit. Both need the same thing underneath, and I said I'd get cache metrics into our provenance records.&lt;/p&gt;

&lt;p&gt;That turned into a bigger job than I expected. Here's what came out of it, what it costs, and the two parts I still haven't worked out.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why store more than the output?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Because the output can't tell you whether something was always broken or broke last Tuesday, whether it's one record or ten thousand, or whether you changed something or the vendor did.&lt;/p&gt;

&lt;p&gt;A &lt;code&gt;total_amount&lt;/code&gt; column has the value. Nothing about where it came from.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where do you record it?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;At the provider, not in the feature.&lt;/p&gt;

&lt;p&gt;I did it in the feature first. It worked fine. But then every new AI feature has to remember to do the same thing, and eventually one won't. So it moved down to where providers get built:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;withProvenance&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;meta&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;startedAt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;performance&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
      &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;recordAiCall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;buildAiCallPayload&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;meta&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;startedAt&lt;/span&gt; &lt;span class="p"&gt;}))&lt;/span&gt;
      &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;recordAiCall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;buildAiCallPayload&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;meta&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;startedAt&lt;/span&gt; &lt;span class="p"&gt;}))&lt;/span&gt;
      &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="nx"&gt;error&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nothing opts in because nothing gets asked.&lt;/p&gt;

&lt;p&gt;One detail: this write goes outside the handler's transaction. If the business transaction rolls back, the call still happened and you still paid for it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What's in the payload?&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"providerId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"anthropic"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"handlerName"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"invoice-extract"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"requestedModel"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"claude-sonnet-5"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"respondedModel"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"claude-sonnet-5-20260514"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"promptVersion"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"3f9a1c0e77b2"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"inputHash"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"9c4e1ab77f30d552"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"latencyMs"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1180&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"usage"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"inputTokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;2140&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"outputTokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;318&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"cacheCreationInputTokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"cacheReadInputTokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1890&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"reportedCostUsd"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.0042&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"stopReason"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"end_turn"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The two cache fields are there because of Valentin's comment. &lt;code&gt;cacheReadInputTokens&lt;/code&gt; against &lt;code&gt;inputTokens&lt;/code&gt; is your hit ratio, and now it's per call rather than a monthly average, so you can key an alert on it per prompt version like he suggested instead of watching one global number.&lt;/p&gt;

&lt;p&gt;Requested and responded model are separate fields. Looks pedantic until a vendor routes you somewhere else and they don't match.&lt;/p&gt;

&lt;p&gt;No prompt text, no output. Only hashes. The log lives forever and prompts are full of customer invoices.&lt;/p&gt;

&lt;p&gt;Failures get a row too, with error kind and HTTP status. Most setups only log the successes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do you version a prompt?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I planned to put a version in the prompt file and bump it by hand. Then I didn't, because hand-maintained versions rot.&lt;/p&gt;

&lt;p&gt;It hashes the stable, model-visible part of the request at runtime instead: system instructions, corpus, and the tool schemas. Twelve hex characters, and nobody bumps anything. Messages stay out on purpose, because if they were in it every call would get its own version and the thing would group nothing.&lt;/p&gt;

&lt;p&gt;The "and the tool schemas" part is newer than this post. It used to hash the cacheable prefix only. Our extraction handler builds its tool from the caller's output schema, so you could change that schema, send the model a demonstrably different request, and &lt;code&gt;promptVersion&lt;/code&gt; wouldn't move. &lt;code&gt;inputHash&lt;/code&gt; saw it, but that one is unique per call, so it identifies without grouping. Which is the missed bump the runtime hash was supposed to make impossible. Found it writing this, wrote an issue, fixed it, merged it before the post went out. &lt;code&gt;toolChoice&lt;/code&gt; went in at the same time, the field Tae flagged in the last thread ;)&lt;/p&gt;

&lt;p&gt;What's left after that is key order, and it's a real one. &lt;code&gt;JSON.stringify&lt;/code&gt; hashes the serialisation, so reordering properties in a tool schema flips the version without changing anything semantically.&lt;/p&gt;

&lt;p&gt;The obvious move is to sort keys and hash a canonical form. We decided against it, and the reason is the interesting bit: the model sees the prompt as serialised text, and property order in a JSON schema affects the order an LLM generates fields in. Two differently sorted schemas are genuinely two different prompts. Sorting them into one bucket would rebuild exactly the missed bump we just removed, only invisibly. So the trade isn't false splits against churn. It's a visible false split against an invisible missed bump, and visible wins.&lt;/p&gt;

&lt;p&gt;The other thing is on purpose. The hash is one-way, so you get &lt;code&gt;3f9a1c0e77b2&lt;/code&gt; and no route back to the prompt. You can tell which calls are affected but not what changed in them. Reconstructing means git plus rehashing against the corpus from that day. The hand-bumped file version would have handled that with &lt;code&gt;git blame&lt;/code&gt;. I didn't think about it until after I'd shipped the hash.&lt;/p&gt;

&lt;p&gt;We have a prompt store with proper revision history sitting in the same codebase, connected to none of this. That's probably the answer and I haven't wired it up.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What does it cost?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Postgres 16, synthetic events, shape checked against real ones from the provider path:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;513 bytes per payload at the median, 517 at p95&lt;/li&gt;
&lt;li&gt;about 1,036 bytes per call with row overhead and the event store's three indexes&lt;/li&gt;
&lt;li&gt;at 10k calls a day that's roughly 296 MB a month, 3.5 GB a year&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Same per-call number at 10k, 100k and 1M rows. Storage isn't where the money goes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can you query it?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is the part where measuring changed my mind.&lt;/p&gt;

&lt;p&gt;"Every call with prompt version X" is what you run when something is wrong. An event store indexes tenant, aggregate type and time. Not the inside of a JSONB payload.&lt;/p&gt;

&lt;p&gt;Scoped to one tenant and a time window: 162 ms over a million events. The index picks ~67,000 candidate rows and the payload filter narrows those to 8,490, so you pay for hauling 67,000 rows out of the heap.&lt;/p&gt;

&lt;p&gt;Globally it's a seq scan, so you add an expression index on the payload field. At a million rows:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;distinct prompt versions&lt;/th&gt;
&lt;th&gt;no index&lt;/th&gt;
&lt;th&gt;with index&lt;/th&gt;
&lt;th&gt;plan&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;115 ms&lt;/td&gt;
&lt;td&gt;281 ms&lt;/td&gt;
&lt;td&gt;ignored it, seq scan&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;50&lt;/td&gt;
&lt;td&gt;107 ms&lt;/td&gt;
&lt;td&gt;61 ms&lt;/td&gt;
&lt;td&gt;bitmap heap scan&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;td&gt;107 ms&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.02 ms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;bitmap heap scan&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Index costs 6.9 MB in all three.&lt;/p&gt;

&lt;p&gt;With eight versions one version is 12.5% of the table, so Postgres scans and is right to. The index just sits there.&lt;/p&gt;

&lt;p&gt;The problem is that day one is when you benchmark this, see nothing, and decide the index isn't worth it. Six months in you've edited prompts a few hundred times, one version is 0.2% of the table, and that same index is a hundred times faster. Add it once you're past single digits.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do I need event sourcing?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;No. An append-only &lt;code&gt;ai_calls&lt;/code&gt; table gets you nearly all of it. We already had the event log so it came out of that for free.&lt;/p&gt;

&lt;p&gt;The bit worth copying either way is where you put the recording.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What would you change before this goes live?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Asking for real. It's built and merged but not out to a live tenant yet, so this is still a good moment to hear that it's wrong.&lt;/p&gt;

&lt;p&gt;Which is also why there are no cost or latency numbers from real traffic in here. We don't have them yet. In two months I will.&lt;/p&gt;

&lt;p&gt;If you've run something like this for a while: what's in your call log that isn't in mine? And where does this fall apart at a volume I haven't hit?&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://github.com/CosmicDriftGameStudio/kumiko-framework/blob/main/docs/reference/ai-call-provenance-benchmark.sql" rel="noopener noreferrer"&gt;benchmark SQL&lt;/a&gt; is in the repo if you'd rather measure your own database than trust mine. It's self-contained, so &lt;code&gt;psql -v rows=1000000 -v prompt_versions=500 -f ai-call-provenance-benchmark.sql&lt;/code&gt; against a throwaway database reproduces the table above.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I build &lt;a href="https://kumiko.rocks" rel="noopener noreferrer"&gt;Kumiko&lt;/a&gt;, an event sourced framework for multi-tenant B2B systems in TypeScript.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>monitoring</category>
      <category>performance</category>
    </item>
    <item>
      <title>Multi-Tenant AI Chat: From Hardcoded Config to BYOK in 4 Steps</title>
      <dc:creator>Marc</dc:creator>
      <pubDate>Mon, 03 Aug 2026 08:35:50 +0000</pubDate>
      <link>https://dev.to/marc_kumiko/multi-tenant-ai-chat-from-hardcoded-config-to-byok-in-4-steps-51aa</link>
      <guid>https://dev.to/marc_kumiko/multi-tenant-ai-chat-from-hardcoded-config-to-byok-in-4-steps-51aa</guid>
      <description>&lt;p&gt;Two tenants, two AI providers, two prompts. Sounds simple, and on day one it is. That's the trap. It stays simple right up until customer number two sends their first "quick question," and eighteen months later you're running a small distributed system to answer it. Here's the honest version of that slide, four stages, each one caused by a real human typing a real request into Slack.&lt;/p&gt;

&lt;h2&gt;
  
  
  Day 1: it just works (he says, foolishly) 😅
&lt;/h2&gt;

&lt;p&gt;Tenant A wants OpenAI. Tenant B wants Claude. Both want their own system prompt. The obvious first version: one config object, one row per tenant. What could possibly go wrong. (Everything. Everything could go wrong. But not yet.)&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;┌─────────────┐     ┌──────────────────┐     ┌─────────────┐
│  Request    │ →   │  TENANT_CONFIG   │ →   │  Provider   │
│  (tenantId) │     │  (hardcoded obj) │     │  SDK call   │
└─────────────┘     └──────────────────┘     └─────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;TENANT_AI_CONFIG&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;tenantA&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;openai&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-5&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;You are terse and technical.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;tenantB&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;anthropic&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;claude-sonnet-5&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;You are friendly. Antworte auf Deutsch.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;handleChat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;userMessage&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;config&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;TENANT_AI_CONFIG&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;provider&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;openai&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nx"&gt;openai&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;anthropic&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;userMessage&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ships in an afternoon. Two tenants, two rows, demo goes great, everyone claps 👏. Put this moment in a frame, it's the calmest the codebase will ever be.&lt;/p&gt;

&lt;h2&gt;
  
  
  Evolution 1: Tenant B wants their own twist
&lt;/h2&gt;

&lt;p&gt;A week in (a week, we didn't even get a full sprint), tenant B messages: "can we change the prompt ourselves, without waiting for a deploy?" Fair ask, they know their users, we don't, and also nobody wants to be the on-call engineer who gets paged to edit a string literal. A hardcoded object can't answer that, it needs a rebuild to change a comma.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;┌─────────────┐     ┌──────────────────┐     ┌─────────────┐
│  Request    │ →   │  tenant_settings │ →   │  Provider   │
│  (tenantId) │     │  (DB row, admin  │     │  SDK call   │
│             │     │   editable)      │     │             │
└─────────────┘     └──────────────────┘     └─────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Config moves from a code constant to a table the tenant's own admin UI can write to.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;getAiConfig&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;row&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tenantSettings&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;findOne&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;tenantId&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;row&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;aiConfig&lt;/span&gt; &lt;span class="c1"&gt;// { provider, model, prompt }&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same shape as before, different source. &lt;code&gt;handleChat&lt;/code&gt; doesn't change at all, it has no idea any of this happened, which is exactly the point of putting the lookup behind one function.&lt;/p&gt;

&lt;h2&gt;
  
  
  Evolution 2: "wait, who changed the prompt?" 🕵️
&lt;/h2&gt;

&lt;p&gt;Self-service is great until it isn't: tenant B's prompt quietly changed last Tuesday, their bot started answering in pirate-speak for reasons nobody can reconstruct, support gets a ticket, and the honest answer is "we have no idea, the database doesn't remember either." A plain DB row just gets overwritten, the past has no representation, it's Ctrl+Z with no undo history.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;┌──────────────┐     ┌────────────────┐     ┌──────────────────┐
│  Admin edits │ →   │  ConfigChanged │ →   │  current config  │
│  the prompt  │     │  event (who,   │     │  = fold(events)  │
│              │     │  when, diff)   │     │                  │
└──────────────┘     └────────────────┘     └──────────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The config becomes event-sourced instead of a mutable row: every change is an event, the current value is a projection over them.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;updateAiConfig&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;patch&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;Partial&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;AiConfig&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;actor&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;emit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;AiConfigChanged&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;patch&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;actor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;at&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;getAiConfig&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;AiConfig&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;events&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;events&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;find&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;AiConfigChanged&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;events&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;reduce&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;cfg&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt;&lt;span class="nx"&gt;cfg&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;patch&lt;/span&gt; &lt;span class="p"&gt;}),&lt;/span&gt; &lt;span class="nx"&gt;DEFAULT_AI_CONFIG&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now "who changed it and when" is a query, not a seance 🔮. &lt;code&gt;handleChat&lt;/code&gt; still hasn't changed, it just calls &lt;code&gt;getAiConfig&lt;/code&gt;, blissfully unaware it's now talking to an event log instead of a table.&lt;/p&gt;

&lt;h2&gt;
  
  
  Evolution 3: BYOK and a usage cap 💸
&lt;/h2&gt;

&lt;p&gt;A bigger tenant shows up, the kind that gets its own Slack channel, with two demands: they want to use &lt;em&gt;their own&lt;/em&gt; OpenAI key (cost control, their own rate limits, their own finance team breathing down their neck), and they want a hard cap on monthly spend so an over-caffeinated intern's script can't turn into a five-figure invoice.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;┌──────────────┐     ┌───────────────────────┐     ┌──────────────┐
│  Request     │ →   │  config.apiKey?       │ →   │  usage &amp;lt; cap?│
│              │     │  (BYOK, encrypted)    │     │  → call      │
│              │     │  else our shared key  │     │  → else 429  │
└──────────────┘     └───────────────────────┘     └──────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;callProvider&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;userMessage&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;config&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;getAiConfig&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;usage&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;getMonthlyUsage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;usageCap&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;usage&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;usageCap&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;UsageCapExceeded&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;apiKey&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;byokApiKey&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;SHARED_API_KEY&lt;/span&gt; &lt;span class="c1"&gt;// BYOK overrides shared key&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;getClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;apiKey&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;reply&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;userMessage&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;recordUsage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;reply&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;totalTokens&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;reply&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two additive fields on the same config, &lt;code&gt;byokApiKey&lt;/code&gt; and &lt;code&gt;usageCap&lt;/code&gt;, and one counter check before the call. No new architecture, no rewrite of the first three stages, no "sorry, we need a full quarter to redesign this."&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern behind the pattern 🧵
&lt;/h2&gt;

&lt;p&gt;Every stage kept &lt;code&gt;getAiConfig(tenantId) → { provider, model, prompt, ... }&lt;/code&gt; as the seam. Storage changed underneath it four times (constant, DB row, event-sourced projection, projection with encrypted secrets) and the call site never noticed, never cared, never even asked. That's the actual lesson: don't design the multi-tenant AI system upfront, design one seam that can absorb whatever the next Slack message throws at it.&lt;/p&gt;

&lt;p&gt;Skipped on purpose: provider fallback, streaming, per-model cost tables. Add those when a tenant actually asks, same as everything above. If a tenant asks for a fifth provider before you've read this sentence, that's not a counterexample, that's Tuesday. 🙃&lt;/p&gt;

</description>
      <category>ai</category>
      <category>multitenancy</category>
      <category>typescript</category>
      <category>prisma</category>
    </item>
    <item>
      <title>We Cut Our AI Pipeline Costs 25% Without Losing Accuracy (and the fix wasn't a cheaper model)</title>
      <dc:creator>Marc</dc:creator>
      <pubDate>Sat, 01 Aug 2026 18:46:54 +0000</pubDate>
      <link>https://dev.to/marc_kumiko/we-cut-our-ai-pipeline-costs-25-without-losing-accuracy-and-the-fix-wasnt-a-cheaper-model-4l5n</link>
      <guid>https://dev.to/marc_kumiko/we-cut-our-ai-pipeline-costs-25-without-losing-accuracy-and-the-fix-wasnt-a-cheaper-model-4l5n</guid>
      <description>&lt;p&gt;Our AI pipeline runs three step kinds (&lt;code&gt;ai.generate&lt;/code&gt;, &lt;code&gt;ai.extract&lt;/code&gt;, &lt;code&gt;ai.classify&lt;/code&gt;), each independently resolving its own provider, model, and prompt revision at run time. The default model is Sonnet, not Opus — and for a while that felt like a compromise, because Opus was the expensive-but-reliable option and Sonnet needed babysitting to hit the same pass rate.&lt;/p&gt;

&lt;p&gt;The fix that closed the gap wasn't a smarter prompt. It was &lt;code&gt;tool_choice&lt;/code&gt; plus a tightened output schema, forcing the model to commit to an answer shape instead of spending tokens hedging its way there. That alone got Sonnet to pass-parity with Opus's prior output at roughly a quarter of the cost, in our eval. A second, separate lever: Anthropic's own recommended &lt;code&gt;effort: "xhigh"&lt;/code&gt; for agentic tasks produced roughly twice the thinking tokens of &lt;code&gt;"high"&lt;/code&gt; for the same pass rate — thinking tokens bill like output tokens, so that's a straight 2x for zero accuracy gain, and it only matters when Opus is used via an explicit override (Sonnet doesn't get adaptive thinking at all). A third, independent lever: &lt;code&gt;max_tokens&lt;/code&gt; defaulted to 16000 out of caution; dropping it to 4000 (still comfortably above what any real step needed) cut effective output cost again, because Anthropic bills against the cap as an upper bound in some failure paths, not only against what the model actually emits.&lt;/p&gt;

&lt;p&gt;None of these three touched the prompt content or the provider. All three came from looking at per-step token usage, which only exists because every step writes an immutable provenance record on completion: prompt revision, provider, model, token usage. Same idea as event sourcing, applied to LLM calls instead of domain writes. You don't trust "the pipeline probably used prompt v3, on whatever model was configured," you have a row that says so — and you can audit six months of "which prompt touched this tenant" without ever loading the actual LLM payloads (also the DSGVO-friendly shape: provenance rows carry no user content, only call metadata).&lt;/p&gt;

&lt;p&gt;The other real trap, found the hard way: adaptive thinking and forced tool-use don't mix. Anthropic 400s with "Thinking may not be enabled when tool_choice forces tool use" the moment both are set, which sounds obvious in hindsight but not when you're setting both because both individually sound like "make the model try harder." The fix has to be automatic — detect a forced &lt;code&gt;tool_choice&lt;/code&gt; and disable adaptive thinking before the request goes out — because leaving it to every caller to remember means every eval run using &lt;code&gt;toolChoice: { type: "any" }&lt;/code&gt; against an Opus override breaks the same way, all at once.&lt;/p&gt;

&lt;p&gt;Separately: prompt caching only pays off if the &lt;code&gt;cache_control&lt;/code&gt; breakpoint sits on the &lt;em&gt;last&lt;/em&gt; static block, with tools and system prompt as one cacheable prefix. Past roughly 16k tokens without it, requests start hitting the SDK's HTTP timeout before the response streams back far enough to matter. Put the breakpoint on something that changes per request and you get a silent cache-miss that looks identical to a hit in the logs while you keep paying full price.&lt;/p&gt;

&lt;p&gt;Full writeup with the pipeline architecture (with diagrams), the provider-resolution code, and the provenance shape: &lt;a href="https://docs.kumiko.rocks/en/guides/ai-pipeline-provenance/" rel="noopener noreferrer"&gt;docs.kumiko.rocks/en/guides/ai-pipeline-provenance&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>typescript</category>
      <category>llm</category>
      <category>architecture</category>
    </item>
    <item>
      <title>We Shipped a Mortgage Calculator Bug Where Every Test Was Green and the Answer Was Still Wrong</title>
      <dc:creator>Marc</dc:creator>
      <pubDate>Sat, 01 Aug 2026 11:48:07 +0000</pubDate>
      <link>https://dev.to/marc_kumiko/we-shipped-a-mortgage-calculator-bug-where-every-test-was-green-and-the-answer-was-still-wrong-25ic</link>
      <guid>https://dev.to/marc_kumiko/we-shipped-a-mortgage-calculator-bug-where-every-test-was-green-and-the-answer-was-still-wrong-25ic</guid>
      <description>&lt;p&gt;A review pass on our credit calculator flagged something that should have been impossible: a&lt;br&gt;
finding tagged HIGH, on a panel with a full test suite, where every individual test passed and&lt;br&gt;
the answer was still wrong. Not "wrong in an edge case" wrong. Wrong for every single user who had&lt;br&gt;
ever configured a Sondertilgung (a lump-sum extra payment on their mortgage) and then looked at&lt;br&gt;
the "what if I invest the difference instead" comparison next to it.&lt;/p&gt;

&lt;p&gt;The bug: the comparison panel always ran the loan projection with an empty additionals list. The&lt;br&gt;
headline result above it, the number the user actually configured and trusted, included their&lt;br&gt;
real Sondertilgungen. So if you had a 20k lump-sum payment scheduled for month 6, the headline&lt;br&gt;
said "your loan is basically gone by year 8" and the panel right next to it, presented as&lt;br&gt;
directly comparable, was quietly running a different loan. Not a rounding difference. A different&lt;br&gt;
payoff date, a different remaining balance, a different verdict on whether investing instead&lt;br&gt;
would have won.&lt;/p&gt;

&lt;p&gt;Every test was green because every test checked one function against one input. Nothing checked&lt;br&gt;
that the two numbers on the same screen were talking about the same loan.&lt;/p&gt;
&lt;h2&gt;
  
  
  Why "pay down debt faster vs. invest the difference" is rigged by default
&lt;/h2&gt;

&lt;p&gt;That specific bug is a symptom of a more general trap. Almost every calculator online that lets&lt;br&gt;
you compare "pay down debt faster" against "invest the difference instead" gets the comparison&lt;br&gt;
itself wrong, independent of any coding bug — because it quietly changes the monthly budget&lt;br&gt;
between the two scenarios and presents the result as a fair fight.&lt;/p&gt;

&lt;p&gt;Say you have a 300k loan at 3.5%, paying 3% annual amortization. Someone suggests: "drop to 1%&lt;br&gt;
amortization, invest the rest." The naive calculator compares:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Scenario A: 3% amortization, no investing, ~1200/month total outflow&lt;/li&gt;
&lt;li&gt;Scenario B: 1% amortization, invest a fixed 200/month, ~950/month total outflow&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Of course B looks better over ten years. You're comparing spending 1200/month against spending&lt;br&gt;
950/month and marveling that the cheaper option built more wealth. It didn't win on insight. It&lt;br&gt;
won because you fed it a smaller number.&lt;/p&gt;

&lt;p&gt;The fix is a structural invariant, not a footnote:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;credit installment (this month) + ETF contribution (this month) = fixed monthly budget
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every month, for both scenarios, and once a loan is paid off in one scenario, its full budget&lt;br&gt;
rolls into the ETF contribution from that point on. You don't get to just stop spending because&lt;br&gt;
the mortgage happened to end early in the model. Once that holds, "total interest paid" stops&lt;br&gt;
being a meaningful metric (it no longer means the same thing across scenarios that pay&lt;br&gt;
differently), and the only honest number left is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;net worth(t) = ETF portfolio(t) - remaining loan balance(t)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;at a fixed horizon, for both scenarios, with identical cash outflow throughout.&lt;/p&gt;

&lt;h2&gt;
  
  
  The invariant was never the bug. Something else was.
&lt;/h2&gt;

&lt;p&gt;Here's the part that made this finding interesting: our budget invariant was correctly&lt;br&gt;
implemented from day one. Every month's ETF contribution is computed as &lt;code&gt;monthlyBudget -&lt;br&gt;
loanRate&lt;/code&gt;, not hardcoded, not assumed. If you'd audited that one property in isolation, you'd have&lt;br&gt;
signed off.&lt;/p&gt;

&lt;p&gt;The actual failure mode was a level up from that: getting the fair-comparison math right and then&lt;br&gt;
silently comparing two different loans anyway. The panel and the headline agreed on the rule&lt;br&gt;
("budget minus installment goes to the ETF") but disagreed on the input ("what installment,&lt;br&gt;
exactly, on what loan"). A structurally correct comparison of the wrong pair of scenarios is still&lt;br&gt;
wrong, and none of the unit tests around the fair-comparison logic could have caught it, because&lt;br&gt;
that logic was never the part that broke.&lt;/p&gt;

&lt;p&gt;The fix threads the same &lt;code&gt;additionals&lt;/code&gt; the headline uses through the comparison panel, so both&lt;br&gt;
numbers on the screen are now guaranteed to describe the same loan.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual lesson, and something to go check right now
&lt;/h2&gt;

&lt;p&gt;Budget-neutrality gets cited (correctly) as the thing naive pay-down-vs-invest calculators get&lt;br&gt;
wrong. It's necessary. It is not sufficient. You also have to pin every other input, identically,&lt;br&gt;
across both scenarios you're comparing, or the comparison quietly starts answering a different&lt;br&gt;
question than the one on the label, and a green test suite will not tell you.&lt;/p&gt;

&lt;p&gt;If you've got a mortgage calculator, a rent-vs-buy tool, or a pay-down-vs-invest comparison&lt;br&gt;
bookmarked somewhere: go feed it a scenario with something extra configured (a lump-sum payment,&lt;br&gt;
an irregular income month, whatever the tool supports), and check whether the comparison panel&lt;br&gt;
still uses it. Most won't tell you either way. Check anyway.&lt;/p&gt;




&lt;p&gt;This is the Tilgung-vs-ETF panel in &lt;a href="https://cashcolt.kumiko.rocks" rel="noopener noreferrer"&gt;cashcolt&lt;/a&gt;, a free&lt;br&gt;
no-signup mortgage calculator hosted in Germany. No tracking, no login, poke at it with your own&lt;br&gt;
numbers.&lt;/p&gt;

</description>
      <category>bug</category>
      <category>debugging</category>
      <category>softwareengineering</category>
      <category>testing</category>
    </item>
    <item>
      <title>Your event store is already your audit log</title>
      <dc:creator>Marc</dc:creator>
      <pubDate>Wed, 01 Jul 2026 07:51:25 +0000</pubDate>
      <link>https://dev.to/marc_kumiko/your-event-store-is-already-your-audit-log-1keo</link>
      <guid>https://dev.to/marc_kumiko/your-event-store-is-already-your-audit-log-1keo</guid>
      <description>&lt;h1&gt;
  
  
  Your event store is already your audit log
&lt;/h1&gt;

&lt;p&gt;Almost every SaaS I've worked on ends up with an &lt;code&gt;audit_log&lt;/code&gt; table. Someone files a compliance ticket — "we need to know who changed what and when" — and a new table appears next to the domain tables. Then the real work starts: writing to it on every mutating endpoint, keeping it in sync, and quietly discovering six months later that three endpoints forgot to log.&lt;/p&gt;

&lt;p&gt;That table is a second source of truth. And second sources of truth drift.&lt;/p&gt;

&lt;h2&gt;
  
  
  What an audit log actually needs
&lt;/h2&gt;

&lt;p&gt;Strip the compliance language away and an audit entry is five fields:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;who&lt;/strong&gt; did it&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;when&lt;/strong&gt; they did it&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;what&lt;/strong&gt; they touched (which entity)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;which action&lt;/strong&gt; it was&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;the delta&lt;/strong&gt; — what actually changed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Plus, for a multi-tenant app: &lt;strong&gt;whose data&lt;/strong&gt; it was, so tenant A can never read tenant B's history.&lt;/p&gt;

&lt;p&gt;Now look at what an event in an event-sourced system carries. Every state change is an appended event with &lt;code&gt;createdBy&lt;/code&gt;, &lt;code&gt;createdAt&lt;/code&gt;, &lt;code&gt;tenantId&lt;/code&gt;, &lt;code&gt;aggregateType&lt;/code&gt; + &lt;code&gt;aggregateId&lt;/code&gt;, &lt;code&gt;type&lt;/code&gt;, and a &lt;code&gt;payload&lt;/code&gt; holding the delta.&lt;/p&gt;

&lt;p&gt;That's the same five fields. The event log already &lt;em&gt;is&lt;/em&gt; the audit trail — append-only, ordered, and impossible to forget to write, because writing the event &lt;em&gt;is&lt;/em&gt; how state changes in the first place. There's no code path that mutates data without producing an event.&lt;/p&gt;

&lt;h2&gt;
  
  
  So don't build the table. Query the log.
&lt;/h2&gt;

&lt;p&gt;If the audit trail is already there, the whole "audit feature" collapses into one privileged read over the events table:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;listQuery&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;defineQueryHandler&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;list&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;schema&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;object&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;aggregateType&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;string&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;optional&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="na"&gt;aggregateId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;uuid&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;optional&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="na"&gt;eventType&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;string&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;optional&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="na"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;string&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;optional&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="na"&gt;from&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;iso&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;optional&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="na"&gt;to&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;iso&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;optional&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="na"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;number&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="k"&gt;default&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;before&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;string&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;optional&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="c1"&gt;// cursor&lt;/span&gt;
  &lt;span class="p"&gt;}),&lt;/span&gt;
  &lt;span class="na"&gt;access&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;roles&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Admin&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;SystemAdmin&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;handler&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;query&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;where&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;query&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;user&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tenantId&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt; &lt;span class="c1"&gt;// tenant-isolated at the WHERE&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;aggregateType&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;where&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;aggregateType&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;aggregateType&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;aggregateId&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="nx"&gt;where&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;aggregateId&lt;/span&gt;   &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;aggregateId&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;eventType&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;     &lt;span class="nx"&gt;where&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;type&lt;/span&gt;          &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;eventType&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;        &lt;span class="nx"&gt;where&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;createdBy&lt;/span&gt;      &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="c1"&gt;// ...time range + cursor omitted for brevity&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;selectMany&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;eventsTable&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;where&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;orderBy&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;col&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;id&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;direction&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;desc&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
      &lt;span class="na"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No table, no projection, no write path, no sync job. The filter surface an audit UI wants — by entity, by actor, by action, by time — is just &lt;code&gt;WHERE&lt;/code&gt; clauses over columns the events already have.&lt;/p&gt;

&lt;h2&gt;
  
  
  The two things you still owe
&lt;/h2&gt;

&lt;p&gt;Reusing the event log doesn't come completely free. Two concerns are real:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Access control.&lt;/strong&gt; The event log is the most sensitive read in the system — it's literally everything that ever happened. Gate it hard (&lt;code&gt;Admin&lt;/code&gt; / &lt;code&gt;SystemAdmin&lt;/code&gt; above) and pin tenant isolation into the &lt;code&gt;WHERE&lt;/code&gt; clause itself, not into application logic that a future refactor can bypass. Cross-tenant peeking should be structurally impossible, not politely discouraged.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. PII.&lt;/strong&gt; If you dump raw event payloads into an audit view, you'll surface fields you didn't mean to. The clean fix is to strip sensitive values &lt;em&gt;at append time&lt;/em&gt; — mark them in the entity definition and never let them into the stored event. Then the audit read physically cannot leak them, because they were never written. Doing it at read time is a filter you'll eventually forget on some new field; doing it at write time is a guarantee.&lt;/p&gt;

&lt;h2&gt;
  
  
  When this doesn't apply
&lt;/h2&gt;

&lt;p&gt;Honesty: this only works if you're actually event-sourced. If your system does in-place &lt;code&gt;UPDATE&lt;/code&gt;s, there's no historical record to query — you genuinely need to &lt;em&gt;start&lt;/em&gt; capturing one, and a dedicated table (or CDC/logical decoding off the WAL) is the pragmatic path. This isn't an argument to adopt event sourcing &lt;em&gt;for&lt;/em&gt; audit; it's an argument that if you already have it, the second table is redundant.&lt;/p&gt;

&lt;p&gt;One caveat even when it fits: event schemas evolve, so your audit reader sees heterogeneous historical payloads. For an audit log that's a feature — you want the exact shape as it was written — but don't mistake it for a clean queryable projection.&lt;/p&gt;

&lt;h2&gt;
  
  
  The point
&lt;/h2&gt;

&lt;p&gt;An audit log isn't a thing you build. It's a &lt;em&gt;view&lt;/em&gt; onto history you're already keeping. If you're appending events, you've been sitting on a complete, tamper-evident audit trail the whole time — the only missing piece was a gated query with the right &lt;code&gt;WHERE&lt;/code&gt; clauses.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This is exactly how the &lt;code&gt;audit&lt;/code&gt; feature works in &lt;a href="https://kumiko.rocks" rel="noopener noreferrer"&gt;Kumiko&lt;/a&gt;, a Bun/TypeScript framework where multi-tenancy, GDPR, and audit are bundled features rather than boilerplate you rewrite per project — one ~40-line query handler, no separate table.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>eventsourcing</category>
      <category>postgres</category>
      <category>architecture</category>
      <category>backend</category>
    </item>
    <item>
      <title>How do you prevent AI-generated code from drifting away from your conventions over time?</title>
      <dc:creator>Marc</dc:creator>
      <pubDate>Sun, 28 Jun 2026 10:52:37 +0000</pubDate>
      <link>https://dev.to/marc_kumiko/how-do-you-prevent-ai-generated-code-from-drifting-away-from-your-conventions-over-time-4b3l</link>
      <guid>https://dev.to/marc_kumiko/how-do-you-prevent-ai-generated-code-from-drifting-away-from-your-conventions-over-time-4b3l</guid>
      <description>&lt;p&gt;We've been generating production features with AI for a while now — auth flows, billing hooks, notification handlers. And we've hit a pattern we don't have a good answer to yet.&lt;/p&gt;

&lt;p&gt;The first feature the AI generates looks great. It reads the codebase, picks up the patterns, and the output looks like something a senior dev wrote.&lt;/p&gt;

&lt;p&gt;The tenth feature? Less so. Small inconsistencies creep in:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A handler that doesn't follow the error-handling convention&lt;/li&gt;
&lt;li&gt;A schema field with a different naming pattern&lt;/li&gt;
&lt;li&gt;A test that checks existence instead of behavior&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of it is wrong. All of it is subtly inconsistent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fixes we've tried
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;AGENTS.md / CLAUDE.md&lt;/strong&gt; — helps, but gets stale and doesn't scale with the codebase&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code review&lt;/strong&gt; — catches it, but defeats some of the speed advantage&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Linting + formatting&lt;/strong&gt; — catches easy stuff, misses semantic drift&lt;/p&gt;

&lt;p&gt;What we haven't solved: giving the AI a "living" representation of your conventions that stays current as the codebase evolves.&lt;/p&gt;

&lt;p&gt;We're building Kumiko — an opinionated SaaS framework — partly as an answer to this. If the framework constrains what's possible, drift has less surface area. But I'm not convinced that fully solves it either.&lt;/p&gt;

&lt;h2&gt;
  
  
  Curious what's actually working for others
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Do you just review AI output carefully and accept some drift?&lt;/li&gt;
&lt;li&gt;Custom guards / linters that encode your conventions?&lt;/li&gt;
&lt;li&gt;Something that auto-generates AGENTS.md from the codebase?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What's your approach?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
