<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sakura Sky</title>
    <description>The latest articles on DEV Community by Sakura Sky (sakurasky).</description>
    <link>https://dev.to/sakurasky</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Forganization%2Fprofile_image%2F13859%2Fe20f8934-bc82-46f2-a495-bdf8a8522a2f.png</url>
      <title>DEV Community: Sakura Sky</title>
      <link>https://dev.to/sakurasky</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sakurasky"/>
    <language>en</language>
    <item>
      <title>The CTO Playbook for Agentic Systems</title>
      <dc:creator>Andrew Stevens</dc:creator>
      <pubDate>Fri, 18 Sep 2026 08:11:01 +0000</pubDate>
      <link>https://dev.to/sakurasky/the-cto-playbook-for-agentic-systems-2ffd</link>
      <guid>https://dev.to/sakurasky/the-cto-playbook-for-agentic-systems-2ffd</guid>
      <description>&lt;p&gt;Today we are publishing &lt;a href="https://www.sakurasky.com/white-papers/cto-playbook-for-agentic-systems/" rel="noopener noreferrer"&gt;The CTO Playbook for Agentic Systems&lt;/a&gt; on our white papers shelf. It runs to 71 pages across nine parts, and it was written by Andrew Stevens, our CTO and CISO (Stevens, 2026), over eleven months of conversations with engineering leaders running agents in production. It is published by Whitepaper Press rather than by us. We are hosting it here because its audience is the one we spend most of our time with.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why now
&lt;/h2&gt;

&lt;p&gt;Most of the engineering organisations we work with have stopped asking whether agents can write production code. The question has moved on to what happens to an organisation once a large share of its output arrives without a human having typed it, and the public telemetry on that has turned uncomfortable.&lt;/p&gt;

&lt;p&gt;Faros Research's 2026 study, drawn from two years of data across 22,000 developers and more than 4,000 teams, reports throughput up on every measure it tracks, and alongside it a 242.7 percent rise in the ratio of production incidents to merged pull requests, bugs per developer up from 9 percent in its 2025 edition to 54 percent, and pull requests merged with no review at all, human or agentic, up 31.3 percent (Faros Research, 2025; Faros Research, 2026). Faros does not read that last figure as anyone deciding to skip oversight. It reads it as reviewers being unable to keep pace.&lt;/p&gt;

&lt;p&gt;We would hold those findings a little loosely, since they are telemetry from organisations running one vendor's platform and DORA's 2025 survey reaches a more optimistic conclusion about whether strong engineering practice protects you (DORA, 2025). But if your own review load has started behaving anything like that, the shape of the problem is already familiar and the answers are not obvious.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the paper works through
&lt;/h2&gt;

&lt;p&gt;Nine parts, each closing with a specific deliverable and a set of questions to put in front of your own leadership team.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Part 2&lt;/strong&gt; treats verification as a staffed capability rather than an overhead, sizes it, and puts a formula behind the cost per 1,000 governed actions. It also works the case most write-ups avoid, which is what an agent programme looks like from the inside once it has stopped paying for itself, and when the right call is to shrink it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Part 3&lt;/strong&gt; names seven engineering functions that need owners, with a level range and a compensation-band anchor for each, then raises the problem underneath the ladder: agents absorb a good deal of the work junior engineers used to build judgement on.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Parts 4 and 5&lt;/strong&gt; cover how to split work between a squad and its agents, and how to make decision rights auditable. Most first attempts collapse into a binary of what an agent may and may not do, which does not survive an incident review. The paper runs two axes instead, and is direct about the tension that creates between a containment target and a break-glass approval.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Part 6&lt;/strong&gt; rebuilds the delivery lifecycle around a different unit of deployment, and includes a sourcing table you can hand to procurement, with a buy-or-build call and an owner for every stage.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Part 7&lt;/strong&gt; is the leadership work: five conversations worth preparing for, and a first-90-days communication runbook where every announcement carries a precondition that has to be true before you make it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Parts 8 and 9&lt;/strong&gt; charter a governance board, put fourteen metrics on one page for it, set out four layers of hard control, and sequence the whole thing across six phases with a gate before scale. The gate has five criteria and runs in both directions, which is the design decision we find most useful in it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you can actually use
&lt;/h2&gt;

&lt;p&gt;Two things make this more than a read.&lt;/p&gt;

&lt;p&gt;The nine deliverables. Every part closes with a named artefact rather than a recommendation, each with an owner, a cadence, and a description of what finished looks like. An agent inventory with named owners. A rehearsed and timed rollback runbook. A signed, dated declaration of which roadmap phase you are actually in and which gate criterion you currently fail. When we are asked to assess an agent programme, these are reliably the documents that turn out not to exist.&lt;/p&gt;

&lt;p&gt;The readiness assessment. Twenty-eight questions in the appendix, scored 1 to 5, describing what a weak answer reveals rather than what a strong one sounds like. It is built to be scored in a board meeting. We expect to be asked to run it, and we would rather organisations ran it themselves first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who it is for
&lt;/h2&gt;

&lt;p&gt;CTOs and VPs of Engineering with agents already shipping code, the directors and principal engineers designing the review and levelling systems underneath that, security architects who own the control architecture behind an autonomy tier, and the executives who hold engineering accountable.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.sakurasky.com/white-papers/cto-playbook-for-agentic-systems/" rel="noopener noreferrer"&gt;Download the playbook&lt;/a&gt;. For the architecture it assumes underneath, &lt;a href="https://www.sakurasky.com/white-papers/trustworthy-agentic-ai-blueprint/" rel="noopener noreferrer"&gt;The Trustworthy Agentic AI Blueprint&lt;/a&gt; is the deeper treatment and &lt;a href="https://www.sakurasky.com/white-papers/gate/" rel="noopener noreferrer"&gt;GATE&lt;/a&gt; is its implementable form. GATE is an open framework Andrew Stevens authors and maintains personally, not a Sakura Sky product.&lt;/p&gt;

&lt;p&gt;If you would rather not work through the twenty-eight questions alone, that is what our &lt;a href="https://www.sakurasky.com/grc/" rel="noopener noreferrer"&gt;Managed GRC&lt;/a&gt; service line is for, and &lt;a href="https://www.sakurasky.com/contact/" rel="noopener noreferrer"&gt;a conversation&lt;/a&gt; is usually the quickest way to establish which of the nine artefacts you are missing.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Disclosure: The CTO Playbook for Agentic Systems is written by Andrew Stevens, CTO and CISO at Sakura Sky, and published by Whitepaper Press, with copyright held by the author. The white paper is a registration download and Sakura Sky receives the registration details. GATE is an open framework authored and maintained by Andrew Stevens personally under CC BY 4.0 and MIT, and is not a Sakura Sky product. The Trustworthy Agentic AI Blueprint linked here is likewise Sakura Sky and Stevens work, and the same weighting applies. Managed GRC Services is a Sakura Sky offering, and Sakura Sky provides advisory and managed services of the kind discussed here. Third-party research is cited as published and was checked on the dates given in the references.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Not legal advice. This article offers general commentary for an engineering leadership audience. The control-to-framework mapping described in the paper is directional and does not imply conformance with any regime or standard. Readers must obtain independent advice on how any of these apply to their circumstances.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;DORA, 2025. &lt;em&gt;State of AI-assisted Software Development.&lt;/em&gt; Google Cloud. Available at: &lt;a href="https://dora.dev/research/2025/dora-report/" rel="noopener noreferrer"&gt;https://dora.dev/research/2025/dora-report/&lt;/a&gt; [Accessed 18 September 2026].&lt;/p&gt;

&lt;p&gt;Faros Research, 2025. &lt;em&gt;The AI Productivity Paradox Report 2025.&lt;/em&gt; 23 July. Faros AI. Available at: &lt;a href="https://www.faros.ai/blog/ai-software-engineering" rel="noopener noreferrer"&gt;https://www.faros.ai/blog/ai-software-engineering&lt;/a&gt; [Accessed 18 September 2026].&lt;/p&gt;

&lt;p&gt;Faros Research, 2026. &lt;em&gt;Ten takeaways from the AI Engineering Report 2026: The Acceleration Whiplash.&lt;/em&gt; 12 April. Faros AI. Available at: &lt;a href="https://www.faros.ai/blog/ai-acceleration-whiplash-takeaways" rel="noopener noreferrer"&gt;https://www.faros.ai/blog/ai-acceleration-whiplash-takeaways&lt;/a&gt; [Accessed 18 September 2026].&lt;/p&gt;

&lt;p&gt;Stevens, A., 2026. &lt;em&gt;The CTO Playbook for Agentic Systems,&lt;/em&gt; Version 1.2. Whitepaper Press. Available at: &lt;a href="https://www.sakurasky.com/white-papers/cto-playbook-for-agentic-systems/" rel="noopener noreferrer"&gt;https://www.sakurasky.com/white-papers/cto-playbook-for-agentic-systems/&lt;/a&gt; [Accessed 18 September 2026].&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agentic</category>
      <category>governance</category>
      <category>strategy</category>
    </item>
    <item>
      <title>A 150B MoE on a Developer Laptop: An Upper Bound on What Faster Storage Buys You</title>
      <dc:creator>Andrew Stevens</dc:creator>
      <pubDate>Wed, 16 Sep 2026 19:45:26 +0000</pubDate>
      <link>https://dev.to/sakurasky/a-150b-moe-on-a-developer-laptop-an-upper-bound-on-what-faster-storage-buys-you-293i</link>
      <guid>https://dev.to/sakurasky/a-150b-moe-on-a-developer-laptop-an-upper-bound-on-what-faster-storage-buys-you-293i</guid>
      <description>&lt;p&gt;Frontier-scale models are being published as open weights faster than most teams have worked out what to do with them. The question I wanted answered was narrow and practical: can an engineer run one of these on the laptop they were issued, and if so, what is actually holding it back?&lt;/p&gt;

&lt;p&gt;The use case is evaluation rather than production. Before anything gets deployed, somebody has to find out whether a model holds a house style across a long document, whether it calls tools without falling over, whether its output is worth the trouble at all. That work belongs on an engineer's own machine, using material that has not been sent anywhere, and it does not need throughput. It needs the model to run.&lt;/p&gt;

&lt;p&gt;Conventional wisdom says it will not, or that it will crawl, and the reason given is always storage. A model this size cannot sit in memory, so the parts you are not using have to stream off a disk, and the disk becomes the wall. The engine used here documents a reference machine where one token takes 24 seconds and states plainly that the bytes have to arrive, so faster silicon will not help.&lt;/p&gt;

&lt;p&gt;That was my hypothesis going in. On a machine with a current NVMe, storage would still dominate, and the number worth measuring was how badly.&lt;/p&gt;

&lt;p&gt;Across a three minute generation the drive spent under eleven seconds actually reading anything. A disk with infinite bandwidth would have moved the result from 1.11 tokens per second to roughly 1.18.&lt;/p&gt;

&lt;p&gt;That figure is a bound, not a benchmark. Your machine will produce different throughput than mine, but the ceiling on what faster storage could win is the part that transfers, and on any drive of this class it is small.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mechanism
&lt;/h2&gt;

&lt;p&gt;A sparsely-activated mixture-of-experts model has a large parameter count and a small activation count. A router picks a handful of experts per token per layer and the rest of the network sits idle for that token.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/JustVugg/colibri" rel="noopener noreferrer"&gt;Colibrì&lt;/a&gt;, an Apache-2.0 inference engine by Vincenzo Fornaro, treats that property as a placement problem (Fornaro, 2026a). Rather than requiring the model to fit in memory, it stages VRAM, RAM and storage as one hierarchy: the dense part of the network stays resident, the routed experts live on disk, and the engine reads the ones the router asks for. It keeps a per-layer cache, records which experts a workload routes to, and pins the hot ones. The engine is one C file per model family with no BLAS and no Python at runtime.&lt;/p&gt;

&lt;p&gt;The project is explicit that this is research rather than a product, and that its scope is wider than storage. Its stated goal covers "model formats, memory hierarchy, storage I/O, placement, scheduling, kernels, speculation, and CPU/GPU overlap", and it ships a fully-resident configuration in which the disk drops out of the decode path entirely. What follows measures one machine at one point on that hierarchy.&lt;/p&gt;

&lt;h2&gt;
  
  
  I ran it on my work laptop
&lt;/h2&gt;

&lt;p&gt;An AMD Ryzen AI 9 365 with 20 logical CPUs and 61 GB of RAM, a single WD Blue SN5100 1 TB NVMe, an integrated Radeon 880M this engine cannot use, and Ubuntu 24.04.&lt;/p&gt;

&lt;p&gt;Nothing about that is unusual for an engineering laptop in 2026, which is why I used it. Benchmarks published on eight H100s are interesting and I cannot act on them. I wanted to know what the machine on my desk would do, and I suspect most people reading this want the same. If you have a current NVMe, enough RAM to hold a dense set plus a cache, and 85 GB of free disk, everything below should reproduce. There is no accelerator anywhere in it, which conveniently removes the variable that usually decides these comparisons before they start.&lt;/p&gt;

&lt;p&gt;Two things about the silicon are worth knowing before anyone reads too much into the CPU figures later. The part mixes performance and compact cores at different sustained clocks, and a laptop chassis holds a lower sustained power limit than a desktop would. Either could inflate CPU time per unit of work, and neither is something I measured.&lt;/p&gt;

&lt;p&gt;The model was DeepSeek V4 Flash, in the &lt;a href="https://huggingface.co/puwaer/DeepSeek-V4-Flash-0731-reap-150b" rel="noopener noreferrer"&gt;REAP-pruned 150B variant published by puwaer&lt;/a&gt;: 150,128,549,111 parameters across 43 layers of 132 experts, routing to six per token, occupying 84.7 GB (puwaer, 2026). Routed experts are native fp4, which is why Colibrì reads the published checkpoint with no conversion step.&lt;/p&gt;

&lt;p&gt;REAP is an expert-pruning method whose authors report near-lossless compression on code generation at fifty percent of experts removed, on models from 20B to 1T parameters (&lt;a href="https://arxiv.org/abs/2510.13999" rel="noopener noreferrer"&gt;Lasby et al., 2025&lt;/a&gt;). I did not evaluate that claim. This post makes no quality argument at all, and the one answer reproduced below contains a factual error about its own architecture, which is worth holding on to.&lt;/p&gt;

&lt;p&gt;Colibrì v1.11.0, built by its own &lt;code&gt;setup.sh&lt;/code&gt; with gcc 13.3.0 at &lt;code&gt;-O3 -march=native -fopenmp -flto&lt;/code&gt;. That build line matters more than it looks for a hand-rolled no-BLAS engine, since &lt;code&gt;-march=native&lt;/code&gt; on Zen 5 is the difference between measuring the workload and measuring your compiler.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measure the storage first
&lt;/h2&gt;

&lt;p&gt;The engine refuses to plan without a storage measurement and marks it required. It ships its own benchmark, and the block size is not arbitrary: each expert's matrices are stored adjacently and read in one call, so the read that matters is several megabytes rather than several kilobytes.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;M&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;~/models/dsv4-reap-150b/model-00003.safetensors
./iobench &lt;span class="nv"&gt;$M&lt;/span&gt; 13 200 1 1
./iobench &lt;span class="nv"&gt;$M&lt;/span&gt; 13 200 4 1
./iobench &lt;span class="nv"&gt;$M&lt;/span&gt; 13 200 16 1
./iobench &lt;span class="nv"&gt;$M&lt;/span&gt; 13 200 4 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With O_DIRECT: 5.67 GB/s at one thread, 6.17 at four, 5.72 at sixteen. Buffered at four threads: 7.43 GB/s, though that figure is inflated by the page cache on a file written minutes earlier and should not be used as a baseline for anything.&lt;/p&gt;

&lt;p&gt;One run per configuration, no variance reported, so the 8% spread across thread counts is not a result. What the numbers do support is that a single thread already reaches within 10% of the best figure observed, which is unsurprising given that a 13 MB application read is split into many device-level requests before it reaches the drive.&lt;/p&gt;

&lt;p&gt;For scale, the reference machine in Colibrì's &lt;a href="https://github.com/JustVugg/colibri/blob/main/docs/glm53-flash.md" rel="noopener noreferrer"&gt;GLM-5.3-Flash documentation&lt;/a&gt; measures 72 MB/s at queue depth 1 and saturates near 200 MB/s, and that drive imposes a floor of 24 seconds per token on that model (Fornaro, 2026b). A 2026 consumer NVMe is roughly eighty times faster at the comparable setting. The floor that dominates that document is a property of one drive, which is worth knowing before anyone reads those figures as the cost of the technique.&lt;/p&gt;

&lt;h2&gt;
  
  
  The plan
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;RAM    39.6 GB budget · 8.8 GB dense · 5.9 GB runtime · 24.9 GB warm experts · cap 43/layer
disk   51.0 GB cold experts · 416.6 GB free
limit  disk expert misses
hit    33% projected expert residency
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two thirds of the expert set stays on disk and the planner names disk misses as the expected constraint. It disabled speculative decoding, on the reasoning that drafting widens the set of experts a token needs, and enabled read and compute overlap.&lt;/p&gt;

&lt;p&gt;That overlap setting is load-bearing for everything that follows, and I will come back to it.&lt;/p&gt;

&lt;p&gt;The self-tuning pass then swept eleven candidates: one OMP team size, three loader-thread counts, two smaller RAM caches, two CUDA paths that a CPU-only run still evaluates, and confirmation runs. None cleared its three percent improvement gate, so the defaults were kept. The profile recorded 0.92 tokens per second at a 66.75 percent hit rate.&lt;/p&gt;

&lt;h2&gt;
  
  
  The run
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;time&lt;/span&gt; ./coli run &lt;span class="nt"&gt;--model&lt;/span&gt; ~/models/dsv4-reap-150b &lt;span class="nt"&gt;--no-think&lt;/span&gt; &lt;span class="nt"&gt;--ngen&lt;/span&gt; 200 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"Write one paragraph explaining what a mixture-of-experts model is."&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two hundred tokens in 180.1 seconds: 1.11 tokens per second, or 0.90 seconds per token. Time to first token 17.2 seconds on a seventeen-token prompt. Expert hit rate 87.6 percent, from 46,665 hits against 6,597 misses across 53,262 selections.&lt;/p&gt;

&lt;p&gt;The answer was coherent and said that mixture-of-experts models activate "typically just two or three" experts. This checkpoint routes to six.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the disk actually cost
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;v4_tokens prompt=17 generated=200 expert_requests=53262 hits=46665 misses=6597 bytes=88197562368
v4_direct reads=4816 flock_reads=0 fallbacks=1781 payload_bytes=60599304192
timing time_to_first_token=17.175s after_first=162.946s
real 3m2.409s / user 46m32.525s / sys 0m22.993s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two counters report bytes and they disagree, so it is worth saying which one this analysis uses and why. &lt;code&gt;v4_direct reads&lt;/code&gt; plus &lt;code&gt;fallbacks&lt;/code&gt; comes to 6,597, which is exactly the miss count, so &lt;code&gt;payload_bytes&lt;/code&gt; at 60.6 GB is the physical I/O attributable to misses: one read per miss, averaging 9.19 MB. The larger 88.2 GB figure in &lt;code&gt;v4_tokens&lt;/code&gt; does not divide into any plausible per-read size and is some other accounting. If you reproduce this, check both.&lt;/p&gt;

&lt;p&gt;So: 60.6 GB delivered by a drive measured at 5.67 GB/s is &lt;strong&gt;10.7 seconds of device time across a 180-second generation&lt;/strong&gt;. Remove the disk entirely and the same work finishes in about 169 seconds, which is 1.18 tokens per second. &lt;strong&gt;The ceiling on infinitely fast storage is about six percent.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The calculation needs no claim about where the other 169 seconds went. It is bytes divided by measured bandwidth, with no timing in it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Halving the cache
&lt;/h2&gt;

&lt;p&gt;A bound is not a diagnosis. If misses are so cheap, the obvious test is to cause a lot more of them and see what happens. So I pinned the cache budget explicitly, dropped the page cache before each run so nothing was read from memory, and halved it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;sh &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s1"&gt;'sync; sysctl -w vm.drop_caches=3'&lt;/span&gt;
&lt;span class="nb"&gt;time&lt;/span&gt; ./coli run &lt;span class="nt"&gt;--model&lt;/span&gt; ~/models/dsv4-reap-150b &lt;span class="nt"&gt;--no-think&lt;/span&gt; &lt;span class="nt"&gt;--ngen&lt;/span&gt; 200 &lt;span class="nt"&gt;--ram&lt;/span&gt; 20 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"Write one paragraph explaining what a mixture-of-experts model is."&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;40 GB budget&lt;/th&gt;
&lt;th&gt;20 GB budget&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;tokens/sec&lt;/td&gt;
&lt;td&gt;1.045&lt;/td&gt;
&lt;td&gt;0.885&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;expert hit rate&lt;/td&gt;
&lt;td&gt;85.5%&lt;/td&gt;
&lt;td&gt;66.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;misses&lt;/td&gt;
&lt;td&gt;7,719&lt;/td&gt;
&lt;td&gt;17,998&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;bytes read&lt;/td&gt;
&lt;td&gt;71.4 GB&lt;/td&gt;
&lt;td&gt;167.5 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;decode&lt;/td&gt;
&lt;td&gt;191.3 s&lt;/td&gt;
&lt;td&gt;226.0 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;user CPU&lt;/td&gt;
&lt;td&gt;49m 10s&lt;/td&gt;
&lt;td&gt;54m 39s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;device time at 5.67 GB/s&lt;/td&gt;
&lt;td&gt;12.6 s&lt;/td&gt;
&lt;td&gt;29.5 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Halving the cache cost 15% of throughput, so misses are not free and the engine's own advice to give it memory is right. My first draft implied memory had stopped paying at this configuration, and that was too strong.&lt;/p&gt;

&lt;p&gt;The interesting part is the accounting. Decode got 34.7 seconds slower. Of that, 16.9 seconds is the extra device time from reading 96 GB more. The user CPU went up by 329 seconds, which spread across 17 threads is about 19 seconds of wall clock. The two together come to 36 seconds against an observed 35, which is close enough to say that &lt;strong&gt;a cache miss costs you roughly as much CPU as it costs you disk&lt;/strong&gt;. Dequantising an fp4 expert, copying it into place, evicting something else and updating the bookkeeping is real work, and it scales with misses just as reads do.&lt;/p&gt;

&lt;p&gt;That reframes what memory is buying. The README explains the benefit through read volume. On this machine about half the benefit arrives as CPU you do not have to spend.&lt;/p&gt;

&lt;p&gt;And the drive still never becomes the constraint. Two and a third times the traffic raised decode time by 18%, and even at the worse setting device time was 29.5 seconds of a 226 second run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ruling out spin-wait
&lt;/h2&gt;

&lt;p&gt;Before claiming any of that CPU time was arithmetic, there is an obvious objection. GNU OpenMP spins by default while waiting at a barrier, so a thread stuck behind a straggler or behind a loader sitting in &lt;code&gt;pread&lt;/code&gt; burns CPU and computes nothing. High user time proves the threads were runnable and nothing more.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;time &lt;/span&gt;&lt;span class="nv"&gt;OMP_WAIT_POLICY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;passive ./coli run &lt;span class="nt"&gt;--model&lt;/span&gt; ~/models/dsv4-reap-150b &lt;span class="nt"&gt;--no-think&lt;/span&gt; &lt;span class="nt"&gt;--ngen&lt;/span&gt; 200 &lt;span class="nt"&gt;--ram&lt;/span&gt; 40 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"Write one paragraph explaining what a mixture-of-experts model is."&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Decode 191.703 seconds against 191.290, user CPU 49m 16s against 49m 10s. Hit rate, bytes and read counts match to three significant figures. Nothing moved.&lt;/p&gt;

&lt;p&gt;So the threads were not spinning at OpenMP barriers. One caveat I will not paper over: this rules out that specific mechanism, and the engine runs its own loader threads with their own synchronisation that &lt;code&gt;OMP_WAIT_POLICY&lt;/code&gt; does not touch. Spinning somewhere else remains possible and I have no instruction counts to exclude it.&lt;/p&gt;

&lt;p&gt;Those two runs also give something I was missing. 191.290 and 191.703 seconds at identical settings is 0.2% apart, so run-to-run variance on this machine is small enough to ignore at the resolution anything here is argued at.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three things that cut the other way
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Prefill was slower per token than decode.&lt;/strong&gt; 17.2 seconds for 17 prompt tokens is 1.01 seconds each, against 0.90 in decode. Prefill is batched and should be faster per token. The likely explanation is a cold sweep through experts, which would make time to first token nearly pure I/O and roughly ten percent of the wall clock genuinely storage-bound.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One short prompt touched three quarters of the model.&lt;/strong&gt; 4,216 distinct experts out of 5,676 in 200 tokens. Whatever the cache is learning, it is not learning that this workload uses a small corner of the network.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Residency and hit rate are different quantities.&lt;/strong&gt; The planner projected 33 percent residency and the run recorded 87.6 percent hit rate. Those are a capacity ratio and a routing outcome, and the gap between them may be a skewed router rather than anything learned. The project lists learned pinning in its own open-hypotheses table, with the note that history "can overfit a prompt" and that the validation it wants is held-out cross-session A/B testing. Running one prompt repeatedly is the overfitting case, not evidence against it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this leaves the project's advice
&lt;/h2&gt;

&lt;p&gt;Colibrì's README says of this model family, in bold, "Give it RAM", and that the expert cache hit rate is what sets tokens per second, making &lt;code&gt;--ram&lt;/code&gt; the single most valuable knob (Fornaro, 2026a).&lt;/p&gt;

&lt;p&gt;The sweep says that is right, and says the stated mechanism is half the story. Memory does set throughput here. It sets it partly by avoiding reads, which is the documented reason, and about equally by avoiding the CPU cost of servicing a miss, which is not mentioned. If you are tuning a machine, that distinction matters, because it means the payoff from more memory does not disappear the moment your drive stops being busy.&lt;/p&gt;

&lt;p&gt;There is a sharper version I got wrong in an earlier draft and want to correct in public. This model is 84.7 GB. A machine with 96 GB or more of usable memory holds all of it, at which point there is no streaming, no cache and no miss rate to improve, and the project publishes exactly that configuration. So the shape of the curve is: memory helps a lot while the hit rate is poor, helps less as it climbs, and then the moment the whole model fits, the streaming machinery and everything it costs in CPU goes away at once.&lt;/p&gt;

&lt;p&gt;For comparison, the project's published ladder for GLM-5.2 at 744B runs from 5.8 to 6.8 tokens per second on six RTX 5090s with full residency, 1.8 on a 128 GB CPU-only desktop, 1.07 on a laptop-class box with an RTX 5070 Ti, and 0.05 to 0.1 on a 25 GB machine reading cold. Its DeepSeek-specific figure is about 1.6 tokens per second at 3k context on an RTX 5080 with two NVMe drives. The 128 GB desktop datapoint is worth sitting with, since it is 62 percent faster than this laptop and the obvious difference is memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this does not say
&lt;/h2&gt;

&lt;p&gt;One machine, one model, one prompt. Four runs, with repeatability of 0.2% between the two at identical settings, which is enough to trust the comparisons and nowhere near a benchmark suite.&lt;/p&gt;

&lt;p&gt;It is also one engine, and that is the gap I would most want closed. The CPU time is real work, but it could be a property of this engine's hand-written kernels rather than of the workload itself. Running a comparable MoE through a second engine on the same box would separate those, and I have not done it. Until someone does, read the compute finding as being about this engine on this machine.&lt;/p&gt;

&lt;p&gt;I also have no instruction counts, no device-level I/O trace, and no memory bandwidth figure, so where the CPU time actually goes is unmeasured. The share attributed to miss servicing comes from the difference between two runs, not from a profile.&lt;/p&gt;

&lt;p&gt;And 1.11 tokens per second is slow. A 600-word answer is around ten minutes, extrapolated from a 200-token run rather than measured. That is workable for queued jobs and unusable for a conversation.&lt;/p&gt;

&lt;h2&gt;
  
  
  So can an engineer do this on their own machine
&lt;/h2&gt;

&lt;p&gt;Yes, at 200 tokens in three minutes, which decides what it is good for.&lt;/p&gt;

&lt;p&gt;Checking whether a model follows a style guide, holds context across a long document, or handles a tool-calling loop does not need throughput. It needs the model, a prompt and some patience. Queue the job, do something else, read the answer. Anything conversational is out.&lt;/p&gt;

&lt;p&gt;The arithmetic that decides whether it will run at all is simple enough to do before installing anything. The dense set has to be resident, the expert cache wants whatever memory is left, and the free disk has to hold the whole checkpoint. This model needed 6.2 GB resident and 85 GB on disk, and 61 GB of system memory left 39.6 GB for the engine to divide. At 20 GB it still worked, 15% slower. On a 32 GB machine there would be very little cache left after the dense set, and the results above suggest that gets painful rather than impossible.&lt;/p&gt;

&lt;p&gt;Storage is the part to worry about least. Any current NVMe is well past the point where the drive is what slows you down, and the two bounds are the argument: even after tripling the miss rate on purpose, device time was 13% of the run. Between a faster drive and more memory, the memory is what moves the number, and about half of what it buys is CPU not spent servicing misses.&lt;/p&gt;

&lt;p&gt;The ceiling is worth naming too. Once the whole model fits in memory there is no streaming, no cache and no miss rate, and this checkpoint is 84.7 GB. Below that line you are trading against the cache. Above it the entire mechanism this post is about stops being relevant.&lt;/p&gt;

&lt;p&gt;I should say what this does to our own position, since Sakura Sky runs managed environments for open-weight models and the measurement above cuts against part of the pitch.&lt;/p&gt;

&lt;p&gt;Weights of this class run on an engineer's laptop. Anyone who has been told that large open models are out of reach without renting hardware has been told something that stopped being true. That makes local evaluation free, and it removes the excuse for not doing it before committing to anything.&lt;/p&gt;

&lt;p&gt;What the laptop does not give you is fifty people using it at once, isolation between them, an audit trail, a recovery path, or an answer in under ten minutes. None of that is affected by anything measured here. What has changed is that nobody has to rent hardware to find out whether the model is any good first. Which for me, was the real win from this.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproducing it
&lt;/h2&gt;

&lt;p&gt;The engine is at &lt;a href="https://github.com/JustVugg/colibri" rel="noopener noreferrer"&gt;github.com/JustVugg/colibri&lt;/a&gt;, with its &lt;a href="https://github.com/JustVugg/colibri/blob/main/docs/quickstart.md" rel="noopener noreferrer"&gt;quick start&lt;/a&gt;, &lt;a href="https://github.com/JustVugg/colibri/blob/main/docs/tuning.md" rel="noopener noreferrer"&gt;tuning notes&lt;/a&gt; and &lt;a href="https://github.com/JustVugg/colibri/blob/main/docs/benchmarks.md" rel="noopener noreferrer"&gt;benchmark protocol&lt;/a&gt; in the same repository. The checkpoint is &lt;a href="https://huggingface.co/puwaer/DeepSeek-V4-Flash-0731-reap-150b" rel="noopener noreferrer"&gt;puwaer/DeepSeek-V4-Flash-0731-reap-150b&lt;/a&gt; on Hugging Face, and the &lt;a href="https://github.com/JustVugg/colibri/blob/main/docs/deepseek-v4.md" rel="noopener noreferrer"&gt;engine notes for that family&lt;/a&gt; cover the flags used here. Neither the engine nor the checkpoint needed a conversion step, so an afternoon is a build, an 85 GB download and a run.&lt;/p&gt;

&lt;p&gt;If you run it, record what the project asks for: hardware, commit, model container, exact command, prompt, cache state, throughput, time to first token, expert hit rate and bytes read. Add the build line and the page cache state, since a warm cache will hand you a number that means nothing. What I still cannot give you is an instruction count or a device-level I/O trace, and both would say more about where the CPU time goes than anything I measured.&lt;/p&gt;

&lt;p&gt;The comparison worth making is not tokens per second between machines. It is the ratio of device time to wall clock on your own, because that tells you whether your next purchase should be a drive, more memory, or neither.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Disclosure: Sakura Sky designs and operates managed data, security and AI platforms for customers, including managed environments for open-weight models, which is a commercial interest in this subject and one the argument above partly cuts against. Colibrì is not a Sakura Sky project and we have no relationship with its author. The REAP-pruned checkpoint was published by a third party and its output quality was not evaluated here. All measurements come from a single machine on 16 September 2026 and are reproducible from the commands shown, which is not the same as being representative.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;Fornaro, V. (2026a) &lt;em&gt;colibrì: run frontier MoE models on hardware you already own&lt;/em&gt;. Available at: &lt;a href="https://github.com/JustVugg/colibri" rel="noopener noreferrer"&gt;https://github.com/JustVugg/colibri&lt;/a&gt; (Accessed: 16 September 2026).&lt;/p&gt;

&lt;p&gt;Fornaro, V. (2026b) &lt;em&gt;GLM-5.3-Flash engine (c/glm53.c)&lt;/em&gt;. Available at: &lt;a href="https://github.com/JustVugg/colibri/blob/main/docs/glm53-flash.md" rel="noopener noreferrer"&gt;https://github.com/JustVugg/colibri/blob/main/docs/glm53-flash.md&lt;/a&gt; (Accessed: 16 September 2026).&lt;/p&gt;

&lt;p&gt;Lasby, M., Lazarevich, I., Sinnadurai, N., Lie, S., Ioannou, Y. and Thangarasa, V. (2025) 'REAP the Experts: Why Pruning Prevails for One-Shot MoE compression', &lt;em&gt;arXiv:2510.13999&lt;/em&gt;. Available at: &lt;a href="https://arxiv.org/abs/2510.13999" rel="noopener noreferrer"&gt;https://arxiv.org/abs/2510.13999&lt;/a&gt; (Accessed: 16 September 2026).&lt;/p&gt;

&lt;p&gt;puwaer (2026) &lt;em&gt;DeepSeek-V4-Flash-0731-reap-150b&lt;/em&gt;. Available at: &lt;a href="https://huggingface.co/puwaer/DeepSeek-V4-Flash-0731-reap-150b" rel="noopener noreferrer"&gt;https://huggingface.co/puwaer/DeepSeek-V4-Flash-0731-reap-150b&lt;/a&gt; (Accessed: 16 September 2026).&lt;/p&gt;

</description>
      <category>ai</category>
      <category>optimization</category>
      <category>architecture</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Tollgate v0.2.3: What Long-Context Pricing Does to a Spend Cap</title>
      <dc:creator>Andrew Stevens</dc:creator>
      <pubDate>Fri, 11 Sep 2026 09:21:15 +0000</pubDate>
      <link>https://dev.to/sakurasky/tollgate-v023-what-long-context-pricing-does-to-a-spend-cap-ia5</link>
      <guid>https://dev.to/sakurasky/tollgate-v023-what-long-context-pricing-does-to-a-spend-cap-ia5</guid>
      <description>&lt;p&gt;Most teams find out what their long-context requests cost when the invoice arrives, and the gap is usually wider than the extra tokens explain. Some current models bill the whole request at a higher rate once the prompt crosses a size threshold, and a spend tool that stores one rate per token class has no way to represent that.&lt;/p&gt;

&lt;p&gt;At the end of August I wrote about &lt;a href="https://www.sakurasky.com/blog/tollgate-launch/" rel="noopener noreferrer"&gt;Tollgate&lt;/a&gt;, an open-source gateway that prices every LLM request by token, reserves the worst-case cost before forwarding, and refuses the request if that reservation would break a budget. Once the provider reports usage, the reservation is settled to the actual figure. That post described v0.1.1. Four releases later, v0.2.3 is out (Stevens, 2026), and most of the work in between went into the pricing rather than into the reserve and settle path itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where a rate per token stops working
&lt;/h2&gt;

&lt;p&gt;This is the part worth checking against your own provider's pricing page, because it is easy to miss and it costs money in one direction only.&lt;/p&gt;

&lt;p&gt;Google prices Gemini 2.5 Pro at $1.25 per million input tokens for prompts up to 200,000 tokens and $2.50 above it, with output going from $10.00 to $15.00 at the same threshold. Gemini 3.1 Pro Preview uses the same 200,000 token threshold and the same 2x input and 1.5x output multiples at different absolute rates (Google, 2026a). The same tiering appears on Vertex AI, which is the Google surface Tollgate proxies (Google, 2026b). What re-rates is the whole request, rather than the excess above the threshold, so a prompt one token over the line costs twice as much on its input leg as the same prompt one token under it.&lt;/p&gt;

&lt;p&gt;This is a property of individual models. Anthropic states that Claude 4.6 and later include the full one million token context window at standard pricing, and that a 900,000 token request is billed at the same per-token rate as a 9,000 token one (Anthropic, 2026). A provider that tiered last year may stop, and a new model family may arrive tiered.&lt;/p&gt;

&lt;p&gt;For a proxy sitting in front of the model, that creates a specific problem. A price, as most systems store one, is one rate per token class. It cannot hold a rate that changes with request size. A proxy that stores "input costs $1.25 per million" and multiplies will under-charge every long request by the tier multiple, without any error to show for it. The ledger then disagrees with the invoice by an amount that looks like a rounding difference right up until somebody runs a long-context workload, at which point the same defect shows up as a several-thousand-dollar variance nobody can account for.&lt;/p&gt;

&lt;p&gt;Migration 0010, in v0.2.2, puts a threshold and separate input and output multiples on the price row, resolved field by field against a deployment default. The tier applies to the reservation as well as to the settlement, so the two cannot drift apart on exactly the largest requests, which are the ones where drift is expensive.&lt;/p&gt;

&lt;p&gt;Tiering stays off until you configure it, per model.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;tollgate admin price &lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--provider&lt;/span&gt; vertex &lt;span class="nt"&gt;--model&lt;/span&gt; gemini-2.5-pro &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--long-context-threshold&lt;/span&gt; 200000 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--long-context-input-permille&lt;/span&gt; 2000 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--long-context-output-permille&lt;/span&gt; 1500
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Defaulting to an uplift would invent a charge and break the meaning of the unit, since a million tokens would stop costing the per-million rate. Until the tier is set, a prompt above the threshold is logged as a possible under-charge rather than re-rated. I would rather leave a gap somebody can see than fill it with a number I made up.&lt;/p&gt;

&lt;h2&gt;
  
  
  What else changed since v0.1.1
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;It speaks the protocols the clients already speak.&lt;/strong&gt; There is a native Anthropic route at &lt;code&gt;POST /v1/messages&lt;/code&gt;, which takes the key as &lt;code&gt;x-api-key&lt;/code&gt; and returns the upstream body unwrapped, and an OpenAI-compatible route at &lt;code&gt;POST /v1/chat/completions&lt;/code&gt; that takes a bearer token. The Anthropic SDK, OpenAI clients, LiteLLM and Google ADK agents route through it by changing a base URL and a key, with no Tollgate SDK to adopt. Where a request does cross protocols, as an OpenAI-shaped request to a Google model does, the cost is metered from the usage the provider reports rather than from the shape of the request, because token accounting is usually the first thing a translation layer loses. Tollgate builds on proxy and cost-control patterns that LiteLLM established (BerriAI, 2026).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Streaming is metered from the stream.&lt;/strong&gt; Usage is read from the provider's own end-of-stream reporting rather than estimated afterwards. A stream that does not finish cleanly is charged its full reservation rather than going unbilled, and the row is recorded as &lt;code&gt;estimated&lt;/code&gt;, with &lt;code&gt;/console/usage&lt;/code&gt; reporting &lt;code&gt;measured_cost&lt;/code&gt; alongside &lt;code&gt;estimated_cost&lt;/code&gt;. Keeping the measured figure separate from the assumed one has been more useful to me than the total, because it tells you how much of a month's spend number rests on an estimate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt-cache classes are priced per class, per model.&lt;/strong&gt; A class left unpriced falls back to a conservative multiple of the input rate, 1x on reads and 2x on writes, and logs a warning naming the model. Several under-charges that had been present from the first release were fixed across v0.1.2 and v0.2.0: Gemini thinking tokens and tool-use prompt tokens were not billed at all, reasoning tokens through Vertex's OpenAI shim were not billed, and Anthropic cached requests were under-counted by up to 99%.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Budgets recover from losing their cache.&lt;/strong&gt; Budget counters live in Valkey for speed and are rebuilt from the Postgres ledger. The reserve script used &lt;code&gt;INCRBY&lt;/code&gt;, which recreates a missing key at the reserve amount, so a flush, a failover to a cold replica or an eviction restarted every budget's period with nothing logged, and left it restarted until somebody bounced the gateway. An absent counter carries no information about the spend that preceded it, and &lt;code&gt;INCRBY&lt;/code&gt; treats it as a counter at zero. Reserve now refuses on an absent counter, rebuilds it from the ledger and retries once, and refuses the request outright if the ledger cannot be reached.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A refusal says whose fault it is.&lt;/strong&gt; &lt;code&gt;x-tollgate-reason&lt;/code&gt; now uses one vocabulary on every route. The native Anthropic path used to emit the SDK's error kind, so a budget refusal on &lt;code&gt;/v1/messages&lt;/code&gt; reported &lt;code&gt;permission_error&lt;/code&gt; while the same refusal on &lt;code&gt;/v1/chat/completions&lt;/code&gt; reported &lt;code&gt;budget_exceeded&lt;/code&gt;, and an alert keyed on the latter would have missed every Claude refusal. v0.2.3 extends that to failures. A pre-flight count that fails is now reported as an upstream error with a &lt;code&gt;502&lt;/code&gt; rather than a backend error with a &lt;code&gt;503&lt;/code&gt;, which sends the operator to the provider instead of to Valkey and Postgres. A count endpoint that rejects a malformed body surfaces as a &lt;code&gt;400&lt;/code&gt; instead of a &lt;code&gt;502&lt;/code&gt;, which SDKs had been retrying three times before anyone saw the real fault.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Privileged changes leave a trail.&lt;/strong&gt; The &lt;code&gt;audit_log&lt;/code&gt; table has existed with append-only triggers since migration 0002, and nothing wrote to it. Key issuance, revocation, budget changes and price changes now each append a row. The principal recorded is the OS user and host the command ran as, which is enough to line a change up against a shell history. It is not evidence that anyone was authenticated, and &lt;code&gt;SECURITY.md&lt;/code&gt; says so.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you are already running v0.1.x
&lt;/h2&gt;

&lt;p&gt;Three of the fixes above change numbers you may have relied on, so they are worth acting on rather than reading past.&lt;/p&gt;

&lt;p&gt;If you ran cache-heavy Anthropic traffic, or Gemini traffic using thinking or tools, on v0.1.1, the recorded cost for that period is low against the invoice, in the Anthropic cache case by as much as 99%. Reconcile that window against the bill rather than trusting the ledger for it.&lt;/p&gt;

&lt;p&gt;If your Valkey instance was flushed, evicted or failed over to a cold replica while a budget period was open, that period restarted at zero and stayed restarted until the gateway was next bounced. Any cap that looked comfortable during such a window was not being applied as configured.&lt;/p&gt;

&lt;p&gt;If you use &lt;code&gt;exact&lt;/code&gt; admission on Vertex, it was refusing requests rather than counting them, because the body it sent carried fields the count endpoint rejects.&lt;/p&gt;

&lt;p&gt;The upgrade path is &lt;code&gt;tollgate admin migrate&lt;/code&gt; before starting the new binary, migration 0010 included, and a review of budgets before you start, because cache-heavy workloads will show higher spend once the classes are priced properly. Finish the rollout before relying on the counter repair. A v0.2.0 instance still recreates a missing counter at the reserve amount, so during a mixed rollout the old behaviour is live on the old instances.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three bugs found by measuring the provider
&lt;/h2&gt;

&lt;p&gt;Three of the fixes in v0.2.3 came out of a script that calls the providers' token-counting endpoints and prints what comes back. None were found by reading the code, and two contradicted a reading of the documentation that looked obvious at the time.&lt;/p&gt;

&lt;p&gt;Vertex &lt;code&gt;:countTokens&lt;/code&gt; rejects &lt;code&gt;safetySettings&lt;/code&gt;, &lt;code&gt;labels&lt;/code&gt;, &lt;code&gt;toolConfig&lt;/code&gt; and &lt;code&gt;cachedContent&lt;/code&gt; with a 400. An SDK-built body routinely carries &lt;code&gt;safetySettings&lt;/code&gt;, so for most callers &lt;code&gt;exact&lt;/code&gt; admission on Vertex refused the request instead of counting it, and the code that built the body looked reasonable. The endpoint accepts and counts &lt;code&gt;contents&lt;/code&gt;, &lt;code&gt;systemInstruction&lt;/code&gt;, &lt;code&gt;tools&lt;/code&gt; and &lt;code&gt;generationConfig&lt;/code&gt;, which is what is now sent.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;generationConfig&lt;/code&gt; being counted matters more than it sounds. A &lt;code&gt;responseSchema&lt;/code&gt; lives inside it and is billed as prompt material. A one-word message counted 1 token bare and 51 tokens with a small four-property schema attached. The schema had been dropped from the count on the reasoning that it describes the output rather than the input. That reasoning is tidy and it is wrong, and it under-reserved every structured-output request on Vertex.&lt;/p&gt;

&lt;p&gt;The third was not a counting bug at all. &lt;code&gt;reqwest&lt;/code&gt; has been built without its default features since the first release, to make the TLS backend an explicit choice, and that took &lt;code&gt;http2&lt;/code&gt; out with it. Every outbound provider request in every release up to now went over HTTP/1.1, with no connection multiplexing, on the busiest path in the product. A dependency's default feature list is not something most of us re-read after the first time we set it.&lt;/p&gt;

&lt;p&gt;Where a provider's behaviour is the ground truth, measure it. Reasoning about which fields cannot possibly affect a token count is how two of those three bugs were written.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it still does not do
&lt;/h2&gt;

&lt;p&gt;The README carries a long limitations list on purpose, and it is the part I would read before the feature list.&lt;/p&gt;

&lt;p&gt;The cap is a reservation, and a settlement can exceed it. Server-injected tool-use tokens, prompts that name a stored context cache, and long-context re-rating can all settle above what was reserved, and under the default &lt;code&gt;fast&lt;/code&gt; admission the input estimate is deliberately rough. What the design aims for is direction rather than precision. It errs upward, because under-charging is the failure that makes a spend control worth nothing.&lt;/p&gt;

&lt;p&gt;Provider-side context cache creation and storage are not visible from the request path. They bill against the cache resource, which no proxy in this position sees.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;/metrics&lt;/code&gt; reports something as of v0.2.3, where it previously exposed a liveness gauge alone: request totals by decision, a cost total, and the configured limit for each budget. A scrape reads process memory only, and does not query the ledger, because an observability path that depends on Postgres and Valkey stops answering during the incident you would be using it to diagnose.&lt;/p&gt;

&lt;p&gt;Streaming under &lt;code&gt;exact&lt;/code&gt; admission is refused on the OpenAI route, and supported on the native Anthropic one. External media by reference is refused on both. The ledger drops partitions older than 90 days by default, which is a retention setting worth looking at before it becomes a gap in your own reporting.&lt;/p&gt;

&lt;p&gt;It is beta and pre-1.0. Interfaces and the schema may still change.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it without infrastructure
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;cargo run &lt;span class="nt"&gt;--&lt;/span&gt; demo
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No Postgres, no Valkey, no cloud credentials. It prints an API key and two budgets, serves a mock provider and the console, and you can watch the per-key meter climb and hard-stop.&lt;/p&gt;

&lt;p&gt;The open-source core remains single-operator and single-tenant, with keys and budgets managed through the CLI and an observe-only console. A Tollgate Enterprise edition will follow for teams running across several tenants and operators, adding web and API management, role-based access control and SSO, budget-change approval workflows, multi-org scoping, and audit export. If that edition is relevant to your environment, a &lt;a href="https://github.com/sakura-sky/tollgate/discussions" rel="noopener noreferrer"&gt;GitHub Discussion&lt;/a&gt; or a note to Sakura Sky is the place to say so.&lt;/p&gt;

&lt;p&gt;The repository is at &lt;a href="https://github.com/sakura-sky/tollgate" rel="noopener noreferrer"&gt;github.com/sakura-sky/tollgate&lt;/a&gt; and the release is &lt;a href="https://github.com/sakura-sky/tollgate/releases/tag/v0.2.3" rel="noopener noreferrer"&gt;v0.2.3&lt;/a&gt;. If you put a real budget in front of real traffic and the ledger disagrees with your invoice, open an issue with the shape of the workload. That is the report I want most, because the failure this release was built around is one you only see on the bill. Please keep the &lt;code&gt;SECURITY.md&lt;/code&gt; inbox for vulnerability reports.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Disclosure: I created Tollgate. It is published as a Sakura Sky open-source project under MIT, with copyright held by me and its contributors, and there is no contributor licence agreement. Sakura Sky intends to offer a commercial Enterprise edition alongside the free core, which is a commercial interest worth stating. Sakura Sky also builds and runs managed LLM gateway environments for customers, including LiteLLM deployments, and continues to do so; Tollgate is a narrower tool for a specific job and is not a migration recommendation for anyone. The provider prices quoted here were checked on the dates given in the references and change often.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;Anthropic (2026) &lt;em&gt;Pricing&lt;/em&gt;. Available at: &lt;a href="https://platform.claude.com/docs/en/about-claude/pricing" rel="noopener noreferrer"&gt;https://platform.claude.com/docs/en/about-claude/pricing&lt;/a&gt; (Accessed: 11 September 2026).&lt;/p&gt;

&lt;p&gt;BerriAI (2026) &lt;em&gt;LiteLLM&lt;/em&gt;. Available at: &lt;a href="https://github.com/BerriAI/litellm" rel="noopener noreferrer"&gt;https://github.com/BerriAI/litellm&lt;/a&gt; (Accessed: 11 September 2026).&lt;/p&gt;

&lt;p&gt;Google (2026a) &lt;em&gt;Gemini Developer API pricing&lt;/em&gt;. Available at: &lt;a href="https://ai.google.dev/gemini-api/docs/pricing" rel="noopener noreferrer"&gt;https://ai.google.dev/gemini-api/docs/pricing&lt;/a&gt; (Accessed: 11 September 2026).&lt;/p&gt;

&lt;p&gt;Google (2026b) &lt;em&gt;Vertex AI pricing for generative AI&lt;/em&gt;. Available at: &lt;a href="https://cloud.google.com/vertex-ai/generative-ai/pricing" rel="noopener noreferrer"&gt;https://cloud.google.com/vertex-ai/generative-ai/pricing&lt;/a&gt; (Accessed: 11 September 2026).&lt;/p&gt;

&lt;p&gt;Stevens, A. (2026) &lt;em&gt;Tollgate v0.2.3&lt;/em&gt;. Available at: &lt;a href="https://github.com/sakura-sky/tollgate/releases/tag/v0.2.3" rel="noopener noreferrer"&gt;https://github.com/sakura-sky/tollgate/releases/tag/v0.2.3&lt;/a&gt; (Accessed: 11 September 2026).&lt;/p&gt;

</description>
      <category>ai</category>
      <category>finops</category>
      <category>rust</category>
      <category>gcp</category>
    </item>
    <item>
      <title>Tollgate: Open-Source Hard Limits for LLM Spend</title>
      <dc:creator>Andrew Stevens</dc:creator>
      <pubDate>Mon, 31 Aug 2026 12:39:10 +0000</pubDate>
      <link>https://dev.to/sakurasky/tollgate-open-source-hard-limits-for-llm-spend-2cah</link>
      <guid>https://dev.to/sakurasky/tollgate-open-source-hard-limits-for-llm-spend-2cah</guid>
      <description>&lt;p&gt;In my experience most teams find out what their LLM workload costs from the invoice. Usage grows, a retry loop misbehaves, a feature ships with a larger context window, and the month closes on a number nobody forecast. Dashboards and budget alerts do not prevent that. They report money that has already left.&lt;/p&gt;

&lt;p&gt;Tollgate turns the budget into an admission decision, taken before the request reaches the model. It is an open-source AI gateway and spend-control proxy, out now as v0.1.1 under MIT (Stevens, 2026).&lt;/p&gt;

&lt;h2&gt;
  
  
  What it does
&lt;/h2&gt;

&lt;p&gt;Tollgate sits between your applications and their LLM providers, currently Anthropic and Google Vertex, both config-gated, plus a built-in mock for testing. There is no OpenAI adapter yet. Each request is authenticated, priced by token, and checked against every budget that applies to it before anything is forwarded. If a hard cap would be exceeded the request is refused and never reaches the provider. Spend is written to a durable, append-only ledger in Postgres, and a read-only web console shows live budget meters, the tokens-to-cost breakdown, and the gateway's own added latency.&lt;/p&gt;

&lt;p&gt;It is written in Rust and deploys to Google Cloud Run. The Terraform module in the repository provisions the full stack and keeps the database URL and API-key pepper in Secret Manager, injected by reference rather than as plaintext environment values.&lt;/p&gt;

&lt;p&gt;A refused request returns &lt;code&gt;402 Payment Required&lt;/code&gt; with an &lt;code&gt;x-tollgate-reason&lt;/code&gt; header. An unroutable or unpriced model returns &lt;code&gt;400&lt;/code&gt;, also with a reason header.&lt;/p&gt;

&lt;p&gt;That is worth unpacking, because the obvious alternative is a trap and the choice made carries its own caveat. &lt;code&gt;429&lt;/code&gt; is the rate-limiting code, and provider SDKs commonly auto-retry it, so a monthly cap answering &lt;code&gt;429&lt;/code&gt; would invite well-behaved clients to keep knocking at a gate that cannot open for weeks. &lt;code&gt;402&lt;/code&gt; avoids that. But &lt;code&gt;402&lt;/code&gt; is reserved rather than standardised, and the x402 payment standard, now a Linux Foundation project, is busy teaching agent clients that a &lt;code&gt;402&lt;/code&gt; means pay and retry (x402 Foundation, 2026). Tollgate advertises no payment challenge and its callers authenticate with an issued key, so x402 middleware has nothing to act on here. If you run agents with payment middleware in the stack, check that interaction before relying on the status code alone. The reason header is there to be read.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it does not do yet
&lt;/h2&gt;

&lt;p&gt;The README carries a Limitations section. Three entries decide whether Tollgate fits a given workload today, so they belong here rather than three clicks away.&lt;/p&gt;

&lt;p&gt;There is no streaming. Requests must be non-streaming, because usage cannot be metered from a partial stream in this release.&lt;/p&gt;

&lt;p&gt;In the default admission mode, requests referencing external media by &lt;code&gt;fileUri&lt;/code&gt;, &lt;code&gt;file_id&lt;/code&gt; or URL are refused with a &lt;code&gt;400&lt;/code&gt;. The token cost of fetched media is not bounded by the size of the request body, so the reservation would be far too low to be safe. Those requests need &lt;code&gt;exact&lt;/code&gt; admission, described below. If your traffic is multimodal, that decides your configuration on day one.&lt;/p&gt;

&lt;p&gt;Prompt-caching cost is approximate. Cache-read and cache-write token classes are not yet priced separately, so on cache-heavy workloads the recorded cost can diverge from the provider's bill. Standard non-caching usage is exact. Metrics are minimal too: &lt;code&gt;/metrics&lt;/code&gt; exposes only a liveness gauge, so Postgres and the console are where spend actually lives.&lt;/p&gt;

&lt;h2&gt;
  
  
  How a request moves through it
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fwww.sakurasky.com%2Fimages%2Fblog%2Fmermaid-diagram-2026-08-30-tollgate-request-lifecycle.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fwww.sakurasky.com%2Fimages%2Fblog%2Fmermaid-diagram-2026-08-30-tollgate-request-lifecycle.png" title="Figure 1. The Tollgate request lifecycle: authenticate, price, reserve, forward, settle, with the hard stop landing before the provider call." alt="Sequence diagram of the Tollgate request lifecycle. A client sends a POST to /v1/{provider} carrying an x-tollgate-key header. Tollgate authenticates the key with HMAC-SHA256, prices the request from tokens to cost, and asks Valkey to reserve the worst-case cost. If the reservation would exceed a hard cap, Valkey denies it and Tollgate returns 402 Payment Required to the client without ever contacting the provider. If the request is within budget, Tollgate forwards it to the provider, receives the response and reported token usage, settles the reservation to the actual cost in Valkey, and returns 200 to the client with the cost breakdown and an x-tollgate-overhead-us header." width="800" height="590"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Pricing happens before the forward, which is what separates an enforceable cap from an advisory one. The worst-case cost of a request is reserved up front as a single atomic check-and-increment against Valkey, and once the provider responds with reported token usage the reservation is settled to the actual figure.&lt;/p&gt;

&lt;p&gt;Four properties hold that together. Money is handled in integer micros with no floating point anywhere on the cost path. Postgres is the system of record and the Valkey counters are only a hot-path cache, reconciled from the ledger at startup in a direction that can raise a counter but never lower it. A request matching no budget at all is denied, so a missing configuration row cannot leave a deployment unenforced without anyone noticing. And a cache failure fails closed, which is the right call for a spend gate but carries a real cost: Tollgate is an in-path dependency, so if Valkey is unreachable your LLM traffic stops with a &lt;code&gt;503&lt;/code&gt; rather than flowing unmetered.&lt;/p&gt;

&lt;p&gt;Budgets resolve by scope: a specific key, a provider, a model, or the mandatory global backstop. Periods are daily, weekly, or monthly, and a request has to clear every scope that applies to it, so a per-key cap and a deployment-wide monthly ceiling operate together.&lt;/p&gt;

&lt;p&gt;Two admission modes decide how strict the cap really is. Output tokens are bounded by the request's own &lt;code&gt;max_tokens&lt;/code&gt; or &lt;code&gt;maxOutputTokens&lt;/code&gt;. Input tokens are counted either by &lt;code&gt;fast&lt;/code&gt;, the default, which estimates from the request body with no extra call and on token-dense input can settle a little over the cap, or by &lt;code&gt;exact&lt;/code&gt;, which makes a pre-flight token-count call to the provider and enforces strictly. &lt;code&gt;exact&lt;/code&gt; costs a round trip, and it adds a second provider dependency: if the count call fails the request fails with a &lt;code&gt;503&lt;/code&gt; rather than falling back to the weaker estimate. If you need the cap strict to the token, that is the trade.&lt;/p&gt;

&lt;h2&gt;
  
  
  What v0.1.1 gives you
&lt;/h2&gt;

&lt;p&gt;Per-key and global token budgets with a hard stop before the provider. Token-level cost accounting, currency-agnostic and integer-precise. Anthropic and Vertex adapters plus the mock. Budget and price changes that apply within the reload interval without a restart. Ledger retention by monthly partition, with partitions created ahead of need and old ones dropped. Per-request gateway latency on an &lt;code&gt;x-tollgate-overhead-us&lt;/code&gt; header, with median and p95 in the console, so the overhead figure for your own traffic is something you read rather than something I assert.&lt;/p&gt;

&lt;p&gt;To see it without provisioning anything, &lt;code&gt;cargo run -- demo&lt;/code&gt; boots an in-memory gateway with the mock provider, prints a key and two budgets, and serves the console on loopback.&lt;/p&gt;

&lt;p&gt;Tollgate builds on proxy and cost-control patterns that LiteLLM established (BerriAI, 2026), implemented in Rust.&lt;/p&gt;

&lt;h2&gt;
  
  
  Open core
&lt;/h2&gt;

&lt;p&gt;The open-source core is single-operator and single-tenant by design. Keys and budgets are managed through the CLI, the console is observe-only with no write endpoints, and authorization on the observe endpoints is flat: any valid key can view the whole deployment's budgets and usage. Enforcement stays per key. A Tollgate Enterprise edition will follow for teams running across many tenants and operators, adding web and API management of keys and budgets, role-based access control and SSO, budget-change approval workflows, multi-org scoping, richer analytics, and audit export.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;The repository is at &lt;a href="https://github.com/sakura-sky/tollgate" rel="noopener noreferrer"&gt;github.com/sakura-sky/tollgate&lt;/a&gt; and the release is &lt;a href="https://github.com/sakura-sky/tollgate/releases/tag/v0.1.1" rel="noopener noreferrer"&gt;v0.1.1&lt;/a&gt;. Run the demo, and open an issue if it breaks in your environment. This is beta and pre-1.0, so interfaces and the schema may still change, and a report from someone who put a real budget in front of real traffic is worth more to me than any amount of design discussion. If the Enterprise edition is relevant to you, open a &lt;a href="https://github.com/sakura-sky/tollgate/discussions" rel="noopener noreferrer"&gt;GitHub Discussion&lt;/a&gt; or contact Sakura Sky. Please keep the &lt;code&gt;SECURITY.md&lt;/code&gt; inbox for vulnerability reports.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Disclosure: I created Tollgate. It is published as a Sakura Sky open-source project under MIT, with copyright held by me and its contributors, and there is no contributor licence agreement. Sakura Sky intends to offer a commercial Enterprise edition alongside the free core, which is a commercial interest worth stating. Sakura Sky also builds and runs managed LLM gateway environments for customers, including LiteLLM deployments, and continues to do so; Tollgate is a narrower tool for a specific job and is not a migration recommendation for anyone.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;BerriAI (2026) &lt;em&gt;LiteLLM&lt;/em&gt;. Available at: &lt;a href="https://github.com/BerriAI/litellm" rel="noopener noreferrer"&gt;https://github.com/BerriAI/litellm&lt;/a&gt; (Accessed: 29 August 2026).&lt;/p&gt;

&lt;p&gt;Stevens, A. (2026) &lt;em&gt;Tollgate v0.1.1&lt;/em&gt;. Available at: &lt;a href="https://github.com/sakura-sky/tollgate/releases/tag/v0.1.1" rel="noopener noreferrer"&gt;https://github.com/sakura-sky/tollgate/releases/tag/v0.1.1&lt;/a&gt; (Accessed: 29 August 2026).&lt;/p&gt;

&lt;p&gt;x402 Foundation (2026) &lt;em&gt;x402: an open standard for internet-native payments&lt;/em&gt;. The Linux Foundation. Available at: &lt;a href="https://x402.org" rel="noopener noreferrer"&gt;https://x402.org&lt;/a&gt; (Accessed: 29 August 2026).&lt;/p&gt;

</description>
      <category>ai</category>
      <category>finops</category>
      <category>rust</category>
      <category>gcp</category>
    </item>
    <item>
      <title>GATE v1.5 Roadmap</title>
      <dc:creator>Andrew Stevens</dc:creator>
      <pubDate>Mon, 31 Aug 2026 08:55:18 +0000</pubDate>
      <link>https://dev.to/sakurasky/gate-v15-roadmap-1k3d</link>
      <guid>https://dev.to/sakurasky/gate-v15-roadmap-1k3d</guid>
      <description>&lt;p&gt;The &lt;a href="https://www.sakurasky.com/blog/gate-roadmap-v1-4/" rel="noopener noreferrer"&gt;v1.4 roadmap&lt;/a&gt; named six things: C20 output validation, break-glass formalisation, the OWASP AISVS and MITRE ATLAS mappings, gate-rust, gate-fuzz, and gate-knowledge. Five shipped in full, although C20 shipped with a streaming constraint I come back to below. gate-fuzz shipped one of its three deliverables, the Python-to-Rust differential harness. The bundle-derived strategy generator and the HTTP-level protocol fuzzer moved to v1.5, and the &lt;a href="https://www.sakurasky.com/blog/gate-v1-4-release/" rel="noopener noreferrer"&gt;v1.4 release post&lt;/a&gt; named the deferral on the day. v1.5 delivers them.&lt;/p&gt;

&lt;p&gt;v1.4 completed the boundary model: identity at instantiation, policy at execution, observation throughout, classification at delivery. v1.5 is about depth. Certification alignment, streaming, a reference implementation someone can actually run, and paying down the debt v1.4 named.&lt;/p&gt;

&lt;p&gt;Here is what is coming.&lt;/p&gt;

&lt;h2&gt;
  
  
  ISO/IEC 42001, mapped with the asymmetry written down
&lt;/h2&gt;

&lt;p&gt;GATE has claimed alignment with ISO/IEC 42001 since the first release, on the strength of a high-level theme table in the Standard Mappings appendix. A theme table is enough to start a conversation with an auditor and not enough to survive one.&lt;/p&gt;

&lt;p&gt;v1.5 replaces it with a per-control mapping from C01 through C20 to 42001 clauses and Annex A controls, researched against the published standard (ISO/IEC, 2023) rather than against secondary summaries. It ships as gate-knowledge documents plus a machine-readable &lt;code&gt;gate-conformance/mappings/iso-42001.yaml&lt;/code&gt;, in the same shape as the AISVS and ATLAS mappings but not with the same guarantee behind it. Those two validate against pinned open upstreams that the repository can carry a snapshot of. 42001 is a paywalled standard, so the yaml pins the edition and carries clause and Annex A identifiers, and the validator checks internal consistency rather than resolving each title against a shipped copy. That is a weaker check and the mapping header will say so.&lt;/p&gt;

&lt;p&gt;The asymmetry document is the part that matters. 42001 is a management system standard: it asks for leadership commitment, a documented AI policy, competence and communication, supplier governance, and lifecycle planning that extends well past runtime. GATE addresses none of that, and it should not pretend to. The document lists each requirement 42001 makes that GATE does not touch, so a team reading the mapping can see the shape of the gap instead of inferring it. GATE conformance does not imply 42001 conformance, and after v1.5 the mapping will say exactly where the difference lies.&lt;/p&gt;

&lt;p&gt;Sakura Sky plans a companion technical note alongside it, for CISOs reading the mapping in a 42001 certification context. It addresses the question the asymmetry document raises, which is what an auditor will still ask for once GATE is in place.&lt;/p&gt;

&lt;h2&gt;
  
  
  Streaming-aware C20
&lt;/h2&gt;

&lt;p&gt;v1.4 ships output validation with a named constraint: streaming is disabled at the &lt;code&gt;high_privilege&lt;/code&gt; tier for any agent whose action matrix can produce a hold or review obligation. You cannot classify a response you have not finished generating, and you cannot un-send a token. Rather than hide that, v1.4 wrote it into the spec as a limitation.&lt;/p&gt;

&lt;p&gt;v1.5 closes it. Chunk-level classification with a sentence-boundary default, buffer-then-deliver semantics per chunk, per-tier classifier latency budgets so the buffering cost is bounded and stated, and a partial-classification event contract so the evidence trail covers a streamed response as well as it covers a single one.&lt;/p&gt;

&lt;p&gt;Buffer-then-deliver is the design decision underneath all of that. A chunk is classified before it reaches the user, so there is nothing to take back. Token-level delivery with retroactive retraction stays out of scope. Once a user has read a token, withdrawing it does not undo the disclosure.&lt;/p&gt;

&lt;p&gt;This one touches every repository. The contract in gate-contracts, the guardrail in gate-policies, a runner sub-check in gate-conformance, and the classification path in both gate-python and gate-rust.&lt;/p&gt;

&lt;h2&gt;
  
  
  gate-fuzz: converting PARTIAL to PASS
&lt;/h2&gt;

&lt;p&gt;The runner reports 9 AUTOMATED and 11 PARTIAL by default, 11 and 9 when bundle stores are configured. Several of those PARTIAL results are not manual by nature; they are PARTIAL because closing them requires executing test scenarios a query tool cannot perform. gate-fuzz was scoped in June to close them, and v1.4 shipped the harness without the closure work. v1.5 ships the two deferred deliverables.&lt;/p&gt;

&lt;p&gt;The first is Hypothesis strategies generated from the signed C09 invariant bundle, so invariants in the bundle produce property tests without anyone writing them by hand, for the invariant types the generator covers. That changes what a passing suite means: it becomes evidence about your bundle rather than evidence about the harness that tests it. The second is the HTTP-level protocol fuzzer, which sends mutation variants at a running Tool Gateway and treats any response that is not a schema rejection or a policy denial as a finding.&lt;/p&gt;

&lt;p&gt;The framing that matters is the one the June roadmap set out and v1.4 did not deliver. With the README mapping each test to the conformance check it closes, running gate-fuzz alongside the runner converts PARTIAL results into verifiable PASS results, leaving in PARTIAL only the checks that need human judgement: controlled drills, process inspections, runbook sign-offs. v1.5 completes that scope, and the runner counts change accordingly.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gate-rust conformance bridge
&lt;/h2&gt;

&lt;p&gt;gate-rust can produce canonical JSON, envelopes, ledger events and signatures that match gate-python byte for byte, but it cannot hand the conformance runner evidence directly. A team running a Rust gateway still needs Python in the evidence path. That is its own commitment in v1.5, and one of the five known issues v1.4 named.&lt;/p&gt;

&lt;p&gt;The bridge goes through either PyO3 or a subprocess boundary, so a single test suite runs over both implementations and produces byte-level proofs at the envelope, ledger, and signing layers. That is what unlocks the envelope-hash and ledger-event-hash byte-parity properties, which account for the four tests currently skipped in the gate-fuzz suite. A Rust gateway then produces runner-compatible evidence without a Python sidecar in the path.&lt;/p&gt;

&lt;h2&gt;
  
  
  gate-adk-reference: show me the code
&lt;/h2&gt;

&lt;p&gt;The most common response to the framework paper is some version of "show me the code". The repositories answer it in pieces: a schema here, a policy bundle there, a library that implements the hashing. Nobody has published the thing those pieces add up to.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;gate-adk-reference&lt;/code&gt; is an end-to-end GATE-governed data analyst agent built on Google's Agent Development Kit (Google, 2026), with &lt;a href="https://github.com/sakura-sky/sql-guard" rel="noopener noreferrer"&gt;sql-guard&lt;/a&gt; as the deterministic policy engine at the SQL tool boundary. Eight controls wired across all four layers. A BigQuery public dataset is the headline scenario, and a documented SQLite path covers air-gapped contexts where a cloud warehouse is not available.&lt;/p&gt;

&lt;p&gt;The scope statement is part of the deliverable. This demonstrates a single integration pattern on one agent framework. It is not hardened, and copying it verbatim into production would be unsafe. That goes in the first screen of the README rather than the last.&lt;/p&gt;

&lt;h2&gt;
  
  
  A public playground
&lt;/h2&gt;

&lt;p&gt;Reading about a policy denial is not the same as watching one happen. v1.5 plans a hosted, sandboxed instance of the reference agent where you can issue a prompt and watch the GATE event stream next to the chat: the policy decision, the quality gate on retrieval, the output classification, each appearing as the agent works.&lt;/p&gt;

&lt;p&gt;Details are still being finalised, so I am not naming a hosting platform, a model, a domain, or a cost envelope here. Those land with the release.&lt;/p&gt;

&lt;h2&gt;
  
  
  AI Catalog and ARD interop
&lt;/h2&gt;

&lt;p&gt;Scoped for v1.5, gated on the ARD specification stabilising. The final in-or-out call is not made.&lt;/p&gt;

&lt;p&gt;When C17 discovers an unenrolled workload, GATE has no way to read an assurance claim the publisher has already made about it. Everything GATE knows about that workload, it has to establish itself.&lt;/p&gt;

&lt;p&gt;Google and a set of industry partners published Agentic Resource Discovery in June, an open specification for publishing, discovering, and verifying AI capabilities across the web, covering MCP servers, A2A agents and OpenAPI tools, built on the Linux Foundation AI Catalog data model (Bu and Krishnan, 2026). Its trust manifest is the cryptographic layer, carrying identity and attestations for a published capability.&lt;/p&gt;

&lt;p&gt;The work is small and bounded: an optional &lt;code&gt;catalog_refs&lt;/code&gt; field on the ABOM and enrolment contracts, a verification helper in gate-policies, a signal recorded on the C17 enrolment decision when a reference verifies, and a conformance check on the integrity of catalog references where they are present. What a verified reference proves is that the publisher controls the domain the catalog is served from, not that the capability is safe to run, so it is recorded as provenance rather than treated as a substitute for the checks C17 already makes.&lt;/p&gt;

&lt;p&gt;The boundary is explicit and I want it stated before the code exists. GATE consumes catalog entries. It does not publish them, and it does not extend the ai-catalog schema. Keeping the discovery layer and the governance layer separable means a change in one does not force a change in the other.&lt;/p&gt;

&lt;h2&gt;
  
  
  Contract debt, named and paid
&lt;/h2&gt;

&lt;p&gt;Three items from the v1.4 known issues, each scoped for closure in v1.5.&lt;/p&gt;

&lt;p&gt;The unified exception register: one &lt;code&gt;exception_record&lt;/code&gt; contract covering the exception surface that C09 break-glass and the C17, C18, and C19 policies all consume, instead of adjacent shapes added one control at a time.&lt;/p&gt;

&lt;p&gt;The multi-approver HITL Decision Record: approvals needing two or more approvers carried by the HITL record itself rather than borrowed from the break-glass path.&lt;/p&gt;

&lt;p&gt;gate-python publication to PyPI, so the reference library installs the way a Python library is expected to install.&lt;/p&gt;

&lt;p&gt;Together with streaming C20, the fuzz completion, and the conformance bridge, that accounts for every item the v1.4 release post listed under "What v1.5 picks up". None of them has been dropped without saying so.&lt;/p&gt;

&lt;h2&gt;
  
  
  The HTML edition becomes canonical
&lt;/h2&gt;

&lt;p&gt;One meta change, and the gap it closes is a maintenance one. The cloud quickstarts and the day-2 runbooks currently live inside the paper, so correcting a quickstart means cutting a release. Page-number citations rot every time the paper is re-laid out.&lt;/p&gt;

&lt;p&gt;From v1.5 the HTML edition at &lt;a href="https://deterministicagents.ai" rel="noopener noreferrer"&gt;deterministicagents.ai&lt;/a&gt; is canonical and the PDF is an automated export of it. The quickstarts and runbooks move to live surfaces, cutting roughly thirty pages from the normative core. Citations point at section anchors rather than page numbers, which stop being stable between builds.&lt;/p&gt;

&lt;h2&gt;
  
  
  Timeline
&lt;/h2&gt;

&lt;p&gt;No dates. Sequencing, which is the part I can actually commit to.&lt;/p&gt;

&lt;p&gt;The framework paper update comes first, then the contracts, then the repositories that consume them, in dependency order. The ISO 42001 mapping is the substantive content workstream and runs alongside the paper rather than behind it. The reference implementation lands before the playground, and the playground goes last because it depends on everything above it.&lt;/p&gt;

&lt;p&gt;Two open questions get settled at kickoff rather than in a later release. The AISVS candidate triage may produce new conformance checks, and if the triage confirms them they join the suite in v1.5. And the canonical-JSON float restriction documented in v1.4 is either the end state or it is not; the alternative is a shared float emitter across the Python and Rust implementations, and that call gets made before the contracts are cut.&lt;/p&gt;

&lt;p&gt;If you are building agentic AI infrastructure and any of the above addresses a gap you have hit in practice, the issues and discussions at &lt;a href="https://github.com/deterministic-agents" rel="noopener noreferrer"&gt;github.com/deterministic-agents&lt;/a&gt; are the right place to say so before the spec is locked.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;GATE is published at &lt;a href="https://deterministicagents.ai" rel="noopener noreferrer"&gt;deterministicagents.ai&lt;/a&gt; under CC BY 4.0 for the documentation and MIT for the code. The strategic companion to this framework is the &lt;a href="https://www.sakurasky.com/white-papers/trustworthy-agentic-ai-blueprint/" rel="noopener noreferrer"&gt;Trustworthy Agentic AI Blueprint&lt;/a&gt;, co-authored with Sakura Sky.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Disclosure: GATE is authored and maintained by me personally rather than by Sakura Sky, and there is no paid tier, hosted version, or commercial product built on it. Agent governance is also my day job at Sakura Sky, which is a commercial interest worth stating. Two Sakura Sky items appear above: sql-guard, the policy engine named in the reference implementation, which is developed by Sakura Sky and released under Apache-2.0, and the planned companion technical note on 42001. Everything described in this post is planned work. None of it has shipped, and scope can change before it does.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;Bu, J. and Krishnan, S. (2026) &lt;em&gt;Announcing the Agentic Resource Discovery specification&lt;/em&gt;, Google Developers Blog, 17 June. Available at: &lt;a href="https://developers.googleblog.com/announcing-the-agentic-resource-discovery-specification/" rel="noopener noreferrer"&gt;https://developers.googleblog.com/announcing-the-agentic-resource-discovery-specification/&lt;/a&gt; (Accessed: 19 August 2026).&lt;/p&gt;

&lt;p&gt;Google (2026) &lt;em&gt;Agent Development Kit (ADK)&lt;/em&gt;. Available at: &lt;a href="https://adk.dev/" rel="noopener noreferrer"&gt;https://adk.dev/&lt;/a&gt; (Accessed: 19 August 2026).&lt;/p&gt;

&lt;p&gt;ISO/IEC (2023) &lt;em&gt;ISO/IEC 42001:2023 Information technology - Artificial intelligence - Management system&lt;/em&gt;. International Organization for Standardization. Available at: &lt;a href="https://www.iso.org/standard/42001" rel="noopener noreferrer"&gt;https://www.iso.org/standard/42001&lt;/a&gt; (Accessed: 19 August 2026).&lt;/p&gt;

&lt;p&gt;Stevens, A. (2026) &lt;em&gt;Governed Agent Trust Environment (GATE) v1.4&lt;/em&gt;. Available at: &lt;a href="https://github.com/deterministic-agents/gate/releases/tag/v1.4" rel="noopener noreferrer"&gt;https://github.com/deterministic-agents/gate/releases/tag/v1.4&lt;/a&gt; (Accessed: 19 August 2026).&lt;/p&gt;

</description>
      <category>ai</category>
      <category>governance</category>
      <category>security</category>
      <category>agentic</category>
    </item>
    <item>
      <title>GATE v1.4: Output Validation, a Rust Companion, and the Conceptual Layer as OKF</title>
      <dc:creator>Andrew Stevens</dc:creator>
      <pubDate>Tue, 18 Aug 2026 19:16:56 +0000</pubDate>
      <link>https://dev.to/sakurasky/gate-v14-output-validation-a-rust-companion-and-the-conceptual-layer-as-okf-338c</link>
      <guid>https://dev.to/sakurasky/gate-v14-output-validation-a-rust-companion-and-the-conceptual-layer-as-okf-338c</guid>
      <description>&lt;p&gt;GATE v1.4 ships today. It adds one control, three standard mappings, and three repositories, taking the Governed Agent Trust Environment from 19 controls to 20 across the same four layers: Identity and Integrity, Runtime Enforcement, Observability and Forensics, and Orchestration and Ecosystem. It is the largest release since I first published the framework in April. Existing C01 through C19 implementations remain compatible.&lt;/p&gt;

&lt;p&gt;The release lives at &lt;a href="https://github.com/deterministic-agents/gate/releases/tag/v1.4" rel="noopener noreferrer"&gt;github.com/deterministic-agents/gate/releases/tag/v1.4&lt;/a&gt; (Stevens, 2026), with the artifacts bundle, SHA256SUMS, a single-file markdown export of the paper, and the &lt;a href="https://github.com/deterministic-agents/gate/releases/download/v1.4/GATE-v1.4.pdf" rel="noopener noreferrer"&gt;141-page PDF&lt;/a&gt;. The paper source is now public at &lt;a href="https://github.com/deterministic-agents/gate-framework-paper" rel="noopener noreferrer"&gt;gate-framework-paper&lt;/a&gt;. The framework home is &lt;a href="https://deterministicagents.ai" rel="noopener noreferrer"&gt;deterministicagents.ai&lt;/a&gt; and the components sit under open licences at &lt;a href="https://github.com/deterministic-agents" rel="noopener noreferrer"&gt;github.com/deterministic-agents&lt;/a&gt;. For the architectural rationale, &lt;a href="https://www.sakurasky.com/blog/gate-launch/" rel="noopener noreferrer"&gt;"GATE: The Missing Infrastructure Layer for Agentic AI"&lt;/a&gt; remains the canonical introduction, and the &lt;a href="https://www.sakurasky.com/blog/gate-roadmap-v1-4/" rel="noopener noreferrer"&gt;v1.4 roadmap post&lt;/a&gt; from June sets out what I said this release would contain.&lt;/p&gt;

&lt;h2&gt;
  
  
  C20 Agent-to-Human Output Validation
&lt;/h2&gt;

&lt;p&gt;GATE governed what an agent could execute at the tool boundary and what it could retrieve at the memory boundary. It said nothing about what the agent delivered at the end. C20 closes that.&lt;/p&gt;

&lt;p&gt;The control sits in Layer 3, Observability and Forensics, alongside C13 intent telemetry and C19 drift monitoring. It performs per-response classification at the delivery boundary. Each final response carries a sensitivity tier, any regulated content categories it touches, a confidence score, and the obligations that follow: &lt;code&gt;redact_fields&lt;/code&gt;, &lt;code&gt;hitl_review&lt;/code&gt;, &lt;code&gt;hold_for_review&lt;/code&gt;. At the &lt;code&gt;high_privilege&lt;/code&gt; tier the control fails closed: a response whose classification matches no explicit pass-through entry in the signed action matrix is held, not delivered, enforced by the bundle schema and a policy guardrail rule. An output classification event is emitted per response, which gives the delivery boundary the same evidence coverage that tool envelopes give the tool boundary.&lt;/p&gt;

&lt;p&gt;That is the fourth boundary in the model: identity at instantiation, policy at execution, observation throughout, classification at delivery. For most deployments the output gap is tolerable. For healthcare, legal advice, financial guidance, and HR decision support, obligations about what may be delivered, to whom, and under what review conditions are as real as obligations about what may be executed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Break-glass becomes a contract
&lt;/h2&gt;

&lt;p&gt;C09 has always required a signed record and dual approval for an invariant-halt override. What it lacked was a machine-verifiable format, so the record lived as a narrative document tied to a ledger event by convention.&lt;/p&gt;

&lt;p&gt;v1.4 adds &lt;code&gt;break_glass_record.schema.json&lt;/code&gt; to gate-contracts, with dual approval, scope binding, and expiry enforced by the schema. The invariant-halt ledger event now carries a &lt;code&gt;break_glass_record_id&lt;/code&gt; referencing a valid record whenever an override was authorised, so a conformance runner can verify that each halt either resolves to one or was never overridden. That distinction previously required manual audit. This is a tightening rather than a new control. C09 stays C09 and the semantics do not change.&lt;/p&gt;

&lt;p&gt;Two other controls tighten alongside it. C17 gains an automated enrolment fast-path backed by a signed policy contract, so discovered workloads can be enrolled without a human in the path when the policy says they qualify. C18 provenance now chains back to a registered source or an approved external feed through two normative URI schemes, which removes the ambiguity in what counted as an acceptable provenance claim.&lt;/p&gt;

&lt;h2&gt;
  
  
  OWASP AISVS, MITRE ATLAS, and NIST SSDF
&lt;/h2&gt;

&lt;p&gt;The paper's standard mappings previously covered NIST AI RMF and ISO/IEC 42001, with EU AI Act alignment carried in the control text. v1.4 adds three more: OWASP AISVS, which reached 1.0 on 24 June and states testable security requirements across the AI system lifecycle (OWASP, 2026); MITRE ATLAS, the technique taxonomy for adversarial attacks on AI systems, whose tactics and techniques the mapping ties to the specific test scenarios C16 requires, so an adversarial validation harness has a threat model behind it rather than a list of categories (MITRE, 2026); and NIST SSDF, the secure development practice set issued under US Executive Order 14028 (Souppaya, Scarfone and Dodson, 2022), mapped as a deliberately narrow intersection: GATE touches SSDF at the artifact-integrity and tool-gateway surfaces (C03, C05) and marks everything else out of scope.&lt;/p&gt;

&lt;p&gt;The mappings are pinned to upstream snapshots (AISVS v1.0, ATLAS content 2026.05) and machine-validated by a validator that ships in the repo: every referenced requirement and technique identifier must resolve against the pinned upstream with a matching title. Identifiers drift between revisions of a standard, and a mapping that has fallen out of date against its upstream is worse than no mapping, because it invites a reader to trust a claim nobody has rechecked.&lt;/p&gt;

&lt;p&gt;The gaps are documented rather than papered over. Training data quality, model lifecycle management, and AI supply chain security sit outside GATE's scope, and the AISVS mapping says so. The June roadmap named these gaps as model evaluation, training data quality, and AI supply chain security; AISVS 1.0 organises the model-evaluation territory under model lifecycle management, and the mapping follows the published chapter names. The ATLAS mapping lists the techniques the framework does not address: the model interior is not a control-plane problem, and mapping to ATLAS does not make it one.&lt;/p&gt;

&lt;p&gt;These mappings are informative. Passing GATE conformance does not imply conformance with AISVS, coverage of ATLAS, or certification against ISO/IEC 42001 or any other standard. They exist so an implementation team can trace a control to the requirement a reviewer will ask about. That is the whole purpose.&lt;/p&gt;

&lt;h2&gt;
  
  
  gate-rust v1.0.0
&lt;/h2&gt;

&lt;p&gt;The Python library is the reference implementation and the right tool for prototyping gateways and building compliance tooling. It is the wrong tool for a high-throughput Tool Gateway written in Rust, where allocation patterns matter and Python interop adds latency.&lt;/p&gt;

&lt;p&gt;gate-rust covers the primitives that have to be fast: canonical JSON, envelope construction, the hash-chained ledger, and ES256 signing and verification. The hash compatibility guarantee is the primary contract. gate-rust and gate-python must produce identical hashes for identical inputs, enforced by shared test vectors and a CI digest lockstep in both repositories. A single failing vector means the implementations disagree and both are wrong until one is fixed.&lt;/p&gt;

&lt;p&gt;gate-rust covers the fast primitives only. Replay recording, schema validation, and the higher-level orchestration stay in Python where they make more sense. Neither library is on a package registry yet. v1.4 defers crates.io publication, so gate-rust installs from a tagged git ref (&lt;code&gt;cargo add gate-rust --git https://github.com/deterministic-agents/gate-rust --tag v1.0.0&lt;/code&gt;), and gate-python publication to PyPI is a v1.5 item. Minimum supported Rust version is 1.78.&lt;/p&gt;

&lt;h2&gt;
  
  
  gate-fuzz v1.0.0
&lt;/h2&gt;

&lt;p&gt;The v1.4 release ships one of three gate-fuzz deliverables originally scoped for the release. What lands is the Python-to-Rust differential harness that enforces byte equivalence between the two implementations at canonical-JSON, signing, and schema-validation layers. The other two deliverables - the bundle-derived Hypothesis strategy generator and the HTTP-level protocol fuzzer against a running Tool Gateway - move to v1.5. Naming the deferred items in the same release where the differential harness ships keeps the roadmap honest against what practitioners find when they clone the repo.&lt;/p&gt;

&lt;h2&gt;
  
  
  gate-knowledge v1.0.0
&lt;/h2&gt;

&lt;p&gt;Google Cloud published Open Knowledge Format v0.1 in June, an open specification for representing curated knowledge as a directory of markdown files with YAML frontmatter, where concepts link to each other with ordinary markdown links and the only required field is &lt;code&gt;type&lt;/code&gt; (McVeety and Hormati, 2026). It defines a file layout and nothing else. There is no SDK to install and nothing to run as a service.&lt;/p&gt;

&lt;p&gt;GATE has the problem OKF was designed for. The control catalogue, the threat model, and the ABOM are all structured knowledge, and until now they lived as YAML, PDF, and HTML. An agent building a GATE-conformant system had to fetch several documents from several sources and assemble the picture itself.&lt;/p&gt;

&lt;p&gt;gate-knowledge publishes the conceptual layer as an OKF v0.1 bundle: 20 control documents, the threat model, the adoption path, an ABOM template a team can fork and populate, typed relationship links between concepts, and a custom validator. The relationship links are the point. C17 links to C04 because discovery feeds the commission lifecycle. C09 links to C05 because invariants evaluate after policy. C19 links to C16 because they address different failure modes at the same tier and must not be merged. An agent navigating the bundle can understand a control's dependency surface before implementing it, without reading the full paper.&lt;/p&gt;

&lt;p&gt;The bundle is the conceptual layer only. gate-contracts remains the normative schema source and gate-conformance remains the check source; nothing in gate-knowledge is authoritative over either.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conformance runner v1.3.0
&lt;/h2&gt;

&lt;p&gt;Check20 joins the suite for the new control. Check17 and Check18 move from PARTIAL to AUTOMATED when bundle stores are configured. A default configuration now returns 9 AUTOMATED and 11 PARTIAL results; a configured one returns 11 and 9. PARTIAL remains a statement about what the runner can see, not a failure: the report names the specific artefact each PARTIAL check still needs from the operator.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is published in v1.4
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/deterministic-agents/gate-contracts" rel="noopener noreferrer"&gt;gate-contracts v1.2.0&lt;/a&gt;&lt;/strong&gt; carries the normative JSON Schema definitions, extended with the C20 output classification event, the break-glass record, the auto-enrolment policy, and the approved feed registry.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/deterministic-agents/gate-policies" rel="noopener noreferrer"&gt;gate-policies v1.2.0&lt;/a&gt;&lt;/strong&gt; carries the OPA/Rego policy and invariant bundles, adding C20 output classification with a fail-closed guardrail, C09 break-glass verification, the C17 auto-enrolment fast-path, and an explicit bundle manifest.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/deterministic-agents/gate-python" rel="noopener noreferrer"&gt;gate-python v1.2.0&lt;/a&gt;&lt;/strong&gt; remains the reference implementation, adding the C20 output module, break-glass verification, strict canonical-JSON hashing, and the cross-language test vectors gate-rust builds against.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/deterministic-agents/gate-conformance" rel="noopener noreferrer"&gt;gate-conformance v1.3.0&lt;/a&gt;&lt;/strong&gt; carries the twenty checks, the report template, the runbooks, the runner, and the three new machine-validated mappings.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/deterministic-agents/gate-rust" rel="noopener noreferrer"&gt;gate-rust v1.0.0&lt;/a&gt;&lt;/strong&gt;, &lt;strong&gt;&lt;a href="https://github.com/deterministic-agents/gate-fuzz" rel="noopener noreferrer"&gt;gate-fuzz v1.0.0&lt;/a&gt;&lt;/strong&gt;, and &lt;strong&gt;&lt;a href="https://github.com/deterministic-agents/gate-knowledge" rel="noopener noreferrer"&gt;gate-knowledge v1.0.0&lt;/a&gt;&lt;/strong&gt; are the three new repositories, described above.&lt;/p&gt;

&lt;p&gt;The framework paper source is public at &lt;a href="https://github.com/deterministic-agents/gate-framework-paper" rel="noopener noreferrer"&gt;gate-framework-paper&lt;/a&gt;, so the Quarto sources, control specifications, and diagram sources are now readable and forkable alongside the published PDF.&lt;/p&gt;

&lt;h2&gt;
  
  
  What v1.5 picks up
&lt;/h2&gt;

&lt;p&gt;The June roadmap scoped gate-fuzz around PARTIAL closure: a README mapping each test to the conformance check it closes, and a suite that converts PARTIAL results into verifiable PASS results. That mapping arrives with the two deferred deliverables in v1.5, so the runner counts above stand until that mapping lands.&lt;/p&gt;

&lt;p&gt;Alongside them, the paper's known issues section names five more: streaming-aware C20 classification, where a response is classified before it is fully formed; a unified exception register, meaning one &lt;code&gt;exception_record&lt;/code&gt; contract for the exception surface that C09 break-glass and the C17, C18, and C19 policies all consume; a multi-approver HITL record, so approvals that need more than one approver are carried by the HITL record itself rather than borrowed from the break-glass path; gate-python publication to PyPI; and a gate-rust conformance bridge so the crate can produce runner-compatible evidence directly.&lt;/p&gt;

&lt;p&gt;That is what v1.5 is scoped around, written down now for the same reason the June roadmap was. If you are building agentic AI infrastructure, deploying agents in a regulated environment, or advising organisations on AI governance, the issues and discussions on each repository are open. A bug report from someone who tried to implement a control is worth more to me than a comment on the paper.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;GATE is published at &lt;a href="https://deterministicagents.ai" rel="noopener noreferrer"&gt;deterministicagents.ai&lt;/a&gt; under CC BY 4.0 for the documentation and MIT for the code. The strategic companion to this framework is the &lt;a href="https://www.sakurasky.com/white-papers/trustworthy-agentic-ai-blueprint/" rel="noopener noreferrer"&gt;Trustworthy Agentic AI Blueprint&lt;/a&gt;, co-authored with Sakura Sky.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Disclosure: GATE is authored and maintained by me personally rather than by Sakura Sky, and there is no paid tier, hosted version, or commercial product built on it. Agent governance is also my day job at Sakura Sky, which is a commercial interest worth stating. The standard mappings described here are informative and were validated against pinned upstream snapshots rather than assessed by a third party.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;McVeety, S. and Hormati, A. (2026) &lt;em&gt;Introducing the Open Knowledge Format&lt;/em&gt;, Google Cloud Blog, 12 June. Available at: &lt;a href="https://cloud.google.com/blog/products/data-analytics/how-the-open-knowledge-format-can-improve-data-sharing" rel="noopener noreferrer"&gt;https://cloud.google.com/blog/products/data-analytics/how-the-open-knowledge-format-can-improve-data-sharing&lt;/a&gt; (Accessed: 18 August 2026).&lt;/p&gt;

&lt;p&gt;MITRE (2026) &lt;em&gt;MITRE ATLAS: Adversarial Threat Landscape for Artificial-Intelligence Systems&lt;/em&gt;. Available at: &lt;a href="https://atlas.mitre.org/" rel="noopener noreferrer"&gt;https://atlas.mitre.org/&lt;/a&gt; (Accessed: 18 August 2026).&lt;/p&gt;

&lt;p&gt;OWASP (2026) &lt;em&gt;Artificial Intelligence Security Verification Standard (AISVS) 1.0&lt;/em&gt;, released 24 June. OWASP Foundation. Available at: &lt;a href="https://owasp.org/www-project-artificial-intelligence-security-verification-standard-aisvs-docs/" rel="noopener noreferrer"&gt;https://owasp.org/www-project-artificial-intelligence-security-verification-standard-aisvs-docs/&lt;/a&gt; (Accessed: 18 August 2026).&lt;/p&gt;

&lt;p&gt;Souppaya, M., Scarfone, K. and Dodson, D. (2022) &lt;em&gt;Secure Software Development Framework (SSDF) Version 1.1: Recommendations for Mitigating the Risk of Software Vulnerabilities&lt;/em&gt;, NIST SP 800-218. National Institute of Standards and Technology. Available at: &lt;a href="https://csrc.nist.gov/pubs/sp/800/218/final" rel="noopener noreferrer"&gt;https://csrc.nist.gov/pubs/sp/800/218/final&lt;/a&gt; (Accessed: 18 August 2026).&lt;/p&gt;

&lt;p&gt;Stevens, A. (2026) &lt;em&gt;Governed Agent Trust Environment (GATE) v1.4&lt;/em&gt;. Available at: &lt;a href="https://github.com/deterministic-agents/gate/releases/tag/v1.4" rel="noopener noreferrer"&gt;https://github.com/deterministic-agents/gate/releases/tag/v1.4&lt;/a&gt; (Accessed: 18 August 2026).&lt;/p&gt;

</description>
      <category>ai</category>
      <category>governance</category>
      <category>security</category>
      <category>agentic</category>
    </item>
    <item>
      <title>Protecting the IP in a Generative World</title>
      <dc:creator>Andrew Stevens</dc:creator>
      <pubDate>Tue, 18 Aug 2026 11:40:02 +0000</pubDate>
      <link>https://dev.to/sakurasky/protecting-the-ip-in-a-generative-world-17hk</link>
      <guid>https://dev.to/sakurasky/protecting-the-ip-in-a-generative-world-17hk</guid>
      <description>&lt;p&gt;A studio found its own back catalogue inside the training data of a public model. Someone had run a handful of its older titles through a detection tool, and the tool reported, with reasonable confidence, that the model had seen them. The legal team moved quickly, because a decades-old library is the company's principal asset and defending it is the job. The engineering team was asked a simpler question and could not answer it. From the studio's own systems, could it prove when each of those works was created, what it was derived from, and whether anyone had ever been licensed to ingest it. The plain answer was no. The studio owned the content and could not produce the provenance.&lt;/p&gt;

&lt;p&gt;That gap is the subject of this post. The studio's real exposure went well beyond a model having trained on its work. Nothing in its own architecture could establish the origin, ownership, and rights history of the things it makes, on demand and in a form an outside party would accept. The wider series argued that trust is something a system produces rather than asserts (see &lt;a href="https://www.sakurasky.com/blog/engineering-underneath-part-6/" rel="noopener noreferrer"&gt;Trust Is an Engineering Output&lt;/a&gt;). For a media business, provenance is the specific shape that takes, and generative models are the forcing function that has finally made its absence expensive. This post works backwards from the discovery: what the old protections were built to do, what protection now demands, and the layer that closes the gap.&lt;/p&gt;

&lt;h2&gt;
  
  
  The back-catalogue discovery
&lt;/h2&gt;

&lt;p&gt;Start with what the studio did have, because it was not nothing. Its digital rights management was intact. Streams were encrypted, playback was licensed, access to the finished assets was controlled, and outbound video carried watermarks. By the standard of the threat that DRM was built for, the studio was well defended. None of it touched the problem in the room.&lt;/p&gt;

&lt;p&gt;What the studio lacked was a record, attached to each asset and maintained by its systems, of when the work was made, who contributed to it, what source material it drew on, and what rights sat over it. That information existed, scattered across contracts in a document store, production notes, and the memories of people who had moved on. It did not exist as data the studio could query. So when the question became "prove this is yours and prove what happened to it," the studio was reduced to reconstructing its own history by hand, which is exactly the position the earlier posts in this programme described as evidence that was never engineered.&lt;/p&gt;

&lt;p&gt;This is the shape most media businesses are in, and it is worth being precise about why it is uncomfortable. Without that record, the studio could not prove ingestion had happened, could not cheaply assert ownership, and could not act on the opt-out mechanisms the law now provides, because it had no machine-readable way to express its rights across a library built over decades. The content was owned. The provenance was not producible. Those are different properties, and only the second one was being tested.&lt;/p&gt;

&lt;h2&gt;
  
  
  What DRM used to mean
&lt;/h2&gt;

&lt;p&gt;Work backwards to how the protections were built, because they were built well for a threat that has changed. Digital rights management was designed to stop one thing: the unauthorised copying and redistribution of a finished asset. Encryption at rest and in transit, licence servers that decide who may play what, access controls, and watermarks that trace a leak back to its source. The protected object was the copy, and the moment of risk was consumption. For piracy, this was the right design, and for piracy it still largely works.&lt;/p&gt;

&lt;p&gt;It was silent, though, on the things that now matter most. Encryption protects the delivery path, so it does block a crawler from ingesting that path, but the training-data problem mostly lives elsewhere: in the trailers, clips, and released titles that circulate in the clear, and in the studio's ability to express its rights and prove its own lineage. DRM protects the outbound stream. It says nothing about how a work may be used for training, nor about the origin story of the work behind the stream.&lt;/p&gt;

&lt;p&gt;This is why DRM modernisation is not a matter of stronger encryption or tighter licence enforcement. The object that now needs protecting is not only the outbound stream. It is the provenance of the work itself, the verifiable account of what it is, when it was made, and what may lawfully be done with it. That account is data, and it has to be engineered as data.&lt;/p&gt;

&lt;h2&gt;
  
  
  What model-era IP protection requires
&lt;/h2&gt;

&lt;p&gt;Once the threat moves from copying to ingestion, content IP protection needs three capabilities that traditional media rights management never had to provide, and the law on both sides of the Atlantic is moving toward them, even as the central US question, whether training counts as fair use, stays unresolved in the courts.&lt;/p&gt;

&lt;p&gt;The first is machine-readable rights reservation. European copyright law gives rightsholders a text-and-data-mining opt-out, but it only functions if the reservation is expressed in a machine-readable form that a crawler can detect (European Parliament and Council, 2019). The AI Act then obliges providers of general-purpose models to put in place a policy to comply with that reservation and to publish a sufficiently detailed summary of the content used to train the model (European Parliament and Council, 2024). An opt-out a studio cannot express at scale, across a whole catalogue, is a right it cannot actually exercise.&lt;/p&gt;

&lt;p&gt;The second is provable ownership. The United States arrives at a related point from the other direction: in a pre-publication report, the US Copyright Office takes the view that assembling a training dataset from copyrighted works implicates the reproduction right, because it involves downloading, storing, and copying those works (U.S. Copyright Office, 2025). Registration and a clean chain of title remain the legal instruments that let a studio sue and recover, so the provenance layer does not replace them. It feeds them, giving the studio a queryable evidence base for what it owned and when, which is the thing a claim or a defence eventually turns on.&lt;/p&gt;

&lt;p&gt;The third is content provenance. The industry has converged on a standard for this, the C2PA content credentials specification, which attaches tamper-evident provenance and edit history to a media asset so its origin travels with it rather than living in a separate database (C2PA, 2026). Generative AI copyright disputes are often, underneath the legal language, arguments about who can prove what about a file, and content credentials are an attempt to make that provable at the level of the asset.&lt;/p&gt;

&lt;p&gt;This is not hypothetical, and music is where it has been tested first. In July 2026 the Munich Regional Court found the AI music service Suno liable for copyright infringement in a case brought by the German collecting society GEMA, holding that specific protected works remained reproducible inside the models and that the outputs reproduced them, and that training carried out in the United States did not place it beyond German law (Music Business Worldwide, 2026). The decision is first-instance and under appeal, but it turns on exactly the question a provenance layer exists to answer: can a rightsholder show that a particular work was used, and that a particular output reproduced it. The studios watching that case are asking whether they could prove the same thing about their own catalogues, and most cannot.&lt;/p&gt;

&lt;h2&gt;
  
  
  The provenance layer
&lt;/h2&gt;

&lt;p&gt;Move forward from the diagnosis and the resolution is an IP-aware data layer, engineered before the dispute rather than assembled during it.&lt;/p&gt;

&lt;p&gt;It does three things, and they are the implementations of the capabilities above. First, it makes provenance a property of every asset: content credentials are attached at the point of creation, and the studio keeps an internal record of each work's creation date, contributors, source and derivation, and attached rights. That record is made tamper-evident through cryptographic signing and hash-linking, so a later alteration is detectable, which is what turns a database into evidence. Second, it expresses machine-readable rights reservation consistently across the whole library, so a model training opt-out is something the studio asserts at scale rather than in principle. Third, it adds monitoring that checks whether the studio's works surface in public datasets or model outputs. That monitoring is probabilistic, the same reasonable-confidence signal the studio started with, so it works as an early-warning system rather than as proof, and describing it that way keeps the claim accurate. Building that layer out of a media company's scattered production and rights data is IP architecture, and it is a large part of what Sakura's &lt;a href="https://www.sakurasky.com/data/" rel="noopener noreferrer"&gt;Data &amp;amp; AI practice&lt;/a&gt; does for content businesses.&lt;/p&gt;

&lt;p&gt;This is where the argument connects back to the rest of the programme. A provenance layer is the media form of evidence as an engineered property: the studio can answer, for any asset, where it came from and what may be done with it, on demand and with the signed record attached. That is data provenance doing the same job in a studio that a hash-linked transaction chain does in a bank, and it turns content provenance from a compliance aspiration into a running feature of the platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happens when this is not engineered
&lt;/h2&gt;

&lt;p&gt;The cost of skipping this is not only the litigation the studio is now in. It is broader, and some of it is opportunity rather than risk.&lt;/p&gt;

&lt;p&gt;Without the layer, a media company cannot readily prove ingestion happened, cannot enforce an opt-out it has no way to express, and cannot assert ownership without a manual reconstruction every time. Each of those is a live exposure as the use of copyrighted works for AI training data becomes contested. The part most businesses miss is on the other side of the ledger. A market for licensing catalogues to model builders is forming: in the same music-industry fight, Warner Music settled its US infringement suit against Suno in late 2025 and signed a licensing partnership, even as Universal and Sony kept litigating (Music Business Worldwide, 2026). Those deals still close on corporate ownership and contractual warranties rather than on asset-level provenance alone, but a content owner that can show clean provenance diligences faster, carries less risk, and negotiates from a stronger position. The same missing layer that leaves the company exposed also leaves value on the table in the market its own content is helping to create.&lt;/p&gt;

&lt;p&gt;The evidence that lets a media business defend its rights or license its catalogue is engineering output produced before the dispute, not a legal artefact produced after it. Building the record that lets a studio prove what it owns, express how it may be used, and show when a boundary was crossed is a media security and evidence problem before it is a legal one, and it is the ground Sakura's &lt;a href="https://www.sakurasky.com/security/" rel="noopener noreferrer"&gt;Security practice&lt;/a&gt; works on with a media company's data and rights teams.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;Coalition for Content Provenance and Authenticity (C2PA), 2026. &lt;em&gt;C2PA Technical Specification, Version 2.4.&lt;/em&gt; Available at: &lt;a href="https://spec.c2pa.org/" rel="noopener noreferrer"&gt;https://spec.c2pa.org/&lt;/a&gt; [Accessed 18 August 2026].&lt;/p&gt;

&lt;p&gt;Music Business Worldwide, 2026. &lt;em&gt;Suno infringed copyright in GEMA case, German court rules.&lt;/em&gt; Reporting the Munich Regional Court first-instance judgment in GEMA v Suno, Case 42 O 763/25, 31 July 2026 (under appeal). Available at: &lt;a href="https://www.musicbusinessworldwide.com/suno-infringed-copyright-in-gema-case-german-court-rules/" rel="noopener noreferrer"&gt;https://www.musicbusinessworldwide.com/suno-infringed-copyright-in-gema-case-german-court-rules/&lt;/a&gt; [Accessed 18 August 2026].&lt;/p&gt;

&lt;p&gt;European Parliament and Council, 2019. &lt;em&gt;Directive (EU) 2019/790 of the European Parliament and of the Council of 17 April 2019 on copyright and related rights in the Digital Single Market and amending Directives 96/9/EC and 2001/29/EC.&lt;/em&gt; Official Journal of the European Union, L 130, 17 May, pp. 92-125. Available at: &lt;a href="https://eur-lex.europa.eu/eli/dir/2019/790/oj" rel="noopener noreferrer"&gt;https://eur-lex.europa.eu/eli/dir/2019/790/oj&lt;/a&gt; [Accessed 18 August 2026].&lt;/p&gt;

&lt;p&gt;European Parliament and Council, 2024. &lt;em&gt;Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act).&lt;/em&gt; Official Journal of the European Union, L 2024/1689, 12 July. Available at: &lt;a href="https://eur-lex.europa.eu/eli/reg/2024/1689/oj" rel="noopener noreferrer"&gt;https://eur-lex.europa.eu/eli/reg/2024/1689/oj&lt;/a&gt; [Accessed 18 August 2026].&lt;/p&gt;

&lt;p&gt;U.S. Copyright Office, 2025. &lt;em&gt;Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version).&lt;/em&gt; United States Copyright Office, Washington, DC, 9 May. Available at: &lt;a href="https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf" rel="noopener noreferrer"&gt;https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf&lt;/a&gt; [Accessed 18 August 2026].&lt;/p&gt;

</description>
      <category>media</category>
      <category>entertainment</category>
      <category>ipprotection</category>
      <category>drm</category>
    </item>
    <item>
      <title>Your System Prompt Is Not a Trust Boundary</title>
      <dc:creator>Andrew Stevens</dc:creator>
      <pubDate>Tue, 18 Aug 2026 09:28:15 +0000</pubDate>
      <link>https://dev.to/sakurasky/your-system-prompt-is-not-a-trust-boundary-2k99</link>
      <guid>https://dev.to/sakurasky/your-system-prompt-is-not-a-trust-boundary-2k99</guid>
      <description>&lt;p&gt;A text-to-SQL agent is one of the easiest AI features to demo and one of the more awkward ones to get through a security review. A non-technical user asks a question in English, a model writes SQL, the warehouse runs it, and a table comes back. The demo takes an afternoon. The review tends to stall on a harder question: what actually stops the agent returning data the person asking is not entitled to see?&lt;/p&gt;

&lt;p&gt;In the reviews I sit in, the first answer is usually a sentence in the system prompt. &lt;em&gt;Never return customer email addresses. Only query the orders table. Do not run anything expensive.&lt;/em&gt; Those sentences are worth writing, and OWASP's own first mitigation for prompt injection is to constrain model behaviour in exactly that way, with specific instructions about the model's role and limits (OWASP, 2025). My argument is not that the sentences are useless. It is that they are the wrong artefact to point at when somebody asks you to demonstrate that the rule holds, because an instruction sitting in the context window is subject to the same forces as every other token in that window.&lt;/p&gt;

&lt;p&gt;Willison, who coined the term prompt injection by analogy with SQL injection, puts the mechanism plainly: LLMs follow instructions in content, and they cannot reliably distinguish the importance of instructions based on where those instructions came from (Willison, 2025). OWASP ranks prompt injection first in its Top 10 for LLM applications and is unusually direct about the outlook, noting that given the stochastic influence at the heart of the way models work, it is unclear whether any fool-proof method of prevention exists (OWASP, 2025).&lt;/p&gt;

&lt;p&gt;In a text-to-SQL agent the injection surface is wider than the user's question. Tool results reach the context, as do column descriptions pulled from the information schema, and a returned row with instructions written into a free-text field. Any of those can carry the argument that talks the model past the rule you wrote.&lt;/p&gt;

&lt;p&gt;My working assumption, and I would not claim it as more than an assumption, is that a rule an LLM enforces is a rule an LLM can be argued out of. If you need to show an auditor that a rule holds, the rule has to live somewhere you can test it and be proven wrong.&lt;/p&gt;

&lt;p&gt;OWASP's own mitigation list says as much further down, recommending deterministic code to validate adherence to expected output, and handling privileged functions in code rather than handing them to the model (OWASP, 2025). That is the design this post is about.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a second model buys you, and what it does not
&lt;/h2&gt;

&lt;p&gt;The reflex fix, once a team accepts that the system prompt is porous, is to add a judge: a smaller LLM that looks at the generated SQL and rules on whether it is safe.&lt;/p&gt;

&lt;p&gt;I think that makes a reasonable detective control and a poor floor. It costs tokens on every query and adds a network round trip to a path users experience as latency, which are the boring objections. The one I care about is that a judge is a probabilistic classifier, and the figure these systems tend to advertise sits somewhere around 95 percent. Willison's line on the guardrail vendor category is that in web application security 95 percent is a failing grade (Willison, 2025), and I would apply the same standard here. A control that is right most of the time can sit above your floor, and I would not build a floor out of one.&lt;/p&gt;

&lt;p&gt;Where a judge does earn its place is on the classes a parser structurally cannot see, and I will get to those below: re-identification through innocuous columns, PII addressed by a string literal inside a JSON payload, aggregations that are each individually fine and jointly identifying. Those are semantic questions, and a grammar has no opinion on them. My preference is to run a judge above a deterministic layer where the budget allows, rather than in place of one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Move the rule to where the SQL is
&lt;/h2&gt;

&lt;p&gt;The SQL an agent produces is a string. Before it reaches the warehouse it is inert, fully inspectable, and, unlike natural language, has a grammar. That is the moment where a rule can be enforced by something that does not take arguments.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;sql-guard&lt;/code&gt; is our attempt at that control. It is a policy engine that sits between the model's output and the warehouse client: it parses the SQL, runs an ordered list of rules against the syntax tree, and returns &lt;code&gt;allow&lt;/code&gt;, &lt;code&gt;confirm&lt;/code&gt; or &lt;code&gt;deny&lt;/code&gt;. There is no LLM in the guard path. It is Apache-2.0, pure Python, and depends on &lt;code&gt;sqlglot&lt;/code&gt; and nothing else. Install &lt;code&gt;agent-sql-guard&lt;/code&gt;, import &lt;code&gt;sql_guard&lt;/code&gt;, Python 3.11 or later (Sakura Sky Engineering, 2026).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;sql_guard&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;PiiDenylist&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;SqlGuard&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;SqlGuardConfig&lt;/span&gt;

&lt;span class="n"&gt;guard&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;SqlGuard&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;SqlGuardConfig&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_settings&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;pii_denylist&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;PiiDenylist&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_mapping&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;columns&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;phone_number&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ssn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;substrings&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;address&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;}),&lt;/span&gt;
    &lt;span class="n"&gt;allowed_tables&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;my-project.analytics.orders&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;dialect&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;bigquery&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="n"&gt;decision&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;guard&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;evaluate_static&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SELECT customer_id, COUNT(*) FROM `my-project.analytics.orders` GROUP BY 1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;denied&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;reason&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Six rules ship by default. Single &lt;code&gt;SELECT&lt;/code&gt; only, so DML, DDL and multi-statement payloads are refused even when buried in a subquery. A PII column denylist. No &lt;code&gt;SELECT *&lt;/code&gt; in any scope. Nothing whose columns the guard cannot enumerate. A table allowlist of fully-qualified names. And a cost cap with three thresholds: auto-execute below $0.10, ask for confirmation between, refuse above a $20.00 hard cap or a 10 GiB bytes-billed ceiling, all of them defaults you can move.&lt;/p&gt;

&lt;p&gt;Parsing rather than pattern-matching is the load-bearing choice. Regex over SQL is famously brittle, and most of the bypasses described below would have been trivial against a regex while still being reachable against an AST walk that was not careful enough. &lt;code&gt;sqlglot&lt;/code&gt; handles the dialect surface (Mao, 2026). BigQuery is the default and the one under heaviest real use. Snowflake, Postgres, Trino, DuckDB, ClickHouse and MySQL have test coverage, and Presto shares Trino's parser. The remainder of &lt;code&gt;sqlglot&lt;/code&gt;'s thirty-plus dialects should work without being battle-worn.&lt;/p&gt;

&lt;p&gt;One caveat belongs here rather than in the limitations section at the end, because it decides whether any of this is a boundary at all. A library the caller has to remember to call is a convention. The host process's warehouse identity is the same whether the guard ran or not, so if another code path in the same process can reach the client directly, the guard is closer to a lint than a control. Putting the credential behind the guarded client, in a separate service account or a proxy the agent process cannot reach around, is what turns the convention into enforcement.&lt;/p&gt;

&lt;p&gt;A parser will not be talked out of a rule. It can still be wrong about what the rule covers, and most of the rest of this post is eight examples of exactly that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Eight ways the PII denylist was bypassed
&lt;/h2&gt;

&lt;p&gt;v0.2.0 exists because an internal adversarial review of a deployed agent found two ways to get denylisted columns past the guard. Fixing those two surfaced six more. The release closes eight, each confirmed against 0.1.1 with a reproducing query before anyone touched a fix, each with a regression test that fails on the old code (Sakura Sky Engineering, 2026). The changelog itemises the work across ten entries, because two of the classes needed separate fixes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Alias laundering.&lt;/strong&gt; Rename a denied column inside a common table expression (CTE), the named subquery a &lt;code&gt;WITH&lt;/code&gt; clause introduces, then project the alias from the outer query. Each CTE gets its own scope, and the rule called a helper that only read the outermost projection list, so the outer select named &lt;code&gt;city&lt;/code&gt; and looked clean.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;WITH&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;billing_city&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;city&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="nv"&gt;`p.d.orders`&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;city&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Derived tables, &lt;code&gt;UNION ALL&lt;/code&gt; arms and multi-hop alias chains through two CTEs all worked the same way.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Inner-scope stars.&lt;/strong&gt; The star rule also ran against the outermost select only, so a &lt;code&gt;SELECT *&lt;/code&gt; inside a CTE body or derived table went through untouched.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stars inside a wrapping construct.&lt;/strong&gt; A separate bug with the same effect. The check inspected the projection's root node, which meant &lt;code&gt;OBJECT_CONSTRUCT(*)&lt;/code&gt; in Snowflake, &lt;code&gt;COLUMNS(*)&lt;/code&gt; in DuckDB, &lt;code&gt;* APPLY(f)&lt;/code&gt; in ClickHouse and &lt;code&gt;ROW(c.*)&lt;/code&gt; in Trino all passed. It is now a deep walk, with &lt;code&gt;COUNT(*)&lt;/code&gt; as the explicit carve-out.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Qualified &lt;code&gt;t.*&lt;/code&gt;.&lt;/strong&gt; In &lt;code&gt;sqlglot&lt;/code&gt; a qualified star parses as a &lt;code&gt;Column&lt;/code&gt; node wrapping a &lt;code&gt;Star&lt;/code&gt;, not as a bare &lt;code&gt;Star&lt;/code&gt;. An &lt;code&gt;isinstance(projection, exp.Star)&lt;/code&gt; check walks straight past it, at the top level as well as anywhere else.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ClickHouse &lt;code&gt;COLUMNS('regex')&lt;/code&gt;.&lt;/strong&gt; Expands to an arbitrary set of columns and parses with no &lt;code&gt;Star&lt;/code&gt; node anywhere in the tree to match on.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;NATURAL JOIN&lt;/code&gt;.&lt;/strong&gt; Joins on whichever columns the two tables happen to share. Without schema introspection the guard cannot rule out a denied column among them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Column names that never become a &lt;code&gt;Column&lt;/code&gt; node.&lt;/strong&gt; &lt;code&gt;sqlglot&lt;/code&gt; parses several positions as a bare &lt;code&gt;Identifier&lt;/code&gt;, so a sweep collecting &lt;code&gt;Column&lt;/code&gt; nodes missed them entirely: &lt;code&gt;JOIN ... USING (email)&lt;/code&gt;, column aliases of the &lt;code&gt;AS g(email)&lt;/code&gt; form, and &lt;code&gt;STRUCT('x' AS email)&lt;/code&gt; field names. The &lt;code&gt;USING&lt;/code&gt; case was a working single-query value oracle.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Aggregates as a blanket exemption.&lt;/strong&gt; Every aggregate function counted as PII-neutralising, so &lt;code&gt;MAX(email)&lt;/code&gt;, &lt;code&gt;MIN(email)&lt;/code&gt;, &lt;code&gt;ARRAY_AGG(email)&lt;/code&gt;, &lt;code&gt;STRING_AGG(email)&lt;/code&gt; and &lt;code&gt;ANY_VALUE(email)&lt;/code&gt; all returned real values. Only aggregates that reduce to a derived statistic qualify now.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Whole-row alias references.&lt;/strong&gt; This one was not in the original report. It was found while fixing the others, and it was the most severe of the eight.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="nv"&gt;`p.d.orders`&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That parses as an ordinary column named &lt;code&gt;c&lt;/code&gt;. The denylist had no denied name to match against, the star check found no star, and the guard auto-executed it. Strictly more powerful than the &lt;code&gt;SELECT *&lt;/code&gt; it had been blocking since day one, and syntactically indistinguishable from selecting a column that happens to be called &lt;code&gt;c&lt;/code&gt;. The fix also covers &lt;code&gt;TO_JSON_STRING(c)&lt;/code&gt;, &lt;code&gt;ARRAY_AGG(c)&lt;/code&gt;, &lt;code&gt;STRUCT(c)&lt;/code&gt;, the unaliased &lt;code&gt;SELECT tbl FROM tbl&lt;/code&gt; form, CTE and derived-table names, and &lt;code&gt;VALUES&lt;/code&gt; and &lt;code&gt;PIVOT&lt;/code&gt; aliases.&lt;/p&gt;

&lt;p&gt;Getting that rule to be usable took more care than getting it to be strict. Denying on a bare name collision would have broken &lt;code&gt;WITH revenue AS (SELECT ..., SUM(x) AS revenue ...) SELECT revenue FROM revenue&lt;/code&gt;, which is a mainstream idiom, so the rule resolves the ambiguity from the AST instead: a table contributes only the name it is addressable by, and a CTE publishes its own output names.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bypasses returned the success state
&lt;/h2&gt;

&lt;p&gt;Most of the eight did not return &lt;code&gt;deny&lt;/code&gt;. They returned &lt;code&gt;confirm&lt;/code&gt;, and &lt;code&gt;confirm&lt;/code&gt; is what static evaluation returns when no rule fires at all.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;rule&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_rules&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;decision&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;rule&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;evaluate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;decision&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;decision&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;GuardDecision&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;outcome&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;GuardOutcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;CONFIRM&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;reason&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Static checks passed; awaiting cost evaluation.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;referenced_tables&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;tuple&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;sorted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tables&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So there was no near-miss and nothing to instrument. Eight queries reached denied columns, and the guard reported that static checks had passed, in the same words it uses for a query that is entirely fine. The whole-row alias went on through the cost gate and executed.&lt;/p&gt;

&lt;p&gt;One thing worth knowing if you are building telemetry on this: &lt;code&gt;GuardOutcome&lt;/code&gt; is shared across the two evaluation phases and carries a different meaning in each. In &lt;code&gt;evaluate_static&lt;/code&gt;, &lt;code&gt;confirm&lt;/code&gt; is the pass. In &lt;code&gt;evaluate_cost&lt;/code&gt; it means &lt;em&gt;ask the user before running this&lt;/em&gt;, and sits between the auto threshold and the hard cap.&lt;/p&gt;

&lt;p&gt;A guardrail that fails closed tends to generate a support ticket and get fixed that afternoon. One that fails open often generates very little, which leaves going and looking as the main way anybody finds out. That asymmetry is my argument for treating adversarial review of this class of component as routine work, and it is why every fixed bypass now carries a test that fails against the version before it. A good share of the tests across those five files exist to keep the eight closed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix broke our own demo
&lt;/h2&gt;

&lt;p&gt;Tightening a guardrail means queries that used to work now do not, and 0.2.0 contains breaking behaviour changes on purpose.&lt;/p&gt;

&lt;p&gt;The clearest casualty was the identity-resolution query bundled with the project as a worked example. It normalises &lt;code&gt;email&lt;/code&gt; and &lt;code&gt;mobile&lt;/code&gt; inside a CTE and projects only &lt;code&gt;COUNTIF&lt;/code&gt; aggregates, an ordinary pattern in analytics that looks careful. It is now denied in both PII modes: the CTE scope projects the denied columns, and &lt;code&gt;COUNTIF(email_norm = 'target')&lt;/code&gt; is itself a value oracle.&lt;/p&gt;

&lt;h2&gt;
  
  
  A WHERE clause is an oracle
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;sql-guard&lt;/code&gt; has two PII modes. &lt;code&gt;"reference"&lt;/code&gt;, the default, denies any mention of a denied column anywhere in the query. &lt;code&gt;"project"&lt;/code&gt; denies only projections, checked across every scope.&lt;/p&gt;

&lt;p&gt;Projection-only checking is the intuitive design and it does not hold, because a denied column in a predicate never appears in the output while still answering a question about itself.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="nv"&gt;`p.d.orders`&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;billing_city&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'Columbus'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run it again with &lt;code&gt;LIKE 'a%'&lt;/code&gt;, then &lt;code&gt;&amp;gt; 'm'&lt;/code&gt;, and you are doing binary search against a value you were never allowed to read. Depending on the cardinality, a handful of queries recovers it. &lt;code&gt;GROUP BY&lt;/code&gt;, &lt;code&gt;HAVING&lt;/code&gt; and &lt;code&gt;ORDER BY&lt;/code&gt; leak the same way. The row count is the channel, and a rule that only inspects projections cannot see it.&lt;/p&gt;

&lt;p&gt;Hence the strict default. If a deployment genuinely needs predicate access to a denied column, &lt;code&gt;pii_mode="project"&lt;/code&gt; is there and the trade-off is stated in the docs. Narrowing the denylist, or pointing the agent at a pre-masked view the denylist does not cover, is usually the better move.&lt;/p&gt;

&lt;p&gt;One nuance is easy to get backwards, and a careful reader of the changelog did: &lt;strong&gt;aggregation is not a safe harbour&lt;/strong&gt;. &lt;code&gt;COUNT&lt;/code&gt;, &lt;code&gt;SUM&lt;/code&gt;, &lt;code&gt;AVG&lt;/code&gt; and their relatives are treated as PII-neutralising only under &lt;code&gt;pii_mode="project"&lt;/code&gt;. Under the default mode no aggregate is exempt, because the whole point of the default is that the agent must not learn the values at all, and &lt;code&gt;COUNT(*) ... WHERE email = ...&lt;/code&gt; learns them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits of a parse-level SQL guard
&lt;/h2&gt;

&lt;p&gt;These limits are structural. I would sooner publish them than have somebody discover them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;PII inside JSON, VARIANT or STRUCT payloads is not covered.&lt;/strong&gt; &lt;code&gt;JSON_VALUE(payload, '$.email')&lt;/code&gt; names only &lt;code&gt;payload&lt;/code&gt;; the field name is a string literal the engine resolves at runtime. Denylist the containing column.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two whole-row reads remain open by design.&lt;/strong&gt; Selecting a STRUCT column whole, and an &lt;code&gt;UNNEST&lt;/code&gt; alias over an array of structs, both return every field without naming one. Neither is distinguishable at parse time from the scalar-array form that is idiomatic and has to stay allowed. Same remedy: denylist the containing column.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Re-identification through non-PII columns is out of scope.&lt;/strong&gt; If &lt;code&gt;uid&lt;/code&gt; maps one-to-one to a person, blocking &lt;code&gt;email&lt;/code&gt; does not stop correlation against an outside dataset.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Side channels remain.&lt;/strong&gt; Row counts, dry-run byte figures and error messages all carry bits about denied values even when every direct reference is refused. The oracle above is the version we closed, and the family is larger than the fix.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;There is no schema introspection.&lt;/strong&gt; If you say a table is allowed, the guard takes your word for it. That constraint is the reason &lt;code&gt;SELECT *&lt;/code&gt; is rejected everywhere instead of reasoned about.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The cost cap bounds a query, not a spend.&lt;/strong&gt; &lt;code&gt;evaluate_cost&lt;/code&gt; is per-call by construction. An agent in a retry loop issuing two thousand queries at nine cents each trips nothing. Cumulative exposure needs warehouse-side maximum-bytes-billed and a budget alert.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is not an authorisation layer.&lt;/strong&gt; Identity, IAM and row-level security sit outside it. It can approve a query that a correctly configured warehouse would have refused on identity grounds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is a boundary only where it is the only path to the credentials.&lt;/strong&gt; Covered above, and in my experience the most common way this control gets downgraded to a suggestion.&lt;/p&gt;

&lt;p&gt;The changelog also carries two known issues we have not fixed and one inconsistency: a top-level &lt;code&gt;EXCEPT DISTINCT&lt;/code&gt; or &lt;code&gt;INTERSECT&lt;/code&gt; is currently rejected as a non-SELECT, which fails closed and is an availability bug rather than a security one; allowlist breaches are under-reported in telemetry built on &lt;code&gt;decision.reason&lt;/code&gt; because that rule runs last; and two spellings of hashed PII are handled inconsistently, in the direction of denial. Anyone evaluating this would do well to read that section alongside the README.&lt;/p&gt;

&lt;h2&gt;
  
  
  Defence in depth, and what the depth sits on
&lt;/h2&gt;

&lt;p&gt;Warehouse-side column-level and row-level security are the durable answer to most of this. Policy tags on sensitive columns, masking policies, row access policies scoped to the calling principal: enforcement that lives with the data, applies to every client, and does not care whether the query came from an agent, a BI tool or somebody's notebook. Where a team can get there, I would push them to.&lt;/p&gt;

&lt;p&gt;The gap &lt;code&gt;sql-guard&lt;/code&gt; fills is between deciding that and having it. Warehouse-side controls need coordinated schema work, a data classification exercise that is usually half-finished, and sign-off from teams who own tables you do not. In the programmes I have watched, that runs to quarters rather than weeks. A denylist and an allowlist in a config file is an afternoon, it gives you a cost cap that column security does not, and it keeps working as a second layer once the warehouse work lands.&lt;/p&gt;

&lt;p&gt;Message-boundary guardrails such as NeMo Guardrails or LangChain's belong in the same picture. They watch intent in the conversation, this watches the query at execution, and the failure modes look different enough to me that running both is usually worth the tokens.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it is
&lt;/h2&gt;

&lt;p&gt;v0.2.0 is on PyPI as &lt;a href="https://pypi.org/project/agent-sql-guard/" rel="noopener noreferrer"&gt;&lt;code&gt;agent-sql-guard&lt;/code&gt;&lt;/a&gt; and the source, changelog and security policy are on &lt;a href="https://github.com/sakura-sky/sql-guard" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;. Two naming notes, because both have caught people. The import is &lt;code&gt;sql_guard&lt;/code&gt; while the distribution is &lt;code&gt;agent-sql-guard&lt;/code&gt;, and the unqualified name &lt;code&gt;sql-guard&lt;/code&gt; on PyPI is an unrelated data-quality package by a different author. 0.2.0 is also the first release published to PyPI at all, because the name collision meant 0.1.x never shipped there.&lt;/p&gt;

&lt;p&gt;The package classifiers say Development Status 4, Beta, which is the accurate description. The whole thing is roughly 1,400 lines across three modules, small enough to read end to end in an afternoon. For a component sitting on a security boundary I would treat that as a feature rather than an apology, and I would rather people read it than take this post's word for anything.&lt;/p&gt;

&lt;p&gt;If you are running an agent that writes SQL against anything sensitive, an hour spent trying to beat your own guard, whatever form it takes, is likely to pay for itself. The eight above are a starting list. If you find something in &lt;code&gt;sql-guard&lt;/code&gt;, please use the disclosure route in &lt;a href="https://github.com/sakura-sky/sql-guard/blob/main/SECURITY.md" rel="noopener noreferrer"&gt;SECURITY.md&lt;/a&gt; instead of a public issue: &lt;code&gt;security@sakurasky.com&lt;/code&gt; or a draft advisory, with the SQL, the config, and the decision you expected against the one you got. Everything that is not a bypass is very welcome in the issue tracker.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Disclosure: &lt;code&gt;sql-guard&lt;/code&gt; is developed and maintained by Sakura Sky and released under Apache-2.0. Sakura Sky uses it in client-facing agent work. There is no paid tier, hosted version or commercial product built on it. The findings described here come from an internal adversarial review of a deployment rather than a third-party security audit.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;OWASP (2025) &lt;em&gt;LLM01:2025 Prompt Injection, OWASP Top 10 for LLM Applications&lt;/em&gt;. OWASP Gen AI Security Project. Available at: &lt;a href="https://genai.owasp.org/llmrisk/llm01-prompt-injection/" rel="noopener noreferrer"&gt;https://genai.owasp.org/llmrisk/llm01-prompt-injection/&lt;/a&gt; (Accessed: 18 August 2026).&lt;/p&gt;

&lt;p&gt;Sakura Sky Engineering (2026) &lt;em&gt;sql-guard: deterministic policy engine for LLM-generated SQL, v0.2.0&lt;/em&gt;. Available at: &lt;a href="https://github.com/sakura-sky/sql-guard" rel="noopener noreferrer"&gt;https://github.com/sakura-sky/sql-guard&lt;/a&gt; (Accessed: 18 August 2026).&lt;/p&gt;

&lt;p&gt;Mao, T. (2026) &lt;em&gt;SQLGlot: no-dependency SQL parser, transpiler, optimizer and engine&lt;/em&gt;. Available at: &lt;a href="https://sqlglot.com/sqlglot.html" rel="noopener noreferrer"&gt;https://sqlglot.com/sqlglot.html&lt;/a&gt; (Accessed: 18 August 2026).&lt;/p&gt;

&lt;p&gt;Willison, S. (2025) &lt;em&gt;The lethal trifecta for AI agents: private data, untrusted content, and external communication&lt;/em&gt;, 16 June. Available at: &lt;a href="https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/" rel="noopener noreferrer"&gt;https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/&lt;/a&gt; (Accessed: 18 August 2026).&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agentic</category>
      <category>security</category>
      <category>python</category>
    </item>
    <item>
      <title>Arkhe: An Ontology Language for AI Systems</title>
      <dc:creator>Andrew Stevens</dc:creator>
      <pubDate>Thu, 23 Jul 2026 10:25:57 +0000</pubDate>
      <link>https://dev.to/sakurasky/arkhe-an-ontology-language-for-ai-systems-3p2o</link>
      <guid>https://dev.to/sakurasky/arkhe-an-ontology-language-for-ai-systems-3p2o</guid>
      <description>&lt;p&gt;Two weeks ago I &lt;a href="https://www.sakurasky.com/blog/own-your-ontology/" rel="noopener noreferrer"&gt;argued&lt;/a&gt; that the ontology layer is where the recurring agentic failures get retired, and that the artifact doing the retiring should live in your repository rather than a vendor's console (Stevens, 2026a). I closed that piece with a demand: insist on a neutral, versioned, reviewable form for your entities, relationships, and permitted actions. The obvious follow-up question: neutral form in what language, exactly?&lt;/p&gt;

&lt;p&gt;So I spent my evenings since building one. &lt;/p&gt;

&lt;p&gt;Arkhe v0.2.0 shipped today: a small, open-source ontology language for AI systems. Arkhe is my personal project, written on my own time, Apache-2.0 licensed, free forever, with no CLA. It is a Sakura Sky product in no sense at all. I am writing about it here because the argument started here.&lt;/p&gt;

&lt;h2&gt;
  
  
  One file, three registers
&lt;/h2&gt;

&lt;p&gt;An Arkhe module is a YAML file in git. It declares the three registers from the previous post: what exists, how it connects, and what may be done, by whom, with what trace.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;entities&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;FinancialModel&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;keys&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;model_id&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
    &lt;span class="na"&gt;properties&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;state&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;draft&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;in_validation&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;validated&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;production&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;retired&lt;/span&gt;&lt;span class="pi"&gt;],&lt;/span&gt; &lt;span class="nv"&gt;initial&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;draft&lt;/span&gt;&lt;span class="pi"&gt;}&lt;/span&gt;
      &lt;span class="na"&gt;last_validated&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;date&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;optional&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;true&lt;/span&gt;&lt;span class="pi"&gt;}&lt;/span&gt;

&lt;span class="na"&gt;links&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;owned_by&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;from&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;FinancialModel&lt;/span&gt;
    &lt;span class="na"&gt;to&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;TradingDesk&lt;/span&gt;
    &lt;span class="na"&gt;cardinality&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;many_to_one&lt;/span&gt;
    &lt;span class="na"&gt;reverse&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;models&lt;/span&gt;

&lt;span class="na"&gt;actions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;grant_production_use&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;target&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;FinancialModel&lt;/span&gt;
    &lt;span class="na"&gt;guard&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;target.status == "validated" &amp;amp;&amp;amp; months_since(target.last_validated) &amp;lt;= &lt;/span&gt;&lt;span class="m"&gt;12&lt;/span&gt;
    &lt;span class="na"&gt;authority&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;head_of_model_risk&lt;/span&gt;
    &lt;span class="na"&gt;audit&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;mandatory&lt;/span&gt;
    &lt;span class="na"&gt;effects&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;target.status&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;production&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A head of model risk can read that file and object to line eleven. That property, reviewability by the person who owns the business concept rather than the pipeline, drives most of the design. Guards are written in CEL, Google's Common Expression Language, because it is small, side-effect free, and already the policy expression language of half the cloud estate (Google, n.d.). The world is closed: entities and actions that are absent from the file are absent from the system, and no inference engine conjures new ones. The philosophical tradition behind the word ontology is acknowledged and then firmly set aside; Gruber's definition, an explicit specification of a conceptualisation, is the whole ambition (Gruber, 1993).&lt;/p&gt;

&lt;p&gt;The other load-bearing decision: the spec is a contract, never a runtime. Arkhe compiles your module into a neutral intermediate form, one tool contract per action, and emitters generate what your systems consume from there. Nothing of Arkhe executes in your serving path. If the project vanished tomorrow, everything it generated for you would keep working, and the YAML would still be yours to compile with something else.&lt;/p&gt;

&lt;h2&gt;
  
  
  What compilation buys you
&lt;/h2&gt;

&lt;p&gt;The compiler pipeline is deliberately dull. Validate the module (structural rules from a published JSON schema, then semantic rules: key references, link endpoints, guard names, effect types). Generate contracts. Emit.&lt;/p&gt;

&lt;p&gt;Two emitters exist today:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The first produces a native Python library of guarded functions: call &lt;code&gt;grant_production_use&lt;/code&gt; on a model that was validated fourteen months ago and you get a refusal that names the failing clause, &lt;code&gt;months_since(target.last_validated) &amp;lt;= 12&lt;/code&gt;, rather than an exception or, worse, a success. An agent wired to those functions inherits the refusal semantics without any prompt engineering. &lt;/li&gt;
&lt;li&gt;The second emitter, new in v0.2, produces an Open Knowledge Format bundle: one markdown concept document per entity, link, and action, generated from the module's annotations, for LLM consumption. &lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I think of the pair as complementary halves: the contracts declare what an agent may do, and the OKF bundle declares what it should know. Both are regenerated from the same file, so they cannot drift apart.&lt;/p&gt;

&lt;p&gt;v0.2 also added synonyms as first-class annotations, so grounding can match the vocabulary people actually use ("desk" for TradingDesk, "sign-off" for grant_production_use) while the canonical names stay stable, and resolved types inline on effects, so downstream consumers stop re-deriving what a write will do.&lt;/p&gt;

&lt;p&gt;The release I am most pleased with, though, is the one users will never invoke: a second implementation. The Rust port of the validator now passes the same frozen golden fixtures as the Python reference. Porting surfaced three behaviours the spec had never actually decided, all inherited silently from Python's YAML library, including the venerable Norway problem, where &lt;code&gt;no&lt;/code&gt; parses as a boolean. All three are now pinned strictly in a public ADR. My conclusion from that exercise is worth its own post: a language with one implementation has a spec by accident.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who it is for
&lt;/h2&gt;

&lt;p&gt;I wrote persona guides for the three audiences I keep meeting in this work, and they map onto the registers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;GRC professionals get controls as structure. &lt;code&gt;authority: head_of_model_risk&lt;/code&gt; and &lt;code&gt;audit: mandatory&lt;/code&gt; are enforced properties with a trace, and a refusal that cites its failing clause is audit evidence in a way that a prompt instruction never will be. The reviewable YAML is the control library; the git history is the change record your auditor keeps asking for.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Data engineers get the semantic layer as code, sitting in the same repository as the schema it describes, validated in CI with finding codes and line numbers like any other artifact. If you have ever maintained business definitions in a wiki while the warehouse drifted underneath them, the appeal of &lt;code&gt;arkhe validate&lt;/code&gt; failing your pipeline should be immediate.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Architects get protocol independence. The tool-contract IR sits at the centre; MCP, OpenAI function schemas, and Google ADK tool definitions are all planned as thin emitters from it, which means the protocol churn of the past two years stops being a rewrite trigger. Choosing Arkhe commits you to a file format, and deliberately to nothing else.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The use case I most want to see built, and intend to build as the flagship module, is an ontology of the AI estate itself: models, agents, tools, datasets, evals, guardrails, and incidents, with actions like promote-to-production guarded on eval results. The industry currently manages its most consequential systems with less structural rigour than it applies to a customer table. That seems worth fixing with the same medicine.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it refuses to be
&lt;/h2&gt;

&lt;p&gt;Arkhe is a language and a compiler. There is no workspace, no operational database, no query engine, no agent runtime. Palantir productised the three registers inside a platform a decade ago (Palantir Technologies, 2026), and several open projects are now circling the same territory with runtimes attached; the prior-art section of the README names them rather than pretending the idea arrived from nowhere. The bet Arkhe makes is narrower than any of them: the durable artifact is the reviewable file and its compiled contracts, and everything stateful belongs to systems you already run. If that bet is wrong, the cost of having tried it is one YAML file you can still parse with anything.&lt;/p&gt;

&lt;p&gt;v0.2.0 is on PyPI (&lt;code&gt;pip install arkhelang&lt;/code&gt;), crates.io (&lt;code&gt;cargo install arkhe&lt;/code&gt;), and GitHub, with the spec sketch, nine ADRs, the golden fixtures, and the persona guides. It is deliberately small. If the previous post convinced you the neutral form should exist, this one is an invitation to argue with a concrete candidate, ideally in the form of an issue with a failing fixture attached.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Disclosure: Arkhe is the author's personal open-source project and is unaffiliated with Sakura Sky; Sakura Sky has no commercial interest in it. The author also maintains &lt;a href="https://deterministicagents.ai/" rel="noopener noreferrer"&gt;GATE&lt;/a&gt;, an open framework for governed agent runtimes, with which Arkhe aims to interoperate on equal terms with any other runtime.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;Google (n.d.) &lt;em&gt;Common Expression Language specification&lt;/em&gt;. Available at: &lt;a href="https://github.com/google/cel-spec" rel="noopener noreferrer"&gt;https://github.com/google/cel-spec&lt;/a&gt; (Accessed: 23 July 2026).&lt;/p&gt;

&lt;p&gt;Gruber, T.R. (1993) 'A translation approach to portable ontology specifications', &lt;em&gt;Knowledge Acquisition&lt;/em&gt;, 5(2), pp. 199-220.&lt;/p&gt;

&lt;p&gt;Palantir Technologies (2026) &lt;em&gt;Ontology overview, Foundry documentation&lt;/em&gt;. Available at: &lt;a href="https://www.palantir.com/docs/foundry/ontology/overview" rel="noopener noreferrer"&gt;https://www.palantir.com/docs/foundry/ontology/overview&lt;/a&gt; (Accessed: 23 July 2026).&lt;/p&gt;

&lt;p&gt;Stevens, A. (2026a) &lt;em&gt;Your ontology is the asset. Stop renting it back.&lt;/em&gt; Sakura Sky, 10 July. Available at: &lt;a href="https://www.sakurasky.com/blog/own-your-ontology/" rel="noopener noreferrer"&gt;https://www.sakurasky.com/blog/own-your-ontology/&lt;/a&gt; (Accessed: 23 July 2026).&lt;/p&gt;

&lt;p&gt;Stevens, A. (2026b) &lt;em&gt;Arkhe: an ontology language for AI systems&lt;/em&gt;. Available at: &lt;a href="https://github.com/arkhelang/arkhelang" rel="noopener noreferrer"&gt;https://github.com/arkhelang/arkhelang&lt;/a&gt; (Accessed: 23 July 2026).&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agentic</category>
      <category>governance</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Identity Is the Bank Now</title>
      <dc:creator>Andrew Stevens</dc:creator>
      <pubDate>Wed, 15 Jul 2026 09:56:08 +0000</pubDate>
      <link>https://dev.to/sakurasky/identity-is-the-bank-now-4bnc</link>
      <guid>https://dev.to/sakurasky/identity-is-the-bank-now-4bnc</guid>
      <description>&lt;p&gt;A customer who has banked with the same institution for years is known to it in a number of different ways:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;To the onboarding system that ran the identity checks, the customer is a passport, an address, and a date on which verification was completed. &lt;/li&gt;
&lt;li&gt;To the fraud engine, the same customer is a pattern of devices, locations, and spending that it scores in real time. &lt;/li&gt;
&lt;li&gt;To the anti-money-laundering system, a risk rating and a stack of alerts, most of them cleared long ago. &lt;/li&gt;
&lt;li&gt;To the app and the contact centre, a name, a photograph, and a service history. &lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Four example systems, four versions of one person, and they do not agree with one another. When that customer moves house, the bank often hears about it three times, because three of the four systems have no idea the other two have already been updated. The customer notices the seams. So, in a different way, do the people who defraud banks for a living.&lt;/p&gt;

&lt;p&gt;Every retail bank has some version of this customer, and the fragmentation is not the failing of one institution. Fraud, financial crime, KYC, and customer experience grew up as separate disciplines, with separate teams, separate budgets, and separate systems, and for a long time that separation held up fine. It holds up much less well now, because the four have collapsed into a single engineering question, and banking identity is the name for the answer. &lt;/p&gt;

&lt;p&gt;Does the bank hold one reliable, current, shared understanding of who its customer is, and can every function that needs it read from and write to that understanding in close to real time? The wider series framed identity as the point where trust stops being a claim and becomes an engineering output (see &lt;a href="https://www.sakurasky.com/blog/engineering-underneath-part-6/" rel="noopener noreferrer"&gt;Trust Is an Engineering Output&lt;/a&gt;). &lt;/p&gt;

&lt;p&gt;Inside a bank, identity is turning into the thing the whole institution runs on.&lt;/p&gt;

&lt;h2&gt;
  
  
  The seamed customer
&lt;/h2&gt;

&lt;p&gt;Look first at what the seams cost, because they are expensive on both sides at once. For the customer, they show up as friction. A repeat verification for a transaction the fraud engine did not recognise, a wait while a legitimate payment is held, a request to re-confirm details the bank already holds, an address updated in one channel while another channel keeps posting to the old one. None of this is catastrophic. All of it, accumulated across millions of customers, is the texture of a bank that feels harder to deal with than it should.&lt;/p&gt;

&lt;p&gt;The other side is where it gets serious. &lt;/p&gt;

&lt;p&gt;The same seams are an attack surface, and financial crime lives in the gaps between the four systems. A synthetic identity, assembled from real and fabricated fragments, can pass onboarding because the KYC system checks documents rather than coherence, and then behave well enough that the fraud engine, which never saw the onboarding anomaly, has no reason to look twice. An account takeover exploits the fact that the fraud engine sees a device change the AML system will never hear about, while the AML system sees a transaction pattern the fraud engine treats as somebody else's problem. No single system holds the whole picture of the customer, so no single system can see the whole picture of the attack. The criminal's real advantage is not sophistication. It is that the bank's view of the customer is split four ways and the criminal's view of that same customer is whole.&lt;/p&gt;

&lt;h2&gt;
  
  
  How identity got fragmented
&lt;/h2&gt;

&lt;p&gt;The fragmentation was built one sensible decision at a time. The KYC system arrived to satisfy onboarding obligations, and those obligations are real: the global standard requires a bank to identify and verify the customer, identify beneficial owners, understand the relationship, and monitor it on an ongoing basis (FATF, 2012). The AML monitoring system was bought separately, often from a different vendor, and keyed on accounts and transactions rather than on people. The fraud engine came in for real-time scoring, keyed on sessions and devices. The customer relationship system was built for service, keyed on a contact record. Each was the right tool for its job, procured by the team that owned that job, on its own timeline.&lt;/p&gt;

&lt;p&gt;What none of them shared was a canonical idea of the customer. Each held its own identifier for the same human being, and nothing tied those identifiers together into one durable customer identity. So the bank ended up able to answer four narrow questions well and the one broad question, who is this customer and what do we currently know about them, not at all. This is the state most KYC architecture is in: strong at the point check, silent on the whole.&lt;/p&gt;

&lt;p&gt;The regulatory current now runs the other way on both sides of the Atlantic, which is worth noticing. In the United States, FinCEN's customer due diligence rule already requires banks to identify and verify the beneficial owners behind their legal-entity customers and to keep that understanding current, and a February 2026 order recast the refresh obligation on a risk basis rather than as a mechanical repeat at every new account (FinCEN, 2016; FinCEN, 2026). In the European Union, the 2024 anti-money-laundering reform moves the bloc onto a single rulebook and stands up a central authority to supervise it, with directly applicable rules on customer due diligence and beneficial ownership designed to make that data consistent across institutions rather than bespoke to each (European Parliament and Council, 2024). Different regimes, one direction of travel, toward shared and coherent identity. A bank whose own AML data cannot be reconciled across its four internal systems is starting that journey a long way back.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a unified identity layer looks like
&lt;/h2&gt;

&lt;p&gt;The fix is a unified identity layer, and it is worth being precise about what that means, because it is not another copy of the data in a warehouse. It is one canonical, current, authoritative record of each customer, built by resolving the fragments the four systems already hold into a single entity. Entity resolution, the work of deciding that this passport, that device pattern, and that contact record all refer to the same real person, is the hard technical core of it, and it is never perfectly clean, which is exactly why it has to be engineered rather than assumed.&lt;/p&gt;

&lt;p&gt;The layer holds the durable answer to who the customer is, and links out to the signals each function produces: the KYC status and its evidence, the AML risk rating and its alerts, the live fraud signals, the service history. The point is not that one team now owns everything. The point is that any function can see the whole customer without having to own all of the customer. The fraud engine can factor in that onboarding flagged something odd. The AML system can see that the fraud engine just watched the customer's device and location change. Onboarding can stop asking for what the bank already holds. The identity layer becomes the shared, low-latency, authoritative view that every other system reads from and contributes to, and building that view out of messy source data is a serious piece of identity engineering rather than a reporting exercise.&lt;/p&gt;

&lt;h2&gt;
  
  
  The security architecture that supports it
&lt;/h2&gt;

&lt;p&gt;Here is the part that a bank has to take seriously before it builds any of this, because getting it wrong turns an asset into a liability. A unified identity layer concentrates the most sensitive data in the entire institution into one place. That is precisely what makes it valuable, and precisely what makes it the highest-value target the bank owns. Unify identity without hardening it and the result is a better-organised honeypot.&lt;/p&gt;

&lt;p&gt;So the unification and the protection are the same project, and banking security has to be designed into the layer from the first day rather than added once it works. That means access that is fine-grained and purpose-limited, so the fraud engine reads the fields it needs for fraud prevention and nothing more, and the contact centre sees a service view that does not expose the full financial-crime picture. It means strong identity for the workloads and the people reaching the layer, not just for the customers described in it. It means tokenising the most sensitive attributes, so that a breach of one consumer does not spill raw identity data. And it means a complete, tamper-evident record of every read and write, both because a supervisor will eventually ask who saw what, and because the bank itself needs to know. Get this right and the concentration of identity is a strength. Get it wrong and it is the worst single point of failure a bank could design. This seam between unifying data and defending it is the ground Sakura's &lt;a href="https://www.sakurasky.com/security/" rel="noopener noreferrer"&gt;Security practice&lt;/a&gt; works on with a bank's identity teams.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changes once it exists
&lt;/h2&gt;

&lt;p&gt;Come back to that customer, and follow what a working identity layer changes for them and for the bank at the same time. The address update lands once and every channel reflects it. The legitimate payment clears because the fraud engine can see the customer is exactly who the rest of the bank already knows them to be. Onboarding to a new product takes minutes because the bank reuses what it verified years ago instead of starting over. The friction that made the bank feel hard to deal with mostly disappears, and it disappears for the same reason the bank gets safer.&lt;/p&gt;

&lt;p&gt;Because the criminal loses the seams. Synthetic identities are harder to sustain when onboarding, fraud, and financial-crime signals resolve to one view that has to stay coherent over time. Account takeover is harder when the device change the fraud engine sees and the transaction pattern the AML system sees are looking at the same customer record rather than two strangers. Fraud prevention and financial-crime detection stop being separate contests the bank fights with half the picture each.&lt;/p&gt;

&lt;p&gt;And the bank gains the property the whole series has been circling. It can answer, for any customer, what it knows and how it knows it, on demand and with the evidence attached. That is trust as an output of the architecture rather than an assertion in a policy, and it turns out to run on a clean identity foundation. A bank that builds one is not a slow bank: Xapo Bank &lt;a href="https://www.sakurasky.com/case-studies/innovation-at-speed-how-xapo-bank-achieved-genai-adoption-in-just-8-weeks/" rel="noopener noreferrer"&gt;reached production on a tightly governed stack in weeks rather than quarters&lt;/a&gt;, because the foundations were engineered once and everything else inherited them.&lt;/p&gt;

&lt;p&gt;The unified identity layer that all four functions read from and write to has to be built out of the fragmented systems a bank already runs, and constructing that resolved, current, authoritative view of the customer is the work Sakura's &lt;a href="https://www.sakurasky.com/data/" rel="noopener noreferrer"&gt;Data &amp;amp; AI practice&lt;/a&gt; does inside financial institutions.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;European Parliament and Council, 2024. &lt;em&gt;Regulation (EU) 2024/1624 of the European Parliament and of the Council of 31 May 2024 on the prevention of the use of the financial system for the purposes of money laundering or terrorist financing (Anti-Money Laundering Regulation).&lt;/em&gt; Official Journal of the European Union, L, 19 June. Available at: &lt;a href="https://eur-lex.europa.eu/eli/reg/2024/1624/oj" rel="noopener noreferrer"&gt;https://eur-lex.europa.eu/eli/reg/2024/1624/oj&lt;/a&gt; [Accessed 10 July 2026].&lt;/p&gt;

&lt;p&gt;Financial Action Task Force, 2012. &lt;em&gt;International Standards on Combating Money Laundering and the Financing of Terrorism and Proliferation: The FATF Recommendations (updated).&lt;/em&gt; Financial Action Task Force, Paris. Available at: &lt;a href="https://www.fatf-gafi.org/en/publications/Fatfrecommendations/Fatf-recommendations.html" rel="noopener noreferrer"&gt;https://www.fatf-gafi.org/en/publications/Fatfrecommendations/Fatf-recommendations.html&lt;/a&gt; [Accessed 10 July 2026].&lt;/p&gt;

&lt;p&gt;Financial Crimes Enforcement Network (FinCEN), 2016. &lt;em&gt;Customer Due Diligence Requirements for Financial Institutions, Final Rule.&lt;/em&gt; 31 CFR Parts 1010, 1020, 1023, 1024 and 1026. US Department of the Treasury. Available at: &lt;a href="https://www.federalregister.gov/documents/2016/05/11/2016-10567/customer-due-diligence-requirements-for-financial-institutions" rel="noopener noreferrer"&gt;https://www.federalregister.gov/documents/2016/05/11/2016-10567/customer-due-diligence-requirements-for-financial-institutions&lt;/a&gt; [Accessed 10 July 2026].&lt;/p&gt;

&lt;p&gt;Financial Crimes Enforcement Network (FinCEN), 2026. &lt;em&gt;Order Granting Exceptive Relief from the Beneficial Ownership Requirements for Legal Entity Customers.&lt;/em&gt; US Department of the Treasury, 13 February. Available at: &lt;a href="https://www.fincen.gov/system/files/2026-02/FinCEN-Order-CCDExceptiveRelief.pdf" rel="noopener noreferrer"&gt;https://www.fincen.gov/system/files/2026-02/FinCEN-Order-CCDExceptiveRelief.pdf&lt;/a&gt; [Accessed 10 July 2026].&lt;/p&gt;

</description>
      <category>financialservices</category>
      <category>banking</category>
      <category>identity</category>
      <category>kyc</category>
    </item>
    <item>
      <title>Migrating a Live Real-Time Communications Platform from AWS to Google Cloud</title>
      <dc:creator>Andrew Stevens</dc:creator>
      <pubDate>Tue, 14 Jul 2026 11:15:08 +0000</pubDate>
      <link>https://dev.to/sakurasky/migrating-a-live-real-time-communications-platform-from-aws-to-google-cloud-4l5f</link>
      <guid>https://dev.to/sakurasky/migrating-a-live-real-time-communications-platform-from-aws-to-google-cloud-4l5f</guid>
      <description>&lt;p&gt;11Sight operates a real-time voice and video engagement platform, the communications backbone behind its AI agents for automotive and hospitality businesses. Calls are the product. An infrastructure migration that takes the platform offline for a weekend was never an option. Sakura Sky's &lt;a href="https://www.sakurasky.com/cloud/" rel="noopener noreferrer"&gt;Cloud practice&lt;/a&gt; moved the platform from AWS to Google Cloud with a phased hybrid strategy that kept it serving live traffic throughout, and held user-facing downtime at the final cutover to under five minutes.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Challenge
&lt;/h3&gt;

&lt;p&gt;11Sight's production environment had grown up on EC2 and RDS: a monolithic web application on VMs, a Jitsi-based conferencing stack with video bridges and recorders, and a PostgreSQL database holding the customer data that every call depends on. Three constraints shaped the engagement:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Live traffic, all the time.&lt;/strong&gt; A real-time communications platform has no quiet maintenance window long enough for a big-bang cutover, so the migration had to run while customers kept making calls.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prove everything before production.&lt;/strong&gt; Cutover timing, VPN latency, and application behaviour on GKE all had to be validated against production-grade infrastructure before any customer traffic depended on them, which made a full rehearsal environment a first-class deliverable of the migration plan.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Modernization over relocation.&lt;/strong&gt; The goal was to land on a cloud-native footing, with Kubernetes where it earned its keep and managed services for state, rather than reproduce the VM-centric estate on new hardware.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The Engineering
&lt;/h3&gt;

&lt;p&gt;We designed a compute-first, three-phase hybrid migration that decoupled the application move from the database move, so each could be validated independently. Four engineering decisions carried the project:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Landing zone before workloads.&lt;/strong&gt; The first deliverable was a Google Cloud foundation built from Sakura Sky's &lt;a href="https://www.sakurasky.com/enclave/" rel="noopener noreferrer"&gt;Enclave&lt;/a&gt; Terraform blueprint: organization structure, IAM groups, centralized logging and monitoring, a Shared VPC, and a Cloud HA VPN linking the AWS and GCP networks. Every subsequent resource was defined in Terraform, in repositories created inside 11Sight's environment from day one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rehearse the whole migration in staging first.&lt;/strong&gt; We provisioned a complete staging environment (GKE for the containerized web application, Cloud SQL for PostgreSQL 16, Memorystore with the Valkey engine, Compute Engine instances behind autoscaling groups for the conferencing workloads) and used it to rehearse the entire migration. That rehearsal validated that VPN latency between the GCP application tier and the AWS database was within production thresholds, and produced a measured downtime estimate of 45 to 90 minutes for the final cutover.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compute first, data second.&lt;/strong&gt; In production, Phase 1 shifted live traffic from the AWS VMs to the GKE load balancer via DNS while all reads and writes continued against AWS RDS over the VPN. Phase 2 enabled logical replication on RDS and ran a continuous Database Migration Service job into Cloud SQL, keeping the two databases in near-real-time sync.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A rehearsed cold cutover.&lt;/strong&gt; Phase 3 was a stop-and-go cutover inside the planned window: drain traffic at the GKE ingress, quiesce the source database, promote the Cloud SQL replica once DMS reported no lag, repoint GKE configuration and secrets, and redeploy. A war room voice bridge kept migration, DevOps, development, and QA leads on one channel, and traffic reopened only after internal health checks passed against the new database.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The Results
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Under 5 minutes of downtime.&lt;/strong&gt; Against a planned 45 to 90 minute maintenance window, actual user-facing downtime at the production cutover was less than five minutes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zero data loss.&lt;/strong&gt; Continuous DMS replication and the no-lag promotion gate meant the cutover moved the database without losing a single write.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A cloud-native production platform.&lt;/strong&gt; The web application now runs containerized on GKE, state lives in managed services, and the conferencing fleet scales behind autoscaling groups instead of hand-tended VMs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A proving ground that outlives the project.&lt;/strong&gt; The rehearsal environment was built to outlast the migration: a permanent, Terraform-defined staging environment where future releases, upgrades, and scaling decisions get validated before they reach customers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Everything as code, owned by the client.&lt;/strong&gt; All infrastructure is defined in Terraform in 11Sight-owned repositories, so there was nothing to hand over at close-out that 11Sight did not already control.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;This is the type of work our &lt;a href="https://www.sakurasky.com/accelerate/" rel="noopener noreferrer"&gt;Accelerate&lt;/a&gt; solution delivers on, with foundations laid by &lt;a href="https://www.sakurasky.com/enclave/" rel="noopener noreferrer"&gt;Enclave&lt;/a&gt;: production capability built inside the client's environment, jointly with their team. &lt;a href="https://www.sakurasky.com/contact/" rel="noopener noreferrer"&gt;Contact us&lt;/a&gt; to scope a similar migration.&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>gcp</category>
      <category>kubernetes</category>
      <category>iac</category>
    </item>
    <item>
      <title>Regulatory Evidence at Machine Speed</title>
      <dc:creator>Andrew Stevens</dc:creator>
      <pubDate>Mon, 13 Jul 2026 09:38:21 +0000</pubDate>
      <link>https://dev.to/sakurasky/regulatory-evidence-at-machine-speed-3i04</link>
      <guid>https://dev.to/sakurasky/regulatory-evidence-at-machine-speed-3i04</guid>
      <description>&lt;p&gt;The request arrived on a Tuesday and gave the compliance team twenty-four hours. The regulator wanted to know why the bank's transaction monitoring system had cleared a particular payment eleven months earlier: which rules fired and which did not, what customer risk score applied at that moment, who reviewed the resulting alert, and what that reviewer actually saw. The system that made the decision was still running, unchanged, and working correctly. The evidence about that one decision was somewhere else entirely. It sat in a queue, behind a request to rebuild a dataset, behind a restore from backup, behind a data engineer who had other work booked. The bank was not being accused of anything. It was being asked to show its working, and it had twenty-four hours to discover whether it could.&lt;/p&gt;

&lt;p&gt;Banks have always been asked for evidence. What has changed is the nature of the request. The old rhythm was periodic, predictable, and aggregate: a return filed on schedule, a report at quarter end, a sample pulled for an inspection booked weeks in advance. The new rhythm is specific, unscheduled, and granular: this decision, this customer, this moment, and show me now. Most banks built their evidence systems for the first rhythm and are now being asked to serve the second. This is the banking form of a tension that runs through most regulated industries, which the wider series has set out as the gap between evidence and speed (see &lt;a href="https://www.sakurasky.com/blog/engineering-underneath-part-3/" rel="noopener noreferrer"&gt;Evidence Versus Speed&lt;/a&gt;). Inside a bank it takes a particular shape, and it has a particular fix.&lt;/p&gt;

&lt;h2&gt;
  
  
  The twenty-four-hour request
&lt;/h2&gt;

&lt;p&gt;What the compliance team actually had to do, in those twenty-four hours, is the most revealing part of the story. The decision they needed to explain was not recorded anywhere as a decision. It had to be reassembled from three systems that had never been designed to be read together.&lt;/p&gt;

&lt;p&gt;The monitoring engine's logs held which rules had evaluated the payment, but those logs rotated after ninety days, so the relevant window had to be restored from backup. The customer risk score was worse. It was recomputed nightly and overwritten each time, so the score that actually applied eleven months ago no longer existed anywhere; it had to be rebuilt by rerunning the scoring logic against archived inputs and hoping the logic had not changed in the meantime. The analyst's review sat in a case management tool with its own retention policy and no link back to the payment except a reference number typed in by hand.&lt;/p&gt;

&lt;p&gt;The team got there, barely, and the answer was correct. The payment had been cleared for good reasons and the bank had done nothing wrong. The cost was two engineers for the better part of a week and an uncomfortable realisation in the room afterwards: nobody was confident they could do it again. The decision itself had been sound. What the bank could not do was demonstrate it on demand, and that is a different kind of risk from getting the decision wrong. Under MAS Notice 626, records must be retained and be retrievable in a usable form (Monetary Authority of Singapore, 2022). Retention that exists in principle but cannot be produced when the regulator asks is not really retention at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  What regulators used to ask for
&lt;/h2&gt;

&lt;p&gt;The reason banks are in this position is that their evidence systems were built, quite rationally, for the requests that used to arrive. Regulatory reporting was a calendar activity. Returns were filed monthly, quarterly, or annually. Inspections were scheduled. The unit of evidence was the aggregate: a capital position, an exposure summary, a count of alerts raised and cleared. Banks built reporting factories to serve that model, with data warehouses, reconciliation processes, sign-off workflows, and a small industry of controls wrapped around the production of the report.&lt;/p&gt;

&lt;p&gt;Even the regulation that pushed hardest on data quality assumed this shape. The Basel Committee's principles for effective risk data aggregation and risk reporting, published in 2013, told banks to be able to aggregate risk data accurately and to trace it, and it remains the reference point for banking data lineage (Basel Committee on Banking Supervision, 2013). But the output it had in mind was still a report, produced on a cycle, for a supervisor who would read it later.&lt;/p&gt;

&lt;p&gt;Under that model, assembling evidence retrospectively was a perfectly sensible strategy, because the ask was predictable and the deadline was known. Banking compliance became organised around producing documents on a schedule. Evidence was a product of the reporting cycle rather than a property of the transaction, and for a long time nobody had reason to notice the difference.&lt;/p&gt;

&lt;h2&gt;
  
  
  What they ask for now
&lt;/h2&gt;

&lt;p&gt;The difference is now impossible to miss, because supervisors have stopped confining themselves to the aggregate. They ask about individual decisions, at short notice, and they expect the bank to trace one end to end.&lt;/p&gt;

&lt;p&gt;The pressure is visible in the supervisors' own output. More than a decade after the Basel principles, the European Central Bank found it necessary to issue a guide pressing banks on effective risk data aggregation and reporting, precisely because so many still cannot demonstrate complete, end-to-end lineage across their data estate (European Central Bank, 2024). The Digital Operational Resilience Act, in application since January 2025, requires financial entities to maintain registers of their ICT arrangements and to evidence their operational resilience continuously rather than annually (European Parliament and Council, 2022). And the five-year retrievability standard in MAS Notice 626 is not satisfied by a backup tape that takes a week to read.&lt;/p&gt;

&lt;p&gt;Put together, these describe a single shift. The deliverable has become the reconstructable chain behind any single decision the bank made, available on request, at something close to the speed the bank operates. The report still gets filed, but it is no longer the thing the supervisor is really testing. Real-time compliance is a slightly misleading phrase, because nobody expects the answer instantly. What is expected is that producing the answer is a query rather than an excavation. A financial services audit is becoming a series of specific questions with specific answers, and regulatory reporting, while it continues, is no longer the whole of the job.&lt;/p&gt;

&lt;h2&gt;
  
  
  Engineering evidence into the flow
&lt;/h2&gt;

&lt;p&gt;Meeting that expectation is an engineering problem, and it has an engineering answer. The evidence has to be emitted by the system as it runs.&lt;/p&gt;

&lt;p&gt;Look again at what those twenty-four hours actually required. Engineers excavated rotated logs from a backup and hoped the restore was complete. They reran scoring logic against archived inputs and hoped the logic had not drifted in eleven months. They tied a case file to a payment through a reference number somebody had typed in by hand. Every step was archaeology. Every step introduced a guess. What the bank finally handed the regulator was a well-argued account of the decision, assembled under deadline by people reconstructing their own system from its debris.&lt;/p&gt;

&lt;p&gt;Engineer the evidence into the flow and the same request lands very differently. At the moment the decision is made, the system writes it down. It records the inputs that fed the rule, the version of the rule and the model that evaluated them, the score they produced, the identity of anyone who touched the alert, and the time it happened, all bound together by a single identifier. It preserves the risk score as it stood that day instead of overwriting it at midnight. It hash-links the record, so any later alteration shows. It makes the record addressable, so one payment resolves to one chain. One approach excavates. The other retrieves.&lt;/p&gt;

&lt;p&gt;This is what engineered compliance means in a bank. The evidence becomes a by-product of operating the system, produced continuously whether or not anyone asks for it. It costs something to build. It costs considerably less than paying for reconstruction every time and never being certain the reconstruction will hold.&lt;/p&gt;

&lt;p&gt;None of that capability lives in the monitoring engine. It lives one layer down, in the data foundation that carries lineage and preserved history underneath every transaction the bank processes, and that layer is precisely what Sakura's &lt;a href="https://www.sakurasky.com/data/" rel="noopener noreferrer"&gt;Data &amp;amp; AI practice&lt;/a&gt; builds inside regulated institutions. A bank standing on that foundation can move quickly: Xapo Bank &lt;a href="https://www.sakurasky.com/case-studies/innovation-at-speed-how-xapo-bank-achieved-genai-adoption-in-just-8-weeks/" rel="noopener noreferrer"&gt;reached production on a tightly governed stack in weeks rather than quarters&lt;/a&gt;, because the controls were engineered into the platform instead of being negotiated afresh for every release.&lt;/p&gt;

&lt;h2&gt;
  
  
  The audit cycle that disappears
&lt;/h2&gt;

&lt;p&gt;The payoff goes well beyond surviving the next twenty-four-hour request. The audit cycle stops being an event in the bank's calendar at all.&lt;/p&gt;

&lt;p&gt;When evidence is a running output, the request that consumed two engineers for a week becomes a query answered in minutes by someone in compliance who does not need to open a ticket. Audit-ready banking means the bank answers without a change freeze, without pulling engineers off delivery, and without the background dread that the answer might not be reproducible. Everything else the bank is trying to build keeps moving while the question is answered.&lt;/p&gt;

&lt;p&gt;The banks that bolt evidence on afterwards pay for it twice. They pay once in reconstruction, and again in the drag on everything else, because a system whose evidence cannot be produced on demand cannot safely be changed quickly. Every release carries an unpriced risk of breaking a chain nobody can currently see. There is a further dividend, too: a bank that can prove exactly what its transaction monitoring did, and why, can afford to tune it more aggressively, because it can demonstrate the effect of the change rather than argue about it. Evidence engineered into the flow is what lets a bank keep moving while it answers.&lt;/p&gt;

&lt;p&gt;The work does not end when the architecture is built. The chain has to hold through every release, every model change, and every new rule, and someone has to be able to prove it still does. Running that discipline inside a bank is what Sakura's &lt;a href="https://www.sakurasky.com/grc/" rel="noopener noreferrer"&gt;GRC service&lt;/a&gt; is for.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;Basel Committee on Banking Supervision, 2013. &lt;em&gt;Principles for effective risk data aggregation and risk reporting.&lt;/em&gt; Bank for International Settlements, Basel. Available at: &lt;a href="https://www.bis.org/publ/bcbs239.htm" rel="noopener noreferrer"&gt;https://www.bis.org/publ/bcbs239.htm&lt;/a&gt; [Accessed 10 July 2026].&lt;/p&gt;

&lt;p&gt;European Central Bank, 2024. &lt;em&gt;Guide on effective risk data aggregation and risk reporting.&lt;/em&gt; ECB Banking Supervision, Frankfurt. Available at: &lt;a href="https://www.bankingsupervision.europa.eu/ecb/pub/pdf/ssm.supervisory_guides240503_riskreporting.en.pdf" rel="noopener noreferrer"&gt;https://www.bankingsupervision.europa.eu/ecb/pub/pdf/ssm.supervisory_guides240503_riskreporting.en.pdf&lt;/a&gt; [Accessed 10 July 2026].&lt;/p&gt;

&lt;p&gt;European Parliament and Council, 2022. &lt;em&gt;Regulation (EU) 2022/2554 of the European Parliament and of the Council of 14 December 2022 on digital operational resilience for the financial sector (Digital Operational Resilience Act).&lt;/em&gt; Official Journal of the European Union, L 333, 27 December, pp. 1-79. Available at: &lt;a href="https://eur-lex.europa.eu/eli/reg/2022/2554/oj" rel="noopener noreferrer"&gt;https://eur-lex.europa.eu/eli/reg/2022/2554/oj&lt;/a&gt; [Accessed 10 July 2026].&lt;/p&gt;

&lt;p&gt;Monetary Authority of Singapore, 2022. &lt;em&gt;MAS Notice 626: Notice to Banks on Prevention of Money Laundering and Countering the Financing of Terrorism.&lt;/em&gt; Monetary Authority of Singapore. Available at: &lt;a href="https://www.mas.gov.sg/regulation/notices/notice-626" rel="noopener noreferrer"&gt;https://www.mas.gov.sg/regulation/notices/notice-626&lt;/a&gt; [Accessed 10 July 2026].&lt;/p&gt;

</description>
      <category>financialservices</category>
      <category>banking</category>
      <category>compliance</category>
      <category>audit</category>
    </item>
  </channel>
</rss>
