<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ken W Alger</title>
    <description>The latest articles on DEV Community by Ken W Alger (@kenwalger).</description>
    <link>https://dev.to/kenwalger</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F15734%2F22d0195e-9fce-4d80-9ae2-3bb416bf8d6f.jpg</url>
      <title>DEV Community: Ken W Alger</title>
      <link>https://dev.to/kenwalger</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/kenwalger"/>
    <language>en</language>
    <item>
      <title>Views Measure Views</title>
      <dc:creator>Ken W Alger</dc:creator>
      <pubDate>Thu, 01 Oct 2026 19:32:00 +0000</pubDate>
      <link>https://dev.to/kenwalger/views-measure-views-4co7</link>
      <guid>https://dev.to/kenwalger/views-measure-views-4co7</guid>
      <description>&lt;p&gt;&lt;em&gt;Nine years in, I finally worked out what else to count.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A writer I follow, &lt;a href="https://dev.to/sylwia-lask"&gt;Sylwia Laskowska&lt;/a&gt;, recently published a post about &lt;a href="https://dev.to/sylwia-lask/the-accidental-blogger-how-i-ended-up-on-dev-5a3f"&gt;accidentally becoming a blogger&lt;/a&gt;. She has been writing on DEV for about a year, and the numbers attached to that year are impressive: hundreds of thousands of views and tens of thousands of followers.&lt;/p&gt;

&lt;p&gt;I have been on DEV for more than nine years. This morning I am at 48,408 total views.&lt;/p&gt;

&lt;p&gt;Before I go any further, I want to be clear about something, because the essay that usually follows a comparison like that is insufferable. Building a large readership for accessible, useful developer writing is genuinely difficult, and doing it in a year is more difficult still. I would like more people to read my work. I can learn a great deal from writers who are better than I am at audience building, topic selection, accessibility, and community participation.&lt;/p&gt;

&lt;p&gt;So this is not a piece about why small numbers are secretly good. It is a piece about what I spent nine years failing to separate.&lt;/p&gt;

&lt;h2&gt;
  
  
  The number that stopped me using one of my metrics
&lt;/h2&gt;

&lt;p&gt;I recently pulled nine years of my own data out of the DEV API, mostly to make a chart of follower growth with my publication dates marked on it, so I could see which posts moved the line.&lt;/p&gt;

&lt;p&gt;The chart was useless, and the reason is instructive.&lt;/p&gt;

&lt;p&gt;In 2026 I gained roughly eighteen thousand followers. My 2026 posts have about fourteen thousand views between them. You cannot acquire eighteen thousand readers from fourteen thousand page loads. The daily follow rate sits around 130 and does not respond to whether I publish anything, and about 37% of the usernames carry auto-generated hex or numeric tails. It is reciprocal-follow farming, it is endemic, and it has nothing to do with me or my writing.&lt;/p&gt;

&lt;p&gt;Here is the cleanest version of it. Since the middle of September, my follower count has gone from 18,148 to 21,048. Over the same fifteen days, my total view count went from 46,774 to 48,408.&lt;/p&gt;

&lt;p&gt;Two thousand nine hundred new followers. One thousand six hundred and thirty-four new views.&lt;/p&gt;

&lt;p&gt;I gained nearly twice as many followers as readers, and a follower is supposed to be a reader who liked something enough to want more. That number had been sitting on my profile for months looking like evidence of something.&lt;/p&gt;

&lt;p&gt;A page view is at least an honest measurement. Somebody loaded the page. That is real information, and I think "vanity metric" is an unfair label if it is taken to mean meaningless.&lt;/p&gt;

&lt;p&gt;The trouble starts when we quietly change the claim from "this post received more views" to "this post was more successful."&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Successful at what?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  A page view is real. It just is not the whole story.
&lt;/h2&gt;

&lt;p&gt;In corporate content the answer to that question is usually explicit. A post might exist to attract someone searching for a problem, introduce a product, move that reader toward a trial, and eventually help create a customer. Ten thousand views with no downstream behaviour may be worth less to that company than five hundred views that put twenty qualified developers into the funnel.&lt;/p&gt;

&lt;p&gt;Personal writing has a funnel too, just a much vaguer one. Someone reads one article, encounters another a month later, starts recognising the name, follows, leaves a substantive comment, references the work elsewhere. The page view was real. It was not necessarily the outcome.&lt;/p&gt;

&lt;p&gt;It also helps to remember that a broadly useful JavaScript tutorial and an article about authority boundaries in AI-generated software are not competing for the same reader. One is relevant to a large share of a developer community. The other starts with a much smaller pool of people who already care about capability security and the distinction between correctness and authority.&lt;/p&gt;

&lt;p&gt;That is not so different from comparing the audience for a popular fantasy novel with the audience for a presidential memoir. Both books can be excellent. Both can do exactly what their authors intended. Their potential readerships are still radically different.&lt;/p&gt;

&lt;p&gt;Audience size is partly a property of the artifact and partly a property of the market around it. That sounds obvious about books. Writers forget it remarkably fast while staring at a dashboard.&lt;/p&gt;

&lt;p&gt;An article can also have a desired consequence without a button under it. Sometimes I want someone to challenge the argument. Sometimes I want a developer to recognise a problem in their own system. Sometimes the entire intended outcome is that a reader leaves with a question they were not asking ten minutes earlier. That still gives me something to evaluate against. It just refuses to appear as &lt;code&gt;conversion=true&lt;/code&gt; in an analytics dashboard.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quality, audience fit, and distribution are different problems
&lt;/h2&gt;

&lt;p&gt;I have come to think of online writing as at least three separate problems.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Writing quality&lt;/strong&gt; is the craft problem. Is the argument coherent? Is the explanation useful? Did I do the research? Is there something here worth another person's time?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Audience fit&lt;/strong&gt; is the relevance problem. How many people where I publish are likely to care about this subject? How much prerequisite knowledge does it demand? Can someone scrolling a feed see immediately why the question matters to them?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Distribution&lt;/strong&gt; is the discovery problem. How does the article reach those people? Search, followers, newsletters, speaking, community participation, platform curation, links from other writers, or some combination.&lt;/p&gt;

&lt;p&gt;Those interact, but they are not interchangeable. An article has to clear all three, and failing any one of them produces the same disappointing number for completely different reasons.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    W["Article"] --&amp;gt; Q{"Quality"}
    Q --&amp;gt;|"weak"| X1["Nobody finishes it"]
    Q --&amp;gt;|"strong"| F{"Audience fit"}
    F --&amp;gt;|"wrong room"| X2["Few people care"]
    F --&amp;gt;|"right room"| D{"Distribution"}
    D --&amp;gt;|"not found"| X3["Nobody sees it"]
    D --&amp;gt;|"found"| R["Readers"]

    classDef gate fill:#FFFFFF,stroke:#166534,color:#14532D,stroke-width:2px;
    classDef miss fill:#FEF2F2,stroke:#991B1B,color:#7F1D1D,stroke-width:2px;
    classDef good fill:#E8F3EE,stroke:#166534,color:#14532D,stroke-width:2px;

    class Q,F,D gate;
    class X1,X2,X3 miss;
    class W,R good;&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;That is the diagnostic value of separating them. Three posts can land at 80 views apiece and need three entirely different responses. A good article can have poor audience fit. An accessible article can have excellent fit and no distribution. A technically modest piece can answer a question a hundred thousand people are asking today. A strong argument can address a question five hundred people know they have.&lt;/p&gt;

&lt;p&gt;For most of my writing life I concentrated almost entirely on the first variable and assumed distribution would sort itself out. Sometimes it did. Frequently it did not. My own numbers say that plainly: about 4% of my traffic comes from search, which for someone whose best-performing historical work is evergreen reference material is a distribution problem rather than a quality one.&lt;/p&gt;

&lt;h2&gt;
  
  
  I did not start with a content strategy
&lt;/h2&gt;

&lt;p&gt;Nine years ago my public technical writing grew out of databases, because databases were the work I was doing and the community I was in. I did not sit down with a personal-brand document. I wrote about what I knew, what I was learning, and what developers were asking about.&lt;/p&gt;

&lt;p&gt;Those articles accumulated into something larger. People began associating my name with certain subjects, questions led to more articles, and some of that work became durable enough to keep attracting readers years later. My database writing eventually included MongoDB's &lt;a href="https://www.mongodb.com/company/blog/building-with-patterns-a-summary" rel="noopener noreferrer"&gt;Building with Patterns&lt;/a&gt;, one of the most heavily visited bodies of content I worked on there.&lt;/p&gt;

&lt;p&gt;That history matters when I wander. If I publish one &lt;a href="https://rust-lang.org/" rel="noopener noreferrer"&gt;Rust&lt;/a&gt; article today, it does not arrive with nine years of association between my name and Rust behind it. That does not mean developers are uninterested in Rust, or that my readers dislike it. It may simply mean I have not given a Rust audience any reason to know who I am.&lt;/p&gt;

&lt;p&gt;Those explanations imply completely different responses. If I wanted Rust to become a real part of my writing, one underperforming article would tell me almost nothing. I would need several useful pieces, participation in that community, and time. If I do not want that, the article stays an interesting experiment.&lt;/p&gt;

&lt;p&gt;"This article performed poorly" is an observation. "There is no audience for me here" is an interpretation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Personal writing gets to discover its strategy
&lt;/h2&gt;

&lt;p&gt;I have spent enough of my career around corporate developer content to know that content strategy matters. A company generally knows why it is publishing: which developers it wants to reach, which capabilities it needs explained, which search terms it wants to own. The strategy should exist before anyone fills the editorial calendar.&lt;/p&gt;

&lt;p&gt;Personal writing is stranger. You can chase a question because it bothered you on Tuesday. You can abandon a series when you have nothing else useful to say. You can spend weeks on a technical experiment and then publish something ridiculous because you started wondering whether all the photographs on your phone technically make it heavier.&lt;/p&gt;

&lt;p&gt;You can also discover the strategy after you have written enough to see the pattern. That has increasingly been my experience. Articles I thought were about AI verification, provenance, missing information, audit evidence, memory, and authority turned out to be different views of the same few questions. I did not design that body of work and then manufacture articles to fill it. The writing is how I found it.&lt;/p&gt;

&lt;p&gt;The conversation that prompted this piece made me realise that is less different from corporate strategy than I assumed. Sylwia described writing mostly by intuition while still making choices about what she wants to be known for, which audiences interest her, which adjacent topics fit, and which opportunities she ignores.&lt;/p&gt;

&lt;p&gt;That is a content strategy. It just has a governance structure of one.&lt;/p&gt;

&lt;p&gt;A company may need content, DevRel, product marketing, SEO, and leadership involved in deciding whether a newly discovered audience matters. A personal writer can notice something in the comments on Tuesday and run the experiment on Thursday. The strategic question is nearly identical. The path from observation to decision is not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Signals are not instructions
&lt;/h2&gt;

&lt;p&gt;This matters because audiences talk back. A recurring question in the comments may reveal an adjacent audience you did not know you had. Search traffic may show people finding an article for a reason you never anticipated. A series may attract platform engineers when you thought you were writing for application developers.&lt;/p&gt;

&lt;p&gt;That is useful information. It is not an order.&lt;/p&gt;

&lt;p&gt;A writer can discover that beginner tutorials have an enormous reachable audience and still decide not to build a body of work around them. A company can discover that a group of users loves a product for an unexpected use case and still decide that market does not fit the strategy.&lt;/p&gt;

&lt;p&gt;Discovering an audience is not the same as deciding to serve it.&lt;/p&gt;

&lt;p&gt;This is where "write more of whatever did best last week" collapses. It is the same loop with one step deleted.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A["Write"] --&amp;gt; B["Publish"]
    B --&amp;gt; C["Signals&amp;lt;br/&amp;gt;views · comments&amp;lt;br/&amp;gt;search terms · who shows up"]
    C --&amp;gt; D{"Is this where&amp;lt;br/&amp;gt;I want to go?"}
    D --&amp;gt;|"yes"| E["Build an audience here"]
    D --&amp;gt;|"no"| F["Note it. Leave it alone."]
    E --&amp;gt; G["Strategy"]
    F --&amp;gt; G
    G --&amp;gt; A

    classDef step fill:#E8F3EE,stroke:#166534,color:#14532D,stroke-width:2px;
    classDef signal fill:#FFFFFF,stroke:#5CA08A,color:#14532D,stroke-width:2px;
    classDef gate fill:#FEF2F2,stroke:#991B1B,color:#7F1D1D,stroke-width:2px;

    class A,B,E,F,G step;
    class C signal;
    class D gate;&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The diamond is the part that gets skipped. Without it the loop still runs, it just runs on autopilot, and the writer ends up somewhere chosen by whatever the feed rewarded in a given week.&lt;/p&gt;

&lt;p&gt;Metrics tell you what happened. Readers reveal opportunities you did not know to look for. Neither one gets to decide what you want the work to become. The feedback loop needs interpretation.&lt;/p&gt;

&lt;h2&gt;
  
  
  I eventually wrote down some rules
&lt;/h2&gt;

&lt;p&gt;Once I could see a body of work forming, it became tempting to turn every passing thought into another strategic article. So I wrote myself a gate.&lt;/p&gt;

&lt;p&gt;For the deliberate part of my technical writing, I now ask whether there is a disputable claim, whether investigating it will put pressure on an actual artifact, whether I have standing through a project or experiment, whether it advances rather than repeats the larger body of work, and whether a reader should do or question something differently afterward. I also want to know what would falsify the claim before I start assembling evidence for it.&lt;/p&gt;

&lt;p&gt;The artifact-pressure test is deliberately hard. If investigating an idea will not change code, a specification, an ADR, a schema, or a demo, it probably does not belong in that stream. The investigation also has to be capable of failing. Building fixtures that encode what I already believe proves very little.&lt;/p&gt;

&lt;p&gt;That produces a shape I have become fond of:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Here is what I thought.&lt;br&gt;
Here is what would have convinced me I was wrong.&lt;br&gt;
Here is what I built.&lt;br&gt;
Here is what happened.&lt;br&gt;
Here is what changed.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those are not my rules for everything. I deliberately keep another lane with almost no gate at all: career observations, language experiments, satire, community responses, project archaeology, and pure curiosity only need to be worth writing.&lt;/p&gt;

&lt;p&gt;A publishing strategy should help me recognise strong work. It should not make me ask permission before being curious.&lt;/p&gt;

&lt;h2&gt;
  
  
  Some of it is luck
&lt;/h2&gt;

&lt;p&gt;There is one variable I cannot put into a strategy with any confidence.&lt;/p&gt;

&lt;p&gt;I can study years of data and conclude that Tuesday at 9:00 a.m. Pacific is the right time to publish. That says nothing about whether it is right for &lt;em&gt;this&lt;/em&gt; article. A major news event may take the morning. Three other posts aimed at the same readers may appear within the hour. A moderator may promote something. A Gem may land. Another writer with a large audience may link to you. None of that says anything new about the quality of the work, and each can transform its distribution.&lt;/p&gt;

&lt;p&gt;Strategy does not eliminate luck. It changes the conditions under which luck operates. You can improve the writing, understand the audience, participate in the community, publish consistently enough that people know you exist, and make the work easy to find.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;You can engineer more opportunities for an outcome without engineering the outcome itself.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And when luck does hand you an unexpected success, strategy comes back. A surprising audience appearing is another signal, not a mandate. You still have to decide whether it points somewhere you want to go.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I actually track now
&lt;/h2&gt;

&lt;p&gt;I still look at views. Pretending otherwise would be silly. If one article gets 5,000 and another gets 40, I want to know why. I just no longer think the first was 125 times more successful.&lt;/p&gt;

&lt;p&gt;The outcomes I care most about do not fit in a platform dashboard. A reader challenges a claim and I change the model. A comment exposes a missing invariant. An experiment breaks the answer I expected to publish. An article changes code, a specification, or an architecture decision.&lt;/p&gt;

&lt;p&gt;That has become much less theoretical lately.&lt;/p&gt;

&lt;p&gt;I published an article that began with a refractometer and a batch of homemade wine, arguing about the difference between a measurement and the state we infer from it. The comments pushed it considerably further: version the correction rule, spend more measurement budget on consequential baselines, decide how conflicting instruments get adjudicated before seeing the readings, distinguish evidence that survives a restart from evidence that dies with the process.&lt;/p&gt;

&lt;p&gt;An article about authority boundaries in AI-generated code did the same thing. Readers pushed on capability lifetime, consumable authority, semantic authority diffs, denied-call telemetry, and who is permitted to modify the authority boundary itself.&lt;/p&gt;

&lt;p&gt;Those comments did not just increase engagement. They changed the model. Those are propagation effects rather than distribution metrics, and I have started tracking them separately: research-induced change, engagement from people with standing outside my field, independent use of a concept, substantive challenges and extensions, and whether writing has started consuming so much time that the projects supplying it have stopped moving.&lt;/p&gt;

&lt;p&gt;My numbers support the split more cleanly than I expected. Across nine years I have 1,372 reactions and 683 comments. Roughly one comment for every two reactions is a strange ratio, and it is concentrated almost entirely in recent work. My 2017 tutorials pulled more than twice the traffic of everything I wrote in 2026 and produced 26 comments in three years. The 2026 essays, with a third of the traffic, have produced hundreds. One of them has a comment thread 35 replies deep.&lt;/p&gt;

&lt;p&gt;One body of work got found and skimmed. The other gets read and argued with. For years I evaluated both with the same number.&lt;/p&gt;

&lt;p&gt;A related consequence: a technically unsuccessful investigation can make a successful article. If I start with a hypothesis, state what would falsify it, build something capable of producing an answer I do not control, and find out my model was wrong, that is useful. Possibly more useful. It tells the reader something, it changes the artifact, and it often exposes a better question than the one I started with.&lt;/p&gt;

&lt;p&gt;I have a line in my current content plan that says missing a publishing day beats manufacturing a weak post. Nine years ago I am not sure I would have been comfortable with that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Nine years later
&lt;/h2&gt;

&lt;p&gt;I am still working this out. I still publish things and wonder whether anyone will care. I still occasionally write something I expect to perform well and watch it vanish. I still publish something on a whim and find it landed exactly where it needed to.&lt;/p&gt;

&lt;p&gt;The difference is that I now have more ways to recognise success when it shows up wearing something other than a large number.&lt;/p&gt;

&lt;p&gt;A successful post might reach 30,000 people. It might produce a conversation that changes the next article. It might expose a flaw in an architecture. It might give someone language for a problem they were already having. It might lead to code. It might turn out to be part of a body of work whose shape I could not see when I wrote the first piece.&lt;/p&gt;

&lt;p&gt;And sometimes it might simply be an article I wanted to write, written well enough that I am still happy to have my name on it years later.&lt;/p&gt;

&lt;p&gt;Views measure views. They are a real measurement of a real thing, and they do not independently measure rigour, usefulness, influence, audience fit, changed behaviour, changed artifacts, or whether the writing moved me toward a better question. Sometimes those correlate. Sometimes they do not. Mine told me almost nothing for nine years, and the twenty-one thousand followers told me less.&lt;/p&gt;

&lt;p&gt;The useful question is not &lt;strong&gt;"Was this post successful?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It is &lt;strong&gt;"What did I want this post to do, and what happened because I published it?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Those are much harder numbers to put on a dashboard. I think they are also the ones worth learning to notice.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Thanks to &lt;a href="https://dev.to/sylwia-lask"&gt;Sylwia Laskowska&lt;/a&gt; for the conversation that prompted this, and for the encouragement to write some of it down.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>career</category>
      <category>writing</category>
      <category>devjournal</category>
      <category>meta</category>
    </item>
    <item>
      <title>The Most Useful Line on Your AI Cost Report Is the One You Can't Explain</title>
      <dc:creator>Ken W Alger</dc:creator>
      <pubDate>Thu, 01 Oct 2026 13:16:00 +0000</pubDate>
      <link>https://dev.to/kenwalger/the-most-useful-line-on-your-ai-cost-report-is-the-one-you-cant-explain-195f</link>
      <guid>https://dev.to/kenwalger/the-most-useful-line-on-your-ai-cost-report-is-the-one-you-cant-explain-195f</guid>
      <description>&lt;p&gt;&lt;em&gt;Attribution, allocation, and why "unknown" belongs in the schema.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This piece grew out of a comment thread on Sarvar Nadaf's &lt;a href="https://dev.to/sarvar_04/per-agent-cost-tracking-for-multi-agent-ai-on-aws-10eg"&gt;Per-Agent Cost Tracking for Multi-Agent AI on AWS&lt;/a&gt;. The schema below was worked out in that conversation, in public, and it is better for it. Where a specific idea came from the exchange, I have tried to say so.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;Most AI cost dashboards answer one question well: how much did this run cost. Total tokens, model spend, per-agent spend, latency, tool usage. Those are real and useful numbers, and for a while they are enough.&lt;/p&gt;

&lt;p&gt;They stop being enough the moment your system becomes a composition. Once a request flows through a retriever, a knowledge graph, three specialist agents, and a supervisor that synthesizes their output, the total tells you almost nothing about what to change. A run can be correct, return HTTP 200, look healthy in every latency-and-errors panel, and still cost forty percent more than an identical run that produced the same answer. The overspend is real. It is just not anywhere you are looking.&lt;/p&gt;

&lt;p&gt;To find it, you have to stop asking where the money was spent and start asking what caused it to be spent. Those are different questions, and the gap between them is the whole subject of this piece.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where a Cost Is Incurred Is Not What Caused It
&lt;/h2&gt;

&lt;p&gt;Consider a retrieval operation that costs $0.004: searching, ranking, fetching. That is the direct cost, and it is easy to attribute. It happened on that span, you can measure it, done.&lt;/p&gt;

&lt;p&gt;Now suppose that retrieval returned 20,000 tokens, and all of them were hydrated into a supervisor's context on the next step. The supervisor then costs $0.009. How much did the retrieval really cost?&lt;/p&gt;

&lt;p&gt;The direct answer is still $0.004. But that is no longer the interesting answer, because the retrieval also caused cost somewhere else. It inflated the supervisor's context, and some portion of that $0.009 exists only because the retrieval handed it too much material. The cost was incurred at the supervisor. It was caused, in part, at the retriever.&lt;/p&gt;

&lt;p&gt;This gives you two distinct dimensions, and a useful cost model has to carry both:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Where the cost was incurred. This is just the span. Directly observed, low ambiguity.&lt;/li&gt;
&lt;li&gt;What caused or contributed to it. This is the interesting axis, and it is the one no aggregate dashboard shows.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A record that captures both might look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;span_id: retrieval-104
direct_cost: $0.0040
primary: retrieval
returned_tokens: 20000

downstream:
  span_id: supervisor-105
  attributed_cost: $0.0021
  caused_by: retrieval-104
  attribution_method: proportional
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The dollars are single-counted. We do not charge the $0.0021 twice. The supervisor genuinely incurred it; the retrieval genuinely contributed to causing it; and the record says both without inventing money. What we have added is lineage: a link from a downstream cost back to the decision that helped produce it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Hard Part Is Honesty About How You Know
&lt;/h2&gt;

&lt;p&gt;Here is the question that breaks naive versions of this: how do we know retrieval-104 actually caused $0.0021 of the supervisor's cost, and not some other amount?&lt;/p&gt;

&lt;p&gt;Sometimes you can measure it. If you have a controlled comparison where the only meaningful change is that retrieval result, the delta is real evidence. Supervisor costs $0.006 without the retrieved material and $0.009 with it, so roughly $0.003 of downstream cost is attributable to that retrieval. That is measured causation, and it is the strongest claim you can make.&lt;/p&gt;

&lt;p&gt;Most production traces do not give you that. In a real run the supervisor is carrying system instructions, conversation state, the outputs of other agents, tool results, and the retrieved material, all at once. There is no clean counterfactual. So you fall back on allocation: split the supervisor's context cost proportionally by the tokens each source contributed. That is a reasonable method. It is not measurement, and the receipt must not pretend it is.&lt;/p&gt;

&lt;p&gt;This is why the single most important field in the whole schema is not a dollar amount. It is this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;attribution_method:
  - measured_delta
  - proportional
  - estimated
  - unknown
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That field is what keeps the entire model honest. It stops a proportional guess from masquerading as measured causation. With it, a line can say "retrieval span 104 contributed an estimated $0.0021 of downstream context cost, allocated proportionally by hydrated token share," and every word in that sentence is defensible, because the method is stated. Without it, the same $0.0021 acquires a precision the evidence never earned.&lt;/p&gt;

&lt;p&gt;Resist collapsing this into a confidence score. A number like &lt;code&gt;confidence: 0.82&lt;/code&gt; feels rigorous and gives you nothing, because now you have a second number whose provenance you have to go investigate. &lt;code&gt;measured_delta&lt;/code&gt;, &lt;code&gt;proportional&lt;/code&gt;, &lt;code&gt;estimated&lt;/code&gt;, and &lt;code&gt;unknown&lt;/code&gt; each tell you why you are entitled to believe the figure. The method is the provenance. A score would hide it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Show the Method Where the Decision Is Made
&lt;/h2&gt;

&lt;p&gt;A natural instinct is to keep the attribution method as drill-down metadata, out of the main view, so the report stays clean. That instinct is wrong, and it is wrong for the same reason aggregate dashboards are wrong: it makes two different claims look equivalent.&lt;/p&gt;

&lt;p&gt;The method belongs inline, next to any attributed cost, with one sensible exception. A directly observed cost carries no ambiguity and needs no method tag:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;retrieval-104   RETRIEVAL   $0.0040
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is nothing to disclose there; it was measured on the span. But the moment a number is attributed rather than observed, the method has to ride along:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;retrieval-104 -&amp;gt; downstream CONTEXT   $0.0021   proportional
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Drop the word &lt;code&gt;proportional&lt;/code&gt; and that $0.0021 visually becomes as solid as the $0.0040 above it, which is a lie of formatting. A report that hides the distinction between what it measured and what it allocated has committed the same sin as the dashboard that only shows a total. If two numbers make materially different claims, the interface must not make them look the same.&lt;/p&gt;

&lt;p&gt;So the main report shows amount, category, direct versus downstream, and method. The drill-down holds the evidence behind the method: hydrated token counts, comparison runs, parent-child span references, the assumptions the allocation rests on. The decision surface stays readable; the receipt is one click away, not dumped into the table.&lt;/p&gt;

&lt;h2&gt;
  
  
  "Unknown" Is Not a Gap in the Accounting
&lt;/h2&gt;

&lt;p&gt;Every honest version of this schema has to allow &lt;code&gt;unknown&lt;/code&gt; as a real value, not a placeholder you feel bad about. And once you sit with it, the &lt;code&gt;unknown&lt;/code&gt; rows turn out to be the most useful rows in the report.&lt;/p&gt;

&lt;p&gt;A high downstream cost with &lt;code&gt;unknown&lt;/code&gt; attribution is not incomplete bookkeeping. It is the system telling you exactly where your observability boundary stops letting you explain its own behavior. It is pointing at the place where you cannot yet answer "what caused this," which is the place most worth instrumenting next. A tidy report with no &lt;code&gt;unknown&lt;/code&gt; rows has usually not achieved understanding. It has hidden its ignorance behind confident allocation.&lt;/p&gt;

&lt;p&gt;There is also a real reason &lt;code&gt;unknown&lt;/code&gt; is sometimes the only honest answer: context is not additive. An extra 5,000 tokens of context does not simply add a proportional slice of cost. It can change caching behavior, alter the reasoning path the model takes, or change how much output the model generates downstream. When that happens, token share and cost share stop mapping to each other cleanly, and any proportional number you report is a polite fiction. In those cases the schema should say &lt;code&gt;unknown&lt;/code&gt; and mean it, rather than allocate a figure it cannot defend.&lt;/p&gt;

&lt;p&gt;Read that way, the report stops being a statement of where you paid and becomes a map of two things at once: what you can explain about your spending, and where your ability to explain it runs out. The second map is the one that tells you what to build.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Actually Costs to Build
&lt;/h2&gt;

&lt;p&gt;The reason this is not a research project is that most of the structure already exists. If your traces are parent-child spans, the causal lineage is physically present already; a downstream cost sits under the decision that produced it. You are not inventing a new tracing mechanism. You are making the attribution semantics explicit on top of a trace you already record.&lt;/p&gt;

&lt;p&gt;Concretely, that is a small number of additions. Stamp a primary category on each span. Add a &lt;code&gt;contributes_to&lt;/code&gt; link and a &lt;code&gt;cause&lt;/code&gt; on spans that produce downstream effects. Add the &lt;code&gt;attribution_method&lt;/code&gt; on any attributed cost. Then roll the report up along two axes, category and direct-versus-downstream, and let &lt;code&gt;unknown&lt;/code&gt; be a first-class row rather than a swept-under one. The intelligence is not in the plumbing. It is in having the honest method field and being willing to publish the &lt;code&gt;unknown&lt;/code&gt; rows.&lt;/p&gt;

&lt;p&gt;One warning from the same conversation that produced all this: how you record and how you attribute are coupled. Change the way spans are emitted and you can silently break the logic that reads them. The defense is the same one that makes the whole model trustworthy, a known-good baseline you compare against, so that when your instrumentation shifts under you, the numbers move and you notice.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Point
&lt;/h2&gt;

&lt;p&gt;We have spent a lot of effort making AI spend visible. Better token counts, per-agent breakdowns, nested traces. All of it answers "how much." Almost none of it answers "why," and "why" is the only version of the question you can act on.&lt;/p&gt;

&lt;p&gt;The move from one to the other is not a bigger dashboard. It is a small, honest schema: separate where a cost was incurred from what caused it, state the method behind every attributed number, and treat the costs you cannot explain as signal rather than embarrassment. Do that, and the report stops telling you what you spent and starts telling you what to fix, including, in the &lt;code&gt;unknown&lt;/code&gt; rows, where to look first.&lt;/p&gt;

&lt;p&gt;Even the cost report, it turns out, needs provenance. Not just the amount and the category, but how sure you are about who to blame. That last column may be the most useful one on the page.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;With thanks to Sarvar Nadaf, whose post and the conversation under it produced this schema, and to the commenters in that thread who pushed on the baseline and the propagation. The receipt is better for the argument.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>observability</category>
      <category>agents</category>
      <category>provenance</category>
    </item>
    <item>
      <title>A wine cellar that remembers what you used to believe</title>
      <dc:creator>Ken W Alger</dc:creator>
      <pubDate>Wed, 30 Sep 2026 22:49:21 +0000</pubDate>
      <link>https://dev.to/kenwalger/a-wine-cellar-that-remembers-what-you-used-to-believe-4i55</link>
      <guid>https://dev.to/kenwalger/a-wine-cellar-that-remembers-what-you-used-to-believe-4i55</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/sanity-2026-09-16"&gt;Sanity Challenge, Path Two: Vibe-Code Something Strange&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Cellar&lt;/strong&gt; is a wine collection that remembers what you used to believe.&lt;/p&gt;

&lt;p&gt;Most cellar apps answer one question: what do I have right now? Cellar answers a harder one. What did my cellar look like on 1 June 1999? What did I think that bottle's drinking window was in 2018, before a later tasting changed my mind? Which bottles were at peak last year while I wasn't paying attention, and are now past saving?&lt;/p&gt;

&lt;p&gt;Those questions rule out the obvious data model. A &lt;code&gt;status&lt;/code&gt; field on a bottle tells you what is true today and nothing about 2004. A &lt;code&gt;drinkUntil&lt;/code&gt; field on a wine tells you what someone currently believes and destroys what they believed before.&lt;/p&gt;

&lt;p&gt;So Cellar stores neither. It stores events and claims:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Acquisitions&lt;/strong&gt; and &lt;strong&gt;consumptions&lt;/strong&gt; are documents with dates, not fields on a bottle. A bottle is consumed as of date T if a consumption event exists on or before T.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Assessments&lt;/strong&gt; are dated, attributed claims about when a wine should be drunk. Nothing overwrites a window. The current window is resolved from the claims that existed at T, by explicit rules: my own tasting note outranks the producer's, which outranks a critic's, and within a tier the most recent claim wins.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything else follows. Move a date control backward and acquisitions disappear, consumptions reverse, and drinking windows change as the controlling claim changes. Move it forward and unopened bottles age through their windows.&lt;/p&gt;

&lt;p&gt;The design rule underneath: &lt;strong&gt;do not store what is true now when you may later need to know what was true then.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There are three surfaces, with different jobs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Sanity Studio&lt;/strong&gt; authors and governs the ledger. Six document types, a custom review queue, two document actions, and a read-only workflow field.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A Sanity App&lt;/strong&gt;, built on the App SDK and running inside Sanity, interprets it. Three views, all computed live from the event log by a pure TypeScript package that knows nothing about Sanity.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A public page&lt;/strong&gt; is that same App without an account. Read-only by construction rather than by convention, since there is no token in the bundle and the dataset's public ACL grants reads only.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What it deliberately does not have is a second CRUD interface. Studio already is one, so adding bottles and recording consumptions happens there. The App interprets; it does not manage.&lt;/p&gt;

&lt;p&gt;Who it is for, honestly: me. But the argument generalizes past wine, which is the reason I built it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Open it with no account: &lt;a href="https://kenwalger.github.io/Cellar/" rel="noopener noreferrer"&gt;kenwalger.github.io/Cellar&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Inside Sanity&lt;/strong&gt; (organization members): &lt;a href="https://www.sanity.io/@opyntsvcl/application/gcxu5htdwn5n9lc15m64sfpp" rel="noopener noreferrer"&gt;The Cellar app&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/bb-9b_9k7WA" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;The same 542-bottle ledger, read at three moments:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;In cellar&lt;/th&gt;
&lt;th&gt;Drinking&lt;/th&gt;
&lt;th&gt;Past window&lt;/th&gt;
&lt;th&gt;Not yet acquired&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1 June 1999&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;536&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;23 September 2026&lt;/td&gt;
&lt;td&gt;248&lt;/td&gt;
&lt;td&gt;166&lt;/td&gt;
&lt;td&gt;33&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;23 July 2035&lt;/td&gt;
&lt;td&gt;248&lt;/td&gt;
&lt;td&gt;62&lt;/td&gt;
&lt;td&gt;182&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;In 1999 the whole cellar is four bottles of 1993 Château Mouton Rothschild, with two already opened. By 2026 it is 248 bottles. Projected to 2035, those same 248 bottles are still there, and 182 of them have gone past window.&lt;/p&gt;

&lt;p&gt;The cellar never loses bottles. It loses opportunities.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Past:&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F02psjha2oavwao358zcz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F02psjha2oavwao358zcz.png" alt="The Cellar at 1 June 1999: 4 bottles in the cellar, 536 not yet acquired" width="634" height="844"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Present:&lt;/strong&gt; &lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsj80i2zypphhazeu8vup.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsj80i2zypphhazeu8vup.png" alt="The Cellar today: 248 bottles in the cellar, 166 drinking, 33 past window" width="633" height="837"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Future:&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhaqgcue8e7wehcyduo4o.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhaqgcue8e7wehcyduo4o.png" alt="Projected to 23 July 2035: still 248 bottles, 62 drinking, 182 past window, with a note that the projection assumes no further acquisitions, consumptions or assessments" width="624" height="856"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Drink Soon lists what closes within twelve months, and every row names the claim that decided it: which tier, whose, and when.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F36jgg3jcsna47d6060cl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F36jgg3jcsna47d6060cl.png" alt="Drink Soon at today's date: 23 bottles across 12 wines, each row showing the source tier, the source name, the assessment date and the number of accepted claims" width="800" height="1034"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Missed Opportunities asks a different question: which bottles were at peak during a period, were never opened, and are past window now.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgzfi2tpk1ql2w64egmfm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgzfi2tpk1ql2w64egmfm.png" alt="Missed Opportunities for 2023: 17 bottles across 9 wines, of which 8 did get opened in time" width="800" height="1053"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;And when nothing was lost, it says why rather than showing an empty panel.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj732v0wsr80r1mz43f1a.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj732v0wsr80r1mz43f1a.png" alt="The empty state for 2019: nothing was lost in this period, with 47 bottles at peak, 32 since opened, and 7 still in hand and not past window" width="800" height="223"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Three views in the App, all computed live from the event log:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cellar health&lt;/strong&gt;, bottle counts by state at any date&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Drink Soon&lt;/strong&gt;, windows closing within twelve months, each row naming which claim decided it&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Missed Opportunities&lt;/strong&gt;, bottles that were at peak during a period, were never opened, and are past window today&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And in Studio, the &lt;strong&gt;review queue&lt;/strong&gt;, where a proposed claim waits for a person.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/kenwalger/Cellar" rel="noopener noreferrer"&gt;github.com/kenwalger/Cellar&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  My Build Process
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Tools:&lt;/strong&gt; Claude Code in a terminal beside WebStorm, with Sanity's MCP server connected so the model could check current documentation rather than recall it. One thing worth stating plainly: the prompts themselves were drafted in a separate planning conversation with Claude in the chat app, where I also made the decisions at each fork before relaying them. Calling this one developer and one agent would be inaccurate.&lt;/p&gt;

&lt;p&gt;The project entered implementation with fourteen planning documents: a content model, a temporal resolution specification, a build plan, a seed data plan, and eleven architecture decision records. None of it was code. That turned out to matter, though not in the way I expected.&lt;/p&gt;

&lt;h3&gt;
  
  
  The pattern that emerged
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;specification
     ↓
agent conflict pass, no code written
     ↓
human decision
     ↓
implementation
     ↓
independent oracle, measurement, or experiment
     ↓
gate only I can close
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every stage started with a prompt that said, roughly: read the specs, check them against the current platform, report where they disagree, and stop. Every stage found something. The count went &lt;strong&gt;up&lt;/strong&gt; as the project went on, not down. Stage 1 found three spec errors. Stage 4 opened with ten conflicts, two of them spec errors.&lt;/p&gt;

&lt;h3&gt;
  
  
  Prompts that worked
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;"Report conflicts and stop before writing code."&lt;/strong&gt; The single highest-value instruction in the project. In Stage 1 it found that I had specified underscore-prefixed projection fields, which Sanity reserves for system fields, and a validation rule requiring an assessment to have been &lt;em&gt;created&lt;/em&gt; in a particular state, which Sanity cannot express because validation sees current state only. Implementing that rule literally would have made the accept transition permanently invalid, breaking the exact workflow the rule existed to protect.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Gates that name what does not count as proof.&lt;/strong&gt; For the App SDK gate I wrote out the path being tested and then listed the ways it could be faked:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Do not satisfy this gate by calling &lt;code&gt;@cellar/core&lt;/code&gt; separately from the App, querying the dataset through another tool, or relying only on the existing test suite. The purpose is to prove the complete path: Content Lake → App SDK → CELLAR_QUERY → toCellarSnapshot → bottleState() → rendered view.&lt;/p&gt;

&lt;p&gt;You cannot see the rendered view, so do not report the gate as passed. Give me the command, the URL, and exactly what I should see.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That last paragraph matters more than the first. Claude Code has no browser, and a model asked to confirm something it cannot observe is under pressure to confirm it anyway.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"State the expected result before deploying."&lt;/strong&gt; For the Functions experiment, the probe had to produce one predetermined log line, written down and approved before anything was deployed. A bundler can resolve an import and still ship a shimmed or stale module; a log line proves nothing on its own.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Oracles, with an explicit instruction not to repair them.&lt;/strong&gt; Expected outputs were generated before implementation and the prompt said to treat them as an oracle rather than as output to be fixed. When a test disagrees, the cheapest move for a model is to adjust the fixture.&lt;/p&gt;

&lt;h3&gt;
  
  
  Prompts that didn't, and where I was wrong
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;I proposed a test that would have written to production.&lt;/strong&gt; Setting up the Functions probe, I suggested scoping the trigger at the production dataset with a filter on a document type that doesn't exist, to avoid creating a scratch dataset. Claude Code declined, correctly: scoping is safe, invoking is not. A document Function can only be invoked by writing a matching document, so my plan would have put two mutations permanently in production's history to avoid a dataset that deletes whole.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I told it to give up on a test I assumed was impossible.&lt;/strong&gt; After finding a test that passed while measuring nothing, I said that if the clause couldn't be isolated, it should fall back to a weaker assertion. It ran the experiment instead and found a fixture, on the realization that &lt;code&gt;now&lt;/code&gt; in this system is a parameter rather than the clock. "As of June 2024, what had I missed that year?" is an ordinary question for a time machine, and it makes the case constructible.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I approved a feature without asking what it cost.&lt;/strong&gt; An extra column showing "3 of 5 at peak, 1 opened in time" came back with a measurement I hadn't requested: it costs about four times what the rest of the view does, because it scans every other bottle of every listed wine. Still 5% of a frame, so nothing changed, but the earlier estimate had been for something narrower than what shipped and it said so.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where the model got stuck
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;It confidently misread an API.&lt;/strong&gt; It reported that Sanity previews cannot follow references and deferred a feature to a later stage. Asked to verify against the docs rather than assert, it found the section demonstrating exactly that, said plainly it had been wrong, and fixed it in one line.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It nearly filed a bug that didn't exist.&lt;/strong&gt; It concluded that installed type definitions contradicted the reference docs about where &lt;code&gt;perspective&lt;/code&gt; lives. They don't: the option is inherited through an interface chain that the reference flattens and the &lt;code&gt;.d.ts&lt;/code&gt; doesn't. It caught this itself and logged the near-miss instead of the finding. That's the more useful record, because a developer who files bugs that turn out to be their own misreading loses credibility fast.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Right conclusion, wrong reasoning.&lt;/strong&gt; It argued for inclusive date comparison on the grounds that strict comparison would reject ordinary windows in the seed data. It wouldn't, because same-year windows normalize to 1 January and 31 December, which are strictly ordered. The conclusion was right anyway. A model that is right for the wrong reason is right only until the case changes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It asserted a gitignore rule instead of checking it.&lt;/strong&gt; A stage list said the public build's output directory was already ignored. It wasn't: a bare &lt;code&gt;dist&lt;/code&gt; pattern matches a directory named &lt;code&gt;dist&lt;/code&gt;, and this one is &lt;code&gt;dist-public&lt;/code&gt;. Six build files were committed, in the same commit whose accompanying document argued that build output does not belong in a repository. The cause is neat: &lt;code&gt;dist-public&lt;/code&gt; was named that way specifically so it could not collide with &lt;code&gt;dist/&lt;/code&gt;, and a name chosen to be distinct from &lt;code&gt;dist&lt;/code&gt; is by construction one that &lt;code&gt;dist&lt;/code&gt; does not match.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It broke an explicit rule twice, and disclosed both times.&lt;/strong&gt; I told it not to run git commands that change the repository. Twice it ran &lt;code&gt;git checkout --&lt;/code&gt; to undo its own uncommitted edits, and twice it reported this unprompted before I found it. The rule now says explicitly that the file tools are the only way to revert anything, including its own work.&lt;/p&gt;

&lt;h3&gt;
  
  
  How I course-corrected
&lt;/h3&gt;

&lt;p&gt;Mostly by making verification harder to fake rather than by writing better instructions. The gates got stricter as the project went on: name the path, exclude the proxies, state the expected numbers in advance, and put the confirmation in the hands of the one participant with eyes.&lt;/p&gt;

&lt;p&gt;The other correction was measurement. Three of the four real bugs in this project were invisible to every check designed for the feature they lived in.&lt;/p&gt;

&lt;h3&gt;
  
  
  The App SDK
&lt;/h3&gt;

&lt;p&gt;The gate was one view rendering live production data inside Sanity, with &lt;code&gt;now&lt;/code&gt; frozen to a date an independent script had already computed. The numbers had to match exactly: hold 45, drinking 166, past window 33, unassessed 4, consumed 294, 542 total. They did, first try, and I confirmed them in a browser because the model couldn't.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;useQuery&lt;/code&gt; subscribes, so publishing an assessment in Studio moves the counts in the App with no synchronization layer. Studio edits the graph, the App reads the graph, GROQ traverses it, Functions react to it, and a pure module decides what it all means.&lt;/p&gt;

&lt;p&gt;Learning the current App SDK was harder than using it. In one session: Sanity's own bundled agent rule supplied a CLI command with a flag the installed CLI rejects; the quickstart named a template the CLI reference doesn't list; and that template installed the SDK a major version behind the docs I was being pointed at. All three are stale vendor guidance rather than product defects, which is a distinction worth making and also a pattern worth reporting.&lt;/p&gt;

&lt;p&gt;The deeper issue was a recurring shape. The documented path is well covered for the common case of rendering and editing documents. Cellar computes an aggregate, and the guidance steered toward &lt;code&gt;useDocuments&lt;/code&gt; and &lt;code&gt;useDocumentProjection&lt;/code&gt;, which for six numbers would mean roughly 1,500 hook instances and round trips instead of one query. Similar story with &lt;code&gt;useState&lt;/code&gt;: the rule correctly warns against holding Content Lake values there, and says nothing about ephemeral view state like a date control. Neither is wrong. But once you step outside the example app, it becomes difficult to tell &lt;strong&gt;not recommended&lt;/strong&gt; from &lt;strong&gt;not documented&lt;/strong&gt; from &lt;strong&gt;not supported&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Workflows
&lt;/h3&gt;

&lt;p&gt;An assessment carries a review state: proposed, accepted, or rejected. Only accepted claims resolve windows. An agent proposes; a person decides; rejected claims stay in the dataset, because deleting them would destroy exactly the record this project exists to keep.&lt;/p&gt;

&lt;p&gt;Two details I'd defend:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The review state is read-only in the form.&lt;/strong&gt; A workflow you can bypass with a radio button is decoration. The two document actions are the only path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An accepted proposal records that a model extracted it.&lt;/strong&gt; Same tier and same author as a hand-written claim, because accepting it makes it yours, but with &lt;code&gt;sourceMethod: extracted&lt;/code&gt; alongside. A system arguing that claims should carry their provenance cannot then be unable to say whether a person wrote this one.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz5x46rrqewsnllsdkckm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz5x46rrqewsnllsdkckm.png" alt="A proposed assessment in the Studio: review state read-only and set to proposed, source method extracted, a reference back to the consumption it came from, and a notes field quoting the phrase the window was drawn from" width="800" height="1364"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The agent half.&lt;/strong&gt; A document action on a consumption sends the tasting note, the date, and the wine to Sanity's Prompt API and writes back a proposed assessment. It abstains when the note carries no temporal signal, and abstention writes no document at all rather than a proposal with empty bounds, because a proposal nobody can accept is noise in a review queue.&lt;/p&gt;

&lt;p&gt;I checked it against twelve real tasting notes from the dataset, asserting direction rather than exact years, since direction is checkable on a non-deterministic reply. Twelve for twelve, and the outputs were better than the pass rate suggests. Every non-abstaining reply quotes the phrase it relied on, which is what makes the claim auditable: you can see which five words produced a window. Confidence tracks how directly the note states the thing, high on "This is the peak" and low on "Fruit starting to dry out at the edges", which is a hedge in my own note. And on "Could have waited another two years" from a 2024 opening it returned 2026, which is arithmetic rather than classification.&lt;/p&gt;

&lt;p&gt;The abstentions are the part I did not expect. They name what the note actually contains: "the note only describes the occasion for opening the bottle and gives no signal about whether the wine was too young, drinking well, or past its best."&lt;/p&gt;

&lt;p&gt;One honest result. Hunting for a proposal that would visibly move the counts turned into a search, because the seed data carries an assessment on nearly every wine, so the cases where an agent's claim wins outright are the leftovers. The first candidate lost on recency within its own tier: a personal claim of mine from 2018 outranked a proposal dated 2017. &lt;strong&gt;The agent's proposals mostly did not change the cellar, because I had usually already assessed the wine myself.&lt;/strong&gt; That is the authority rule working, and editing the seed data to get better footage was considered for about a line, then rejected. It is the dataset three oracles were built against.&lt;/p&gt;

&lt;p&gt;I verified the workflow against a staging copy of production rather than in a unit test. Two claims written in, neither moving a single count while proposed. One rejected, still moving nothing, though accepting it would have moved five bottles and emptied a wine out of Missed Opportunities. One accepted, moving exactly the five bottles predicted a day earlier.&lt;/p&gt;

&lt;h3&gt;
  
  
  What I cut, and why
&lt;/h3&gt;

&lt;p&gt;One planned piece did not get built: a Function maintaining &lt;code&gt;derived&lt;/code&gt; fields on wines and bottles, so the current state of each would be stored rather than computed.&lt;/p&gt;

&lt;p&gt;The reason is that a projection is a cache, and the thing it would cache costs 0.07 milliseconds for all 542 bottles. Measured, not assumed. The App reads the event log directly and recomputes on every date change, which is under half a percent of a frame budget.&lt;/p&gt;

&lt;p&gt;The more interesting reason is that a projection cannot do the thing this application exists to do. A stored state is true as of one date, which ADR 0012 forced into the open: any clock-dependent projection has to store the date it was computed for, beside the value. Nothing precomputed answers "what was the cellar on 1 June 1999" unless you precompute every date, and that is not a cache, it is a different data structure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What would change at scale&lt;/strong&gt;, since "it is fast enough for 542 bottles" is not an architecture:&lt;/p&gt;

&lt;p&gt;The arithmetic is not what breaks first, and the numbers say so. The ledger is 211.9 KB today, about 400 bytes per bottle:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Today&lt;/th&gt;
&lt;th&gt;x100&lt;/th&gt;
&lt;th&gt;1.5M bottles&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Query payload&lt;/td&gt;
&lt;td&gt;211.9 KB&lt;/td&gt;
&lt;td&gt;20.7 MB&lt;/td&gt;
&lt;td&gt;572.6 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Full state tally&lt;/td&gt;
&lt;td&gt;0.07 ms&lt;/td&gt;
&lt;td&gt;~7 ms&lt;/td&gt;
&lt;td&gt;~200 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The fetch becomes impractical somewhere around fifty thousand bottles, two orders of magnitude before the scan starts dropping frames. A projection would not have helped with the half that breaks first. The fix is scoping the query, by producer or vintage or a search, and a scoped query already existed in the codebase before the Function was proposed.&lt;/p&gt;

&lt;p&gt;Where a Function would genuinely earn its place is a different job. A projection is not only a cache, it is also what makes state &lt;strong&gt;queryable&lt;/strong&gt;. Expressing authority-then-recency resolution in GROQ is possible, and my own spec describes it as possible and unreadable, needing a rewrite the first time a tier is added. So "show me every bottle past its window" would mean a second implementation of the load-bearing rule, in a second language, in a place the tests do not reach. That is why two of the Studio polish items were blocked on this Function, and it is a reason that applies at 542 bottles, not only at a million.&lt;/p&gt;

&lt;p&gt;I still cut it, because nothing in the demo reads those fields and the writing was not done. But "we did not need the cache" is the smaller half of the answer.&lt;/p&gt;

&lt;h3&gt;
  
  
  Getting it in front of someone without an account
&lt;/h3&gt;

&lt;p&gt;This took four prompts to settle and produced the finding I'd most want Sanity to see.&lt;/p&gt;

&lt;p&gt;A deployed App SDK app cannot be shown to anyone outside your organization. The Dashboard gates on membership, and a changelog entry lists cross-organization visibility as a bug that was &lt;em&gt;fixed&lt;/em&gt;. Hosting the app statically doesn't help either: &lt;code&gt;AuthBoundary&lt;/code&gt; redirects a logged-out visitor to a login page before a single query is issued. Asked whether any supported configuration satisfies that boundary without a user token, the model enumerated the entire typed auth surface rather than sampling it. Every route resolves to a user token. The SDK recognizes two environments, a Dashboard iframe and a Studio, and a static host is not one of them.&lt;/p&gt;

&lt;p&gt;There is exactly one typed field that would make a static page render by asserting it is a Studio. It was recorded as a near-miss rather than recommended, on three documented grounds including that it asserts something false.&lt;/p&gt;

&lt;p&gt;So the public page drops the App SDK and reads the public dataset with a plain client. That touches one file, because everything downstream takes a &lt;code&gt;Cellar&lt;/code&gt; object and knows nothing about Sanity, and the views are shared rather than forked.&lt;/p&gt;

&lt;p&gt;Then the last surprise. &lt;strong&gt;A public dataset is exempt from authentication, not from CORS.&lt;/strong&gt; An unauthenticated read from a &lt;code&gt;github.io&lt;/code&gt; origin returns 403 until the origin is allowlisted. Every anonymous read this project had done went through curl, and curl sends no &lt;code&gt;Origin&lt;/code&gt; header, so no CORS check had ever applied. Three sessions had verified the dataset was readable using a tool that structurally cannot test the thing that mattered. The failure would have been invisible until the page went live, because localhost was already allowlisted.&lt;/p&gt;

&lt;p&gt;Three distinct access-control mechanisms, then, to get one page working: dataset ACL, organization membership, and origin allowlisting. None of them is discoverable from the surface the other two are configured on.&lt;/p&gt;

&lt;h3&gt;
  
  
  The answer was in the installed code, repeatedly
&lt;/h3&gt;

&lt;p&gt;A pattern worth naming on its own, because it decided four separate questions.&lt;/p&gt;

&lt;p&gt;Two doc pages contradict each other about whether the Prompt action requires a &lt;code&gt;schemaId&lt;/code&gt;. Session 8 concluded that one of them is wrong and that reading more carefully would not reveal which. The installed type file reveals which in about a minute: &lt;code&gt;PromptRequestBase&lt;/code&gt; declares four members, &lt;code&gt;schemaId&lt;/code&gt; is not among them, and it is required four times on the other action types in the same file. The troubleshooting page is not merely wrong, it is backwards. Omitting the field is correct; including it is the type error.&lt;/p&gt;

&lt;p&gt;That finding deleted a step I was about to take. A schema deploy had sat on the plan for five days as a production write, purely as insurance against a contradiction the type file settles.&lt;/p&gt;

&lt;p&gt;The same thing happened with authentication. The HTTP reference says every endpoint requires bearer auth and says nothing about which clients satisfy it. The answer turned out to be in the AI Assist custom field actions guide, which is not a page about authentication and which demonstrates the pattern twice in complete examples with no token anywhere. The authoritative page created the difficulty; the page that resolves it does so incidentally, by example, with no sentence anywhere stating the fact. That removed the second planned production write.&lt;/p&gt;

&lt;p&gt;Add the App SDK auth enumeration and the Functions bundling probe and it is four times in two weeks that the shipped code answered a question the documentation could not. The documentation is not bad. It is describing a platform moving faster than it can.&lt;/p&gt;

&lt;h3&gt;
  
  
  Two more findings I'd hand to Sanity
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Functions bundling is documented as a project-wide choice and isn't.&lt;/strong&gt; The docs describe TypeScript Functions in pnpm workspaces and non-TypeScript npm projects. This repo is npm plus TypeScript, which is neither. Rather than read more, I deployed a throwaway probe with one predetermined expected result. It worked first try, and inspecting the bundle explained why: the CLI inlined and tree-shook the local workspace package while externalizing the registry dependency. It decides per dependency, not per project. A private workspace package never has to be installable, only resolvable at build time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A Function that does nothing gets an editor token.&lt;/strong&gt; The probe constructed no client, ran no queries and wrote nothing. Deployment still provisioned a robot token with editor role and no expiry.&lt;/p&gt;

&lt;h3&gt;
  
  
  What the AI actually contributed
&lt;/h3&gt;

&lt;p&gt;Not mainly speed. Probably faster, but that's the least interesting result.&lt;/p&gt;

&lt;p&gt;It found errors in my specifications. It introduced its own bugs and caught some of them. It used mutation testing twice, unprompted, to prove that tests fail when the code they cover is removed. It refused to fabricate data when the ledger had no column for a link it was asked to infer, and that refusal exposed that one of my own verification scripts was passing on a proxy.&lt;/p&gt;

&lt;p&gt;The conclusion I'd draw is narrower than "specs make AI-generated code trustworthy":&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Specifications make disagreement observable. Independent witnesses make that disagreement useful.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Sometimes the witness is a CSV generated before implementation. Sometimes a second algorithm, a property test over 17,167 inputs, a benchmark, or the contents of a deployed bundle. Sometimes it's me reading six numbers off a screen, because the model cannot see them.&lt;/p&gt;

&lt;p&gt;Four examples of why that matters more than green tests:&lt;/p&gt;

&lt;p&gt;A date bug in the time control got January 31 to March 1 wrong. &lt;strong&gt;All five of my verification dates passed under the buggy version.&lt;/strong&gt; The application worked, the screenshots were right, and every date I planned to demo was fine. A property test across every slider position found it.&lt;/p&gt;

&lt;p&gt;A test asserting that proposed claims change nothing passed while measuring nothing, because its fixture would have lost on recency whether it was proposed or accepted. What caught it was the companion assertion that accepting must change at least one state. Auditing the rest of the suite for that shape found two more.&lt;/p&gt;

&lt;p&gt;Then the same idea from the other direction. A CI check grepping the built bundle for leaked credentials fired on every clean build, because its pattern matched a legitimate export name from &lt;code&gt;@sanity/client&lt;/code&gt;. A check that always fails gets waved through and certifies exactly as little as a check that can never fail, and it's worse, because a false positive doesn't look broken. It looks diligent.&lt;/p&gt;

&lt;p&gt;And spec-first has its own failure mode, which I ran into late. A polish list drafted from the build plan turned out to have three items already done. A plan describes work that was scheduled, not work that is left.&lt;/p&gt;

&lt;p&gt;The finding I keep coming back to is a different shape. A normalization rule written on day one, for a reason that had nothing to do with the interface, turned every drinking window into whole years. Weeks later that rule made a horizon slider useless, because the count could only change once a year. A stage after that, it made the obvious period for Missed Opportunities structurally empty, all year, every year. No document connects those three facts, and no amount of reading would have found them. Measuring the data did, twice, in about fifteen minutes each.&lt;/p&gt;

&lt;p&gt;Specifications don't only prevent errors. They reach forward and constrain decisions nobody anticipated making.&lt;/p&gt;

&lt;h2&gt;
  
  
  Beyond the cellar
&lt;/h2&gt;

&lt;p&gt;Wine makes this a pleasantly low-stakes problem. The pattern applies anywhere "what was true then?" matters alongside "what is true now?"&lt;/p&gt;

&lt;p&gt;A museum needs to know which attribution it accepted when an object was exhibited, not only which one it holds today. An HR system needs to evaluate an action against the policy in force when it happened, not the policy as amended since.&lt;/p&gt;

&lt;p&gt;The model underneath is the same in both: events record what happened, dated claims record what was believed, provenance records where those claims came from, and resolution decides what was knowable at a given moment.&lt;/p&gt;

&lt;p&gt;It is not free. You can never answer "what is the drinking window?" without also answering "as of when?", every read costs a resolution pass instead of a field lookup, and nobody can simply set a value. Most systems should store current state, because most systems are only ever asked about now. The decision worth making deliberately is whether historical truth is a requirement or an accident.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sanity Project Details
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Project ID:&lt;/strong&gt; &lt;code&gt;aos9nze5&lt;/code&gt; · &lt;strong&gt;Dataset:&lt;/strong&gt; &lt;code&gt;production&lt;/code&gt; (public)&lt;/p&gt;

&lt;p&gt;The dataset is world-readable, so the content model can be inspected directly. Six document types: &lt;code&gt;producer&lt;/code&gt;, &lt;code&gt;wine&lt;/code&gt;, &lt;code&gt;bottle&lt;/code&gt;, &lt;code&gt;acquisition&lt;/code&gt;, &lt;code&gt;consumption&lt;/code&gt;, &lt;code&gt;assessment&lt;/code&gt;, plus a registered &lt;code&gt;varietal&lt;/code&gt; object type. 1,645 documents.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;About the data.&lt;/strong&gt; The cellar is loosely based on a real one. The producers, appellations, and club memberships are real, and some of the history is too, including the 1993 Mouton bought in 1996 and opened twice in the nineties. Everything evaluative is invented. Drinking windows, scores, critic notes, and most tasting notes exist for the demo, and critic assessments are attributed to publications that do not exist. Nothing here should be read as a factual claim about any wine, and nothing attributed to a named producer reflects anything they have actually said.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agent Session
&lt;/h2&gt;

&lt;p&gt;The curated session is here: &lt;a href="https://dev.to/agent_sessions/claude-code-session-qv3oup"&gt;Ten conflicts before a line of code&lt;/a&gt;. It is a link rather than an embed: the transcript contains 383 &lt;code&gt;@&lt;/code&gt; characters, in package names and file paths, and the editor counts each as a user mention against a limit of ten, so the post would not save.&lt;/p&gt;

&lt;p&gt;The full record is in the repository either way: fifteen session logs under &lt;a href="https://github.com/kenwalger/Cellar/blob/main/docs/friction-logs/README.md" rel="noopener noreferrer"&gt;docs/friction-logs/&lt;/a&gt;, each with the prompts and outputs verbatim, plus what was found and what it cost. If you only read one, read session 3, where ten conflicts came back before a line of code was written.&lt;/p&gt;




&lt;p&gt;Addendum: the post's own argument, pointed at the post&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/izgorodin"&gt;A reader&lt;/a&gt; noticed something after publication that I had not, and it is the best catch of the project, so it belongs here rather than in a footnote.&lt;/p&gt;

&lt;p&gt;The two questions at the top of this piece are not the same kind of question. What did my cellar look like on 1 June 1999 asks when things happened. What did I think that bottle's window was in 2018 asks when the ledger learned it. One date control answers both only while a claim starts counting on the day it is dated.&lt;/p&gt;

&lt;p&gt;In the seeded data, it always does. The import wrote all 161 assessments as accepted, with &lt;code&gt;assessedA&lt;/code&gt;t as their only date, so every date the demo can be driven to gives the same answer either way. The demo is accurate. It is accurate by accident of the import, not by design.&lt;/p&gt;

&lt;p&gt;The agent workflow is where the two come apart. A proposal is dated the day after the bottle was opened, often years ago, and acceptance happens now. On acceptance, the resolver treats the claim as having counted since that past date, so an acceptance today can change what the cellar says about 2018. The Accept dialog warns about precisely this. The resolver has no way to represent the alternative.&lt;/p&gt;

&lt;p&gt;The reason is that &lt;code&gt;reviewState&lt;/code&gt; is stored current state with no history, which is the exact shape this post spends two thousand words arguing against. A claim either counts or it does not, and nothing records when that became true. It survived twelve architecture decision records because the field reads as workflow rather than as domain data.&lt;/p&gt;

&lt;p&gt;The fix is one field: write &lt;code&gt;acceptedAt&lt;/code&gt; alongside the transition, backfill it to &lt;code&gt;assessedAt&lt;/code&gt; for the imported claims, and resolve on &lt;code&gt;assessedAt &amp;lt;= T &amp;amp;&amp;amp; acceptedAt &amp;lt;= T&lt;/code&gt;. One slider still answers both questions, because T then means the state of knowledge at T rather than what today's knowledge says about T.&lt;/p&gt;

&lt;p&gt;I am not shipping it mid-judging, and it is recorded as &lt;a href="https://github.com/kenwalger/Cellar/blob/main/docs/ADRs/0013-acceptance-time-is-not-recorded.md" rel="noopener noreferrer"&gt;ADR 0013&lt;/a&gt;. A project whose thesis is to write down what you got wrong should not stop doing that the moment it ships.&lt;/p&gt;

</description>
      <category>sanitychallenge</category>
      <category>devchallenge</category>
      <category>ai</category>
      <category>sanity</category>
    </item>
    <item>
      <title>Your Metric Is Not Your State</title>
      <dc:creator>Ken W Alger</dc:creator>
      <pubDate>Tue, 29 Sep 2026 16:13:00 +0000</pubDate>
      <link>https://dev.to/kenwalger/your-metric-is-not-your-state-2lfl</link>
      <guid>https://dev.to/kenwalger/your-metric-is-not-your-state-2lfl</guid>
      <description>&lt;p&gt;&lt;em&gt;What a refractometer taught me about observability&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.kenwalger.com/blog/ai/mcp/the-field-agent-sovereign-vineyard-mcp/?utm_source=dev_to&amp;amp;utm_medium=article&amp;amp;utm_campaign=systems_under_fermentation" rel="noopener noreferrer"&gt;In July&lt;/a&gt;, I wrote about a farmer walking a vineyard row at dawn, crushing a grape onto a refractometer prism, and typing "13.5 Brix" into a chat window. The agent on the other end didn't care how the number arrived. To the Digital Scribe, I wrote, a number is just a number.&lt;/p&gt;

&lt;p&gt;In the vineyard, that's a defensible position. Fresh grape juice is exactly what a refractometer is built to read.&lt;/p&gt;

&lt;p&gt;This fall, I've been pointing the same kind of instrument at juice that's fermenting. It turns out the agent should have cared.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three batches and a prism
&lt;/h2&gt;

&lt;p&gt;Over the last several weeks I've started three small batches: a blackberry wine that's now bulk aging, a pear wine that began fermenting on September 24, and a fireweed honey mead that got its yeast two days later. They're one-gallon batches, which matters for this story, because at that size every sample you pull is wine you don't get to drink and oxygen you've let in.&lt;/p&gt;

&lt;p&gt;So I track them with a refractometer. A few drops of liquid on a glass prism, close the cover, hold it up to the light, read a number off the scale. The number is Brix, roughly the percentage of dissolved sugar. It's fast, cheap, and costs a few drops per reading. During the blackberry's primary fermentation I took readings twice a day, every time I punched down the cap of fruit.&lt;/p&gt;

&lt;p&gt;And once fermentation starts, the refractometer is wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  The instrument isn't lying, exactly
&lt;/h2&gt;

&lt;p&gt;A refractometer measures how much light bends as it passes through a liquid. Dissolved sugar bends light, so in fresh juice the amount of bending is a good proxy for the amount of sugar. The scale on the instrument is calibrated on that assumption.&lt;/p&gt;

&lt;p&gt;Fermentation breaks the assumption. Yeast converts sugar into alcohol, and alcohol bends light too. A few days into primary, the refractometer is reporting the combined effect of the sugar that remains and the alcohol that's been produced, and presenting all of it as sugar. The reading comes out higher than the real sugar content. Taken at face value, it says the fermentation is further behind than it actually is.&lt;/p&gt;

&lt;p&gt;To get something usable, you run each reading through a correction formula. Most home winemakers use calculators built on formulas originally developed for brewing, and every one of them needs the same extra input: the original reading, taken before the yeast went in.&lt;/p&gt;

&lt;p&gt;Here's what that looked like on the blackberry:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Day&lt;/th&gt;
&lt;th&gt;Raw refractometer (Brix)&lt;/th&gt;
&lt;th&gt;Corrected value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;08/29/2026&lt;/td&gt;
&lt;td&gt;20.7&lt;/td&gt;
&lt;td&gt;20.7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;08/30/2026&lt;/td&gt;
&lt;td&gt;18.4&lt;/td&gt;
&lt;td&gt;16.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;08/30/2026&lt;/td&gt;
&lt;td&gt;17.0&lt;/td&gt;
&lt;td&gt;14.3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;08/31/2026&lt;/td&gt;
&lt;td&gt;13.2&lt;/td&gt;
&lt;td&gt;8.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;08/31/2026&lt;/td&gt;
&lt;td&gt;10.8&lt;/td&gt;
&lt;td&gt;4.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;09/01/2026&lt;/td&gt;
&lt;td&gt;8.7&lt;/td&gt;
&lt;td&gt;0.7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;09/01/2026&lt;/td&gt;
&lt;td&gt;8.3&lt;/td&gt;
&lt;td&gt;0.1&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Nothing is malfunctioning. The instrument is doing exactly what it was built to do. What changed is the system it's pointed at, and a number that was a reliable proxy in one state became a misleading one in the next.&lt;/p&gt;

&lt;h2&gt;
  
  
  Neither instrument measures sugar
&lt;/h2&gt;

&lt;p&gt;The traditional alternative is a hydrometer, a weighted glass float. The deeper it sinks, the less dense the liquid. Sugar makes liquid denser, so falling specific gravity means sugar is being consumed.&lt;/p&gt;

&lt;p&gt;Except the hydrometer doesn't measure sugar either. It measures density, and alcohol is less dense than water, so as fermentation proceeds the alcohol pulls the reading down on its own. A finished dry wine routinely reads below 1.000, lower than plain water, which would be impossible if the hydrometer were really a sugar meter.&lt;/p&gt;

&lt;p&gt;So the refractometer measures how light bends and the hydrometer measures how heavy the liquid is. Neither measures what I actually care about: how much sugar is left, whether the yeast is still working, whether this wine is done. I infer those. The instruments give me evidence.&lt;/p&gt;

&lt;p&gt;Even a steady reading is ambiguous. The traditional sign that fermentation has finished is the same reading several days running. But a fermentation that has stalled also produces the same reading several days running. Mead is notorious for this: honey has very little natural buffering, the pH drifts down as fermentation proceeds, and the yeast can slow or stop well short of dry. A flat line can mean &lt;em&gt;done&lt;/em&gt; or &lt;em&gt;stuck&lt;/em&gt;, and the instrument can't tell you which. You tell them apart with other evidence: where the line flattened, what the pH is doing, what you expected to see.&lt;/p&gt;

&lt;p&gt;That distinction between evidence and state sounds pedantic until you notice it's one we get wrong in software constantly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your dashboard has the same problem
&lt;/h2&gt;

&lt;p&gt;CPU utilization isn't health. A service at 15% CPU can be deadlocked, and one at 90% can be doing exactly what you want. An HTTP 200 isn't correctness; it tells you the server returned a response, not that the response was right. A green health check tells you the health check endpoint answered. Model confidence isn't truth; it's a number the model produced about its own output.&lt;/p&gt;

&lt;p&gt;Each of these started as a reasonable proxy under a particular set of conditions. Then the system changed, a new failure mode appeared, or someone started optimizing the number itself, and the proxy drifted away from the state it was meant to represent. The dashboard kept rendering it with exactly the same authority.&lt;/p&gt;

&lt;p&gt;That's the refractometer problem. The metric was calibrated for one state of the system and is now being read in another, and nothing on the display tells you so. And like the flat fermentation line, a steady metric can mean two opposite things: a quiet error rate might mean a healthy service, or one that stopped receiving traffic an hour ago.&lt;/p&gt;

&lt;p&gt;The failure isn't collecting the metric. It's treating the metric as the state instead of as evidence about the state.&lt;/p&gt;

&lt;h2&gt;
  
  
  A reading needs its history
&lt;/h2&gt;

&lt;p&gt;Back to that correction formula. To interpret today's reading, you need the reading from before fermentation began. Without it, the current number is close to uninterpretable. The same Brix value could describe a fermentation that has barely started or one that's nearly finished, depending entirely on where it began.&lt;/p&gt;

&lt;p&gt;The meaning of a measurement depends on its lineage.&lt;/p&gt;

&lt;p&gt;This is where I think observability practice most often falls short. We store the value and the timestamp and call it a record. What we usually drop is everything needed to interpret it later: which instrument produced it, under what conditions, what calibration assumptions it carried, what baseline it should be read against, and what state the system was believed to be in at the time.&lt;/p&gt;

&lt;p&gt;My fermentation log keeps the raw reading and the corrected value side by side. The corrected value is what I act on. The raw value is what I re-derive from if I later learn the correction was off, find a better formula, or discover I misread the original. A corrected value on its own is a conclusion with its evidence thrown away.&lt;/p&gt;

&lt;p&gt;That's a pattern worth stealing. Store the observation separately from the interpretation. Keep the provenance that makes the observation meaningful. Let state be something you derive, with its evidence attached, rather than something you overwrite. It's the same idea behind the forensic receipt work I've been doing: a claim should carry what it was based on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Watching isn't free
&lt;/h2&gt;

&lt;p&gt;One more thing the one-gallon batches have taught me. I use a refractometer because a hydrometer needs a much larger sample, and in a small carboy that sample is a real fraction of the batch, and pulling it lets oxygen in. The act of measuring changes the system being measured.&lt;/p&gt;

&lt;p&gt;It also means deciding how often to look. The blackberry had to be punched down twice a day anyway, so twice-daily readings were free. The pear and the mead don't need that kind of handling, so I've settled on a reading every couple of days: often enough to catch a stall, rare enough to leave the batch alone. That's a measurement policy, and it means my log has deliberate gaps in it. Those gaps are their own story, and I'll come back to them.&lt;/p&gt;

&lt;p&gt;Software has the same trade-off: tracing overhead, probes that add latency, sampling that shifts timing. Usually the effect is small. Sometimes it isn't, and how often to look is a design decision with costs on both sides.&lt;/p&gt;

&lt;h2&gt;
  
  
  Evidence, then state
&lt;/h2&gt;

&lt;p&gt;Here's what I've taken from a few weeks of squinting at a prism.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Separate what you observed from what you concluded.&lt;/strong&gt; "The refractometer read 9 Brix" and "fermentation is about two-thirds done" are different kinds of statement. They belong in different places, and they should be allowed to disagree.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep the raw reading.&lt;/strong&gt; Corrections, normalizations, and aggregations are interpretations. If you keep only the interpreted value, you can never revisit the interpretation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Record the conditions that make a reading meaningful.&lt;/strong&gt; Instrument, baseline, calibration assumptions, the believed state of the system. A metric without its context is a number waiting to be misread.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Treat state as a derived claim.&lt;/strong&gt; "Fermentation is complete" isn't a reading. It's a conclusion drawn from several readings over time, plus whatever evidence rules out "stuck," and it should be traceable back to all of it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Notice when your proxy stops being a proxy.&lt;/strong&gt; Every metric was calibrated for some range of system behavior. When the system leaves that range, the metric keeps reporting with exactly the same confidence.&lt;/p&gt;

&lt;p&gt;The agent in July was right about one thing: it shouldn't matter whether a number comes from a clipboard or a five-thousand-dollar probe. Where it was wrong was in believing the number could travel without its history.&lt;/p&gt;

&lt;p&gt;The pear wine is still in primary as I write this. The raw refractometer reading says it has a long way to go. The corrected number says otherwise, and the corrected number is only trustworthy because I wrote down what the juice read before the yeast went in.&lt;/p&gt;

&lt;p&gt;The number on the instrument isn't the system. It never was.&lt;/p&gt;

</description>
      <category>observability</category>
      <category>programming</category>
      <category>ai</category>
      <category>devops</category>
    </item>
    <item>
      <title>Code Review Is Not an Authority Boundary</title>
      <dc:creator>Ken W Alger</dc:creator>
      <pubDate>Tue, 29 Sep 2026 14:28:19 +0000</pubDate>
      <link>https://dev.to/kenwalger/code-review-is-not-an-authority-boundary-3dfc</link>
      <guid>https://dev.to/kenwalger/code-review-is-not-an-authority-boundary-3dfc</guid>
      <description>&lt;p&gt;&lt;em&gt;AI can generate the implementation. Your architecture still has to decide what that implementation is allowed to do.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I was reading a &lt;a href="https://dev.to/james_anderson_h/the-7-walls-javascript-hits-and-how-webassembly-gets-past-them-3khk"&gt;good post&lt;/a&gt; about the walls JavaScript hits and how WebAssembly gets past them, and I got stuck on one of them: running code you did not write and do not trust.&lt;/p&gt;

&lt;p&gt;The framing we normally reach for is "how do I stop this extension from misbehaving?" That question already concedes something. It assumes we have handed the extension access to things worth abusing, and that our remaining defense is its good conduct.&lt;/p&gt;

&lt;p&gt;There is a different question underneath it. What is this extension actually entitled to have?&lt;/p&gt;

&lt;p&gt;I wanted to know whether that distinction survives contact with running code, so I built a small capability-mediated tool host to find out. It is about 200 lines of Python rather than a Wasm runtime, because the thing worth testing is the host interface, and every environment that mediates untrusted code has one. Everything below that reports behaviour is output I actually got, including the part where my own preferred answer turned out to be wrong.&lt;/p&gt;

&lt;h2&gt;From Trust to Authority&lt;/h2&gt;

&lt;p&gt;For most of software's history we have put enormous weight on code review. Someone writes code, someone else reads it, tests run, static analysis checks it, and eventually we decide it is trustworthy enough to merge, deploy, or execute.&lt;/p&gt;

&lt;p&gt;That model made sense when software was expensive to produce and the people writing it were known participants in the process. AI changes the economics. Code is cheap to generate now, and an agent can write a function, produce a plugin, or assemble an entire implementation in minutes. The question is less whether we can produce the code and more what that code is entitled to do once we run it.&lt;/p&gt;

&lt;p&gt;Those are different problems. Code review, however good, only solves one of them.&lt;/p&gt;

&lt;h2&gt;This Has a Name, and It Is Older Than Most of Us&lt;/h2&gt;

&lt;p&gt;None of the architecture here is my invention. What I am describing is capability-based security, and it has a long, well-developed literature.&lt;/p&gt;

&lt;p&gt;Dennis and Van Horn described capabilities in 1966. Mark Miller named the &lt;a href="http://www.erights.org/talks/no-sep/secnotsep.pdf" rel="noopener noreferrer"&gt;principle of least authority&lt;/a&gt; and built the E language around object capabilities. Capsicum brought a capability model to FreeBSD. seL4 has a machine-checked proof of its capability enforcement. Deno shipped a permission model where filesystem and network access are grants rather than defaults. &lt;a href="https://bytecodealliance.org/articles/WASI-0.2" rel="noopener noreferrer"&gt;WASI Preview 2&lt;/a&gt; and the Component Model are capability-oriented by design, which is exactly why the Wasm conversation keeps circling this.&lt;/p&gt;

&lt;p&gt;The core idea is consistent across all of them. Instead of giving code broad access and expecting it to behave, the execution environment grants specific capabilities: this directory, this socket, this operation and not that one. The implementation does not promise to stay inside its authority. It never receives the authority in the first place.&lt;/p&gt;

&lt;p&gt;So the architecture is not new. What is new is that we are about to produce vastly more code whose author we cannot interview.&lt;/p&gt;

&lt;h2&gt;Sandboxing Constrains Execution, Not Authority&lt;/h2&gt;

&lt;p&gt;A sandbox sounds reassuring. Put untrusted code in an isolated environment and prevent it from reaching the rest of the system.&lt;/p&gt;

&lt;p&gt;Suppose I have a perfectly sandboxed module. It cannot escape its runtime and cannot access anything the host does not explicitly expose. Excellent. Now suppose the host exposes:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;read_customer_records()
write_customer_records()
delete_customer_records()
send_email()
issue_refund()
export_database()
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;The sandbox is working exactly as designed, and the module is wildly over-privileged. We have constrained &lt;em&gt;where&lt;/em&gt; the code executes without constraining &lt;em&gt;what it is authorized to affect&lt;/em&gt;. The host interface has quietly become the real trust boundary.&lt;/p&gt;

&lt;p&gt;The Wasm people say as much themselves. All I/O in a module goes through its imports and exports, which means a module's view of the outside can be virtualized entirely by controlling what those imports are linked to. That is a precise description of the control surface, and also an admission that the surface is where the decisions live.&lt;/p&gt;

&lt;p&gt;That is the eighth wall hiding inside the untrusted-code one. Not "can I run code I do not trust," but "can I make what that code is allowed to do explicit, minimal, enforceable, and inspectable?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sandboxing constrains execution. Capability grants constrain authority.&lt;/strong&gt; You need both, and only one of them is commonly discussed.&lt;/p&gt;

&lt;h2&gt;Two Agents, One Test Suite&lt;/h2&gt;

&lt;p&gt;Here is the experiment. Two implementations reconcile invoices against payments and produce a report. &lt;code&gt;agent_a&lt;/code&gt; reads invoices and reads payments. &lt;code&gt;agent_b&lt;/code&gt; does the same reconciliation and also decides that a shortfall is worth refunding and that someone ought to be told.&lt;/p&gt;

&lt;p&gt;Both produce the identical report. The test suite cannot tell them apart:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;The test suite both implementations have to satisfy:

  PASS  agent_a produces the expected reconciliation report
  PASS  agent_b produces the expected reconciliation report
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Nothing in &lt;code&gt;agent_b&lt;/code&gt; is a bug. A reviewer reading it would find defensible code. The refund is small, the alert is reasonable, the logic is sound. It is simply reaching for a great deal more authority to achieve the same output, and no assertion about the output can detect that, because the difference is not in the output.&lt;/p&gt;

&lt;p&gt;Now run both against a least-privilege manifest:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;agent_a:
  completed
  ALLOW  invoices.read      read_invoices
  ALLOW  payments.read      read_payments

agent_b:
  stopped: issue_refund requires refunds.issue, which was not granted
  ALLOW  invoices.read      read_invoices
  ALLOW  payments.read      read_payments
  DENY   refunds.issue      issue_refund
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;The runtime caught it. No reviewer was involved, no test asserted anything about refunds, and that DENY line is a durable record of an authority boundary being enforced rather than a claim that somebody looked and did not see a problem.&lt;/p&gt;

&lt;p&gt;One caveat on what that manifest actually bounds, before anyone reads the demo as stronger than it is. Miller and Shapiro distinguish permission from authority: permission is what the access graph grants you directly, while authority is what you can cause, including indirectly through other components you are permitted to talk to. My host controls permission. A component granted &lt;code&gt;invoices.read&lt;/code&gt; that can also call something else holding &lt;code&gt;payments.write&lt;/code&gt; has more authority than its manifest suggests. Real capability systems address this by making the reference graph itself the thing you reason about. Two hundred lines of Python do not.&lt;/p&gt;

&lt;p&gt;That is the argument in fourteen lines of output. &lt;strong&gt;Why can an invoice-reconciliation program issue a refund?&lt;/strong&gt; Perhaps today's implementation does not. Perhaps review confirms it never calls the refund API. Those are observations about one implementation, and the implementation is the part we just made disposable.&lt;/p&gt;

&lt;h2&gt;Disposable Implementation, Durable Authority&lt;/h2&gt;

&lt;p&gt;If implementation is cheap, the code becomes increasingly disposable. We can regenerate it, replace it, or ask a different model for an alternative.&lt;/p&gt;

&lt;p&gt;Some things cannot be disposable. The contract describing correct behavior is one. The authority granted to the implementation is another. The manifest above survives every regeneration of the component beneath it, which is precisely why it is worth writing down and the code is not.&lt;/p&gt;

&lt;p&gt;This is already live in the agent tooling most of us are wiring up. An MCP server hands an agent a set of tools. That tool list is a capability grant, whether or not anyone has written it down as one, and the people building these servers keep arriving at the same requirements independently: per-tool audit records, tenant isolation, dynamic registration. Those are capability-system features reached from the practical end.&lt;/p&gt;

&lt;h2&gt;"We Tried This. It Became Permission Fatigue."&lt;/h2&gt;

&lt;p&gt;Here is the objection I would raise if someone showed me a capability manifest, and it is a strong one.&lt;/p&gt;

&lt;p&gt;We have deployed manifest-based capability declaration at planetary scale already. Android permissions. Browser extension manifests. iOS entitlements. The result was not security. It was users tapping Allow on everything, reviewers rubber-stamping manifests they did not read, and developers requesting broad permissions because narrow ones generated support tickets. Capability systems have a long history of being architecturally correct and operationally ignored.&lt;/p&gt;

&lt;p&gt;Two things are different here, and neither is optimism about human diligence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The grantor is not a person.&lt;/strong&gt; In the consumer model a human decides at install time, under time pressure, with no context and every incentive to proceed. Here a policy decides. Policies do not get fatigued, do not want the app to work right now, and can be versioned, tested, and audited.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The requester is regenerable.&lt;/strong&gt; This is the part earlier systems could not use. When a shipped application hits a too-narrow permission, the permission widens, because rewriting the app is expensive and the user is waiting. When a generated component hits a too-narrow capability, regenerating the component is cheap. The pressure that historically pushed grants wider now pushes implementations to fit the grant instead.&lt;/p&gt;

&lt;p&gt;That inversion is the actual reason this might work now, and it is a consequence of implementation becoming disposable rather than of anyone becoming more disciplined.&lt;/p&gt;

&lt;h2&gt;The Model Should Not Grant Its Own Permissions&lt;/h2&gt;

&lt;p&gt;If the same model generates the implementation, proposes its tests, decides which tools it needs, and determines its own permissions, we have built a very cooperative security model. The agent says: here is the code I wrote, here are the tests demonstrating it works, and here are the permissions I determined I require. All three may be reasonable. All three come from the same information lineage.&lt;/p&gt;

&lt;p&gt;I have made &lt;a href="https://www.kenwalger.com/blog/ai/verification-bottleneck-ai-generated-software?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=code_review_authority_boundry" rel="noopener noreferrer"&gt;a version of this argument&lt;/a&gt; before about tests, which I described then as not letting the student grade the exam. Authority is the same problem one layer over, and it is the more consequential one. A wrong test produces a wrong belief. A wrong grant produces a wrong consequence.&lt;/p&gt;

&lt;p&gt;A stronger system keeps some decisions outside that lineage. The model proposes that it needs read access to invoices. A policy determines that invoice reconciliation permits &lt;code&gt;invoices.read&lt;/code&gt;. The runtime grants that capability while leaving &lt;code&gt;payments.write&lt;/code&gt; and &lt;code&gt;refunds.issue&lt;/code&gt; unavailable. An audit record shows what authority was actually granted at execution.&lt;/p&gt;

&lt;p&gt;The model can participate without owning the boundary.&lt;/p&gt;

&lt;h2&gt;Where My Own Answer Broke&lt;/h2&gt;

&lt;p&gt;That leaves the hard part. Somebody has to write the policy, and I wanted to know whether the obvious shortcut works.&lt;/p&gt;

&lt;p&gt;The shortcut is observation. Run the component, record which capabilities it actually reaches for, and derive a least-privilege manifest from what you saw. It is the standard move, it is roughly what audit-mode tooling does across the industry, and it is what I assumed I would end up recommending.&lt;/p&gt;

&lt;p&gt;So I ran it. Discovery mode on &lt;code&gt;agent_a&lt;/code&gt; produces exactly what you would hope:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;Discovery mode: allow everything, record what was actually reached for.

  ALLOW  invoices.read      read_invoices
  ALLOW  payments.read      read_payments

Derived manifest:

component: invoice-reconciler
capabilities:
  invoices: {read: true, write: false}
  payments: {read: true, write: false}
  refunds: {issue: false}
  network: {external: false}
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Tight, minimal, derived from real behaviour rather than a guess. Then I ran the same component against an input the discovery run never saw: a credit note, which legitimately requires writing an adjusting payment record.&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;  stopped: write_payment requires payments.write, which was not granted
  ALLOW  invoices.read      read_invoices
  ALLOW  payments.read      read_payments
  DENY   payments.write     write_payment
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;The component is correct. The manifest is wrong. It forbids a legitimate path because nothing exercised that path on the day we happened to be watching.&lt;/p&gt;

&lt;p&gt;This is not a bug in my implementation, and tightening the derivation would not fix it. Observation tells you what a component did, not what it may need. A derived manifest is a lower bound on required authority presented as an upper bound on granted authority, and those are different claims. In production that failure arrives as an outage on the rare path, which is exactly how capability systems earn their reputation as obstacles and then get widened until they mean nothing.&lt;/p&gt;

&lt;p&gt;I do not have a clean answer, and I am not the first to fail to have one. Miller, Tulloh and Shapiro named this in 2004: treating security as a separate concern has not bridged the gap between principle and practice, they argue, because it operates without knowledge of what constitutes least authority, and only when requests are made can we determine how much authority is adequate. My credit note is that sentence with a stack trace attached. Twenty years on the shortcut still does not work, which suggests the diagnosis was right rather than that I picked an unlucky example.&lt;/p&gt;

&lt;p&gt;What I have is a demonstration that the answer I expected to give is wrong, which is worth more than the recommendation would have been.&lt;/p&gt;

&lt;p&gt;The partial signals that survive: a declared interface implies a floor. Denied-capability telemetry tells you where a grant was too tight, if somebody is reading it. And the credit-note path suggests capability sets belong with the specification rather than the trace, because whoever knew credit notes existed knew it before any code ran.&lt;/p&gt;

&lt;p&gt;That is the &lt;a href="https://www.kenwalger.com/blog/ai-engineering/contract-discovery-bottleneck?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=code_review_authority_boundry" rel="noopener noreferrer"&gt;same conclusion I keep reaching&lt;/a&gt; from other directions. Once implementation is cheap and verification is automatable, the scarce work is discovering what the definition of correct should contain. Capability discovery is that problem wearing different clothes.&lt;/p&gt;

&lt;h2&gt;Correctness and Authority Are Orthogonal&lt;/h2&gt;

&lt;p&gt;The two-agent run makes this concrete. Code can be correct and over-privileged, as &lt;code&gt;agent_b&lt;/code&gt; is. Code can be incorrect and tightly constrained. The authority boundary limits one class of consequence without establishing correctness.&lt;/p&gt;

&lt;p&gt;A word about that heading, since the paper linked earlier is called "Why Security Is Not a Separable Concern" and appears to say the opposite. Miller's argument is that you cannot bolt security on afterwards as its own layer, because knowing what least authority means requires knowing what the program is for. Mine is narrower: evidence that an implementation is correct is silent about what that implementation may affect. He is describing where the work belongs in a design. I am describing what a passing test proves. Both land in the same place, which is that somebody who understands the task has to make the authority decision deliberately.&lt;/p&gt;

&lt;p&gt;A perfectly implemented image-resizing plugin should not reach payroll records. A buggy image-resizing plugin with no filesystem, network, or unrelated API access may cause considerably less damage than a correct implementation carrying unnecessary privileges.&lt;/p&gt;

&lt;p&gt;Verification asks whether an implementation satisfies its behavioral contract. Authority asks what consequences that implementation is permitted to create. Tests do not establish least privilege, and least privilege does not establish correctness. We need evidence for both, and I now have a fourteen-line proof that one kind of evidence is silent about the other.&lt;/p&gt;

&lt;h2&gt;Code Review Still Matters&lt;/h2&gt;

&lt;p&gt;None of this makes code review obsolete. Generated code contains logic errors, security vulnerabilities, bad dependency choices, race conditions, and plenty else worth finding.&lt;/p&gt;

&lt;p&gt;But code review is evidence about implementation. It should not be mistaken for enforcement of authority. If a system's safety depends on a reviewer noticing every dangerous operation an implementation could perform, we have made human attention part of the security boundary.&lt;/p&gt;

&lt;p&gt;A capability boundary gives us another layer. The reviewer can miss something, the model can misunderstand something, the implementation can contain behaviour nobody anticipated, and if the component was never granted the capability required to produce a particular consequence, that consequence remains unavailable.&lt;/p&gt;

&lt;p&gt;That is a considerably stronger guarantee than "we reviewed the code and did not see it doing that."&lt;/p&gt;

&lt;h2&gt;What Survives Regeneration?&lt;/h2&gt;

&lt;p&gt;Suppose an agent writes a component today. Tomorrow we regenerate it with a different model. Next month we replace the implementation entirely. What should remain stable?&lt;/p&gt;

&lt;p&gt;The code does not need to. The tests probably should. The behavioral contract certainly should. The authority boundary absolutely should.&lt;/p&gt;

&lt;p&gt;Which suggests a future where the most important artifacts around software are not implementation artifacts at all. They are declarations of what correctness means and what consequences an implementation is entitled to create.&lt;/p&gt;

&lt;p&gt;The model can write the code. It can even propose the capabilities it believes the code needs. But the system executing that code should have an independent answer to a much more important question: &lt;strong&gt;what is this implementation allowed to do?&lt;/strong&gt;&lt;/p&gt;





&lt;p&gt;&lt;em&gt;The host is about 200 lines of Python with one dependency. Four scenarios: &lt;code&gt;correctness&lt;/code&gt;, &lt;code&gt;authority&lt;/code&gt;, &lt;code&gt;discover&lt;/code&gt;, &lt;code&gt;verify&lt;/code&gt;. The last one is the interesting one, because it is where the obvious answer breaks. &lt;a href="https://gist.github.com/kenwalger/910173d9b8c23fdef5c1450b7fa43bbd" rel="noopener noreferrer"&gt;Gist on GitHub&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>webassembly</category>
      <category>ai</category>
      <category>architecture</category>
    </item>
    <item>
      <title>I Pulled Nine Years of My Own Dev.to Data. The Numbers Were Not What I Expected.</title>
      <dc:creator>Ken W Alger</dc:creator>
      <pubDate>Thu, 24 Sep 2026 13:38:00 +0000</pubDate>
      <link>https://dev.to/kenwalger/i-pulled-nine-years-of-my-own-devto-data-the-numbers-were-not-what-i-expected-37ac</link>
      <guid>https://dev.to/kenwalger/i-pulled-nine-years-of-my-own-devto-data-the-numbers-were-not-what-i-expected-37ac</guid>
      <description>&lt;p&gt;&lt;em&gt;There is an API. It will tell you things about your writing that the dashboard will not.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I have been publishing on Dev.to since April 2017. There is a four-year hole in the middle where I posted almost nothing, and then a return in March 2026 that has produced eighty-six posts in six months.&lt;/p&gt;

&lt;p&gt;That shape turns out to be useful. Two distinct bodies of work, on the same account, separated by a gap long enough that the platform itself changed underneath them. A natural experiment I did not set out to run.&lt;/p&gt;

&lt;p&gt;What I wanted was simple: a chart of follower growth with my article publication dates overlaid, to see whether particular posts moved the line. What I got instead was a fairly uncomfortable education in which of my numbers mean anything.&lt;/p&gt;

&lt;p&gt;The API made all of it possible, and almost nobody seems to use it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Yes, There Is an API
&lt;/h2&gt;

&lt;p&gt;Base URL is &lt;code&gt;https://dev.to/api&lt;/code&gt;. You generate a key at &lt;strong&gt;Settings → Extensions → DEV Community API Keys&lt;/strong&gt;. Most read endpoints for your own data want that key in an &lt;code&gt;api-key&lt;/code&gt; header.&lt;/p&gt;

&lt;p&gt;There is one detail that will silently ruin your afternoon:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;headers&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;api-key&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="c1"&gt;# Omit this and you get v0 responses. No error. No warning.
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;accept&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;application/vnd.forem.api-v1+json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user-agent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;my-analytics-script/1.0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Without the &lt;code&gt;accept&lt;/code&gt; header you are served the older v0 serializer. Nothing fails. The shapes are just quietly different, and you will spend twenty minutes wondering why a documented field is missing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trick One: The Default Page Size Is Not the Maximum
&lt;/h2&gt;

&lt;p&gt;The followers endpoint is documented as returning 80 per page. I have a little over eighteen thousand followers, which is 227 round trips.&lt;/p&gt;

&lt;p&gt;The v1 pagination range actually runs to 1000. The 80 is a default, not a ceiling:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;batch&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;API_ROOT&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/followers/users&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;params&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;page&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;page&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;per_page&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nineteen requests instead of 227. Thirty-one seconds instead of however long 227 polite requests would have taken.&lt;/p&gt;

&lt;p&gt;A Forem instance can cap &lt;code&gt;per_page&lt;/code&gt; lower through an environment variable, so do not hardcode the assumption. Ask for the maximum, then learn the real stride from what comes back:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;page&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;batch&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;requested&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="c1"&gt;# Either the server capped us or that is the entire list.
&lt;/span&gt;    &lt;span class="c1"&gt;# Either way, this is the real page size.
&lt;/span&gt;    &lt;span class="n"&gt;page_size&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;batch&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Rate limiting is real and it is not gentle. Back off on 429, sleep between pages, and checkpoint your merged results to disk every few pages. A long first pull that dies on page 190 should not discard the previous 189.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trick Two: Follower Dates Reconstruct History You Never Recorded
&lt;/h2&gt;

&lt;p&gt;This is the single most useful thing in the API and it is easy to miss.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;/api/followers/users&lt;/code&gt; returns a &lt;code&gt;created_at&lt;/code&gt; on every follower: the date that person started following you. You do not need to snapshot your follower count daily and wait six months to accumulate a time series. One pull reconstructs the entire curve retroactively, back to your first follower.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;collections&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Counter&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;date&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timedelta&lt;/span&gt;

&lt;span class="n"&gt;dates&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sorted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;date&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fromisoformat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;created_at&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][:&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;followers&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;cumulative&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;dates&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;  &lt;span class="c1"&gt;# the growth curve, free
&lt;/span&gt;
&lt;span class="n"&gt;weekly&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Counter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nf"&gt;timedelta&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;days&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;weekday&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;dates&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmwk5ndgtrdno0g21wo0i.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmwk5ndgtrdno0g21wo0i.png" alt="Line chart of cumulative Dev.to followers from 2017 to 2026. The line is nearly flat until March 2026, then rises steeply to about 18,000 by September. Dotted vertical markers show article publication dates. A dashed second line excluding auto-generated usernames tracks noticeably lower" width="799" height="433"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;There is a catch worth stating plainly, because it took me a moment to see it. You only receive people who &lt;em&gt;currently&lt;/em&gt; follow you. Anyone who followed and later unfollowed has vanished from the dataset. So the curve is "current followers by acquisition date," not your follower count as it stood on any given day. It increases monotonically by construction and can never show you a decline that actually happened.&lt;/p&gt;

&lt;p&gt;For working out which posts drove acquisition, fine. For anything about retention, actively misleading.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trick Three: The Analytics Endpoints Exist and Are Barely Documented
&lt;/h2&gt;

&lt;p&gt;There is a whole analytics family that most third-party tooling ignores:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/api/analytics/totals               lifetime views, reactions, comments
/api/analytics/historical           daily series over a date range
/api/analytics/past_day             hourly, last 24 hours
/api/analytics/referrers            where the traffic came from
/api/analytics/follower_engagement  follower growth over time
/api/analytics/dashboard            totals + history + top posts, bundled
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All accept &lt;code&gt;article_id&lt;/code&gt; to scope to a single post. Two things to know.&lt;/p&gt;

&lt;p&gt;The responses nest. They are not flat integers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"page_views"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"total"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;246454&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"average_read_time_in_seconds"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;306&lt;/span&gt;&lt;span class="p"&gt;}}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the historical data degrades as you go back. For my 2017 posts the endpoint returns weekly buckets rather than daily rows, and in aggregate accounts for only about 15% of those posts' lifetime views. For 2019 it covers 95%. For 2026, 100%.&lt;/p&gt;

&lt;p&gt;That matters more than it sounds. I computed a "half-life" for each post, meaning days from publication until it had earned half its views to date. My first attempt confidently reported that several 2017 posts had half-lives around 3,000 days. They do not. The endpoint simply does not remember most of what those posts earned, and dividing a remembered fraction produces a precise, authoritative, meaningless number.&lt;/p&gt;

&lt;p&gt;The fix is a coverage gate:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;lifetime&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;article&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;page_views_count&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;tracked&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;daily_series&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;values&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;

&lt;span class="c1"&gt;# A series accounting for 15% of a post's views will still yield a
# confident half-life. It will be an artifact of what the endpoint
# retained, not of how the post aged.
&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;lifetime&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;tracked&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;lifetime&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mf"&gt;0.8&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;continue&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Eighty-two of my 130 posts survive that gate. The median half-life among them is four days.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0454dyhys5xcdhyw0txj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0454dyhys5xcdhyw0txj.png" alt="Line chart showing cumulative percentage of views earned against days since publication, for twelve posts. Most curves rise almost vertically in the first few days and then flatten. A dotted horizontal line marks the fifty percent level, which most curves cross within a week." width="799" height="433"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Trick Four: Some Metadata Is in the Markdown, Not the JSON
&lt;/h2&gt;

&lt;p&gt;I write multi-part series. None of my posts came back with a &lt;code&gt;collection_id&lt;/code&gt;, which is the field you would reach for to group them.&lt;/p&gt;

&lt;p&gt;The series are right there in the Dev.to UI. The article serializer just does not include the field.&lt;/p&gt;

&lt;p&gt;But &lt;code&gt;/api/articles/me/published&lt;/code&gt; returns &lt;code&gt;body_markdown&lt;/code&gt;, front matter and all, and the series name is sitting in it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;FRONT_MATTER_SERIES&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;^series:\s*(.+?)\s*$&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;M&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;series_name&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;article&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;article&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;body_markdown&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lstrip&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;startswith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;---&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="n"&gt;parts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;---&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;parts&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="n"&gt;match&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;FRONT_MATTER_SERIES&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;parts&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;match&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;group&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="sh"&gt;'"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;match&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;General lesson: when a field you expect is absent from the JSON, check whether the source document came back too. It often did.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft14vx8c6cd2nhtdbnvux.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft14vx8c6cd2nhtdbnvux.png" alt="Line chart plotting lifetime views against part number for five article series. Every line except one descends from part one to part three, losing roughly half to two thirds of its readers." width="799" height="433"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Trick Five: Comments Arrive Pre-Threaded
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;/api/comments?a_id={id}&lt;/code&gt; returns comments as a tree, with each comment's replies nested in &lt;code&gt;children&lt;/code&gt;. The structure is doing analytical work for free, and flattening it while preserving depth takes about eight lines:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;flatten&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;nodes&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;depth&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;node&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;nodes&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;[]:&lt;/span&gt;
        &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;depth&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;depth&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;username&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;node&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;{}).&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;username&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;created_at&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;node&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;created_at&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;strip_html&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;node&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;body_html&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt;
        &lt;span class="p"&gt;})&lt;/span&gt;
        &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;extend&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;flatten&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;node&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;children&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;[],&lt;/span&gt; &lt;span class="n"&gt;depth&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Depth is the metric that matters. Comment &lt;em&gt;count&lt;/em&gt; cannot distinguish eight people each saying "great article" from two people arguing with you for four rounds. Maximum thread depth can. One of my posts has a thread 35 levels deep. That is not a comment section, it is a sustained argument, and no count-based metric would have told me it happened.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqtds0jj2xyp33wxotq7r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqtds0jj2xyp33wxotq7r.png" alt="Scatter plot with comment count on the horizontal axis and maximum thread depth on the vertical. Most points cluster low and left. A few outliers sit high on the vertical axis, indicating deep back-and-forth threads rather than many separate comments." width="799" height="433"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;There is no endpoint for &lt;em&gt;creating&lt;/em&gt; comments, incidentally. Replying still requires the browser. Probably deliberate.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Data Actually Said
&lt;/h2&gt;

&lt;p&gt;Here is where the exercise stopped being a programming problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My follower count is not a readership number.&lt;/strong&gt; I gained roughly 18,000 followers in 2026. My 2026 posts have 14,300 total views between them. You cannot acquire eighteen thousand followers from fourteen thousand views. The daily rate sits at a median of 136 with no meaningful response to whether I published anything, and about 37% of the usernames carry auto-generated-looking hex or numeric tails. That is reciprocal-follow farming, it is endemic, and it has nothing to do with me. It does mean the chart I originally set out to build could never have answered the question I was asking.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My most-viewed work is nine years old and no longer being read.&lt;/strong&gt; My 2017 output was MicroPython, NodeMCU, and MongoDB tutorials. Forty-four posts from 2017 to 2019 pulled 32,474 views. Eighty-six posts in 2026 have pulled 14,300. But the daily series tells the other half: my 5,472-view NodeMCU post has had zero views in the last ninety days. Nineteen of my 130 posts are at zero for the quarter. Lifetime counters never decrease, which makes an archive look alive long after it has stopped breathing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The engagement numbers invert completely.&lt;/strong&gt; Those big old tutorials run about 2 reactions per thousand views. My 2026 essays run 56 to 72, with one at 71.6 reactions and 63.8 comments per thousand. Across the eras: 44 old posts drew 26 comments total. 86 new posts have drawn 606.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5skop7lyops5cmmz55tz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5skop7lyops5cmmz55tz.png" alt="Scatter plot with lifetime views on a logarithmic horizontal axis and reactions per thousand views on the vertical. Older high-traffic posts sit far right and near the bottom. Newer low-traffic posts sit left and high. Red markers indicate posts with no views in the last ninety days." width="799" height="433"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So one body of work got found and skimmed. The other gets read and argued with. They are different products, and I had been evaluating both with the same number.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Four percent of my traffic is Google.&lt;/strong&gt; Direct or unknown is 74.5%. Internal Dev.to is 19%. For someone whose best-performing historical content is evergreen reference material, that is the number I find hardest to look at.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Would Tell Someone Starting This
&lt;/h2&gt;

&lt;p&gt;Pull the followers once and cache them, because the dates reconstruct years of history you never thought to record. Gate every analytics computation on data coverage, because a partial series will hand you a confident wrong answer rather than an error. Read thread depth instead of comment count. And check &lt;code&gt;body_markdown&lt;/code&gt; before concluding a field does not exist.&lt;/p&gt;

&lt;p&gt;Mostly, though: separate the metrics that measure distribution from the metrics that measure whether anyone cared. Views, follower counts, and impressions are the first kind. Comment depth, reply rates, and the fact that the same four people keep showing up in your threads are the second.&lt;/p&gt;

&lt;p&gt;I spent nine years assuming the first kind was the scoreboard. The API took an evening to tell me otherwise.&lt;/p&gt;

</description>
      <category>python</category>
      <category>api</category>
      <category>devjournal</category>
      <category>datascience</category>
    </item>
    <item>
      <title>Memory Is a System, Not a Prompt: Putting the Stack to Work</title>
      <dc:creator>Ken W Alger</dc:creator>
      <pubDate>Tue, 22 Sep 2026 13:54:00 +0000</pubDate>
      <link>https://dev.to/kenwalger/memory-is-a-system-not-a-prompt-putting-the-stack-to-work-3mba</link>
      <guid>https://dev.to/kenwalger/memory-is-a-system-not-a-prompt-putting-the-stack-to-work-3mba</guid>
      <description>&lt;p&gt;&lt;em&gt;Part 9 of the Building the AI Memory Stack series&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This series began with a deceptively simple observation: a context window is not memory.&lt;/p&gt;

&lt;p&gt;That idea evolved into an architectural model for building AI systems that remember, explain, verify, and restore information responsibly. Along the way we explored the purpose of the Context Window, the role of Active Working Memory, the permanence of Durable Memory, the accountability provided by the Reasoning Ledger, the trust established through Write-Side Custody, the cryptographic evidence supplied by Forensic Receipts, the restoration performed by Context Hydration, and the economic realities introduced by architectural taxes.&lt;/p&gt;

&lt;p&gt;Those individual ideas are not isolated concepts. Together they form a coherent system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Six Architectural Responsibilities
&lt;/h2&gt;

&lt;p&gt;The AI Memory Stack is not six storage technologies layered on top of one another.&lt;/p&gt;

&lt;p&gt;It is six architectural responsibilities every trustworthy AI system must answer.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Responsibility&lt;/th&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;th&gt;Sovereign Concept&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Execution&lt;/td&gt;
&lt;td&gt;What can I see right now?&lt;/td&gt;
&lt;td&gt;Context Window&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Assembly&lt;/td&gt;
&lt;td&gt;What matters right now?&lt;/td&gt;
&lt;td&gt;&lt;a href="https://sovereignplatform.dev/terms/active-working-memory.html?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=building_the_ai_memory_stack&amp;amp;utm_content=post_9_memory_is_a_system" rel="noopener noreferrer"&gt;Active Working Memory&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Preservation&lt;/td&gt;
&lt;td&gt;What should survive?&lt;/td&gt;
&lt;td&gt;&lt;a href="https://sovereignplatform.dev/terms/durable-memory.html?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=building_the_ai_memory_stack&amp;amp;utm_content=post_9_memory_is_a_system" rel="noopener noreferrer"&gt;Durable Memory&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Explanation&lt;/td&gt;
&lt;td&gt;Why did this happen?&lt;/td&gt;
&lt;td&gt;&lt;a href="https://sovereignplatform.dev/terms/reasoning-ledger.html?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=building_the_ai_memory_stack&amp;amp;utm_content=post_9_memory_is_a_system" rel="noopener noreferrer"&gt;Reasoning Ledger&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Integrity&lt;/td&gt;
&lt;td&gt;Should this become truth?&lt;/td&gt;
&lt;td&gt;&lt;a href="https://sovereignplatform.dev/terms/write-side-custody.html?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=building_the_ai_memory_stack&amp;amp;utm_content=post_9_memory_is_a_system" rel="noopener noreferrer"&gt;Write-Side Custody&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Evidence&lt;/td&gt;
&lt;td&gt;Can I prove it?&lt;/td&gt;
&lt;td&gt;&lt;a href="https://sovereignplatform.dev/terms/forensic-receipt.html?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=building_the_ai_memory_stack&amp;amp;utm_content=post_9_memory_is_a_system" rel="noopener noreferrer"&gt;Forensic Receipt&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Each responsibility exists because it solves a different problem. Durable Memory cannot explain why a decision was made. A Reasoning Ledger cannot prove its records have not been modified. A Context Window cannot preserve institutional knowledge. Together they create a system that is substantially more capable than any individual component.&lt;/p&gt;

&lt;h2&gt;
  
  
  Teaching Order Versus Runtime
&lt;/h2&gt;

&lt;p&gt;Throughout this series we intentionally walked down through the stack, beginning with the familiar idea of prompts before gradually uncovering the deeper responsibilities beneath them.&lt;/p&gt;

&lt;p&gt;A production system operates in the opposite direction.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    WSC["Write-Side Custody"] --&amp;gt; RL["Reasoning Ledger"]
    RL --&amp;gt; DM["Durable Memory"]
    RL --&amp;gt; FR["Forensic Receipt"]
    DM --&amp;gt; HY["Context Hydration"]
    HY --&amp;gt; AWM["Active Working Memory"]
    AWM --&amp;gt; CW["Context Window"]
    CW --&amp;gt; MI["Model Inference"]

    style WSC fill:#1D9E75,stroke:#11644a,stroke-width:2px,color:#fff
    style RL fill:#1D9E75,stroke:#11644a,stroke-width:2px,color:#fff
    style FR fill:#1D9E75,stroke:#11644a,stroke-width:2px,color:#fff
    style DM fill:#BA7517,stroke:#834f0c,stroke-width:2px,color:#fff
    style HY fill:#6B7280,stroke:#374151,stroke-width:2px,color:#fff
    style AWM fill:#BA7517,stroke:#834f0c,stroke-width:2px,color:#fff
    style CW fill:#BA7517,stroke:#834f0c,stroke-width:2px,color:#fff
    style MI fill:#BA7517,stroke:#834f0c,stroke-width:2px,color:#fff&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;&lt;a href="" class="article-body-image-wrapper"&gt;&lt;img&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;New information first passes through &lt;a href="https://sovereignplatform.dev/terms/write-side-custody.html?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=building_the_ai_memory_stack&amp;amp;utm_content=post_9_memory_is_a_system" rel="noopener noreferrer"&gt;Write-Side Custody&lt;/a&gt;, where it is sieved and evaluated before becoming institutional knowledge. The reasoning behind the decision is recorded in the &lt;a href="https://sovereignplatform.dev/terms/reasoning-ledger.html?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=building_the_ai_memory_stack&amp;amp;utm_content=post_9_memory_is_a_system" rel="noopener noreferrer"&gt;Reasoning Ledger&lt;/a&gt;, that record is sealed with a &lt;a href="https://sovereignplatform.dev/terms/forensic-receipt.html?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=building_the_ai_memory_stack&amp;amp;utm_content=post_9_memory_is_a_system" rel="noopener noreferrer"&gt;Forensic Receipt&lt;/a&gt;, and the verified knowledge is committed to &lt;a href="https://sovereignplatform.dev/terms/durable-memory.html?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=building_the_ai_memory_stack&amp;amp;utm_content=post_9_memory_is_a_system" rel="noopener noreferrer"&gt;Durable Memory&lt;/a&gt;. Later, &lt;a href="https://sovereignplatform.dev/terms/context-hydration.html?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=building_the_ai_memory_stack&amp;amp;utm_content=post_9_memory_is_a_system" rel="noopener noreferrer"&gt;Context Hydration&lt;/a&gt; reconstructs only the verified information required for the current task, placing it into Active Working Memory before it reaches the Context Window.&lt;/p&gt;

&lt;p&gt;Teaching order optimizes understanding.&lt;/p&gt;

&lt;p&gt;Runtime order optimizes execution.&lt;/p&gt;

&lt;h2&gt;
  
  
  From Architecture to Implementation
&lt;/h2&gt;

&lt;p&gt;Throughout this series the discussion has intentionally remained architecture-first. None of these responsibilities require a particular programming language, framework, model provider, or database.&lt;/p&gt;

&lt;p&gt;Architectures should outlive implementations.&lt;/p&gt;

&lt;p&gt;Implementations are how architectures prove themselves.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;Sovereign SDK&lt;/strong&gt; is intended to be a reference implementation of the Sovereign Systems Specification rather than a single monolithic AI framework. Instead of hiding every responsibility behind one package, it decomposes the architecture into focused components that mirror the boundaries defined by the specification.&lt;/p&gt;

&lt;p&gt;That distinction matters.&lt;/p&gt;

&lt;p&gt;The SDK is not the specification.&lt;/p&gt;

&lt;p&gt;It is one implementation of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Responsibilities Become Components
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://github.com/kenwalger/sovereign-sdk" rel="noopener noreferrer"&gt;reference SDK&lt;/a&gt; organizes these responsibilities along the information lifecycle, from the moment an observation is captured to the moment it informs reasoning. Each package owns one boundary in that progression.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Stage&lt;/th&gt;
&lt;th&gt;Package&lt;/th&gt;
&lt;th&gt;Status&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Capture&lt;/td&gt;
&lt;td&gt;&lt;code&gt;sovereign-sdk-sensor&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Active&lt;/td&gt;
&lt;td&gt;Capture observations at the Point of Genesis&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Classification&lt;/td&gt;
&lt;td&gt;&lt;code&gt;sovereign-sdk-edge&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Active&lt;/td&gt;
&lt;td&gt;Route, classify, and preserve locality&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Optimization&lt;/td&gt;
&lt;td&gt;&lt;code&gt;sovereign-sdk-sieve&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Active&lt;/td&gt;
&lt;td&gt;Reduce Prose Tax while preserving meaning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Evidence&lt;/td&gt;
&lt;td&gt;&lt;code&gt;sovereign-sdk-ledger&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Active&lt;/td&gt;
&lt;td&gt;Generate immutable Forensic Receipts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Governance&lt;/td&gt;
&lt;td&gt;&lt;code&gt;sovereign-sdk-airlock&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Active&lt;/td&gt;
&lt;td&gt;Govern outbound boundary crossings&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Memory&lt;/td&gt;
&lt;td&gt;&lt;code&gt;sovereign-sdk-vault&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Planned&lt;/td&gt;
&lt;td&gt;Long-term memory custody and retention&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The responsibilities this series named map onto that lifecycle. The write boundary the series called Write-Side Custody runs across the capture, classification, and optimization stages, where observations are validated and stripped of Prose Tax before they are trusted. Forensic Receipts are minted by the ledger. Outbound governance runs through the airlock. Durable Memory is the vault, the single piece still on the roadmap. Assembly, explanation, and hydration compose these packages in the application layer above them.&lt;/p&gt;

&lt;p&gt;You don't have to take the architecture on faith. The Sovereign Memory Demo, the flagship reference implementation, shows the stack running end to end: audited institutional memory retrieval with local, tamper-evident ledger custody. You can browse it and the demonstrations that follow it from the &lt;a href="https://sovereignplatform.dev/demos/index.html" rel="noopener noreferrer"&gt;demos index&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;As the SDK grows, packages may change names, implementations will evolve, and new capabilities will emerge. The architectural responsibilities, however, remain stable. A Rust implementation, a Go implementation, or a Java implementation could all faithfully implement the same architecture while looking completely different internally.&lt;/p&gt;

&lt;p&gt;That is one of the primary goals of the Sovereign Systems Specification: separating enduring architectural ideas from temporary implementation details.&lt;/p&gt;

&lt;p&gt;Earlier in this series we observed that a context window is not memory. The SDK is the architectural consequence of that observation. Rather than trying to solve every problem inside the prompt, it distributes responsibilities across specialized components that preserve, explain, verify, govern, and restore information throughout its lifecycle.&lt;/p&gt;

&lt;h2&gt;
  
  
  Memory Is Infrastructure
&lt;/h2&gt;

&lt;p&gt;The central argument of this series has never been that prompts are unimportant.&lt;/p&gt;

&lt;p&gt;Prompts matter.&lt;/p&gt;

&lt;p&gt;Models matter.&lt;/p&gt;

&lt;p&gt;Retrieval matters.&lt;/p&gt;

&lt;p&gt;What this series argues is that they are only part of a larger system.&lt;/p&gt;

&lt;p&gt;Thinking in terms of &lt;a href="https://sovereignplatform.dev/terms/memory-as-infrastructure.html?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=building_the_ai_memory_stack&amp;amp;utm_content=post_9_memory_is_a_system" rel="noopener noreferrer"&gt;Memory as Infrastructure&lt;/a&gt; changes the design conversation. Instead of asking how to fit more information into a prompt, we begin asking how information should be accepted, preserved, verified, restored, governed, and ultimately communicated back to a model.&lt;/p&gt;

&lt;p&gt;Those are architectural questions.&lt;/p&gt;

&lt;p&gt;And architectural questions tend to outlive technology cycles.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final Thoughts
&lt;/h2&gt;

&lt;p&gt;Every generation of software eventually discovers that the difficult problem is not computation.&lt;/p&gt;

&lt;p&gt;It is memory.&lt;/p&gt;

&lt;p&gt;Databases changed how applications remembered.&lt;/p&gt;

&lt;p&gt;Version control changed how teams remembered.&lt;/p&gt;

&lt;p&gt;Observability changed how operators remembered.&lt;/p&gt;

&lt;p&gt;Artificial intelligence is forcing us to rethink how intelligent systems remember.&lt;/p&gt;

&lt;p&gt;Larger context windows are one answer.&lt;/p&gt;

&lt;p&gt;Better memory architecture is another.&lt;/p&gt;

&lt;p&gt;This series has argued that the second answer will ultimately matter more.&lt;/p&gt;

&lt;p&gt;Prompts are temporary.&lt;/p&gt;

&lt;p&gt;Memory endures.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The future of AI belongs to systems that remember well.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>llm</category>
      <category>agents</category>
    </item>
    <item>
      <title>Frameworks Are Institutional Memory</title>
      <dc:creator>Ken W Alger</dc:creator>
      <pubDate>Thu, 17 Sep 2026 17:07:00 +0000</pubDate>
      <link>https://dev.to/kenwalger/frameworks-are-institutional-memory-b3l</link>
      <guid>https://dev.to/kenwalger/frameworks-are-institutional-memory-b3l</guid>
      <description>&lt;p&gt;&lt;em&gt;I stored passwords in plaintext on a floppy disk in the 1980s. Recently, a coding agent made a version of the same mistake.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;When I was running bulletin board systems in the 1980s, password management could be remarkably straightforward.&lt;/p&gt;

&lt;p&gt;A user created an account. They chose a password. The BBS needed to determine whether the password they entered later was the same one they had chosen.&lt;/p&gt;

&lt;p&gt;So you stored it.&lt;/p&gt;

&lt;p&gt;In plaintext.&lt;/p&gt;

&lt;p&gt;On a floppy disk.&lt;/p&gt;

&lt;p&gt;Hard drives were expensive and hardly universal in the hobbyist computing world I inhabited, so there was nothing strange about application data living on floppies. The important thing was that the software could retrieve the password and compare it against whatever the user typed the next time they called in.&lt;/p&gt;

&lt;p&gt;Something like this was perfectly understandable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;KEN|swordfish
SARAH|password123
MIKE|enterprise
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There was, of course, a slight problem.&lt;/p&gt;

&lt;p&gt;The person running the BBS could read everybody's passwords.&lt;/p&gt;

&lt;p&gt;So people started solving that. Maybe the password should not be stored exactly as the user entered it. A substitution cipher or some other homegrown transformation could at least stop someone from casually opening the user file and reading every credential.&lt;/p&gt;

&lt;p&gt;Problem solved.&lt;/p&gt;

&lt;p&gt;Well. One problem solved.&lt;/p&gt;

&lt;p&gt;The transformation might be reversible. Two users choosing the same password produced the same stored value. A compromised file exposed everyone at once. Users reused those passwords elsewhere. Password recovery introduced an entirely separate collection of problems.&lt;/p&gt;

&lt;p&gt;Each solution exposed another question.&lt;/p&gt;

&lt;p&gt;Over the following decades, "store a password" accumulated an extraordinary amount of engineering knowledge.&lt;/p&gt;

&lt;p&gt;Today, saying an application needs authentication describes far more than comparing one string against another. It means hashing, salts, password reset, expiring tokens, single-use tokens, session invalidation, account enumeration, brute-force protection, cookie attributes, CSRF, rate limiting, and a long list of other concerns depending on the threat model.&lt;/p&gt;

&lt;p&gt;A developer can still build all of that from scratch.&lt;/p&gt;

&lt;p&gt;The more interesting question is whether they should.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Code You Don't Know You Need
&lt;/h2&gt;

&lt;p&gt;The first thing developers learn about frameworks is that they save us from writing code.&lt;/p&gt;

&lt;p&gt;That is true, and it undersells them badly.&lt;/p&gt;

&lt;p&gt;A mature framework does not merely contain code somebody else wrote. It contains decisions somebody else already had to make, edge cases somebody else already hit, vulnerabilities somebody else already discovered, and fixes somebody else already learned were necessary.&lt;/p&gt;

&lt;p&gt;A mature framework is partly &lt;strong&gt;institutional memory encoded as software&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Take authentication. The happy path is not conceptually difficult:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User submits credentials
        ↓
Find account
        ↓
Check password
        ↓
Create session
        ↓
Return success
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A relatively inexperienced developer can understand that flow and implement a version of it.&lt;/p&gt;

&lt;p&gt;Production authentication is not hard because the happy path is hard. It is hard because of everything surrounding the happy path.&lt;/p&gt;

&lt;p&gt;How should the password be stored? What happens after repeated failed attempts? Can an attacker determine whether a username exists by comparing error messages? What happens to existing sessions after a password change? How long should a reset token remain valid? What happens after it is used? Can it be replayed? What attributes belong on the session cookie?&lt;/p&gt;

&lt;p&gt;The developer writing authentication for the first time does not necessarily know to ask any of that.&lt;/p&gt;

&lt;p&gt;A mature framework often does.&lt;/p&gt;

&lt;p&gt;That is a very different kind of value from saving keystrokes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Abstractions Change Who Has to Think
&lt;/h2&gt;

&lt;p&gt;Software has always progressed partly by taking things that required specialized knowledge and putting them behind abstractions.&lt;/p&gt;

&lt;p&gt;Assembly gave me extraordinary control over what the processor did. It was also a tremendous pain. Higher-level languages let us express more sophisticated ideas without manually considering every instruction. C retained substantial control over memory and execution while being far more productive. Later languages and runtimes provided increasingly sophisticated abstractions for memory management, type safety, concurrency primitives, and networking. Libraries and frameworks moved the boundary again. Managed platforms moved it again.&lt;/p&gt;

&lt;p&gt;None of those layers made the layer underneath disappear.&lt;/p&gt;

&lt;p&gt;Your Python application still causes a processor to execute machine instructions. Your Django application still communicates over networks. Your application on a managed platform still runs on computers somewhere.&lt;/p&gt;

&lt;p&gt;What the abstraction changes is &lt;strong&gt;who has to think about those things routinely&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Control and Responsibility Arrive Together
&lt;/h2&gt;

&lt;p&gt;Developers sometimes discuss control as though it were an unqualified good.&lt;/p&gt;

&lt;p&gt;C gives me more control over memory. Running directly on infrastructure gives me more control than a managed platform. Writing SQL gives me more control than an ORM. Building against browser automation primitives gives me more control than a higher-level testing abstraction.&lt;/p&gt;

&lt;p&gt;All of those statements can be true.&lt;/p&gt;

&lt;p&gt;But control has a twin that gets mentioned far less often: &lt;strong&gt;responsibility&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;If you control memory allocation, you are responsible for memory allocation. If you control your infrastructure, you are responsible for configuring, securing, observing, updating, and troubleshooting it. If you control every selector in an end-to-end test suite, you own every selector when the application changes.&lt;/p&gt;

&lt;p&gt;You could write your inventory API in C. Manage the networking, parse HTTP, handle allocation, implement routing, serialization, authentication, database connectivity, concurrency, and error handling yourself. You would have an enormous amount of control, and you would have volunteered for an enormous number of problems that Flask, FastAPI, Django, Rails, Spring, and others have spent years making uninteresting.&lt;/p&gt;

&lt;p&gt;The trade is not:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;abstraction versus control&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is closer to:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;less control and less responsibility versus more control and more responsibility&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Sometimes the additional control is worth the additional responsibility. Sometimes it is not.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A["Problem"] --&amp;gt; B{"Own the complexity?"}
    B --&amp;gt;|"Yes"| C["More Control"]
    C --&amp;gt; D["More Engineering Responsibility"]
    B --&amp;gt;|"No"| E["Use an Abstraction"]
    E --&amp;gt; F["Accept Its Opinions &amp;amp; Constraints"]
    D --&amp;gt; G["Build Your Differentiating Work"]
    F --&amp;gt; G

    classDef question fill:#FFFFFF,stroke:#166534,color:#14532D,stroke-width:2px;
    classDef control fill:#FEF2F2,stroke:#991B1B,color:#7F1D1D,stroke-width:2px;
    classDef abstract fill:#E8F3EE,stroke:#166534,color:#14532D,stroke-width:2px;
    classDef outcome fill:#F0FDF4,stroke:#5CA08A,color:#14532D,stroke-width:2px;

    class A,G outcome;
    class B question;
    class C,D control;
    class E,F abstract;&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Using the abstraction does not make you less of an engineer. Building the lower layer yourself does not automatically make you a better one. The skill is recognizing which problems deserve your attention, while understanding what you surrendered to have the others solved for you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Opinions Are Part of the Product
&lt;/h2&gt;

&lt;p&gt;"Opinionated" sometimes gets used as a criticism.&lt;/p&gt;

&lt;p&gt;It certainly can be one. If a framework's opinions conflict with your application's fundamental requirements, you will spend more time fighting it than benefiting from it.&lt;/p&gt;

&lt;p&gt;But an opinion is also a decision you do not have to make. A framework with an established approach to authentication, migrations, forms, routing, validation, sessions, and project structure has removed that many questions from your team's agenda, and mature defaults usually carry years of accumulated experience with them.&lt;/p&gt;

&lt;p&gt;The other side of the bargain is that somebody else has partially decided what your options look like.&lt;/p&gt;

&lt;p&gt;Using MongoDB with Django was a long-running example. Django's data layer was built around assumptions that fit relational databases particularly well, and getting MongoDB underneath it was possible through various third-party approaches, though "interesting" would be a charitable description of that experience. You were making one architectural choice while using a framework whose opinions had been designed around another.&lt;/p&gt;

&lt;p&gt;What happened next supports the argument better than the complaint did. &lt;a href="https://github.com/mongodb/django-mongodb-backend" rel="noopener noreferrer"&gt;MongoDB shipped an official Django backend&lt;/a&gt; in public preview in early 2025, and it now handles embedded models, queryable encryption, geospatial lookups, and most of what the third-party packages struggled with. The framework's opinions were not overturned. They were extended, by people willing to do the work of reconciling the document model with Django's assumptions.&lt;/p&gt;

&lt;p&gt;That took years, and it is exactly the process this whole post is about. The accumulated knowledge is the product. It just accumulates slowly.&lt;/p&gt;

&lt;p&gt;So before adopting any framework, do not look only at the complexity it removes. Ask what decisions it makes on your behalf, what assumptions are embedded in those decisions, and how painful things become when your application needs something outside them.&lt;/p&gt;

&lt;p&gt;You can dislike the choices. You can replace some of them. You can decide the framework is wrong for you. What matters is understanding the bargain.&lt;/p&gt;

&lt;h2&gt;
  
  
  Every Platform Sells You Both
&lt;/h2&gt;

&lt;p&gt;Managed platforms made this visible in a way frameworks alone did not.&lt;/p&gt;

&lt;p&gt;Before them, deploying an application meant owning a considerable amount of infrastructure knowledge: somewhere to run it, a way to deploy it, process management, configuration, networking, logging, databases, scaling, and monitoring. Managed platforms did not make any of that cease to exist. They changed who had to own it.&lt;/p&gt;

&lt;p&gt;What is easy to miss is that this was never a choice between vendors. It is a choice available inside every vendor. AWS will happily sell you App Runner, Elastic Beanstalk, Lambda, or Amplify, and it will just as happily sell you EC2 instances, a VPC, load balancers, autoscaling groups, and a pile of IAM policies to assemble yourself. Google Cloud offers Cloud Run and App Engine alongside Compute Engine and GKE. Azure has App Service and Container Apps on one side, virtual machines and AKS on the other.&lt;/p&gt;

&lt;p&gt;Same provider. Same workload. Radically different amounts of responsibility.&lt;/p&gt;

&lt;p&gt;The opinionated paths get you running quickly and constrain what you can do. The assembled paths let you build nearly anything and hand you an operational surface that keeps expanding for as long as the system exists. Most real architectures contain both, which is usually the right answer.&lt;/p&gt;

&lt;p&gt;The constraint is real in either direction. If you need behavior outside the managed model, the abstraction becomes limiting, and there are plenty of workloads where owning the pieces makes more sense. That is not a failure of the abstraction. It is the trade the abstraction offered.&lt;/p&gt;

&lt;p&gt;One asymmetry is worth pricing in advance: the managed path is cheap to enter and expensive to leave. Moving off it means rebuilding the machinery the platform was quietly providing, usually under time pressure, usually at the exact moment the constraint became intolerable. That is not an argument against starting there. It is an argument for knowing which constraint would force the move before you have to make it.&lt;/p&gt;

&lt;p&gt;For most teams the question was never whether their engineers could build and operate infrastructure. It was whether infrastructure was what their customers were paying them to be good at.&lt;/p&gt;

&lt;h2&gt;
  
  
  Spend Your Complexity Budget Wisely
&lt;/h2&gt;

&lt;p&gt;I think of this as a &lt;strong&gt;complexity budget&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Every team has finite attention. There are only so many engineers, so many hours, and so many systems a group can deeply understand and maintain. Complexity spends that budget.&lt;/p&gt;

&lt;p&gt;Build your own authentication and you have spent some of it on authentication. Operate your own infrastructure and you have spent some of it on infrastructure. Build a custom persistence layer and you have spent some of it on persistence.&lt;/p&gt;

&lt;p&gt;Any of those can be excellent investments when the capability differentiates your product or your requirements genuinely demand the control. Dan McKinley's "&lt;a href="https://boringtechnology.club/" rel="noopener noreferrer"&gt;Choose Boring Technology&lt;/a&gt;" makes a neighboring argument with innovation tokens: you get about three, so spend them where they matter. The complexity budget is the same instinct pointed at what you build rather than what you adopt.&lt;/p&gt;

&lt;p&gt;"We could build it ourselves" does not establish anything. Engineers can build lots of things.&lt;/p&gt;

&lt;p&gt;The better question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Is this where we want to spend our complexity budget?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Learn What's Underneath Anyway
&lt;/h2&gt;

&lt;p&gt;None of this argues against learning fundamentals. Quite the opposite.&lt;/p&gt;

&lt;p&gt;Understanding HTTP makes you better at using a web framework. Understanding SQL makes you better at using an ORM. Understanding memory makes you better at diagnosing what a runtime is doing.&lt;/p&gt;

&lt;p&gt;Every abstraction eventually leaks. When it does, knowing what is underneath tells you whether you are looking at a bug, a limitation, a bad assumption, or the consequence of a trade you made a year ago.&lt;/p&gt;

&lt;p&gt;But there is a difference between &lt;strong&gt;understanding a layer&lt;/strong&gt; and &lt;strong&gt;taking responsibility for operating it&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I do not need to fabricate a processor to understand how one works. I do not need to write an HTTP server in C to understand HTTP. I do not need to implement my own password hashing to understand why plaintext was a bad idea.&lt;/p&gt;

&lt;p&gt;Sometimes understanding a problem thoroughly is precisely why you decide not to implement it yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Password Problem Never Really Went Away
&lt;/h2&gt;

&lt;p&gt;Recently I had a coding agent implement password reset for a small application.&lt;/p&gt;

&lt;p&gt;It built the feature in about five minutes. The email arrived. The link opened. The password changed. The user logged in with the new one.&lt;/p&gt;

&lt;p&gt;Everything worked.&lt;/p&gt;

&lt;p&gt;Except the reset link worked a second time.&lt;/p&gt;

&lt;p&gt;The implementation satisfied the obvious behavior while missing the invariant underneath it: once the reset succeeds, the token has to become useless.&lt;/p&gt;

&lt;p&gt;That felt familiar.&lt;/p&gt;

&lt;p&gt;Forty years ago the question was:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Does the password let the right person log into the BBS?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Yes.&lt;/p&gt;

&lt;p&gt;Then somebody asked:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can the sysop read everyone's password?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Oh.&lt;/p&gt;

&lt;p&gt;Years later:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Is the stored password protected?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Yes.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can the stored representation be reversed or cheaply cracked?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Oh.&lt;/p&gt;

&lt;p&gt;And now:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Does the password reset feature work?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Yes.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can I reuse the token?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Oh.&lt;/p&gt;

&lt;p&gt;The technologies change. The pattern does not. We implement the requirements we know about, and experience introduces us to the ones we did not.&lt;/p&gt;

&lt;p&gt;That is why mature abstractions matter. They carry some of that experience forward so the next developer inherits the lesson without personally reliving it. Somebody already hit the edge case. Somebody already found the vulnerability. Somebody already spent three days on the concurrency bug that only appears on Tuesdays.&lt;/p&gt;

&lt;p&gt;Which leaves an open question I do not think we have answered yet.&lt;/p&gt;

&lt;p&gt;A mature framework carries its institutional memory in its decisions. Sometimes the reasons stay attached through issues, commits, documentation, and design discussions. Often they do not, and only the resulting constraint survives. Either way, somebody encountered the problem and changed the system because of it.&lt;/p&gt;

&lt;p&gt;That is the part worth noticing. The single-use token check runs whether or not anyone remembers why it was added. The lesson is enforced rather than remembered.&lt;/p&gt;

&lt;p&gt;A model trained on a very large corpus of code has absorbed something that resembles institutional memory. But it learned from artifacts produced after those decisions were made, rather than from the decisions themselves. It can reproduce what surviving code looks like without inheriting anything that enforces why one implementation survived and another did not.&lt;/p&gt;

&lt;p&gt;Which might explain why my agent wrote a flawless happy path and dropped the invariant.&lt;/p&gt;

&lt;p&gt;Happy paths are abundant. Invariants live in the decisions.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>beginners</category>
    </item>
    <item>
      <title>The Hidden Taxes of Prompt-Only AI</title>
      <dc:creator>Ken W Alger</dc:creator>
      <pubDate>Tue, 15 Sep 2026 14:03:00 +0000</pubDate>
      <link>https://dev.to/kenwalger/the-hidden-taxes-of-prompt-only-ai-24lo</link>
      <guid>https://dev.to/kenwalger/the-hidden-taxes-of-prompt-only-ai-24lo</guid>
      <description>&lt;p&gt;&lt;em&gt;Part 8 of the Building the AI Memory Stack series&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Over the past seven articles we've built an architecture that treats &lt;a href="https://sovereignplatform.dev/terms/memory-as-infrastructure.html?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=building_the_ai_memory_stack&amp;amp;utm_content=post_8_hidden_taxes" rel="noopener noreferrer"&gt;memory as infrastructure&lt;/a&gt; rather than as an oversized prompt. We've separated execution from assembly, preservation from explanation, trust from proof, and finally showed how verified knowledge returns to active reasoning through &lt;a href="https://sovereignplatform.dev/terms/context-hydration.html?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=building_the_ai_memory_stack&amp;amp;utm_content=post_8_hidden_taxes" rel="noopener noreferrer"&gt;Context Hydration&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Now it's time to ask a different question.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What does all of that cost?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Every AI system pays for memory. The only question is &lt;strong&gt;where&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Many systems choose to pay almost every cost inside the prompt itself. As context windows grow larger, it becomes tempting to treat them as an infinitely expandable memory system. If the model forgets something, add more documents. If retrieval misses context, increase the top-k value. If the answer is incomplete, make the prompt longer.&lt;/p&gt;

&lt;p&gt;That approach works surprisingly well, until it doesn't.&lt;/p&gt;

&lt;p&gt;The cost isn't limited to API pricing. Large prompts consume attention, increase latency, complicate orchestration, and force the model to separate important information from noise. The result is an architectural bill that grows long before the invoice from your model provider does.&lt;/p&gt;

&lt;h2&gt;
  
  
  Capacity Is Not Communication
&lt;/h2&gt;

&lt;p&gt;One of the recurring themes throughout this series has been that storage and communication are different problems.&lt;/p&gt;

&lt;p&gt;A library may contain every book ever written, but that doesn't mean every book belongs on your desk while solving today's problem. Likewise, &lt;a href="https://sovereignplatform.dev/terms/durable-memory.html?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=building_the_ai_memory_stack&amp;amp;utm_content=post_8_hidden_taxes" rel="noopener noreferrer"&gt;Durable Memory&lt;/a&gt; can preserve years of organizational knowledge without requiring every byte of it to enter today's Context Window.&lt;/p&gt;

&lt;p&gt;The purpose of architecture is deciding &lt;strong&gt;what should move&lt;/strong&gt;, &lt;strong&gt;when it should move&lt;/strong&gt;, and &lt;strong&gt;what it costs to move it&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Introducing the Tax Model
&lt;/h2&gt;

&lt;p&gt;The Sovereign Systems Specification describes these recurring costs as architectural taxes. They are not bugs. They are the predictable costs of moving, storing, validating, retrieving, and communicating information through an AI system.&lt;/p&gt;

&lt;p&gt;Some taxes are unavoidable.&lt;/p&gt;

&lt;p&gt;Others are self-inflicted.&lt;/p&gt;

&lt;p&gt;Good architecture minimizes the second category.&lt;/p&gt;

&lt;p&gt;The taxes that bear most directly on memory are these.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;a href="https://sovereignplatform.dev/terms/prose-tax.html?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=building_the_ai_memory_stack&amp;amp;utm_content=post_8_hidden_taxes" rel="noopener noreferrer"&gt;Prose Tax&lt;/a&gt;
&lt;/h2&gt;

&lt;p&gt;Every explanation has a cost.&lt;/p&gt;

&lt;p&gt;Humans naturally communicate in paragraphs. Models consume tokens. The more words required to express an idea, the more attention the model must allocate before it can begin reasoning.&lt;/p&gt;

&lt;p&gt;High-information-density representations, such as structured records, schemas, identifiers, and references, often communicate the same meaning with a fraction of the prompt budget.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;a href="https://sovereignplatform.dev/terms/context-tax.html?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=building_the_ai_memory_stack&amp;amp;utm_content=post_8_hidden_taxes" rel="noopener noreferrer"&gt;Context Tax&lt;/a&gt;
&lt;/h2&gt;

&lt;p&gt;Every additional token competes for attention.&lt;/p&gt;

&lt;p&gt;Context windows have grown dramatically, but attention remains finite. As more information enters the prompt, genuinely important information must compete with increasingly irrelevant material.&lt;/p&gt;

&lt;p&gt;Bigger windows increase capacity.&lt;/p&gt;

&lt;p&gt;They do not guarantee better focus.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;a href="https://sovereignplatform.dev/terms/retrieval-tax.html?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=building_the_ai_memory_stack&amp;amp;utm_content=post_8_hidden_taxes" rel="noopener noreferrer"&gt;Retrieval Tax&lt;/a&gt;
&lt;/h2&gt;

&lt;p&gt;Searching for information is not free.&lt;/p&gt;

&lt;p&gt;Embedding generation, vector searches, re-ranking, filtering, serialization, and prompt assembly all consume compute and latency before the model has produced a single token of useful work.&lt;/p&gt;

&lt;p&gt;As argued earlier in this series, retrieval should support memory, not replace it.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;a href="https://sovereignplatform.dev/terms/observer-tax.html?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=building_the_ai_memory_stack&amp;amp;utm_content=post_8_hidden_taxes" rel="noopener noreferrer"&gt;Observer's Tax&lt;/a&gt;
&lt;/h2&gt;

&lt;p&gt;Every measurement has a cost.&lt;/p&gt;

&lt;p&gt;Telemetry, debugging information, traces, evaluation artifacts, and compliance records are essential for production systems. Left unchecked, however, they begin competing with operational workloads for compute, storage, and engineering attention.&lt;/p&gt;

&lt;p&gt;Observability is infrastructure.&lt;/p&gt;

&lt;p&gt;It should not become interference.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;a href="https://sovereignplatform.dev/terms/ingestion-tax.html?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=building_the_ai_memory_stack&amp;amp;utm_content=post_8_hidden_taxes" rel="noopener noreferrer"&gt;Ingestion Tax&lt;/a&gt;
&lt;/h2&gt;

&lt;p&gt;The cheapest place to improve information quality is before information enters the system.&lt;/p&gt;

&lt;p&gt;Poorly structured data generates downstream costs forever. Duplicate records, inconsistent schemas, missing provenance, and unverifiable observations all create future work for retrieval pipelines, prompt assembly, and reasoning itself.&lt;/p&gt;

&lt;p&gt;Every bad write compounds.&lt;/p&gt;

&lt;p&gt;Every good write pays dividends.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fiscal Architecture
&lt;/h2&gt;

&lt;p&gt;Viewed individually, these taxes seem manageable.&lt;/p&gt;

&lt;p&gt;Viewed together, they become an architectural discipline.&lt;/p&gt;

&lt;p&gt;None of these taxes exist in isolation. Attempts to reduce one often increase another. Expanding a prompt may reduce retrieval work while increasing Context Tax. Adding more telemetry may improve observability while increasing Observer's Tax. The goal isn't minimizing a single tax; it's balancing the entire system.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tax&lt;/th&gt;
&lt;th&gt;What it charges for&lt;/th&gt;
&lt;th&gt;Lowered by&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Prose Tax&lt;/td&gt;
&lt;td&gt;Meaning expressed in more tokens than it needs&lt;/td&gt;
&lt;td&gt;Structured records over paragraphs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context Tax&lt;/td&gt;
&lt;td&gt;Irrelevant tokens competing for finite attention&lt;/td&gt;
&lt;td&gt;Hydrating only what the task needs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Retrieval Tax&lt;/td&gt;
&lt;td&gt;Search, embedding, and re-ranking before any output&lt;/td&gt;
&lt;td&gt;Higher memory quality, less searching&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Observer's Tax&lt;/td&gt;
&lt;td&gt;Telemetry and traces competing with real work&lt;/td&gt;
&lt;td&gt;Bounded, purposeful observability&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ingestion Tax&lt;/td&gt;
&lt;td&gt;Poor structure and missing provenance at the write&lt;/td&gt;
&lt;td&gt;Verified, structured writes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Organizations often spend months optimizing prompts while ignoring the systems that create those prompts. Yet the largest savings usually come from improving memory quality, reducing unnecessary movement, and preserving information in forms that are inexpensive to hydrate later.&lt;/p&gt;

&lt;p&gt;Traditional software architecture optimizes CPU, memory, network bandwidth, and storage.&lt;/p&gt;

&lt;p&gt;AI architecture adds another economic dimension: &lt;em&gt;attention&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Every architectural decision ultimately affects how much attention the system spends producing useful reasoning.&lt;/p&gt;

&lt;p&gt;In other words, the cheapest token is often the one that never needed to exist.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Goal Isn't Zero Tax
&lt;/h2&gt;

&lt;p&gt;Every system pays taxes.&lt;/p&gt;

&lt;p&gt;A trustworthy system willingly pays some of them.&lt;/p&gt;

&lt;p&gt;Hashes must be computed. Receipts must be signed. Memory must be verified before it is restored. Good engineering accepts these costs because they purchase integrity, explainability, and confidence.&lt;/p&gt;

&lt;p&gt;The goal is not eliminating cost.&lt;/p&gt;

&lt;p&gt;The goal is paying the &lt;strong&gt;right&lt;/strong&gt; costs in the &lt;strong&gt;right&lt;/strong&gt; places.&lt;/p&gt;

&lt;h2&gt;
  
  
  Looking Ahead
&lt;/h2&gt;

&lt;p&gt;The final article brings everything together.&lt;/p&gt;

&lt;p&gt;We'll revisit the AI Memory Stack from the perspective of runtime execution rather than teaching order, showing how data actually flows through the architecture and mapping each responsibility to the &lt;a href="https://github.com/kenwalger/sovereign-sdk" rel="noopener noreferrer"&gt;Sovereign SDK&lt;/a&gt;. By the end, the stack should feel less like a collection of concepts and more like a blueprint that can be implemented today.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>architecture</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Living a Colorful Life in a Black-and-White ATS World</title>
      <dc:creator>Ken W Alger</dc:creator>
      <pubDate>Mon, 14 Sep 2026 13:48:00 +0000</pubDate>
      <link>https://dev.to/kenwalger/living-a-colorful-life-in-a-black-and-white-ats-world-2m0o</link>
      <guid>https://dev.to/kenwalger/living-a-colorful-life-in-a-black-and-white-ats-world-2m0o</guid>
      <description>&lt;h3&gt;
  
  
  A career has edges. The schema only stores nodes.
&lt;/h3&gt;

&lt;p&gt;I spent years in commercial kitchens and over a decade running a one-man construction business before I ever shipped software professionally. Here is what those years look like inside an applicant tracking system:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sous Chef&lt;/strong&gt;, Illahe Hills Country Club, 1996-1997&lt;br&gt;
&lt;strong&gt;Owner&lt;/strong&gt;, Alger Construction, 2006-2018&lt;/p&gt;

&lt;p&gt;Two rows. Six fields. Everything that made those years worth putting on a resume in the first place lives in the space between them, and there is no field for the space between them.&lt;/p&gt;

&lt;p&gt;That is not a bug in any particular vendor's product. It is what happens when a career gets stored.&lt;/p&gt;

&lt;h2&gt;
  
  
  Compression Is the Point, and Compression Has a Cost
&lt;/h2&gt;

&lt;p&gt;A resume is already a lossy representation of a working life. An applicant tracking system compresses it again into titles, employers, dates, skills, and years of experience. Recruiter searches narrow it further. Screening criteria narrow it again. By the time a hiring manager sees a candidate, a great deal of what made that person worth meeting has been discarded by systems doing exactly what they were built to do.&lt;/p&gt;

&lt;p&gt;Hiring at scale requires structure. Nobody sensible is arguing otherwise.&lt;/p&gt;

&lt;p&gt;But compression is only lossless when the thing being compressed matches the schema. Software engineer becomes senior software engineer becomes staff software engineer. The titles line up. The keywords line up. The progression survives the trip through the pipeline nearly intact.&lt;/p&gt;

&lt;p&gt;For a career that moved sideways, the same pipeline behaves very differently. It preserves the least interesting facts and drops the reason those facts belong on the same page.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Resume Contains the Nodes. Something Has to Explain the Edges.
&lt;/h2&gt;

&lt;p&gt;Running a construction business teaches you schedules, budgets, vendors, dependencies, unhappy customers, and what it costs to discover a problem after the wall is closed. Working a station on a busy line teaches you sequencing, preparation, and how to communicate when six things are failing at once and none of them are your fault. Managing a distributed technical team a decade later draws on both, whether or not the job description has a checkbox for it.&lt;/p&gt;

&lt;p&gt;Systems fail. Priorities collide. Resources are finite. People misunderstand each other. Plans meet reality. Learning to operate under those conditions is transferable, and it is the actual content of a nonlinear career.&lt;/p&gt;

&lt;p&gt;A person reading a resume can sometimes reconstruct those connections. A parser has a harder problem, because almost everything it has been given is categorical. It is very good at representing what fits in a field. It has nowhere to put "running a small business taught me something about operational risk that later made me better at managing engineers."&lt;/p&gt;

&lt;p&gt;The information exists. The schema has no column for it.&lt;/p&gt;

&lt;p&gt;This is why the quiet disappearance of the cover letter is more consequential than it looks. I understand why nobody misses them. Most were three paragraphs about being passionate about synergizing solutions in a fast-paced environment, and they earned their reputation. But the format was carrying a load that nothing replaced. The more nonlinear the candidate, the more of their value lives in relational information, and the application form keeps getting narrower precisely where that information needs to go.&lt;/p&gt;

&lt;p&gt;We tell people to bring their whole selves to work while steadily reducing the bandwidth available to transmit one through the front door.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Strongest Objection Is Not Efficiency
&lt;/h2&gt;

&lt;p&gt;The honest counterargument to everything above is not that narrative is slow to read. It is that narrative is where bias lives.&lt;/p&gt;

&lt;p&gt;Structured screening was partly a correction. Unstructured hiring, the kind that runs on interesting backgrounds and good conversations and a hiring manager's instinct about fit, has a well documented tendency to favor candidates who resemble the person doing the evaluating. "This person's unusual path is fascinating" has historically meant "this person reminds me of me." Fields, rubrics, and consistent criteria were a response to a real failure, not an accident of software convenience.&lt;/p&gt;

&lt;p&gt;So the argument cannot be that we should bring back the cover letter and read it with an open heart. That is a request to reopen a door that was closed for a reason.&lt;/p&gt;

&lt;p&gt;The argument is narrower, and I think it survives the objection: the correction discarded a category of information instead of evaluating it consistently. Those are different problems with different fixes. Consistency is achievable for prose. Ask every candidate the same question. Score the answers against the same rubric. Read them at the same stage. That is structured data that happens to arrive as sentences.&lt;/p&gt;

&lt;p&gt;What we built instead was a process that solved the bias problem by deleting the input.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Keyword Arms Race Is a Symptom of the Deletion
&lt;/h2&gt;

&lt;p&gt;The predictable response has been optimization on both sides. Candidates tailor resumes to postings. Tools score keyword alignment. Generative models rewrite bullets in the vocabulary of the job description, while recruiters deploy increasingly sophisticated matching and ranking on the other end.&lt;/p&gt;

&lt;p&gt;None of that is unreasonable in isolation. If a company says "developer enablement" and my resume says "developer education," making that equivalence explicit helps humans and machines alike.&lt;/p&gt;

&lt;p&gt;Taken far enough, though, the process gets strange. One model rewrites a person's career into the vocabulary most likely to satisfy another model deciding whether a human being should ever see it. We have built a translation layer between two parties who could have had a conversation.&lt;/p&gt;

&lt;p&gt;And the translation runs in one direction only. Optimizing for machine legibility strips out exactly the information companies claim to be looking for. Adaptability is messy. Cross-disciplinary experience is messy. Pivots are messy. That messiness is usually where the signal is.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Hiring Process Is a Statement of What an Organization Preserves
&lt;/h2&gt;

&lt;p&gt;It is easy to blame the ATS, because it is the visible machinery. It is also the least interesting part of the story.&lt;/p&gt;

&lt;p&gt;Someone decides which fields matter. Someone writes the screening questions. Someone marks a qualification as required rather than preferred. Someone chooses how recruiters query the pool, and whether a hiring manager sees fifty candidates, ten, or only the five that most closely resemble the profile used to write the posting. And &lt;a href="https://www.kenwalger.com/blog/career/hiring-process-http-status-codes/?utm_source=kenwalger_blog&amp;amp;utm_medium=blog&amp;amp;utm_campaign=colorful_life" rel="noopener noreferrer"&gt;someone decides you will never be told which of those things happened to you&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The resulting process is not neutral. It is a record of what an organization decided was worth keeping.&lt;/p&gt;

&lt;p&gt;A process that preserves titles, years, and exact skills while discarding explanation will reliably favor careers legible through titles, years, and exact skills. A process that rewards similarity to a predefined profile will filter out candidates whose value comes from unusual combinations, because unusual combinations are by definition dissimilar.&lt;/p&gt;

&lt;p&gt;For plenty of roles that is fine. Nobody wants an airline hiring pilots because the career story was compelling.&lt;/p&gt;

&lt;p&gt;It is harder to defend when the same organization says it is looking for adaptable leaders, systems thinkers, creative problem solvers, and people who challenge established assumptions. Those are all contextual qualities. Context is the first thing lossy compression throws away.&lt;/p&gt;

&lt;h2&gt;
  
  
  We Do Not Need Better Resume Optimizers
&lt;/h2&gt;

&lt;p&gt;I have not solved hiring, and candidates still have to operate inside the system that exists. Make transferable skills explicit. Describe outcomes in the language of the role. Find a human being when you can. A nonlinear career does not become legible because we insist it deserves to be.&lt;/p&gt;

&lt;p&gt;But another generation of tools that help candidates compress themselves more efficiently is not a fix. It is better compliance with the thing causing the problem.&lt;/p&gt;

&lt;p&gt;The smaller fix is available today and costs almost nothing. One field, asked of every applicant, scored the same way: &lt;em&gt;what connects these experiences, and what did you carry from one to the next?&lt;/em&gt; Two hundred words. Read before the rejection, not after the offer. It is consistent, it is auditable, and it gives relational information somewhere to exist inside a structured process rather than outside it.&lt;/p&gt;

&lt;p&gt;If an organization genuinely believes that unconventional backgrounds and adaptable thinkers create value, its hiring process needs a place to put them. Not in the culture page. In the schema, where decisions actually get made.&lt;/p&gt;

&lt;p&gt;Otherwise, "bring your whole self to work" carries an undocumented prerequisite:&lt;/p&gt;

&lt;p&gt;First, make sure your whole self fits in the fields.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;P.S. I have also prepared my &lt;a href="https://www.kenwalger.com/resumes/Ken_Alger_Curriculum_Vitae_Latine.pdf" rel="noopener noreferrer"&gt;curriculum vitae in Latin&lt;/a&gt;. Thirty-five years of experience, zero recoverable keywords, and one very specific classics-major-turned-recruiter out there who will open it and say:&lt;/em&gt; &lt;strong&gt;hic est candidatus quem exspectabam&lt;/strong&gt;.&lt;/p&gt;

</description>
      <category>career</category>
      <category>hiring</category>
      <category>discuss</category>
      <category>jobhunting</category>
    </item>
    <item>
      <title>The Contract Discovery Bottleneck</title>
      <dc:creator>Ken W Alger</dc:creator>
      <pubDate>Thu, 10 Sep 2026 16:44:00 +0000</pubDate>
      <link>https://dev.to/kenwalger/the-contract-discovery-bottleneck-48jb</link>
      <guid>https://dev.to/kenwalger/the-contract-discovery-bottleneck-48jb</guid>
      <description>&lt;p&gt;&lt;em&gt;AI can generate the code. We can verify the behavior. But who decides&lt;br&gt;
what correct means?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I &lt;a href="https://www.kenwalger.com/blog/ai/verification-bottleneck-ai-generated-software/" rel="noopener noreferrer"&gt;wrote recently&lt;/a&gt; about a coding agent that built me a password reset&lt;br&gt;
flow with a reset link that worked more than once.&lt;/p&gt;

&lt;p&gt;The bug survived because nobody had written down that a reset link&lt;br&gt;
should be single use. It was obvious right up until it wasn't.&lt;/p&gt;

&lt;p&gt;My argument was that as AI makes implementation cheaper, verification&lt;br&gt;
becomes the bottleneck. The feature request said "build password reset."&lt;br&gt;
The agent built password reset. The happy path worked. The tests passed.&lt;br&gt;
The implementation looked finished.&lt;/p&gt;

&lt;p&gt;What nobody had asked was whether the same reset link should work twice.&lt;/p&gt;

&lt;p&gt;So I added an independently written behavioral specification. The agent&lt;br&gt;
implemented against it. The verifier rejected the reusable token. The&lt;br&gt;
agent fixed the implementation. The verifier passed it.&lt;/p&gt;

&lt;p&gt;That seemed like a useful pattern:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fwww.kenwalger.com%2Fblog%2Fwp-content%2Fuploads%2F2026%2F09%2Fmermaid-diagram-2026-09-10-081104.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fwww.kenwalger.com%2Fblog%2Fwp-content%2Fuploads%2F2026%2F09%2Fmermaid-diagram-2026-09-10-081104.png" alt="Flowchart showing three sequential stages: a specification leads to an implementation, which leads to independent verification." width="506" height="555"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Then I published the article, and the comments started finding things my&lt;br&gt;
specification didn't say. That exposed a harder problem.&lt;/p&gt;

&lt;h2&gt;The specification wasn't finished either&lt;/h2&gt;

&lt;p&gt;One reader asked what would happen if two password reset requests using&lt;br&gt;
the same token arrived at the same time.&lt;/p&gt;

&lt;p&gt;I hadn't tested that. My test covered sequential reuse:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fwww.kenwalger.com%2Fblog%2Fwp-content%2Fuploads%2F2026%2F09%2Fmermaid-diagram-2026-09-10-081312.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fwww.kenwalger.com%2Fblog%2Fwp-content%2Fuploads%2F2026%2F09%2Fmermaid-diagram-2026-09-10-081312.png" alt="Sequence diagram showing a user using a password reset token successfully, the server marking the token consumed, and a second attempt with the same token being rejected." width="800" height="647"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;But concurrent reuse is different:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fwww.kenwalger.com%2Fblog%2Fwp-content%2Fuploads%2F2026%2F09%2Fmermaid-diagram-2026-09-10-081509.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fwww.kenwalger.com%2Fblog%2Fwp-content%2Fuploads%2F2026%2F09%2Fmermaid-diagram-2026-09-10-081509.png" alt="Sequence diagram showing two concurrent password reset requests. Request A validates the token and the server reports it unused. Request B validates the same token before A has consumed it, and the server again reports it unused. Both requests then consume the token and both resets succeed, producing two successful resets from a single link." width="800" height="705"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Both requests check the token while it is still unused. Both proceed.&lt;br&gt;
If validation and consumption are not a single atomic operation, "single&lt;br&gt;
use" can still produce two successful resets.&lt;/p&gt;

&lt;p&gt;The original invariant was incomplete.&lt;/p&gt;

&lt;p&gt;That doesn't make the specification useless. It makes the specification&lt;br&gt;
provisional.&lt;/p&gt;

&lt;p&gt;The interesting part is where the new knowledge goes.&lt;/p&gt;

&lt;p&gt;Once somebody discovers that "single use" also means competing attempts&lt;br&gt;
cannot both succeed, that should stop being knowledge held by the person&lt;br&gt;
who noticed it. It belongs in the durable definition of correct&lt;br&gt;
behavior.&lt;/p&gt;

&lt;p&gt;The specification changes.&lt;/p&gt;

&lt;p&gt;Which means the loop is really closer to this:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fwww.kenwalger.com%2Fblog%2Fwp-content%2Fuploads%2F2026%2F09%2Fmermaid-diagram-2026-09-10-081649.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fwww.kenwalger.com%2Fblog%2Fwp-content%2Fuploads%2F2026%2F09%2Fmermaid-diagram-2026-09-10-081649.png" alt="Flowchart showing a cycle. Human intent produces a provisional specification, which leads to an implementation, then to independent verification. Verification surfaces a newly discovered invariant, which feeds back into the specification, and the cycle repeats." width="676" height="971"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That is messier than the first diagram.&lt;/p&gt;

&lt;p&gt;It is also much closer to engineering.&lt;/p&gt;

&lt;h2&gt;Separate tests can share the same mistake&lt;/h2&gt;

&lt;p&gt;Another reader described an integration builder where an agent wrote&lt;br&gt;
both a connector and the tests for that connector.&lt;/p&gt;

&lt;p&gt;Everything passed.&lt;/p&gt;

&lt;p&gt;Both were wrong.&lt;/p&gt;

&lt;p&gt;The connector and its tests encoded the same incorrect assumption about&lt;br&gt;
OAuth token refresh. The mistake only surfaced when a customer's token&lt;br&gt;
expired during a live session.&lt;/p&gt;

&lt;p&gt;The implementation and test suite were separate artifacts. They were not&lt;br&gt;
independent in the way that mattered.&lt;/p&gt;

&lt;p&gt;They shared an assumption.&lt;/p&gt;

&lt;p&gt;That distinction matters because "independent verification" can sound&lt;br&gt;
like an organizational property:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;different file&lt;/li&gt;
&lt;li&gt;different test suite&lt;/li&gt;
&lt;li&gt;different agent&lt;/li&gt;
&lt;li&gt;different step in the pipeline&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of those necessarily provides independence.&lt;/p&gt;

&lt;p&gt;If the implementation and verifier derive their definition of correct&lt;br&gt;
behavior from the same incomplete prompt, they can agree perfectly and&lt;br&gt;
still be wrong.&lt;/p&gt;

&lt;p&gt;The student is no longer literally grading the same exam.&lt;/p&gt;

&lt;p&gt;Two students have simply studied from the same incorrect answer key.&lt;/p&gt;

&lt;h2&gt;Generating more tests doesn't discover the missing rule&lt;/h2&gt;

&lt;p&gt;Another commenter asked whether property-based testing or giving an&lt;br&gt;
agent an adversarial security persona might do a better job uncovering&lt;br&gt;
these unstated constraints.&lt;/p&gt;

&lt;p&gt;I think both are interesting, but they expose the same boundary.&lt;br&gt;
Property-based testing can explore a stated invariant extremely well.&lt;/p&gt;

&lt;p&gt;If I tell a framework:&lt;/p&gt;

&lt;blockquote&gt;
  A successfully consumed reset token must never produce another
  successful reset.
&lt;/blockquote&gt;

&lt;p&gt;it can generate combinations and sequences I would never think to&lt;br&gt;
hand-author.&lt;/p&gt;

&lt;p&gt;But it cannot tell me that single use was a requirement if nobody&lt;br&gt;
expressed it.&lt;/p&gt;

&lt;p&gt;An adversarial agent has a similar problem. Asking a model to "try to&lt;br&gt;
break this" may produce better tests than asking it to "write tests for&lt;br&gt;
this feature." But if the adversary shares the same context, model&lt;br&gt;
assumptions, and incomplete understanding of the requirement, how&lt;br&gt;
independent is it really?&lt;/p&gt;

&lt;p&gt;The question starts shifting from &lt;em&gt;who writes the tests?&lt;/em&gt; to a more&lt;br&gt;
difficult one: &lt;strong&gt;where does the definition of correct behavior come&lt;br&gt;
from?&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;The verifier can be wrong too&lt;/h2&gt;

&lt;p&gt;One of the most interesting examples in the discussion came from a&lt;br&gt;
verification harness rather than generated application code.&lt;/p&gt;

&lt;p&gt;A capability test timed out.&lt;/p&gt;

&lt;p&gt;The harness recorded the result as a failure.&lt;/p&gt;

&lt;p&gt;But a timeout did not establish that the capability failed. It&lt;br&gt;
established that the harness did not obtain a result within the allotted&lt;br&gt;
time.&lt;/p&gt;

&lt;p&gt;Those are different claims.&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;FAILED
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;and&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;NOT TESTED
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;are not interchangeable.&lt;/p&gt;

&lt;p&gt;The verifier had turned an observation failure into an assertion about&lt;br&gt;
capability.&lt;/p&gt;

&lt;p&gt;That's a useful warning for any architecture built around deterministic&lt;br&gt;
verification: deterministic does not mean correct.&lt;/p&gt;

&lt;p&gt;A verifier can enforce the wrong invariant with absolute consistency.&lt;/p&gt;

&lt;p&gt;So can a specification.&lt;/p&gt;

&lt;p&gt;The goal isn't to replace an unreliable agent with an infallible&lt;br&gt;
verifier. There is no infallible verifier.&lt;/p&gt;

&lt;p&gt;The goal is to make the definition of correctness explicit enough that&lt;br&gt;
it can be inspected, challenged, tested, and revised independently of&lt;br&gt;
the implementation.&lt;/p&gt;

&lt;h2&gt;So who writes the contract?&lt;/h2&gt;

&lt;p&gt;This was the question that pushed the argument furthest for me.&lt;/p&gt;

&lt;p&gt;If humans have to write complete behavioral specifications before agents&lt;br&gt;
can implement anything, haven't we simply moved the bottleneck back to&lt;br&gt;
humans?&lt;/p&gt;

&lt;p&gt;Probably.&lt;/p&gt;

&lt;p&gt;And worse, the concurrency example demonstrates that humans don't&lt;br&gt;
necessarily know the complete specification beforehand either.&lt;/p&gt;

&lt;p&gt;So "humans write the contract" isn't much of an answer.&lt;/p&gt;

&lt;p&gt;An agent could propose it.&lt;/p&gt;

&lt;p&gt;That sounds circular at first. If the agent proposes the implementation&lt;br&gt;
and proposes the contract, aren't we back to the student grading the&lt;br&gt;
exam?&lt;/p&gt;

&lt;p&gt;Only if proposing the contract and accepting the contract are the same&lt;br&gt;
operation.&lt;/p&gt;

&lt;p&gt;They don't have to be.&lt;/p&gt;

&lt;p&gt;An agent might generate a candidate operating contract:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;reset token:
  may be used once
  competing attempts cannot both succeed
  expires after N minutes
  cannot authorize a different account
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;A human, another system, or some combination can then challenge that&lt;br&gt;
much smaller artifact.&lt;/p&gt;

&lt;p&gt;The question being reviewed becomes:&lt;/p&gt;

&lt;blockquote&gt;
  Is this an adequate definition of correct behavior?
&lt;/blockquote&gt;

&lt;p&gt;rather than:&lt;/p&gt;

&lt;blockquote&gt;
  Is this entire implementation correct?
&lt;/blockquote&gt;

&lt;p&gt;That doesn't solve the trust problem, but it reduces its surface area.&lt;br&gt;
Reviewing four lines is a different activity than reviewing four hundred.&lt;br&gt;
One is a conversation about intent. The other is an audit.&lt;/p&gt;

&lt;p&gt;But this runs straight back into the answer key problem.&lt;/p&gt;

&lt;p&gt;If the same model that will implement the feature also proposes the&lt;br&gt;
contract, they share assumptions. An agent that doesn't know single use&lt;br&gt;
matters won't propose single use as an invariant. It will produce a&lt;br&gt;
confident, well-formatted contract with the same hole in it, and now the&lt;br&gt;
hole has been written down and approved.&lt;/p&gt;

&lt;p&gt;So accepting a contract has to do more than approve it. It has to&lt;br&gt;
introduce something the proposing agent didn't have.&lt;/p&gt;

&lt;p&gt;That might be a person who has debugged this class of bug before. It&lt;br&gt;
might be a genuinely different model, though I'm unsure how much&lt;br&gt;
independence that buys. It might be a checklist derived from past&lt;br&gt;
incidents, which is really institutional memory in a form an agent can&lt;br&gt;
read. For a reset token, somebody's list somewhere already says: single&lt;br&gt;
use, expiry, no account substitution, no concurrent success, session&lt;br&gt;
invalidation.&lt;/p&gt;

&lt;p&gt;The value comes from the independence of the source, not from the&lt;br&gt;
ceremony of the review.&lt;/p&gt;

&lt;p&gt;That may be a more tractable thing to build tooling around than&lt;br&gt;
verification itself.&lt;/p&gt;

&lt;h2&gt;Maybe verification isn't the deepest bottleneck&lt;/h2&gt;

&lt;p&gt;This is where the comments changed my framing.&lt;/p&gt;

&lt;p&gt;I started with:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fwww.kenwalger.com%2Fblog%2Fwp-content%2Fuploads%2F2026%2F09%2Fmermaid-diagram-2026-09-10-081827.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fwww.kenwalger.com%2Fblog%2Fwp-content%2Fuploads%2F2026%2F09%2Fmermaid-diagram-2026-09-10-081827.png" alt="Flowchart showing the implementation bottleneck leading to the verification bottleneck, with a dashed arrow to a third stage labeled contract discovery, marked as an open question." width="798" height="118"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I'm less sure that's where it stops.&lt;/p&gt;

&lt;p&gt;Once implementation is cheap and verification is increasingly&lt;br&gt;
automatable, the harder problem may become discovering the invariants&lt;br&gt;
worth verifying.&lt;/p&gt;

&lt;p&gt;Call it contract discovery.&lt;/p&gt;

&lt;p&gt;The requirement says:&lt;/p&gt;

&lt;blockquote&gt;
  Reset my password.
&lt;/blockquote&gt;

&lt;p&gt;Somebody has to discover:&lt;/p&gt;

&lt;blockquote&gt;
  The link works once.
&lt;/blockquote&gt;

&lt;p&gt;Then:&lt;/p&gt;

&lt;blockquote&gt;
  Two concurrent attempts cannot both succeed.
&lt;/blockquote&gt;

&lt;p&gt;Then perhaps:&lt;/p&gt;

&lt;blockquote&gt;
  The token cannot authorize a different account.
  
  A token issued before another successful reset may no longer be valid.
  
  A reset invalidates existing sessions.
&lt;/blockquote&gt;

&lt;p&gt;Some of those are product decisions. Some are security properties. Some&lt;br&gt;
are implementation-independent behavioral invariants. Some may not apply&lt;br&gt;
at all.&lt;/p&gt;

&lt;p&gt;The difficult work is deciding which ones belong to the definition of&lt;br&gt;
correct.&lt;/p&gt;

&lt;p&gt;AI can help propose them.&lt;/p&gt;

&lt;p&gt;Property-based testing can explore them.&lt;/p&gt;

&lt;p&gt;Deterministic systems can enforce them.&lt;/p&gt;

&lt;p&gt;Production incidents will unfortunately discover some of them for us.&lt;/p&gt;

&lt;p&gt;But none of those eliminates the need to decide which claims actually&lt;br&gt;
define correctness.&lt;/p&gt;

&lt;h2&gt;This gets harder when agents start acting&lt;/h2&gt;

&lt;p&gt;There is another reason I think this matters beyond generated code:&lt;br&gt;
agents don't just write things anymore. They call things.&lt;/p&gt;

&lt;p&gt;An agent calls an API. The response is &lt;code&gt;200&lt;/code&gt;. The agent moves on.&lt;/p&gt;

&lt;p&gt;But a &lt;code&gt;200&lt;/code&gt; says the request was processed. It doesn't say the&lt;br&gt;
constraint the agent's plan depended on was enforced. Maybe the call&lt;br&gt;
timed out after the write succeeded, so the retry performed the effect&lt;br&gt;
twice. Maybe the operation was legitimate the first time and should have&lt;br&gt;
been rejected the second.&lt;/p&gt;

&lt;p&gt;That second one should look familiar. It's the reset link, one layer&lt;br&gt;
out.&lt;/p&gt;

&lt;p&gt;A bad implementation leaves an artifact somebody can inspect later.&lt;/p&gt;

&lt;p&gt;A bad tool call already happened.&lt;/p&gt;

&lt;p&gt;It sent the email. Charged the card. Revoked the access. Posted the&lt;br&gt;
message.&lt;/p&gt;

&lt;p&gt;There is no diff to read.&lt;/p&gt;

&lt;p&gt;This is where the contract-discovery problem becomes more consequential.&lt;br&gt;
The system needs some definition of what the agent is permitted to cause&lt;br&gt;
and what evidence would establish that the intended effect actually&lt;br&gt;
happened.&lt;/p&gt;

&lt;p&gt;I don't think I have the architecture for that yet, but one boundary is&lt;br&gt;
becoming clearer:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The specification can be agent-readable without being agent-owned.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The agent should be able to see the invariant. Withholding the&lt;br&gt;
requirement only makes the work guesswork.&lt;/p&gt;

&lt;p&gt;But the agent shouldn't be able to quietly redefine the invariant when&lt;br&gt;
satisfying it becomes inconvenient.&lt;/p&gt;

&lt;p&gt;Whatever accepts, stores, and evaluates the contract needs some&lt;br&gt;
independence from the reasoning that produced the implementation or&lt;br&gt;
action.&lt;/p&gt;

&lt;p&gt;Where that boundary belongs remains a harder question.&lt;/p&gt;

&lt;h2&gt;The specification is durable because it can change&lt;/h2&gt;

&lt;p&gt;Calling the specification a durable artifact can sound like calling it&lt;br&gt;
an immutable one.&lt;/p&gt;

&lt;p&gt;I don't mean that.&lt;/p&gt;

&lt;p&gt;A durable specification should change when we learn something about what&lt;br&gt;
correct behavior actually requires.&lt;/p&gt;

&lt;p&gt;What makes it durable is that the knowledge survives the implementation&lt;br&gt;
that taught us the lesson.&lt;/p&gt;

&lt;p&gt;The reset implementation may be rewritten next month.&lt;/p&gt;

&lt;p&gt;The framework may change.&lt;/p&gt;

&lt;p&gt;The agent may change.&lt;/p&gt;

&lt;p&gt;The database may change.&lt;/p&gt;

&lt;p&gt;But once we've established that two competing reset attempts cannot both&lt;br&gt;
succeed, that invariant should survive all of them.&lt;/p&gt;

&lt;p&gt;The same applies to an integration. Once a production failure teaches us&lt;br&gt;
what token refresh must guarantee, that knowledge should not remain&lt;br&gt;
attached to the incident report or the engineer who debugged it.&lt;/p&gt;

&lt;p&gt;It should become part of what "correct connector" means.&lt;/p&gt;

&lt;p&gt;The implementation may be disposable. The accumulated definition of&lt;br&gt;
correctness is not.&lt;/p&gt;

&lt;h2&gt;I still don't think this is solved&lt;/h2&gt;

&lt;p&gt;There are plenty of uncomfortable questions left.&lt;/p&gt;

&lt;p&gt;How independent does a verifier have to be?&lt;/p&gt;

&lt;p&gt;Can two agents using different prompts but the same underlying model&lt;br&gt;
provide meaningful independence?&lt;/p&gt;

&lt;p&gt;Who accepts an agent-proposed contract?&lt;/p&gt;

&lt;p&gt;How do you distinguish a genuine product invariant from an&lt;br&gt;
implementation detail that shouldn't survive the current code?&lt;/p&gt;

&lt;p&gt;What happens when two valid invariants conflict?&lt;/p&gt;

&lt;p&gt;How do contracts evolve without quietly weakening previous guarantees?&lt;/p&gt;

&lt;p&gt;And how do we verify effects in external systems where state is delayed,&lt;br&gt;
partially observable, or distributed?&lt;/p&gt;

&lt;p&gt;I don't have good answers to all of those. That's partly why I don't&lt;br&gt;
think the answer is simply "write better tests." The tests are&lt;br&gt;
downstream of the harder question.&lt;/p&gt;

&lt;h2&gt;Write down what you mean by correct&lt;/h2&gt;

&lt;p&gt;The original password-reset bug happened because a rule existed in&lt;br&gt;
someone's head and nowhere else.&lt;/p&gt;

&lt;p&gt;The comments on that experiment showed the next problem: writing down&lt;br&gt;
one rule doesn't mean you've found all the others.&lt;/p&gt;

&lt;p&gt;That's fine. The specification doesn't have to arrive complete. It has&lt;br&gt;
to provide somewhere for discovered invariants to go, and that somewhere&lt;br&gt;
has to be a place with a history: versioned, reviewable, and attached to&lt;br&gt;
the behavior rather than to the incident that revealed it.&lt;/p&gt;

&lt;p&gt;Maybe an agent proposes them. Maybe a human notices them. Maybe&lt;br&gt;
property-based testing exposes them. Maybe an independent reviewer asks&lt;br&gt;
the annoying question nobody else asked. And sometimes production will&lt;br&gt;
teach us the expensive way.&lt;/p&gt;

&lt;p&gt;The important part is that each discovery makes the durable definition&lt;br&gt;
of correct behavior better.&lt;/p&gt;

&lt;p&gt;AI is making it remarkably cheap to turn an instruction into working&lt;br&gt;
code.&lt;/p&gt;

&lt;p&gt;Verification asks whether the code did what we said. Contract discovery&lt;br&gt;
asks whether we said enough. I'm starting to think that's the harder&lt;br&gt;
problem.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>programming</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Why User Provisioning Matters for Enterprise AI</title>
      <dc:creator>Ken W Alger</dc:creator>
      <pubDate>Thu, 10 Sep 2026 12:44:00 +0000</pubDate>
      <link>https://dev.to/kenwalger/why-user-provisioning-matters-for-enterprise-ai-534m</link>
      <guid>https://dev.to/kenwalger/why-user-provisioning-matters-for-enterprise-ai-534m</guid>
      <description>&lt;p&gt;&lt;em&gt;Offboarding a user is not enough when the credentials they leave behind can still spend money and invoke tools.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;An orphaned SSO account is inert. Nobody logs in, nobody clicks anything, and eventually someone notices it during an access review.&lt;/p&gt;

&lt;p&gt;An orphaned AI credential is different.&lt;/p&gt;

&lt;p&gt;A &lt;a href="https://dev.to/kenwalger/ai-access-control-for-enterprise-ai-turning-policy-into-runtime-enforcement-5bkk"&gt;virtual key&lt;/a&gt; can still have a budget, provider and model access, and perhaps permission to invoke MCP tools against other systems. The person it belonged to may have left six months ago while the credential continues doing exactly what it was authorized to do, with nobody attached to it and nobody wondering why its usage still appears on the bill.&lt;/p&gt;

&lt;p&gt;That is why user provisioning for enterprise AI is not simply another implementation of the joiner, mover, leaver problem IAM teams have managed for decades. The identity lifecycle is familiar. The consequences of getting it wrong are not.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Should Enterprise AI Provision New Users?
&lt;/h2&gt;

&lt;p&gt;The joiner case is the easy one, which is precisely why it should not get much attention.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://dev.to/kenwalger/rbac-for-ai-governing-the-ai-control-plane-e59"&gt;previous post&lt;/a&gt; looked at how roles and access profiles connect human authorization to runtime policy. With &lt;a href="https://docs.getbifrost.ai/enterprise/user-provisioning" rel="noopener noreferrer"&gt;Bifrost user provisioning&lt;/a&gt;, an identity-provider group can map a user to a role, the role supplies its default access profile, and that profile issues governed access when the user arrives. Providers, models, budgets, rate limits, and tool access have already been decided. The new employee inherits policy instead of inventing it.&lt;/p&gt;

&lt;p&gt;That is the ideal state: nobody files a ticket for an API key or remembers which model Marketing may use. Organizational intent was encoded before the employee arrived, so provisioning mechanically applies a decision already made.&lt;/p&gt;

&lt;p&gt;Joiners are the flattering part of lifecycle management. Movers are where exceptions accumulate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Are Role Changes Dangerous for Enterprise AI Access?
&lt;/h2&gt;

&lt;p&gt;Imagine &lt;a href="https://dev.to/kenwalger/ai-governance-for-enterprise-ai-why-governance-comes-before-the-gateway-2hm1"&gt;the acquisition from the first post&lt;/a&gt; six months later. One engineer initially joins the migration team, later moves to the platform group, and eventually transfers into an internal agent project.&lt;/p&gt;

&lt;p&gt;Each move is legitimate. Each requires new access. The dangerous question is what happened to the old access.&lt;/p&gt;

&lt;p&gt;Role-based provisioning can replace assignments that originated from a role default. But a profile assigned directly to a user represents an explicit exception, so it survives the role change. The system can distinguish a role-derived assignment from a direct assignment; it cannot determine whether the reason behind that exception still exists.&lt;/p&gt;

&lt;p&gt;Combine that with the &lt;a href="https://dev.to/kenwalger/rbac-for-ai-governing-the-ai-control-plane-e59"&gt;role behavior from the previous post&lt;/a&gt;: when someone holds multiple roles, permissions resolve upward rather than to their intersection. The result is a familiar enterprise failure with an AI-specific blast radius. People gain access at every transition and shed it at none.&lt;/p&gt;

&lt;p&gt;Three teams into someone's tenure, their effective authorization may resemble the union of all three jobs they have held. That can mean access to models approved for a previous team, a larger budget, different logs, or tools their current role has no reason to invoke.&lt;/p&gt;

&lt;p&gt;Movers therefore need more than synchronization. They need reconciliation. Bifrost's &lt;a href="https://docs.getbifrost.ai/enterprise/user-provisioning" rel="noopener noreferrer"&gt;user provisioning model&lt;/a&gt; reflects role and team changes when users are reconciled against the identity provider, but lifecycle management still has to account for exceptions outside those inherited defaults. Automation is good at applying known rules and poor at deciding whether an exception created fourteen months ago still has a business justification.&lt;/p&gt;

&lt;p&gt;The question is not simply, "What access should this person's new role add?" It is also, "What access no longer has a reason to exist?" Who owns this key, why does this direct assignment still stand, and would the organization grant the same access today if asked fresh? Those are governance questions rather than provisioning mechanics, and no synchronization job will answer them.&lt;/p&gt;

&lt;p&gt;In practice, reconciliation is a scheduled comparison rather than an event handler. Once a quarter, list every access profile assignment that did not come from a role default, along with who created it, when, and against which stated justification. Anything without a justification on record is the finding. Anything with one is a question for the person who owns that team, and the question is not whether the access is being used but whether it would be granted again today.&lt;/p&gt;

&lt;p&gt;That distinction matters because usage is a misleading signal here. An unused grant is not evidence of safety; it is a capability sitting idle until an incident or an automation reaches for it. A budget that has never been spent still authorizes the spending.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Should Happen to AI Access When a User Leaves?
&lt;/h2&gt;

&lt;p&gt;Leavers expose the difference between identity management and AI lifecycle management most clearly.&lt;/p&gt;

&lt;p&gt;The obvious response to a departing employee is deletion: remove the account, revoke credentials, clean up the objects, and leave the system tidy. &lt;a href="https://dev.to/kenwalger/ai-access-control-for-enterprise-ai-turning-policy-into-runtime-enforcement-5bkk"&gt;The second post in this series&lt;/a&gt; discussed why an expired virtual key should instead fail closed without disappearing. During offboarding, that distinction becomes critical.&lt;/p&gt;

&lt;p&gt;By the time someone leaves, the system may hold an identity, roles, access-profile assignments, a virtual key, budget counters, model permissions, rate limits, tool grants, and a record of requests made under that authority. Some must stop working immediately. Some must survive.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Event&lt;/th&gt;
&lt;th&gt;Identity&lt;/th&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;th&gt;Access profile&lt;/th&gt;
&lt;th&gt;Virtual key&lt;/th&gt;
&lt;th&gt;MCP credentials&lt;/th&gt;
&lt;th&gt;Record&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Joins&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Created on first login&lt;/td&gt;
&lt;td&gt;From IdP group&lt;/td&gt;
&lt;td&gt;Role default applies&lt;/td&gt;
&lt;td&gt;Issued automatically&lt;/td&gt;
&lt;td&gt;Authorized per user&lt;/td&gt;
&lt;td&gt;Created&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Moves&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Unchanged&lt;/td&gt;
&lt;td&gt;Replaced&lt;/td&gt;
&lt;td&gt;Role defaults swap; direct assignments persist&lt;/td&gt;
&lt;td&gt;Reissued&lt;/td&gt;
&lt;td&gt;Persist separately&lt;/td&gt;
&lt;td&gt;Retained&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Leaves&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Deactivated&lt;/td&gt;
&lt;td&gt;Removed&lt;/td&gt;
&lt;td&gt;Copy retained&lt;/td&gt;
&lt;td&gt;Fails closed&lt;/td&gt;
&lt;td&gt;Revoked separately&lt;/td&gt;
&lt;td&gt;Preserved&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The last row is the architecture: &lt;strong&gt;fails closed, record preserved.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Suppose Security investigates an event six months later. The useful questions are historical: which models could this person reach, what budget applied, did they have MCP tool access, and was that access inherited or directly assigned?&lt;/p&gt;

&lt;p&gt;Deleting lifecycle objects because the employee left makes those questions harder precisely when their answers matter most. This is a &lt;a href="https://sovereignplatform.dev/terms/write-side-custody.html" rel="noopener noreferrer"&gt;write-side custody&lt;/a&gt; problem as much as an identity problem. The system should preserve enough state to explain what happened rather than reconstructing it later from an IdP, billing export, deleted credential, and today's configuration.&lt;/p&gt;

&lt;p&gt;Offboarding therefore has two obligations that pull against each other: end the authority, and preserve the evidence. A leaver should become incapable of causing new actions without becoming invisible to the historical record.&lt;/p&gt;

&lt;p&gt;That distinction matters most with tools, and there is a second object hiding here. Where an MCP server uses per-user OAuth, the departing employee holds credentials that are separate from their virtual key and have to be revoked separately. Bifrost exposes a sessions view for inspecting and revoking those per-user MCP credentials, which is easy to miss in an offboarding checklist designed before agentic AI entered the enterprise stack.&lt;/p&gt;

&lt;p&gt;It is worth being precise about why that is three operations rather than one, because they fail independently.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Deactivating the identity&lt;/strong&gt; stops someone logging in. It does not reach anything already issued under that identity. An offboarding process that ends here has closed the front door and left every key that was cut from it in circulation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Invalidating the virtual key&lt;/strong&gt; stops the requests. This is the operation most teams think of as revocation, and it is the one that actually severs runtime authority. It should fail closed while remaining inspectable, for the reasons above.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Revoking per-user MCP credentials&lt;/strong&gt; is the one most likely to be missed, because the credential is not entirely yours. Per-user OAuth means the employee authorized an external service directly, and the resulting grant lives at that service. Invalidating their virtual key stops requests flowing through the gateway. It does not, on its own, tell the external provider that the authorization behind those requests should end. That grant has to be revoked where it lives, which is why a sessions view exists as a separate surface rather than as a checkbox on the user record.&lt;/p&gt;

&lt;p&gt;The practical consequence is that "we deactivated their account" answers only the first of three questions, and the third one is the one nobody thinks to ask.&lt;/p&gt;

&lt;p&gt;The exposure is conditional rather than automatic, and the condition is the interesting part. By default Bifrost does not execute tool calls on its own; a model returns suggestions and the application has to call the execution endpoint explicitly. But a workload running in &lt;a href="https://docs.getbifrost.ai/mcp/overview" rel="noopener noreferrer"&gt;agent mode&lt;/a&gt; with auto-approval configured has no such pause. A forgotten SaaS login is dormant until somebody uses it. An orphaned key attached to an autonomous workload is already being used, and the human whose authority justified those tool grants left two quarters ago.&lt;/p&gt;

&lt;p&gt;An orphaned SSO account is inert.&lt;/p&gt;

&lt;p&gt;An orphaned virtual key is an actor.&lt;/p&gt;

&lt;p&gt;And there is a question sitting underneath all of this that the series has not answered. The record survives the person, but six months later, what does that record actually prove, and who should be allowed to read it? Removing access is only half of offboarding. The other half is being able to explain what happened while that access still existed.&lt;/p&gt;

&lt;p&gt;One practical distinction before designing around any of this. &lt;a href="https://getmax.im/githubdevto" rel="noopener noreferrer"&gt;Bifrost itself is open source&lt;/a&gt;, including the gateway and the core governance around virtual keys, budgets, routing, and MCP tool filtering. The user provisioning and identity lifecycle discussed here are Enterprise capabilities, and the &lt;a href="https://docs.getbifrost.ai/overview" rel="noopener noreferrer"&gt;documentation&lt;/a&gt; is the better place to check that boundary than a blog post.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Happens to AI Credentials That Never Had a User?
&lt;/h2&gt;

&lt;p&gt;There is one uncomfortable limit to this model: lifecycle management assumes a lifecycle.&lt;/p&gt;

&lt;p&gt;None of it helps with an identity the system never knew about. A contractor handed a key directly. A shared service account nobody quite owns. A credential minted for a pipeline in 2024 by someone who has since left. These may be the most consequential actors in the environment, because they run continuously, spend consistently, and often hold broad permissions on the theory that automation is inconvenient when it fails. Unlike a person, they never transfer teams and never trigger an offboarding workflow. You cannot deprovision an identity that never existed.&lt;/p&gt;

&lt;p&gt;The only workable answer is to give them the lifecycle events they will never generate on their own. Every non-human credential needs a named human owner, not a team alias, and a review date on the calendar. The owner is who gets asked when the credential appears in an incident. The review date is the substitute for the HR event that will never arrive. Expiry helps here more than it does for people: a workload credential that has to be renewed annually forces someone to state, once a year, that the thing still needs to exist. Most of the credentials that worry security teams would not survive that question.&lt;/p&gt;

&lt;p&gt;That leaves enterprise AI with a harder question than joiners, movers, and leavers: how do you govern actors that were never people in the first place? An agent does not join a team, does not sit in a reorganization, and does not have a last day. It has an owner who may or may not remember creating it, a budget nobody reviews, and a set of tool grants that made sense on the afternoon they were issued.&lt;/p&gt;

&lt;p&gt;That is where the lifecycle problem stops being about users and starts being about workloads.&lt;/p&gt;




&lt;p&gt;This article was commissioned by the &lt;a href="https://www.getmaxim.ai/bifrost" rel="sponsored nofollow noopener noreferrer"&gt;Bifrost team&lt;/a&gt;. The link to the &lt;a href="https://sovereignplatform.dev?utm_source=devto&amp;amp;utm_medium=footer&amp;amp;utm_campaign=ai-control-plane" rel="noopener noreferrer"&gt;Sovereign Systems Specification&lt;/a&gt; points to my own work. The architectural perspective and conclusions expressed here are my own.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>architecture</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
