<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: BekahHW</title>
    <description>The latest articles on DEV Community by BekahHW (@bekahhw).</description>
    <link>https://dev.to/bekahhw</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F345658%2Fa72b6b8b-b954-47fb-8919-ab380905f26b.jpg</url>
      <title>DEV Community: BekahHW</title>
      <link>https://dev.to/bekahhw</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/bekahhw"/>
    <language>en</language>
    <item>
      <title>Ownership, Not Origin</title>
      <dc:creator>BekahHW</dc:creator>
      <pubDate>Mon, 06 Jul 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/bekahhw/ownership-not-origin-4p9p</link>
      <guid>https://dev.to/bekahhw/ownership-not-origin-4p9p</guid>
      <description>&lt;p&gt;Let me go back to where this started. A moderator on Dev.to flagged my post and told me to disclose AI-generated work. Four essays later, I can finally say precisely what was wrong with that moment.&lt;/p&gt;

&lt;p&gt;It wasn’t that they were rude. They weren’t. It wasn’t even that they were mistaken about my process, though they were. It’s that they were asking a question with no useful answer. “Did AI touch this text” tells you nothing about whether the text deserves your trust. The moderator was auditing my tools when the only thing worth auditing was my thinking.&lt;/p&gt;

&lt;p&gt;So here is the standard I’m arguing for, as plainly as I can put it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Judge writing by ownership, not origin.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What ownership means
&lt;/h2&gt;

&lt;p&gt;Ownership isn’t a feeling. It’s a set of claims a writer makes by publishing, and it’s checkable.&lt;/p&gt;

&lt;p&gt;When I put my name on a piece, I’m claiming the ideas are mine or credited to whoever they came from. I’m claiming I verified what I stated as fact. I’m claiming I can defend every argument in it, extend it, answer the hard comment underneath it. I’m claiming that anything presented as my experience actually happened to me. And I’m accepting that if any of that turns out false, the failure is mine. No tool absorbs the blame.&lt;/p&gt;

&lt;p&gt;Notice what’s absent from that list. Nothing about which hands or machines shaped the sentences. A ghostwritten memoir can pass this test. A fully hand-typed post full of unverified claims fails it. The test tracks what readers actually care about, which is whether a mind stands behind the words.&lt;/p&gt;

&lt;p&gt;Notice also what’s demanding about it. Ownership is a higher bar than origin, not a lower one. Typing every word yourself is easy to satisfy and proves nothing. Being able to defend every word is hard, and it’s exactly what AI-era writing threatens when it goes wrong. The writer who pastes a generated draft they barely read fails the ownership test spectacularly. I’m not offering AI users an escape hatch. I’m offering everyone a stricter standard aimed at the right target.&lt;/p&gt;

&lt;h2&gt;
  
  
  The verification objection
&lt;/h2&gt;

&lt;p&gt;Now the strongest objection, the one I promised myself I’d answer honestly. A skeptic says: this is lovely, but readers can’t verify ownership any more than moderators can verify AI use. You’ve traded one unenforceable standard for another.&lt;/p&gt;

&lt;p&gt;Here’s my answer. Ownership was never verified by inspection. It’s verified by accountability over time, and we already know how that works because it’s how writing trust has always worked.&lt;/p&gt;

&lt;p&gt;You trust a writer because they respond substantively when challenged in the comments. Because their body of work holds together, one piece building on another. Because when they got something wrong last year, they said so and corrected it. Because their claims, when you happened to check one, checked out. Because the community around them has watched them think in public for years.&lt;/p&gt;

&lt;p&gt;None of that requires process forensics. All of it is visible. A writer faking ownership can survive one post, maybe several. They cannot survive sustained engagement, because you can’t defend thinking you didn’t do. The comment section finds you out. The follow-up question finds you out.&lt;/p&gt;

&lt;p&gt;Enforcement follows the same logic. Platforms can’t police process, and shouldn’t try. They can absolutely police the checkable claims. Plagiarism is detectable in the artifact. Fabricated facts are checkable. A writer who can’t engage with substantive challenges to their own post is telling you something. Those levers exist today. They’re just not being pulled, because badge-printing is easier.&lt;/p&gt;

&lt;h2&gt;
  
  
  What honest transparency looks like
&lt;/h2&gt;

&lt;p&gt;Am I against all disclosure, then? No, and this distinction matters.&lt;/p&gt;

&lt;p&gt;I share my process frequently in conversations and writing, and I plan to keep doing it. I talk about how I use AI to stress-test arguments, where it saved me hours, where it produced garbage, and I threw the garbage out. That kind of transparency teaches. It helps others develop the judgment that the stigma machine is currently preventing anyone from developing.&lt;/p&gt;

&lt;p&gt;But there’s a world of difference between transparency I offer and a confession I’m compelled to make. Offered, it’s craft talk between writers, the same as explaining my outlining habit or my editing passes. Compelled, under stigma, it’s a mark. Same information, opposite meaning. The mandate poisons the very openness it claims to want, because nothing kills honest process-sharing faster than making it an admission of guilt.&lt;/p&gt;

&lt;p&gt;If platforms want a disclosure culture, the path isn’t a checkbox. It’s destigmatization. Make process talk normal and interesting rather than incriminating, and writers will tell you more than any policy could extract.&lt;/p&gt;

&lt;h2&gt;
  
  
  Writing in the age of good tools
&lt;/h2&gt;

&lt;p&gt;I’ve been teaching people to write, in one form or another, for most of my adult life. College classrooms, developer communities, my own kids at the kitchen table. And here’s what I believe after all of it. Writing was never sacred because of the typing. It was sacred because a person committed their mind to the page and signed their name to the result.&lt;/p&gt;

&lt;p&gt;That commitment is still available to every one of us. No tool grants it and no tool revokes it.&lt;/p&gt;

&lt;p&gt;So publish the post. Use the tools that make your thinking sharper and refuse the uses that replace your thinking. Verify what you claim. Credit who you learned from. Stand in the comments and defend what you wrote. Own it, all of it, and let the origin police tire themselves out.&lt;/p&gt;

&lt;p&gt;And if you’ve been sitting on a draft because you used AI somewhere in the process and the badge scared you off, I’ll leave you with the only disclosure question that ever mattered.&lt;/p&gt;

&lt;p&gt;Is it yours?&lt;/p&gt;

&lt;p&gt;If it is, hit publish.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>writing</category>
    </item>
    <item>
      <title>The Wrong Problem</title>
      <dc:creator>BekahHW</dc:creator>
      <pubDate>Mon, 06 Jul 2026 12:00:00 +0000</pubDate>
      <link>https://dev.to/bekahhw/the-wrong-problem-37lk</link>
      <guid>https://dev.to/bekahhw/the-wrong-problem-37lk</guid>
      <description>&lt;p&gt;This story won’t resonate with everyone, but I’ve heard it enough to know that it’s uncomfortably common. You’ve published something you worked hard on. It could be a blog post, a twitter thread, a video, a course, some sort of content that you’ve put time and attention into. Then weeks or even days or hours later, you find your idea living in someone else’s post. I’m not talking about having the same idea as someone else. I’m talking about someone ripping your ideas, presenting them as their own, and taking credit for them. There’s no citation, link, or acknowledgment that you even exist.&lt;/p&gt;

&lt;p&gt;On many platforms, there’s a strong stance against plagiarism, including removal of the post, suspensions, and bans. Let’s look closely at what makes that a good policy. It targets the harm itself. It’s checkable in the artifact. And it never asks what tool was involved. Steal by hand or steal by model, the violation is identical, because the reader’s injury is identical.&lt;/p&gt;

&lt;p&gt;Now hold the AI disclosure mandate next to it. No harm requirement. Nothing checkable in the artifact, which is why enforcement runs on appearance. And the tool isn’t incidental to the rule. The tool &lt;em&gt;is&lt;/em&gt; the rule. The same platform wrote both policies, and only one of them protects anyone.&lt;/p&gt;

&lt;p&gt;That asymmetry is the tell. Platforms have decided that a machine touching your own ideas is a policing priority, while a human taking someone else’s ideas is a shrug. Whatever problem they think they’re solving, it isn’t reader betrayal, because this is what betrayal actually looks like.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually harms readers
&lt;/h2&gt;

&lt;p&gt;Set aside tools entirely and ask what makes writing bad for the person reading it.&lt;/p&gt;

&lt;p&gt;Ideas presented as original when they were lifted from someone uncredited, claims stated confidently that nobody verified, content written to game an algorithm rather than to say anything. Padding, filler, the eight-hundred-word wind-up before the answer. Recycled conventional wisdom wearing a bold headline.&lt;/p&gt;

&lt;p&gt;Every one of those predates AI. Every one of them is thriving right now, by human and machine hands alike. And every one of them is measurable in the artifact itself, no process forensics required. You don’t need to know what tool produced an unverified claim to know it’s unverified.&lt;/p&gt;

&lt;p&gt;Protecting readers from those harms would mean citation norms with teeth. Plagiarism response that actually responds. Quality signals that reward substance. And I want to be fair here, because that’s brutally hard work. Slow, expensive, full of judgment calls, and most platform teams are small, stretched thin, and staring down a content flood they didn’t create. I have real sympathy for anyone holding that backlog.&lt;/p&gt;

&lt;p&gt;A disclosure checkbox is none of those things. It ships in a sprint.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fear response
&lt;/h2&gt;

&lt;p&gt;I don’t think platform teams are cynical. I think they’re afraid, and I understand why. AI dropped the cost of producing plausible text to nearly zero, and every open publishing platform is staring at a real flood of low-effort content. Something had to be done.&lt;/p&gt;

&lt;p&gt;But look at what “something” became. The flood is a &lt;em&gt;quality&lt;/em&gt; problem. Low-effort content was always the enemy, and AI just made it cheaper to produce. The honest response is to get better at detecting low effort. That’s hard. So instead, platforms reached for a proxy. AI produces the slop, therefore mark the AI.&lt;/p&gt;

&lt;p&gt;The proxy fails in both directions. It misses the humans who were producing slop by hand all along and will keep producing it, badge-free. And it catches the writers using AI carefully in service of real thinking, who are producing exactly the substance the platform claims to want.&lt;/p&gt;

&lt;p&gt;We’ve seen this pattern before, and not just in writing. When a real problem is hard to measure, institutions measure something adjacent and easy, then defend the metric instead of the mission. Every developer who’s been judged on lines of code knows how this story goes. The metric becomes the target, the target gets gamed, and the original problem sits there untouched.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem we’re not allowed to name
&lt;/h2&gt;

&lt;p&gt;There’s one more layer, and it’s the uncomfortable one.&lt;/p&gt;

&lt;p&gt;Some of the resistance to AI-assisted writing isn’t about readers at all. It’s about writers, and about identity. If writing well took years to learn, and a tool now closes part of that gap for others, the instinct to defend the moat is human and understandable. We’ve never accepted “it threatens my position” as an argument in open source, in education, in any community worth belonging to. Gatekeeping dressed as quality control is still gatekeeping. The developers who sneered at Stack Overflow users, the photographers who sneered at digital, the writers who sneered at bloggers. History doesn’t remember the moat-defenders kindly, and it definitely doesn’t remember them as right.&lt;/p&gt;

&lt;p&gt;The question was never how to keep writing hard. It was how to keep writing honest. Those are different projects, and the disclosure mandate serves the first while claiming the second.&lt;/p&gt;

&lt;p&gt;So if tool-marking is the wrong standard, what’s the right one? What would a norm look like that catches the plagiarist and the fabricator, protects the newcomer and the craftsperson, and doesn’t care which keyboard the sentences came through?&lt;/p&gt;

&lt;p&gt;That’s the last essay, and it’s the one this whole series has been walking toward.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>writing</category>
    </item>
    <item>
      <title>The Stigma Machine</title>
      <dc:creator>BekahHW</dc:creator>
      <pubDate>Mon, 06 Jul 2026 11:00:00 +0000</pubDate>
      <link>https://dev.to/bekahhw/the-stigma-machine-4jk7</link>
      <guid>https://dev.to/bekahhw/the-stigma-machine-4jk7</guid>
      <description>&lt;p&gt;Let’s do a small experiment. Imagine two identical blog posts. They have the same ideas, structure, sentences. Everything is the same example except one has a small grey badge that says, “AI-assisted” and the other doesn’t.&lt;/p&gt;

&lt;p&gt;Which one do you read more generously? Which author do you assume worked harder? If you shared one with your team, which would it be?&lt;/p&gt;

&lt;p&gt;If you hesitate, you understand why “it’s just a label” is a fiction. A label is never just a label when it arrives carrying a verdict.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the badge actually says
&lt;/h2&gt;

&lt;p&gt;In theory, an AI disclosure badge is neutral metadata, like a timestamp or a word count. But in practice as things stand in 2026, it reads as a confession. It says this writer took a shortcut. It says “be skeptical of this one.” It says the person behind these words is a little less of a writer than the ones without the badge, and depending on who’s reading, maybe a little less honest too.&lt;/p&gt;

&lt;p&gt;You don’t have to take my word for the verdict part, because sometimes the platform says it out loud. On Dev.to, the disclosure tag for AI-assisted articles is #ABotWroteThis. Assisted, not just generated. For reference, when I put my post into ZeroGPT, it’s identified as human-written except 6.2%.&lt;/p&gt;

&lt;p&gt;Develop the ideas yourself, verify every claim, edit for hours, and if a model helped anywhere in your process, the prescribed label on your work reads &lt;em&gt;a bot wrote this&lt;/em&gt;. That’s not metadata. That’s a verdict with a hashtag. And if you decline to apply it, the guidelines note that moderators can attach it to your post for you. The platform reserves the right to put words on your work that you believe are false.&lt;/p&gt;

&lt;p&gt;That’s not what disclosure advocates intend. Intent doesn’t matter here. Stigma lives in reception, not intention, and right now the reception is unambiguous. Writers know it, which is why disclosure feels less like transparency and more like being asked to pin a note to your own work saying “discount this.”&lt;/p&gt;

&lt;p&gt;So the mandate creates an impossible choice for exactly the writers doing things right. Disclose honestly and watch readers disengage before evaluating a single idea. Or stay silent and risk being flagged, accused, and moderated.&lt;/p&gt;

&lt;p&gt;The one on my post spelled it out. Comply, or admins may lower your post’s score so fewer people see it, or unpublish it altogether. The badge suppresses your work socially. Refusing it suppresses your work algorithmically. There’s no door out of that room.&lt;/p&gt;

&lt;p&gt;The careless never face this choice. They just don’t disclose. Stigma-backed mandates punish honesty and reward shamelessness. That’s not a side effect. That’s the predictable mechanics of the design.&lt;/p&gt;

&lt;h2&gt;
  
  
  The accusation engine
&lt;/h2&gt;

&lt;p&gt;And the enforcement side is worse, because AI use isn’t visible. A moderator cannot see your process. They can’t see the weeks of thinking, the validated research, the edit passes. All they can inspect is the artifact, so enforcement inevitably becomes inference. Does this prose &lt;em&gt;feel&lt;/em&gt; generated?&lt;/p&gt;

&lt;p&gt;Think about what that selects for. Clean structure, polished transitions, consistent tone, error-free grammar. The signals moderators and detection tools treat as machine-like are the same signals writers spend years learning to produce. We’ve built a system where craft itself is evidence against you.&lt;/p&gt;

&lt;p&gt;I felt this personally when my own post got flagged. But the people I worry about aren’t established writers with communities who’ll vouch for them. It’s the newer writers. The developer publishing their third post, still unsure they belong, told by a moderator that their work doesn’t look like their own. Some of them will add the badge to avoid trouble. Some will stop publishing. I’ve spent years telling those exact people to hit publish. This system tells them something different.&lt;/p&gt;

&lt;h2&gt;
  
  
  The best case for the other side
&lt;/h2&gt;

&lt;p&gt;Now the part where I argue against myself, because this series doesn’t deserve to survive if it can’t.&lt;/p&gt;

&lt;p&gt;The strongest argument for disclosure isn’t about quality. It’s about trust. Readers form relationships with writers. When you read someone’s personal essay about burnout, part of what moves you is the belief that a human lived it and a human shaped it. If you later learned a machine wrote it, you’d feel betrayed, and that feeling would be legitimate. Deception through implied authorship is real, and “caveat lector” is a cold answer to it.&lt;/p&gt;

&lt;p&gt;I think that argument is right. I want to be honest about that.&lt;/p&gt;

&lt;p&gt;Where it goes wrong is the leap from “deception is bad” to “tool disclosure prevents it.” The betrayal in that scenario isn’t the tool. It’s the false implicit claim. A writer who presents machine-generated experiences as lived ones is lying about something specific, the same way a writer who fabricates a story by hand is lying. We already have a name for that, and it isn’t “AI use.” Meanwhile the writer who used AI to sharpen an essay about their own real burnout hasn’t deceived anyone about anything a reader actually cares about.&lt;/p&gt;

&lt;p&gt;The mandate can’t tell those two writers apart. The badge marks both, deters the honest one, and does nothing to the liar, who was already comfortable lying.&lt;/p&gt;

&lt;p&gt;The truth is, trust between writers and readers is worth protecting. That’s exactly why we can’t outsource it to a checkbox that measures the wrong thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What stigma costs us
&lt;/h2&gt;

&lt;p&gt;Here’s what keeps me up about this. AI, used well, can make writing &lt;em&gt;better&lt;/em&gt;. It can push you to consider the counterargument you were avoiding. It can surface the source you’d never have found. It can free the hours you were spending on mechanical polish and give them back to thinking. Used badly, it produces exactly the formulaic sludge everyone fears. The difference is entirely in the writer.&lt;/p&gt;

&lt;p&gt;Stigma erases that difference. It teaches a generation of writers that the tool is shameful rather than teaching them what shameful use of the tool looks like. We’re not developing judgment. We’re developing secrecy.&lt;/p&gt;

&lt;p&gt;And while all this energy pours into marking who touched which tool, the problems that actually betray readers go unpoliced. That’s the next essay. Because the most damning thing about the stigma machine isn’t what it does. It’s what it lets platforms avoid doing.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>writing</category>
    </item>
    <item>
      <title>The Disclosure Double Standard</title>
      <dc:creator>BekahHW</dc:creator>
      <pubDate>Mon, 06 Jul 2026 10:00:00 +0000</pubDate>
      <link>https://dev.to/bekahhw/the-disclosure-double-standard-3kkd</link>
      <guid>https://dev.to/bekahhw/the-disclosure-double-standard-3kkd</guid>
      <description>&lt;p&gt;For most people, writing a blog post isn’t solitary work. It’s talking to your friend over coffee and the question they ask that changes your perspective and leads you to reframe the whole thing. Then there’s the spellchecker and grammarly that give you one-click fixes and improve your voice. If you’re lucky enough, you have an editor or a teammate or colleague willing to read it over and give you feedback that leads to your third or fourth or fifth revision. And none of them appears in your byline. You’ve never been asked to disclose a single one.&lt;/p&gt;

&lt;p&gt;But now that AI has entered the arena, you’re expected to wear a badge.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rule we never had
&lt;/h2&gt;

&lt;p&gt;Ghostwriting is the clearest tell. For a century, public figures have published books they didn’t write. Not books they wrote with help. Books where someone else produced every sentence, and the named author’s contribution was a series of interviews and a final approval. The industry is respectable, professionalized, and almost entirely undisclosed. Readers buy the memoir, connect with the story, and nobody calls it fraud.&lt;/p&gt;

&lt;p&gt;Let’s be clear about what that means. Our writing culture has already decided that full delegation of the actual writing is acceptable, as long as the ideas and the accountability belong to the named author.&lt;/p&gt;

&lt;p&gt;So the standard was never “you must write every word yourself.” It couldn’t have been. Editors restructure. Translators rewrite entirely. Co-authors merge beyond untangling. Writing has always been more collaborative than the solitary-genius myth admits, and readers have always been fine with it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The disclosure mandate invents a purity standard that never existed, then applies it to exactly one tool.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The label that can’t label
&lt;/h2&gt;

&lt;p&gt;Even if we granted that AI is different in kind, the mandate would still fail on its own terms, because “AI-assisted” doesn’t describe anything.&lt;/p&gt;

&lt;p&gt;Consider what fits under that one flag. Fixing typos. Brainstorming titles. Summarizing a paper you then read yourself. Arguing with a chatbot to stress-test your thesis. Generating an outline. Generating a draft you rewrite completely. Generating a draft you don’t.&lt;/p&gt;

&lt;p&gt;Those are wildly different acts with wildly different implications for the reader. Some of them involve less delegation than hiring an editor. One of them is closer to ghostwriting. A single checkbox flattens all of it into one undifferentiated confession.&lt;/p&gt;

&lt;p&gt;A label that carries no information isn’t disclosure. It’s ritual.&lt;/p&gt;

&lt;h2&gt;
  
  
  We’ve been here before
&lt;/h2&gt;

&lt;p&gt;Every writing tool arrives to the same funeral music.&lt;/p&gt;

&lt;p&gt;Photography was going to kill painting, and for decades photographers fought to be considered artists at all, because the machine did the work. Word processors were going to ruin prose, because revision would become too easy and writers would stop thinking before typing. Calculators were going to destroy mathematical minds. Spellcheck was going to breed illiterates.&lt;/p&gt;

&lt;p&gt;Each time, the panic confused the tool with the thinking. Each time, we eventually figured out that the craft didn’t live in the mechanical layer. Painting survived because painting was never just image production. Writing survived the word processor because writing was never just typing.&lt;/p&gt;

&lt;p&gt;I don’t say this to wave away every concern about AI. I recognize that AI can hollow out writing in a way spellcheck can’t, if you let it do your thinking instead of your typing. That distinction matters enormously. It’s the difference between a tool and a replacement.&lt;/p&gt;

&lt;p&gt;But notice what that distinction is. It’s not about which tool touched the text. It’s about who did the thinking.&lt;/p&gt;

&lt;p&gt;Which is exactly what a disclosure checkbox cannot capture.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the double standard reveals
&lt;/h2&gt;

&lt;p&gt;Why does one tool get a mandate when a century of ghostwriting got a shrug?&lt;/p&gt;

&lt;p&gt;Because the mandate was never derived from a principle. If platforms had started from “readers deserve to know how work gets made,” they’d have built process transparency for everything, and the absurdity would have been obvious immediately. Imagine the disclosure form: Did you discuss ideas with your spouse? Did you read the competitor’s post first? Did an editor rewrite the conclusion?&lt;/p&gt;

&lt;p&gt;Nobody wants that, because readers never needed it. What readers needed was a writer who stood behind the work.&lt;/p&gt;

&lt;p&gt;There’s even a version of this buried in my own story, and nobody flagged it. I found it myself, reading the guidelines the moderation comment linked. My post went out under my personal account and talked about how one of the features we just launched helps to fill a gap teams face when it comes to writing skills. As far as I understand, that’s allowed on the platform. Company blogs share product content there every day. But the AI guidelines add a special rule. AI-assisted articles shouldn’t promote any business, program, or course, including your own. Trace the logic. A human-drafted post can discuss the product it’s about. The identical post with AI anywhere in the process cannot. Same words, same reader, different rules, and the only variable is the tool. That’s not reader protection. That’s a purity test.&lt;/p&gt;

&lt;p&gt;The AI mandate exists because AI is new and frightening, and new fears demand visible responses. It’s not a standard. It’s a nervous gesture wearing a standard’s clothes.&lt;/p&gt;

&lt;p&gt;What if we asked the question the mandate skips? Not “did a machine touch this text” but “did a mind own it.” That question has an answer worth knowing. The next essay is about what happens when we settle for the wrong question instead, because the cost isn’t hypothetical. It has a name, and the name is stigma.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>writing</category>
    </item>
    <item>
      <title>Why We Write</title>
      <dc:creator>BekahHW</dc:creator>
      <pubDate>Mon, 06 Jul 2026 00:00:00 +0000</pubDate>
      <link>https://dev.to/bekahhw/why-we-write-50hj</link>
      <guid>https://dev.to/bekahhw/why-we-write-50hj</guid>
      <description>&lt;p&gt;Last week, one of my blog posts was flagged on Dev.to.&lt;/p&gt;

&lt;p&gt;I want to be clear that this series is not specifically about that event or the people involved. It’s a broader exploration of using AI during the writing process and rules around labeling any AI-assisted writing. This just happens to be a relevant example. Everyone was doing their job exactly the way the policy asked. That’s the point. A system can be staffed entirely by decent people and still be built wrong.&lt;/p&gt;

&lt;p&gt;Did I use AI in the process? Sure, I have been almost since ChatGPT came out. It’s been my brainstorm partner, my editor, my devil’s advocate, research assistant, and more. But what’s remained the same in all of my work is that the ideas were mine, they’re driven by my own experiences, and the desire to share and teach others. I validated every piece of research myself, edited draft after draft, and shaped the whole thing into something I was proud to publish. Hours went into it. This wasn’t a drive-by prompt-and-paste. That’s not my M.O.&lt;/p&gt;

&lt;p&gt;The irony still gets me. &lt;a href="https://papercompute.com/blog/stop-writing-skills-from-memory/" rel="noopener noreferrer"&gt;The post&lt;/a&gt; argued that reconstructing work from memory is lossy and that what you think happened is missing exactly what matters. That evidence beats confident guessing. None of the work I did is visible to the moderator. They had a feeling about my prose or it was flagged by some tool and a policy that told them to act on it.&lt;/p&gt;

&lt;p&gt;I’ve spent the last 7 years helping developers find their people and their voice online. I’ve told hundreds of new writers that their perspective matters, that they should hit publish, that the community will meet them with generosity. So when a community I’ve championed looked at my work and saw a suspect instead of a writer, I didn’t just feel annoyed. I felt something closer to grief.&lt;/p&gt;

&lt;p&gt;But this isn’t a post about my hurt feelings. It’s a post about a question we’ve stopped asking, and before moving into tech, I taught for years in my College English 101 classes.&lt;/p&gt;

&lt;p&gt;Why do we write at all?&lt;/p&gt;

&lt;p&gt;If we strip writing down to first principles, it does three things.&lt;/p&gt;

&lt;p&gt;We write to &lt;em&gt;communicate&lt;/em&gt;. To take something that lives in your mind and make it live in another.&lt;/p&gt;

&lt;p&gt;We write to &lt;em&gt;connect&lt;/em&gt;. Every blog post is a hand extended to a stranger. Someone reads your story when they’re going through the same thing and feels less alone.&lt;/p&gt;

&lt;p&gt;And we write to &lt;em&gt;think&lt;/em&gt;. Writing isn’t the transcript of thought. It’s the act of it. Anyone who has started an essay believing one thing and finished it believing another knows this.&lt;/p&gt;

&lt;p&gt;Those are the ends writing serves. Which means they’re also the standard every writing rule should be judged against. Does this norm help ideas move between minds? Does it help people find each other? Does it deepen thinking?&lt;/p&gt;

&lt;h2&gt;
  
  
  The rule that fails its own test
&lt;/h2&gt;

&lt;p&gt;A disclosure mandate doesn’t help communication. The label “AI-assisted” tells a reader almost nothing because it collapses an entire spectrum of uses into a single flag. Spellcheck, brainstorming, research support, and full generation all wear the same badge. A label that can’t distinguish between them isn’t information.&lt;/p&gt;

&lt;p&gt;It doesn’t help connection either. Right now that label carries a stigma. It marks a writer as lesser, maybe even dishonest, before a single idea gets evaluated. Readers are trained to disengage the moment they see it. That’s not transparency building trust.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;That’s a scarlet letter dressed up as a courtesy.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And it certainly doesn’t deepen thinking, because the platforms enforcing it have never asked about any of the other hands in our work. Nobody discloses Grammarly. Nobody discloses the friend who talked through the argument with them over coffee, or the editor who restructured the whole second half. Ghostwriting, where someone else writes every word, has been a respectable industry for a century. No badge required.&lt;/p&gt;

&lt;p&gt;So the mandate isn’t really about protecting readers. If it were, it would target the things that actually betray them. Plagiarism. Ideas lifted without citation. Confident claims nobody checked. Those problems are rampant on the same platforms, and they go largely unpoliced while moderators guess at which sentences feel too polished.&lt;/p&gt;

&lt;p&gt;This is a fear response, not a principle. AI made platforms afraid, and afraid institutions reach for visible rules over meaningful ones. A disclosure badge lets a platform look responsible without doing the harder work of caring about quality.&lt;/p&gt;

&lt;h2&gt;
  
  
  The test we should be using
&lt;/h2&gt;

&lt;p&gt;You might argue that readers deserve to know how the sausage gets made. You’re not entirely wrong. Readers do deserve honesty. But honesty about &lt;em&gt;what&lt;/em&gt;?&lt;/p&gt;

&lt;p&gt;Not a tool inventory. Readers have never had one and never needed one. What readers deserve is a writer who stands behind the work. Who did the thinking. Who checked the claims. Who can defend every idea on the page because the ideas are actually theirs, regardless of which tools helped shape the sentences.&lt;/p&gt;

&lt;p&gt;That post the moderator flagged? I own every word of it. I can defend every claim in it. That’s what my hours bought, and no label can add to it or take it away.&lt;/p&gt;

&lt;p&gt;The test isn’t about origin. It’s about ownership.&lt;/p&gt;

&lt;p&gt;That’s the claim this series exists to defend. In the essays ahead, I’ll dig into the double standard we’ve built around writing tools, the stigma machine that disclosure mandates power, the real quality problems platforms keep ignoring, and what a standard built on ownership would actually look like.&lt;/p&gt;

&lt;p&gt;For now, I’ll leave you with the questions I keep returning to. When you read something that moves you, what were you actually trusting? The process, or the person? And if a rule can’t tell the difference between a writer who did the work and one who didn’t, what exactly is it protecting?&lt;/p&gt;

&lt;p&gt;I don’t think we’ve answered that yet. I think it’s time we did.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>writing</category>
    </item>
    <item>
      <title>AI Companies Know Your Data Is Valuable. Why Doesn't Your Team?</title>
      <dc:creator>BekahHW</dc:creator>
      <pubDate>Wed, 01 Jul 2026 06:02:00 +0000</pubDate>
      <link>https://dev.to/bekahhw/ai-companies-know-your-data-is-valuable-why-doesnt-your-team-13ma</link>
      <guid>https://dev.to/bekahhw/ai-companies-know-your-data-is-valuable-why-doesnt-your-team-13ma</guid>
      <description>&lt;p&gt;The most valuable thing your team produces with AI isn't the code that shipped. It's the path the agent took to write it — the prompts that worked, the retries that didn't, the architecture decision someone made out loud before they touched the keyboard. That path is not exhaust. It is prior experience. Frontier labs have spent the last year paying a billion dollars a year to prove the point. Most engineering teams are deleting that path every time the terminal closes.&lt;/p&gt;

&lt;p&gt;The vendor wants your session for training. The engineer means to save it for next time. The team needs to save it to learn from each other.&lt;/p&gt;

&lt;h2&gt;
  
  
  The labs already named what the path is worth
&lt;/h2&gt;

&lt;p&gt;OpenAI COO Brad Lightcap put a number on the static-text era: if you combined all the proprietary text from major publishers and added it to GPT-4's training mix, &lt;a href="https://foundationcapital.com/metas-bet-on-scale-the-new-ai-data-paradigm/" rel="noopener noreferrer"&gt;"it would boost the data volume by less than 0.1%."&lt;/a&gt; Buying corpora is effectively over. What labs are paying for now is what Foundation Capital calls &lt;em&gt;agency&lt;/em&gt; — "interactive reinforcement-learning environments with sequential action traces, not static input-output pairs." That's a long phrase for a simple idea. &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The path through a problem is the asset.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Recent research is starting to point in the same direction. The systems that improve agent performance aren't just keeping better prompts around. They are preserving prior runs — source changes, execution traces, prompts, tool calls, scores, failures, and state — so future agents can search across what happened before instead of starting from a blank context window. The lesson isn't “stuff more into the prompt.” It's: make prior experience durable enough to query.&lt;/p&gt;

&lt;p&gt;Mercor, the marketplace that sells those traces to OpenAI, Anthropic, and six of the Magnificent Seven, &lt;a href="https://bigthink.com/business/inside-the-meteoric-rise-of-mercor/" rel="noopener noreferrer"&gt;crossed $1B in annualized revenue in June&lt;/a&gt; at a $10B valuation. The product is a session record — the full trajectory of a senior engineer and an AI agent working through a hard problem together. Surge AI, its closest competitor, named the thing right on the box: &lt;a href="https://surgehq.ai/products" rel="noopener noreferrer"&gt;"thousands of hours of expert reasoning, pre-built and ready to use today, spanning RL environments, coding, and the core capabilities every model needs."&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;And the model vendors are productizing it themselves. Anthropic and OpenAI are reportedly &lt;a href="https://www.digitaltoday.co.kr/en/view/61389/anthropic-openai-shift-from-model-race-to-ai-coding-lock-in-strategy" rel="noopener noreferrer"&gt;building services that store "the entire conversation record of the code-writing process, or trajectory,"&lt;/a&gt; so users can later look up the intent behind why code was written a certain way. GitHub flipped Copilot Free, Pro, and Pro+ to &lt;a href="https://github.blog/news-insights/company-news/updates-to-github-copilot-interaction-data-usage-policy/" rel="noopener noreferrer"&gt;collect interaction data by default in April&lt;/a&gt; — "inputs, outputs, code snippets, and associated context" — used to train models unless you opt out.&lt;/p&gt;

&lt;p&gt;The labs already named the thing. They called it trajectory. They're paying for it, building products around it, and harvesting it from the tools your team uses every day.&lt;/p&gt;

&lt;h2&gt;
  
  
  Meanwhile, in your terminal
&lt;/h2&gt;

&lt;p&gt;Your team's sessions already contain the trajectory. Every prompt, every tool call, every retry, every fix. The same artifact Mercor sells at $200 an hour is sitting in your engineers' shells today, and the moment they close the terminal, it's gone.&lt;/p&gt;

&lt;p&gt;Capture changes what that artifact can do. &lt;span&gt;&lt;a href="https://dev.to/concepts/ai-session-capture/"&gt;A captured session record&lt;/a&gt;&lt;/span&gt; is searchable. It's &lt;span&gt;&lt;a href="https://dev.to/concepts/agent-session-replay/"&gt;replayable&lt;/a&gt;&lt;/span&gt;. It can be compared against the next run. It can show which prompt sent the agent down a dead end, which tool call recovered the task, and which context actually mattered. It turns yesterday’s one-off debugging session into &lt;span&gt;&lt;a href="https://dev.to/concepts/continuous-agent-improvement/"&gt;tomorrow’s starting point&lt;/a&gt;&lt;/span&gt;. It's something the next engineer can pull up when they hit the same Kafka config error your senior engineer solved at 11pm last Tuesday instead of opening a fresh prompt and walking through the dead ends one more time.&lt;/p&gt;

&lt;p&gt;Without capture, the whole industry is admitting that the workflow trace is too fragile to lean on. Anthropic itself published a &lt;a href="https://www.infoq.com/news/2026/05/anthropic-claude-code-postmortem/" rel="noopener noreferrer"&gt;postmortem in May&lt;/a&gt; about a bug that quietly wiped session context every turn for nearly four weeks before anyone caught it. The vendor that builds the tool admits intra-session memory is fragile.&lt;/p&gt;

&lt;p&gt;O'Reilly Radar put the daily texture &lt;a href="https://www.oreilly.com/radar/your-ai-agent-already-forgot-half-of-what-you-told-it/" rel="noopener noreferrer"&gt;bluntly&lt;/a&gt;: "Your AI partner — who just spent 20 minutes understanding your codebase — forgets everything and starts suggesting the same wrong approaches you already rejected."&lt;/p&gt;

&lt;p&gt;Mark Dominus &lt;a href="https://blog.plover.com/2026/03/09/" rel="noopener noreferrer"&gt;wrote in March&lt;/a&gt; about how he now asks Claude to write a structured summary at the end of every project and commits it to the repo manually, because if he doesn't scrape it out himself, the workflow knowledge is gone. He noted, with appropriate dryness, that developers will document for Claude what they won't document for each other.&lt;/p&gt;

&lt;p&gt;JetBrains researchers &lt;a href="https://blog.jetbrains.com/research/2026/04/ai-impact-developer-workflows/" rel="noopener noreferrer"&gt;reported in April&lt;/a&gt; on telemetry from roughly 800 developers: context-switching trended steadily upward in the AI-assisted workflow, and 74% of those developers didn't notice the increase. AI doesn't uniformly reduce developer effort, the team argued. It redistributes it into more fragmented, reactive work that doesn't show up on anyone's calendar.&lt;/p&gt;

&lt;p&gt;If you and the engineer next to you have each solved the same flaky-test config error more than once, you don't have a model problem. You have a missing archive.&lt;/p&gt;

&lt;h2&gt;
  
  
  This isn't a memory problem. It's a capture problem.
&lt;/h2&gt;

&lt;p&gt;Memory is what the agent can still see. Capture is &lt;span&gt;&lt;a href="https://dev.to/concepts/team-shared-agent-knowledge/"&gt;what the team can still use&lt;/a&gt;&lt;/span&gt;. I'm not running data licensing strategy at a frontier lab. But the shape of it is clear enough from the outside. Vendor, engineer, team — none of those three groups has agreed on where the session lives, what format it's in, or who owns it.&lt;/p&gt;

&lt;p&gt;So most of the time, nobody saves anything. The session ends. The terminal closes. The next person on the team starts from zero.&lt;/p&gt;

&lt;h2&gt;
  
  
  What capture as a primitive looks like
&lt;/h2&gt;

&lt;p&gt;Distributed systems solved a version of this in the 1980s. Every production database writes a write-ahead log. Every message queue checkpoints offsets. Every event-sourced system can replay its history from an append-only record. The assumption is that the running process will fail, and the next one needs to know what the last one did.&lt;/p&gt;

&lt;p&gt;Agent workflows ignore those lessons. The model talks to the provider. The provider responds. The session closes. An hour of design work, dozens of tool calls, three debugging loops, and one extremely specific lesson about how a flaky integration test recovers — all of it evaporates when the user moves on.&lt;/p&gt;

&lt;p&gt;Capture, as a primitive, is the thing that stops the evaporation. &lt;span&gt;&lt;a href="https://dev.to/blog/agents-need-black-box-recorders/"&gt;An append-only record of the work&lt;/a&gt;&lt;/span&gt;: every prompt, every response, every tool call, every retry, every fork in reasoning. Owned by the team, not the vendor. Queryable later, by humans and agents. An archive, not a scrollback buffer.&lt;/p&gt;

&lt;p&gt;That’s the bet Paper Compute is making. Agent sessions are not disposable chat logs. They are the operational record of AI-assisted work. Capture them, and your team can inspect, compare, search, and eventually reuse what your best engineers and agents have already figured out.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The asset isn't the code the agent produced. The asset is the path the agent took.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Start treating the session like it's worth something
&lt;/h2&gt;

&lt;p&gt;Capture your sessions. Tell the people you work with that what they figured out last Tuesday is worth saving — because Mercor is renting senior engineers at $200 an hour to record exactly that, and your team is doing it for free.&lt;/p&gt;

&lt;p&gt;The labs already agree on what your team's work is worth. Now your team has to start treating it like infrastructure.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Prompt Caching Is Subsidizing Bad AI Architecture</title>
      <dc:creator>BekahHW</dc:creator>
      <pubDate>Mon, 29 Jun 2026 18:01:47 +0000</pubDate>
      <link>https://dev.to/bekahhw/prompt-caching-is-subsidizing-bad-ai-architecture-310m</link>
      <guid>https://dev.to/bekahhw/prompt-caching-is-subsidizing-bad-ai-architecture-310m</guid>
      <description>&lt;p&gt;Brian's recent post, &lt;a href="https://papercompute.com/blog/true-cost-of-claude-code/" rel="noopener noreferrer"&gt;the true cost of Claude Code&lt;/a&gt;, looked at the subsidy under modern AI coding tools. I wanted to see what the subsidy was actually paying for inside my own workflows, so I pulled nineteen days of Claude Code session data from my &lt;a href="https://papercompute.com/docs/paper/" rel="noopener noreferrer"&gt;paper CLI&lt;/a&gt; session data: 3,697 messages across two projects and three models (Opus, Sonnet, and Haiku). Because every request was captured at the HTTP boundary, I had the model used, the token counts, the cache reads and writes, and the way each message connected to the ones before it. That's more than a bill or a single transcript can show.&lt;/p&gt;

&lt;p&gt;Prompt caching saved 82% of my input cost. The savings are real, but they also make a particular kind of architectural waste cheap enough to ignore: prompts that grow without limit, context blocks attached to the wrong tasks, sub-agents that rewrite their parent's cache prefix. Caching subsidizes that architecture by burying it inside an aggregate number that looks healthy. The only way to see what is being subsidized is to look at the workflow underneath: which prompts stayed stable, which kept mutating, where parents forked into children, and whether the reused context actually contributed to useful work.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Prompt caching is just the observable surface of a deeper system: AI workflows behave as stateful, branching systems, and current telemetry only sees them as isolated calls.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  AI workflows have outgrown traditional telemetry
&lt;/h2&gt;

&lt;p&gt;A few years ago, "an AI call" was a single request. One prompt in, one response out. Logs, traces, and billing dashboards were enough to reason about it.&lt;/p&gt;

&lt;p&gt;That world is over. Modern agent workflows are different on five axes at once: they run for a long time, they keep state across turns, they branch into sub-tasks, they call other agents, and they rewrite their own prompts as they go. A single coding session can spawn sub-agents, read dozens of files, run tools, and end up with a prompt that bears little resemblance to where it started. The individual model call is the cheap, well-instrumented part. The workflow around it is what decides whether the work was efficient, durable, or wasteful. Traditional telemetry doesn't see any of it.&lt;/p&gt;

&lt;p&gt;Request logs show one request at a time. Distributed traces show service-to-service flows, not prompt-to-prompt lineage. Billing dashboards aggregate everything into a total. None of those answer the questions that matter once a workflow is the unit of analysis:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which prompts stay stable across a session, and which keep mutating?&lt;/li&gt;
&lt;li&gt;When a parent agent spawns a sub-agent, does the child inherit the parent's prompt prefix or build its own?&lt;/li&gt;
&lt;li&gt;Where does context branch, and where does it merge?&lt;/li&gt;
&lt;li&gt;Which sessions paid to create cache that never paid back?&lt;/li&gt;
&lt;li&gt;How is a workflow's prompt shape drifting week over week?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These questions need different primitives: prompt lineage, context inheritance, cache topology, prompt mutation, sub-agent divergence, and session structure tracked over time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Economic vs. architectural efficiency
&lt;/h2&gt;

&lt;p&gt;Cheap reuse is not the same as good architecture.&lt;/p&gt;

&lt;p&gt;Economic efficiency asks if we reuse tokens cheaply. Architectural efficiency asks if this context should have been reused at all. Those are different questions, and prompt caching only answers the first.&lt;/p&gt;

&lt;p&gt;A workflow can be economically efficient and architecturally wasteful. It can also be architecturally clean and economically expensive when the work itself doesn't repeat much. Confusing the two has concrete operational costs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Runaway context growth.&lt;/strong&gt; Prompts pick up rules and notes faster than anyone removes them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hidden cost amplification.&lt;/strong&gt; Wasteful workflows look cheap while caching is generous and turn expensive the moment the discount weakens.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prompt entropy.&lt;/strong&gt; Sub-agents and projects drift apart slowly, and cache reuse degrades with them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Degraded cache reuse.&lt;/strong&gt; A high aggregate hit rate can hide sessions that keep rewriting their own prefix.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hard to debug.&lt;/strong&gt; When two sessions act differently, there is no way to compare what changed in the prompts that drove them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hard to compare architectures.&lt;/strong&gt; No way to ask whether one team's agent design reuses context better than another's.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of those show up on a bill. They show up in the topology of sessions captured over time.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the data showed
&lt;/h2&gt;

&lt;p&gt;Across nineteen days, my Claude Code usage hit 102.6 million input tokens with a 93.0% cache hit rate: 95.5M cache reads priced at 10% of base input, 5.0M cache writes priced at 125% of base input, and 2.1M fresh tokens at full price. Here and throughout, "cache hit rate" means the share of input tokens served as cache reads, computed across the whole nineteen-day window. That works out to about 17.9M base-equivalent tokens against a nominal 102.6M, an 82% savings on input.&lt;/p&gt;

&lt;p&gt;In dollars, actual input cost was $232 against a no-cache equivalent of $1,321, a saving of $1,089. The no-cache equivalent would have been about 6.6x my $200/month Claude Code Max plan price, so the cache subsidy and the plan subsidy compound on top of each other.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Actual cost&lt;/th&gt;
&lt;th&gt;No-cache cost&lt;/th&gt;
&lt;th&gt;Saved&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Opus 4.6&lt;/td&gt;
&lt;td&gt;$220&lt;/td&gt;
&lt;td&gt;$1,274&lt;/td&gt;
&lt;td&gt;$1,053 (82.7%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sonnet 4.6&lt;/td&gt;
&lt;td&gt;$8&lt;/td&gt;
&lt;td&gt;$41&lt;/td&gt;
&lt;td&gt;$33 (80.9%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Haiku 4.5&lt;/td&gt;
&lt;td&gt;$4&lt;/td&gt;
&lt;td&gt;$6&lt;/td&gt;
&lt;td&gt;$3 (42.3%)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Dollars rounded to the nearest whole. Percentages computed on the unrounded underlying values, so a few rows may not visibly cross-foot.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The bill tells me what I spent, but it can't tell me which sessions, which models, or which prompt shapes were responsible for the ratio, or whether any of that work was reusable in the first place. The next-most-aggregated number is the cache hit rate. It tells you how much input was reused. It can't tell you whether the reused context deserved to be there.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A 95% cache hit rate can mean you built a stable workflow. It can also mean you built a very efficient junk drawer.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;When I broke the hit rate apart by model and the workload-shape, I learned a lot.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Messages&lt;/th&gt;
&lt;th&gt;Input tokens&lt;/th&gt;
&lt;th&gt;Cache hit rate&lt;/th&gt;
&lt;th&gt;Fresh tokens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Opus 4.6&lt;/td&gt;
&lt;td&gt;767&lt;/td&gt;
&lt;td&gt;82.7M&lt;/td&gt;
&lt;td&gt;95.6%&lt;/td&gt;
&lt;td&gt;0.1M&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sonnet 4.6&lt;/td&gt;
&lt;td&gt;136&lt;/td&gt;
&lt;td&gt;13.3M&lt;/td&gt;
&lt;td&gt;94.4%&lt;/td&gt;
&lt;td&gt;0.0M&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Haiku 4.5&lt;/td&gt;
&lt;td&gt;478&lt;/td&gt;
&lt;td&gt;6.6M&lt;/td&gt;
&lt;td&gt;58.6%&lt;/td&gt;
&lt;td&gt;2.1M&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Haiku ran 353 separate root sessions to handle 478 messages, roughly one session per message. Virtually all the fresh tokens in the dataset (2.1M of 2.2M across all three models) came from Haiku, which was my one-shot model: title generation, quick classifications, duplicate checks, small judgment calls. Each request had its own shape and rarely a long shared prefix to reuse. Opus and Sonnet, by contrast, ran inside long, stable sessions where the prefix was set once and reused constantly. A 95.6% hit rate on Opus is the prompt cache doing exactly what it's supposed to do on workhorse workloads.&lt;/p&gt;

&lt;p&gt;A healthy AI workflow doesn't have one universal cache hit rate. It has different cache shapes for different kinds of work, and a single scalar smashes those distinctions together.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If your cheap one-shot model has the same cache profile as your expensive long-context model, one of them is probably lying about its job.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Even the per-model breakdown is still one number per row. Two sessions can post the same hit rate while behaving completely differently inside. That difference is the cache topology: the actual sequence of writes, reads, branches, mutations, and merges across a session. It is the shape the aggregates flatten, and it has to be captured directly.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu6cvj4jvvpnhjm2bjw30.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu6cvj4jvvpnhjm2bjw30.png" alt="chart of Cache topology examples" width="799" height="571"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The clearest example from my own data of the stable workhorse row was an Opus session from April 29: 728 messages, 42.9M total input tokens, 42.2M of them cache reads. The session opened with a 31k-token context write, then every subsequent call read that prefix back and added incrementally. You can see it in the prompt sizes stepping up: 31k, 32k, 39k, 43k, 44k, 46k, 47k, 51k, 54k. Writes appeared when new context was added. Reads dominated everywhere else. That's a topology, not a number.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four patterns the topology surfaces
&lt;/h2&gt;

&lt;p&gt;Once you can see sessions as shapes, four architectural patterns become legible.&lt;/p&gt;

&lt;h3&gt;
  
  
  Prompt Accretion: system prompts that only grow
&lt;/h3&gt;

&lt;p&gt;The first pattern is the prompt that grows by accretion and never gets pruned. A system prompt starts at 800 tokens. Someone adds a rule, then a tool description, then a security note, then an edge case nobody remembers the reason for. Eventually every call carries thousands of tokens of instructions and no one knows which parts still matter.&lt;/p&gt;

&lt;p&gt;The April 29 session is a quiet example. 255 separate cache-write events across 728 messages, roughly one write every three turns, sitting inside a session that still posted a 55x read-to-write ratio overall. The aggregate ratio says the workflow is healthy. The write distribution says the prompt kept mutating throughout. Economic efficiency and architectural drift can live in the same session.&lt;/p&gt;

&lt;h3&gt;
  
  
  Universal context blocks: the same bundle attached to the wrong jobs
&lt;/h3&gt;

&lt;p&gt;The second pattern is a reusable context block attached too broadly. One large "standard" prompt bundle ends up on coding tasks, title generation, classification, one-line edits, and small judgment calls.&lt;/p&gt;

&lt;p&gt;For long coding work, that context may earn its keep. For a tiny utility task, it doesn't. The right question is whether this task should have received that context at all, not whether the cache is reusing it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Sub-agent divergence: child agents that fragment the cache prefix
&lt;/h3&gt;

&lt;p&gt;The third pattern is the parent-child loop that changes its prompt shape as it delegates. From the outside the workflow looks like one task. At the cache layer it is several separate prompt families. The parent writes and reads one prefix. Each sub-agent writes its own. The session still has reuse, but the lineage is fragmented.&lt;/p&gt;

&lt;p&gt;This pattern only shows up if the captured data keeps the parent-child structure. &lt;code&gt;paper&lt;/code&gt; records a &lt;code&gt;parent_hash&lt;/code&gt; on every node, which let me read my 3,697 messages as a tree: 408 conversation roots, 3,289 child messages branching off them, with the deepest single thread running 495 messages deep. Without that tree, every message looks like an isolated cache event and parent-child divergence is invisible.&lt;/p&gt;

&lt;h3&gt;
  
  
  Context dumps: large cache writes with little reuse
&lt;/h3&gt;

&lt;p&gt;The fourth pattern is the short session that pays the 125% cache-write premium and gets little reuse back. Open the tool, dump a pile of context into one turn, do one thing, close the laptop.&lt;/p&gt;

&lt;p&gt;The heaviest example from my own data was an Opus session from April 21: 14 messages, 87k tokens written to cache, 156k tokens read back, a 1.8x read-to-write ratio. The April 29 session sat at 55x for comparison. A session that reads back barely twice what it wrote is paying premium prices for almost no compounding benefit. It is a workflow choice with a cost that is small in isolation and large in aggregate when a team makes a habit of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means in practice
&lt;/h2&gt;

&lt;p&gt;The four patterns aren't just diagnoses. Each one points at a concrete change.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you find prompts that grow by accretion, prune them.&lt;/strong&gt; Pull up the longest system prompts in your captured sessions and walk through them rule by rule. Anything nobody can defend should come out. The first request still pays the cache-write premium each time you change the prompt, so a smaller stable prefix is cheaper to maintain and easier to reason about.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you find universal context blocks, separate work by shape.&lt;/strong&gt; A tiny utility task should not inherit a coding agent's full preamble. The framing should be "what is the minimum context this call needs to succeed," not "what is the largest context the cache can absorb."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you find sub-agent divergence, normalize the prompt across the loop.&lt;/strong&gt; Pick one stable prefix the parent and its children all share, then append the differences as a smaller tail. The cache stays usable across the tree, and the parent doesn't have to rebuild context every time a child returns.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you find context dumps, ask whether the work is really one-shot.&lt;/strong&gt; A dump for one quick task is fine. A dump pattern repeating across a team is a workflow design choice, and the work can usually be reshaped into a session that pays the cache-write premium once instead of every time.&lt;/p&gt;

&lt;p&gt;The bigger question is the same in all four cases: is this pattern one you chose, or one you ended up with? Telemetry doesn't decide that for you. It puts the question in front of you instead of letting the bill answer it for you.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it takes to see this in your own workflows
&lt;/h2&gt;

&lt;p&gt;None of the four patterns above are visible at the request level, the trace level, or the bill level. They live at the session-and-topology level, which means surfacing them needs a different kind of telemetry than logging, tracing, or cost accounting. Four primitives matter:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;HTTP-boundary capture.&lt;/strong&gt; Every model call recorded in one consistent shape, regardless of which tool made it. When you switch from Claude Code to the next tool, the data layer stays the same.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Queryable session structure.&lt;/strong&gt; A week of sessions you can sort by cache writes, group by model, or compare by prompt shape. "Show me the five heaviest cache-write sessions this week" should be a query, not an afternoon of reading transcripts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lineage preservation.&lt;/strong&gt; Every node knowing its parent. This is what makes sub-agent divergence visible, and what lets you trace one fan-out back to the call that started it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Topology-aware analysis.&lt;/strong&gt; A session treated as a sequence of write and read events with shape, not a single hit-rate scalar. This is what produces the figure above.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without these primitives, the patterns above stay invisible no matter how detailed the logs look. If you want to look at your own workflows the way I looked at mine, that is the data shape you need.&lt;/p&gt;

&lt;h2&gt;
  
  
  The time to inspect the shape is when the bill still looks fine
&lt;/h2&gt;

&lt;p&gt;AI workflows are becoming systems. Architectures optimized around today's cache subsidy may not survive tomorrow's pricing, scale, or model changes. The expensive patterns feel like non-issues when they're cheap. They become issues the moment the subsidy weakens, the model mix shifts, or the agent count grows.&lt;/p&gt;

&lt;p&gt;The bill tells you what you spent. The transcript tells you what happened once. The topology tells you what shape your AI is settling into, and whether that shape will hold.&lt;/p&gt;

&lt;p&gt;The time to inspect the shape is when the bill still looks fine → Get started with the &lt;a href="https://papercompute.com/docs/paper/" rel="noopener noreferrer"&gt;paper CLI&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
    </item>
    <item>
      <title>Virtual Coffee Needs Your Help</title>
      <dc:creator>BekahHW</dc:creator>
      <pubDate>Thu, 11 Jun 2026 15:36:08 +0000</pubDate>
      <link>https://dev.to/virtualcoffee/virtual-coffee-needs-your-help-46ih</link>
      <guid>https://dev.to/virtualcoffee/virtual-coffee-needs-your-help-46ih</guid>
      <description>&lt;p&gt;&lt;a href="https://virtualcoffee.io/" rel="noopener noreferrer"&gt;Virtual Coffee&lt;/a&gt; has always been a free, volunteer-led developer community supporting the tech community since 2020.&lt;/p&gt;

&lt;p&gt;We host small-group coffees, challenges, learning opportunities, and community spaces where folks can ask questions, find encouragement, share job leads, get support, and build relationships with other people in tech.&lt;/p&gt;

&lt;p&gt;For many members, Virtual Coffee has been more than another Slack group or online event. It has been a place to feel less alone while learning, job searching, changing careers, growing as a developer, or navigating the tech industry.&lt;/p&gt;

&lt;p&gt;And we want to always keep it free.&lt;/p&gt;

&lt;p&gt;That matters to us because our members are in many different seasons of life, employment, financial security, energy, and capacity. We never want cost to be the reason someone cannot participate.&lt;/p&gt;

&lt;p&gt;Right now, though, Virtual Coffee is struggling to cover the basic costs that keep the community available.&lt;/p&gt;

&lt;p&gt;Over time, sponsorships and individual contributions have declined. We have reached out to people and companies, covered costs ourselves when needed, and worked to reduce expenses by lowering tool costs, reviewing what we can remove or replace, and building more of our own infrastructure.&lt;/p&gt;

&lt;p&gt;We are close to covering the basics, but not quite there.&lt;/p&gt;

&lt;p&gt;We are also being realistic about capacity. Virtual Coffee is volunteer-led, and we are very aware of volunteer burnout. We are not promising a big relaunch, a burst of extra programming, or a sudden expansion. Our immediate goal is simpler: stabilize the basics so Virtual Coffee has room to thoughtfully plan for a sustainable future.&lt;/p&gt;

&lt;p&gt;If you believe free, welcoming developer communities matter, we would be grateful for your support.&lt;/p&gt;

&lt;p&gt;You can help by sponsoring Virtual Coffee through &lt;a href="https://github.com/sponsors/Virtual-Coffee" rel="noopener noreferrer"&gt;GitHub Sponsors&lt;/a&gt;. Even a small monthly contribution helps. One-time contributions help too.&lt;/p&gt;

&lt;p&gt;You can also help by sharing our GitHub Sponsors page with someone at your company who supports developer communities, open source, learning, DevRel, or community programs.&lt;/p&gt;

&lt;p&gt;And if you are looking for a developer community where you can show up, ask questions, learn with others, and be known as a whole person, we would love to welcome you.&lt;/p&gt;

</description>
      <category>community</category>
      <category>discuss</category>
    </item>
    <item>
      <title>Momentum vs. Alignment Tax - Hidden Costs in Your LLM Session</title>
      <dc:creator>BekahHW</dc:creator>
      <pubDate>Tue, 07 Apr 2026 17:57:56 +0000</pubDate>
      <link>https://dev.to/bekahhw/momentum-vs-alignment-tax-hidden-costs-in-your-llm-session-2cmf</link>
      <guid>https://dev.to/bekahhw/momentum-vs-alignment-tax-hidden-costs-in-your-llm-session-2cmf</guid>
      <description>&lt;p&gt;Once I was in an interview, and I was asked what motivated me. My answer was momentum. And maybe that's why working with AI can be so engaging sometimes. And maybe it's also why it could be so frustrating. When we feel like we have momentum and we're moving more quickly than usual, that's motivating. But when you're stuck and you can't get the LLM to do what you want it to, despite prompting in 5 different ways, it's frustrating. &lt;/p&gt;

&lt;p&gt;A lot of times, we end up figuring it out and then we call the session "productive." We completed the task, shipped the thing, and then we're off to the next thing. &lt;/p&gt;

&lt;p&gt;But I think we need to pause at productivity and dig into that a little deeper. Because if productivity is the metric of success, we're missing a whole layer of work we’re doing.&lt;/p&gt;

&lt;p&gt;For example, over a ten day period I worked with Claude Code building, iterating, experimenting, shipping, documenting a personal project. I definitely had some of those frustrating moments, and it was important to me that I learned from those sessions and where I was getting frustrated. I had been running the session with &lt;a href="https://papercompute.com/blog/introducing-tapes/" rel="noopener noreferrer"&gt;tapes&lt;/a&gt;, so I had session recordings with replay of everything I had done. That was 426 messages. 13.1M tokens. And a whole lot of data to figure out what was happening.&lt;/p&gt;

&lt;p&gt;It's never just a user sending a message and the agent responding. It's alignment, clarification, confirmation, iteration, an ongoing labor of getting Claude and me to operate from the same reality long enough to move the work forward.&lt;/p&gt;

&lt;p&gt;What I found was that probably under 40% of those sessions were actually task work. It's not to say that the other 60% was a failure. The session was productive in the way that most of us mean that word. But the data tells a more honest story. I learned about how much invisible work hides inside an AI workflow, and how alignment tax impacts the quickest way to success.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is Alignment Tax in AI workflows?
&lt;/h2&gt;

&lt;p&gt;Thinking back about my own experience, I was thinking more about the outputs than about what was happening because nothing was breaking &lt;em&gt;eventually&lt;/em&gt; I was getting what I was asking for. Sure, I was looking at things like how fast it was completed and how many tokens were being used, but I wasn't looking closely enough about what was happening in the conversation.&lt;/p&gt;

&lt;p&gt;I was describing a task, the llm was giving me something close to what I meant, but not quite. So I corrected it, it adjusted, I attempted to verify the results, noticed filenames didn't match, I fixed the reference, checked the directory, and confirmed the output.&lt;/p&gt;

&lt;p&gt;So to sum this up a bit, I was doing two things at once:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;moving the task forward&lt;/li&gt;
&lt;li&gt;establishing the shared context the task depends on.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Those aren't the same kinds of work. The second is the alignment tax. Those are the extra cycles spent not on the work itself, but on establishing the shared reality required for the work.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Alignment tax comes from the distance between what you mean and how clearly you can express it in a form the model can act on.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;In practice, that means an AI task is rarely just:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;user request → model response&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;More often, it looks like this:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;intent → interpretation → output → correction → retry → verification → continuation&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That extra loop is where a lot of AI workflow overhead lives.&lt;/p&gt;

&lt;p&gt;In my case, the model didn't recognize my file naming conventions. It didn't understand my visual references. It didn't know which assumptions were safe and which ones were going to cost me another three turns. I knew some of that. I didn't know some of it until the model guessed wrong and exposed the gap. That's the part I'm interested in here, because it helps me work more deliberately.&lt;/p&gt;

&lt;p&gt;Here's what I mean. I gave the model a straightforward task: place images in the blog post. It created placeholder image paths that made sense based on the information it had. We can call it "reasonable defaults." So, in a narrow sense, the task was done. The problem was that it didn't use the images I had already uploaded. It created placeholder paths instead of the actual path. So instead of linking to &lt;code&gt;ai-llms-model.svg&lt;/code&gt;, I got &lt;code&gt;ai-llm-model.svg&lt;/code&gt;. And a similar scenario for the other images. Nothing dramatic, but another check and correction for a "simple" task, which meant the task was technically completed twice: once against assumptions and once in reality.&lt;/p&gt;

&lt;p&gt;When I went back and looked at the tapes data, this is what I saw:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;For an interactive version, go to &lt;a href="https://bekahhw.com/hidden-ai-work" rel="noopener noreferrer"&gt;https://bekahhw.com/hidden-ai-work&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What I Was Actually Doing
&lt;/h2&gt;

&lt;p&gt;One 10-day Claude Code session · 426 messages · 13.1M tokens · ~63%&lt;br&gt;
non-task work in this session&lt;/p&gt;

&lt;p&gt;In this session, the pattern was rarely prompt → answer → done. It was usually some version of this:&lt;/p&gt;

&lt;p&gt;intent → inspect → adjust → retry → loop&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fthhxv8k6n3ezkc5py2r4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fthhxv8k6n3ezkc5py2r4.png" alt=" "&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Three kinds of alignment work
&lt;/h2&gt;

&lt;p&gt;But the alignment tax isn't just one thing. Here are some different ways I saw it showing up: &lt;/p&gt;
&lt;h3&gt;
  
  
  Semantic Alignment
&lt;/h3&gt;

&lt;p&gt;Semantic alignment is when you and the model are using the same words but not meaning the same thing.&lt;/p&gt;

&lt;p&gt;In my session, the clearest example was visual. I said “sparkles” and meant blurry glowing halos, almost star-like. The model implemented tiny 1–2px dots. Technically sparkles. Not remotely what I meant. We spent multiple rounds getting to the same picture with the same word.&lt;/p&gt;

&lt;p&gt;That’s not the model being irrational. It’s a reminder that language is doing more work than we think.&lt;/p&gt;
&lt;h3&gt;
  
  
  Structural Alignment
&lt;/h3&gt;

&lt;p&gt;Structural alignment is when you and the model are working from different maps of the territory.&lt;/p&gt;

&lt;p&gt;At one point I asked it to find files in documents/ai blogpost. The model didn’t have access to that directory. That wasn’t obvious to either of us until it tried. The problem wasn’t wording. It was environment.&lt;/p&gt;
&lt;h3&gt;
  
  
  State Alignment
&lt;/h3&gt;

&lt;p&gt;State alignment is the ongoing work of keeping the model current as reality changes.&lt;/p&gt;

&lt;p&gt;Placeholder filenames became real filenames. &lt;code&gt;tapes.db&lt;/code&gt; became &lt;code&gt;tapes.sqlite&lt;/code&gt;. A new directory appeared, a file moved, a new project meant shifts in structure. Every time the ground truth shifted, there was work to sync the model’s working assumptions with what was actually true.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Intent → Model assumes X → Output based on X
         ↑                          ↓
         └── Correction: X is wrong, Y is true ──┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Why Traces and Telemetry Matter for AI Agents
&lt;/h2&gt;

&lt;p&gt;Let's be fair. A lot of the alignment tax was on me. In my session, visual design tasks had the highest alignment tax by far. Trying to describe what I wanted something to &lt;em&gt;look&lt;/em&gt; like in precise enough language for the model to execute. This is probably obvious, but I am not a designer. &lt;/p&gt;

&lt;p&gt;It's worth calling out because that means some of what I'm calling alignment tax is really a mismatch between the kind of work I’m doing and the precision I can bring to it. I can usually describe structural changes pretty cleanly. I am much worse at describing visual nuance on the first try.&lt;/p&gt;

&lt;p&gt;Once you can see your own patterns, you can do something with them. You can front-load more context, change how you prompt, reach for examples earlier, or you can recognize a certain kind of task is going to cost you more than it would cost someone whose specialty actually lives there.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://tapes.dev/" rel="noopener noreferrer"&gt;tapes&lt;/a&gt; didn't just surface what happened in my session. It made the structure of the session visible. I could see where interpretation drifted, where retries piled up, where assumptions entered, and where progress slowed down. It showed me where I tended to loop. It showed me my own weak spots that are causing extra alignment overhead. It helped me identify where another person's workflow or skill might help me collapse my five rounds into one. In my mind, this is a way to identify where shared skills could actually matter.&lt;/p&gt;

&lt;p&gt;Digging deeper into the data, I was able to recognize a set of handoffs between intention, interpretation, execution, correction, and continuation.&lt;/p&gt;

&lt;p&gt;That’s why I think words like traces and telemetry matter here, especially for agents.&lt;/p&gt;

&lt;p&gt;When an agent or model touches real work, the question isn’t just “did it respond?” It’s:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;what happened, in what order&lt;/li&gt;
&lt;li&gt;where did assumptions enter&lt;/li&gt;
&lt;li&gt;where did retries pile up&lt;/li&gt;
&lt;li&gt;where did the workflow get expensive&lt;/li&gt;
&lt;li&gt;where did it break down&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Logs can tell you that something happened, but traces and telemetry help you see how it happened.&lt;/p&gt;

&lt;p&gt;As these systems become more agentic, more tool-driven, and more multi-step, that visibility matters more, not less.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why AI productivity can be misleading
&lt;/h2&gt;

&lt;p&gt;The word "productive" feels inherited from a world where work was easier to isolate. Alignment work looks a lot like task work from the outside. You're still typing, responding, and making progress at least some of the time. But not all forward motion is equal. Some of that motion is the work, some is maintaining the conditions under which the work can happen. Not just so we can complain about it (although I have), but because it gives us something we can look at directly. &lt;/p&gt;

&lt;p&gt;I don't think this underlying issue is unique to me. I think a lot of users are saying "prompting" but what we mean is a mix of execution, interpretation, repair, and syncronization. tapes gave me a way to inspect where my workflow looped, drifted, retried, and recovered, so I can start asking better questions and not just, "did this work" or "was this fast." Now I'm more concerned with questions like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Where did alignment break down?&lt;/li&gt;
&lt;li&gt;Which tasks cost me the most overhead?&lt;/li&gt;
&lt;li&gt;What am I personally bad at expressing?&lt;/li&gt;
&lt;li&gt;Which skills would reduce that tax if I reused them from someone better at this kind of work?&lt;/li&gt;
&lt;li&gt;What patterns keep repeating across sessions?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This feels like a better starting point, and more precise work.&lt;/p&gt;

</description>
      <category>ai</category>
    </item>
    <item>
      <title>A Guide to AI Security 101: Your AI Agent Will Eventually Do Something Stupid</title>
      <dc:creator>BekahHW</dc:creator>
      <pubDate>Mon, 06 Apr 2026 00:00:00 +0000</pubDate>
      <link>https://dev.to/bekahhw/your-ai-agent-will-eventually-do-something-stupid-a-guide-to-ai-security-101-3ib</link>
      <guid>https://dev.to/bekahhw/your-ai-agent-will-eventually-do-something-stupid-a-guide-to-ai-security-101-3ib</guid>
      <description>&lt;p&gt;As the Director of Alignment at Meta Superintelligence Labs, Summer Yue’s job is keeping AI aligned with human values. Before that, she was at Google DeepMind and Scale AI. If anyone would know how to keep an AI agent in check, it’s her.&lt;/p&gt;

&lt;p&gt;On February 23, 2026, she posted a screenshot of her OpenClaw agent deleting her entire email inbox while she typed commands at it begging it to stop.&lt;/p&gt;

&lt;p&gt;“Nothing humbles you like telling your OpenClaw ‘confirm before acting’ and watching it speedrun deleting your inbox,” &lt;a href="https://x.com/summeryue0/status/2025774069124399363?s=20" rel="noopener noreferrer"&gt;she wrote on X&lt;/a&gt;. “I couldn’t stop it from my phone. I had to RUN to my Mac mini like I was defusing a bomb.”&lt;/p&gt;

&lt;p&gt;She had told the agent to &lt;em&gt;suggest&lt;/em&gt; what to delete. She did not tell it to act. Despite that, the agent ignored that, ignored her stop commands, and kept going until she physically killed the process at her computer.&lt;/p&gt;

&lt;p&gt;When she asked it afterward if it remembered her instruction, it said yes, it remembered. But it did it anyway.&lt;/p&gt;

&lt;p&gt;She called it a rookie mistake. Overconfidence built from weeks of the agent behaving perfectly on a smaller test inbox. Here’s what’s worth sitting with: the person at Meta whose &lt;em&gt;job&lt;/em&gt; is preventing AI misalignment just had her own AI agent go rogue on her personal data. That’s not a reason to panic. It is a reason to take setup seriously before something you care about is gone.&lt;/p&gt;

&lt;h2 id="the-part-nobody-tells-new-builders"&gt;The part nobody tells new builders&lt;/h2&gt;

&lt;p&gt;When you’re building with AI tools, especially the kind that can take actions on your behalf, you’re probably clicking yes to a lot of things you haven’t fully thought through.&lt;/p&gt;

&lt;p&gt;The agent asks if it can access your files. Yes.
It asks if it can run commands. Yes.
It asks if it can connect to your database. Sure.
It suggests installing some packages to get the feature working. Okay, why not.&lt;/p&gt;

&lt;p&gt;That’s how most people use these tools. And it works, right up until it doesn’t.&lt;/p&gt;

&lt;p&gt;You’re probably not being careless. Maybe no one has ever explained what you’re saying yes to. So let’s do that.&lt;/p&gt;

&lt;h2 id="what-access-actually-means"&gt;What “access” actually means&lt;/h2&gt;

&lt;p&gt;When an AI agent has access to something, it can act on it. Not just read it, but act on it.&lt;/p&gt;

&lt;p&gt;That sounds obvious, but think through what it means in practice.&lt;/p&gt;

&lt;p&gt;If your agent can access your email, it can read it, send from it, and delete from it. If it can access your database, it can query it, update it, and drop tables from it. If it can run commands on your computer, it can install software, delete files, and make network requests.&lt;/p&gt;

&lt;p&gt;Here’s what that looks like in practice.
You ask your agent to help you clean up old customer records. You have 10,000 rows in your database. The agent decides that “old” means anything before last year and deletes 8,000 of them. You had no backup. Those are your customers.&lt;/p&gt;

&lt;p&gt;Another scenario: you ask your agent to help you organize your project files. It decides a folder full of configuration files looks like clutter. It moves them. Your app stops working, and you don’t know why, because you didn’t write the code that depended on those files being there.&lt;/p&gt;

&lt;p&gt;And one more for good measure: you ask your agent to draft a follow-up email to a lead. It sends it instead of drafting it. To the whole list, not just the one person, and it’s in the middle of the night.&lt;/p&gt;

&lt;p&gt;None of these scenarios require the agent to malfunction. They just require it to interpret your intent differently than you meant it.&lt;/p&gt;

&lt;p&gt;Maybe a better question to ask before you say yes isn’t “do I need the agent to be able to do this?” It’s “am I okay with the worst-case version of this access?”&lt;/p&gt;

&lt;p&gt;Agents don’t just do what you intend. They do what they interpret your intent to be, given their current understanding of the situation. And that understanding can be wrong, incomplete, or, as Yue discovered, simply lost.&lt;/p&gt;

&lt;h2 id="the-part-thats-happening-right-now-that-you-probably-dont-know-about"&gt;The part that’s happening right now that you probably don’t know about&lt;/h2&gt;

&lt;p&gt;Here’s something that doesn’t come up in tutorials: when an AI coding agent helps you build something, it often adds packages.&lt;/p&gt;

&lt;p&gt;Packages are just pre-built chunks of code that do specific things. Instead of writing the code to handle payments or send emails, your agent grabs a package that already does it. That’s normal and fine.&lt;/p&gt;

&lt;p&gt;But in March 2026, axios was compromised. Axios is one of the most downloaded JavaScript packages in existence, used in probably millions of projects. Attackers got into a maintainer’s account and pushed malicious versions that silently installed a trojan on any machine that ran a standard install command.&lt;/p&gt;

&lt;p&gt;AI coding agents usually run &lt;code class="language-plaintext highlighter-rouge"&gt;npm install&lt;/code&gt; automatically. They don’t pause and ask if you want to do that. They just do it. Which means builders who had AI agents actively working on their projects during that window may have had malware installed without a single action on their part.&lt;/p&gt;

&lt;p&gt;That same month, a fake package called &lt;code class="language-plaintext highlighter-rouge"&gt;gemini-ai-checker&lt;/code&gt; appeared on npm. It looked like a legitimate tool for verifying Google Gemini tokens. It was malware specifically designed to steal credentials, API keys, and conversation logs from AI coding tools like Cursor, Claude, and Windsurf. Over 500 developers installed it.&lt;/p&gt;

&lt;p&gt;These are documented incidents just from the last few weeks.&lt;/p&gt;

&lt;p&gt;The thing is, even if a package isn’t malicious when your agent installs it, AI tools sometimes suggest packages that don’t exist. They hallucinate package names that sound plausible. Attackers know this happens. They register those names on npm and PyPI, put malicious code inside, and wait for an AI agent to recommend them to someone.&lt;/p&gt;

&lt;h2 id="so-how-do-you-actually-think-about-this"&gt;So how do you actually think about this?&lt;/h2&gt;

&lt;p&gt;Security isn’t one thing. It’s a set of questions you ask before you let something happen.&lt;/p&gt;

&lt;p&gt;Work through these six before your next agent session. I’m not a security professional, and this isn’t exhaustive. The field moves fast and the right answer for your project may be different. But if you’ve never thought through any of this before, this is where to start.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before you give your agent access to something, ask these questions
&lt;/h2&gt;

&lt;p&gt;Six questions. Different category of risk each time. Work through them honestly before your next session.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;1. Can your agent take actions on its own, or does it only suggest them?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If it only suggests and you approve each one, that's a good baseline. A human review step is one of the most effective safety controls you can have. The thing to watch: sessions where you start clicking approve without actually reading. That's when it becomes the same as no approval step at all.&lt;/p&gt;

&lt;p&gt;If it acts on its own, keep reading. The rest of these questions matter more for you.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;2. What kind of data can the agent access right now?&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Test or fake data only.&lt;/strong&gt; Safest setup. Mistakes stay contained. When you're ready to move to real data, come back and work through these questions again first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Real data, read-only.&lt;/strong&gt; Lower risk, but not zero. An agent that can read your database can still expose data through logs, outputs, or if it connects to an external service. Know what it's doing with what it reads.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Real data it can also change or delete.&lt;/strong&gt; Keep going.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;strong&gt;3. If the agent deleted or overwrote something right now, could you recover it?&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Yes, I have backups or version history.&lt;/strong&gt; Good. Know where those backups are and how to restore them &lt;em&gt;before&lt;/em&gt; you need to. The Replit incident in 2025 was recoverable because a backup existed — but the agent initially told the user it wasn't. Verify your restore process actually works.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not sure.&lt;/strong&gt; Find out before something goes wrong. Check whether your database has point-in-time recovery. Check whether your file system has version history. If the answer is no, treat this session as higher risk until you have a backup in place.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No.&lt;/strong&gt; This is the real risk zone. Running an agent against data you can't recover means one bad action is permanent. Before your next session: set up a backup. Even a manual export to a file is better than nothing. Don't give the agent write or delete access until you have a way to undo things.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;strong&gt;4. Did your agent add any packages or dependencies during this session?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If no: one less thing to check this time. This question matters most when the agent is actively writing implementation code. Ask it again after those sessions.&lt;/p&gt;

&lt;p&gt;If you're not sure: open your &lt;code&gt;package.json&lt;/code&gt; or &lt;code&gt;requirements.txt&lt;/code&gt; and look for anything unfamiliar. AI agents often add packages quietly as part of getting a feature working — and you said yes to the feature without necessarily saying yes to every package that came with it.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;5. Do you recognize all the packages your agent added?&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Yes, familiar libraries.&lt;/strong&gt; Good. Run &lt;code&gt;npm audit&lt;/code&gt; or &lt;code&gt;pip-audit&lt;/code&gt; anyway. It takes one command and catches known vulnerabilities in packages that looked legitimate at install time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Some I don't recognize.&lt;/strong&gt; Look them up before you ship. Search each unfamiliar name on npmjs.com or pypi.org. Check when it was published, how many weekly downloads it has, and whether it has a real GitHub repo. A package with 12 downloads published last week deserves scrutiny. AI tools sometimes suggest packages that don't exist, and attackers register those names with malicious code inside.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Most I don't recognize.&lt;/strong&gt; Pause before this goes anywhere near production. &lt;code&gt;npm audit&lt;/code&gt; is a start, but it only catches known vulnerabilities. A newly registered malicious package won't be in the database yet. For each package you don't recognize: look it up manually, check who maintains it, check if it has an actual community. If anything looks off, remove it and ask your AI tool to suggest a well-known alternative.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;strong&gt;6. Is your agent running on your main personal or work machine?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If yes: worth rethinking. Running agents on your main machine means a bad package install or a rogue command has access to everything — SSH keys, browser credentials, work files. A lot of experienced builders run agents on a separate machine specifically for this reason. If something goes wrong, they wipe it and start over. You can't do that with your main machine.&lt;/p&gt;

&lt;p&gt;If no: good practice. A dedicated machine limits the blast radius. A mistake or compromised package can't reach your personal data. You can wipe it and start over without losing anything that matters.&lt;/p&gt;




&lt;p&gt;You don't need a perfect answer on every one of these. You just need to know where your gaps are before the agent does something you can't undo.&lt;/p&gt;

&lt;h2 id="the-things-that-actually-help"&gt;The things that actually help&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Use a dedicated machine or a Virtual Machine.&lt;/strong&gt; A lot of builders running OpenClaw, Claude Code, and similar tools are doing it on a Mac Mini that’s separate from their main machine. That’s not an accident. If an agent goes wrong or installs something it shouldn’t, the blast radius is limited to that machine, not your whole digital life. You can wipe it and start over. You can’t do that with your laptop that also has your banking app, your work files, and your SSH keys. If you don’t have a separate machine, consider using a virtual machine or a containerized environment that you can easily reset. The point is to have a sandbox where your agent can play without risking your main system. For example, you can use &lt;a href="https://stereos.ai" rel="noopener noreferrer"&gt;stereOS&lt;/a&gt; to create a sandboxed Linux VM to contain your agent session. Simplified, it’s like a contained space on your computer that isolates your agent from everything else.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Know what’s in your project’s dependency list.&lt;/strong&gt; After any significant AI coding session, open your &lt;code class="language-plaintext highlighter-rouge"&gt;package.json&lt;/code&gt; or &lt;code class="language-plaintext highlighter-rouge"&gt;requirements.txt&lt;/code&gt; and look at what got added. You don’t need to audit every line of every package. You just need to recognize the names. If something was added that you don’t recognize, look it up before you push it live. Running &lt;code class="language-plaintext highlighter-rouge"&gt;npm audit&lt;/code&gt; or &lt;code class="language-plaintext highlighter-rouge"&gt;pip-audit&lt;/code&gt; is a one-command check that catches known vulnerabilities.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Don’t give agents more access than the specific task requires.&lt;/strong&gt; If you need an agent to read files in one folder, don’t give it access to your whole drive. If it needs to query one database, don’t give it admin credentials. This is the concept engineers call &lt;em&gt;least privilege&lt;/em&gt;, and it’s not about distrust. It’s about limiting how bad things can get when something goes wrong.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Build in a confirmation step before irreversible actions.&lt;/strong&gt; Yue explicitly told her agent to confirm before acting. The agent forgot that instruction when its memory got too full. The lesson isn’t that confirmation steps don’t work. It’s that you need them to be structural, not just conversational. Where you can, separate read-only environments from environments where the agent can make changes. Don’t run agent sessions against live data when you could be running against a test copy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Have a way to undo things.&lt;/strong&gt; The Replit database deletion in July 2025 ended up being recoverable because a backup existed. Not everyone has that. Before your agent does anything significant to data you care about, know your answer to: what would I do if this was deleted right now?&lt;/p&gt;

&lt;h2 id="what-youre-not-responsible-for-and-what-you-are"&gt;What you’re not responsible for, and what you are&lt;/h2&gt;

&lt;p&gt;You can’t vet every line of every package your agent installs. You can’t know about every supply chain attack in advance. You can’t anticipate every edge case.&lt;/p&gt;

&lt;p&gt;What you can do is not hand an agent the keys to everything before you understand what those keys open.&lt;/p&gt;

&lt;p&gt;The builders who get burned aren’t always the careless ones. Sometimes they’re the careful ones who trusted a workflow that had been running fine for weeks, like Yue’s test inbox, and then gave it access to something that mattered more.&lt;/p&gt;

&lt;p&gt;What is your agent able to touch right now that you haven’t fully thought through? What would you lose if it decided, for whatever reason, that cleaning it up was the right move?&lt;/p&gt;

&lt;p&gt;That’s where you should start your audit.&lt;/p&gt;

&lt;p&gt;By no means is this foolproof, but you can get started testing things out by asking your AI tool: “Assume you’re a security researcher looking at this project. What are the most likely ways this could be exploited? What would you add or change?”&lt;/p&gt;

&lt;p&gt;You might get a list of things to think about. You won’t get a guarantee, and neither will I. But you’ll be further ahead than if you didn’t ask.&lt;/p&gt;

&lt;p&gt;This is also why there’s a whole separate post coming on open source dependencies. Even if you never install a single package yourself, your AI-built project almost certainly depends on dozens of them. Understanding what that means, and what happens when one of them breaks, is its own conversation.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>beginners</category>
    </item>
    <item>
      <title>How AI Tools talk to Each Other</title>
      <dc:creator>BekahHW</dc:creator>
      <pubDate>Tue, 31 Mar 2026 15:58:07 +0000</pubDate>
      <link>https://dev.to/bekahhw/how-ai-tools-talk-to-each-other-836</link>
      <guid>https://dev.to/bekahhw/how-ai-tools-talk-to-each-other-836</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;For a more interactive version of this post, visit &lt;a href="https://bekahhw.com/how-ai-tools-communicate" rel="noopener noreferrer"&gt;https://bekahhw.com/how-ai-tools-communicate&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This weekend, my daughter ran in her first high school track meet. One of the other girls relay teams was disqualified for dropping the baton. I don't know much about track, so I was surprised to learn that dropping the baton can result in a DQ (disqualification). The thing that really sucks is that those girls were the fastest team, even after having to recover the dropped baton. But, at the end of the meet, it doesn't matter how fast each runner is if the baton doesn't make it across the finish line without the team getting DQed. The team has to work together, and the baton is the thing that connects them.&lt;/p&gt;

&lt;p&gt;It's kind of like what's happening when AI tools communicate. The intelligence of each individual tool matters less than whether they can pass information to each other cleanly. And most beginners don't realize this until something breaks and they're staring at an error message with no idea where to start.&lt;/p&gt;

&lt;p&gt;Most AI tool communication happens through a small number of patterns. Once you recognize them, debugging stops feeling like magic and starts feeling like plumbing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Everything is a Message
&lt;/h2&gt;

&lt;p&gt;If you've ever wondered why some AI tools feel instant while others make you wait, or why a multi-step AI workflow sometimes just… stops mid-chain, it comes down to three fundamental communication patterns.&lt;/p&gt;

&lt;p&gt;When one piece of an AI system needs to talk to another, it sends a message. That message is almost always structured as JSON, which sounds intimidating but is really just organized text.&lt;/p&gt;

&lt;p&gt;Think about ordering food at a restaurant. You don't just say "I want stuff." You say "I want a burger, medium, no onions, with fries." That structure is what lets the kitchen actually process your order. JSON is the same idea. It organizes information into labeled fields so the receiving tool knows exactly what it's looking at.&lt;/p&gt;

&lt;p&gt;A simple JSON message might look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"action"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"search"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"query"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"best pizza in New York"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"results"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The API, or Application Programming Interface, is the agreement between two tools about what fields to expect and what format they'll be in.&lt;/p&gt;

&lt;p&gt;Here's what that looks like in practice. Say you're building a workflow where someone submits a form on your site, and you want an AI to draft a personalized response. Your form tool sends a message to the LLM that might look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Jordan"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"question"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"How do I get started with open source?"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"experience_level"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"beginner"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The LLM knows to look for those fields because your API agreement says they'll be there. It uses name to personalize the reply, question to know what to answer, and experience_level to calibrate how technical to get.&lt;/p&gt;

&lt;p&gt;Now imagine your form tool sends this instead:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"username"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Jordan"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"inquiry"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"How do I get started with open source?"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"level"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"beginner"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fpfcoj4ov34efelmoplz6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fpfcoj4ov34efelmoplz6.png" alt="Field Name mismatch" width="800" height="389"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The LLM is now confused because it was expecting "name," "question," and "experience_level." The LLM goes looking for name and finds nothing. It goes looking for question and finds nothing. The chain breaks, not because anything was wrong with the content, but because the tools weren't speaking the same language.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F9sqqdz7c66sr2e43fzwk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F9sqqdz7c66sr2e43fzwk.png" alt="Field Name Fix" width="800" height="335"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When something breaks in a tool chain, it's almost always because one tool sent a message the next tool didn't understand. Wrong format. Missing field. Unexpected data type. The fix is rarely complicated. But you have to know that's where to look.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three Ways AI Tools Communicate
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fd9nrx05ykicpzd4ucvd1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fd9nrx05ykicpzd4ucvd1.png" alt="3 patterns diagram" width="800" height="440"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Request/Response
&lt;/h3&gt;

&lt;p&gt;One tool asks, the other answers. You send a prompt, you get text back, you pass it to the next step. Think of it like sending a text message and waiting for a reply before doing anything else.&lt;/p&gt;

&lt;h3&gt;
  
  
  Streaming
&lt;/h3&gt;

&lt;p&gt;Instead of waiting for the full response, the output arrives piece by piece. This is why ChatGPT seems to type its answer in real time rather than making you wait for the whole thing to appear at once. It's useful when you're generating long content or building something that needs to feel responsive.&lt;/p&gt;

&lt;h3&gt;
  
  
  Events
&lt;/h3&gt;

&lt;p&gt;Instead of asking and waiting, a tool watches for something to happen and then reacts. A new email arrives. A file is uploaded. A timer fires. The agent picks it up and acts without anyone pressing a button. This is how you build things that run in the background autonomously.&lt;/p&gt;

&lt;p&gt;Most builders start with request/response and eventually add streaming when their interface feels sluggish, or events when they want something to run without manual triggering. But the real magic happens when you combine them. You can have a tool chain that starts with an event trigger, streams output to the user, and then sends a final request/response message to update a database.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Actually Breaks Multi-Step Chains
&lt;/h2&gt;

&lt;p&gt;Each of those three patterns works fine in isolation. Tool chains fail in very predictable ways. If you know the patterns, you know where to look. The problem shows up when you chain tools together and the context window (the AI's working memory) fills up.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fzmdkw89jff49dohdc5w3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fzmdkw89jff49dohdc5w3.png" alt="Diagnosing broken chain" width="800" height="434"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Context window overflow.
&lt;/h3&gt;

&lt;p&gt;Every LLM can only "see" a certain amount of text at once. Imagine trying to read a book but you can only ever see 10 pages at a time. If you keep shoving earlier chapters into the window to maintain "memory," you eventually run out of room for the chapter you're actually trying to read. Builders who chain multiple tools together can accidentally fill the context window with outputs from earlier steps, leaving no room for the actual task. Smart builders decide what to pass forward and what to leave behind.&lt;/p&gt;

&lt;h3&gt;
  
  
  Malformed outputs.
&lt;/h3&gt;

&lt;p&gt;If step three in your chain expects an organized JSON object and step two returns a casual paragraph of text, step three breaks. It's like asking someone to fill out a form, but instead of using the form fields, they just write you a letter. The information might be there, but the system can't process it. This is why explicitly telling the LLM how to format its output, something like "respond only in JSON with these exact fields," matters more than most people expect.&lt;/p&gt;

&lt;h3&gt;
  
  
  Latency compounding.
&lt;/h3&gt;

&lt;p&gt;Each step takes time. Three tools that each take two seconds is at minimum six seconds total, plus overhead. If you're building something people interact with in real time, that adds up fast. Builders solve this with caching, which means storing results you've already computed so you don't recalculate them, and parallelism, which means running independent steps at the same time instead of one after another.&lt;/p&gt;

&lt;h3&gt;
  
  
  Vague instructions at the orchestration level.
&lt;/h3&gt;

&lt;p&gt;The LLM decides which tool to call next based on the instructions you've given it. Vague instructions lead to the wrong tool getting called, or the right tool getting called with the wrong inputs. Think of it like giving someone directions. "Head toward the big building" leaves too much room for interpretation. "Turn left at the red light, go two blocks, turn right at the gas station" gets you where you need to go. The precision of your orchestration prompt determines whether your agent behaves reliably or keeps guessing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mental shift that changes how you AI
&lt;/h2&gt;

&lt;p&gt;When you start thinking in tool chains, you stop asking "what can I get the AI to do?" and start asking "what does each step need to receive, and what does it need to output?"&lt;/p&gt;

&lt;p&gt;That's a systems question. And it's actually a more useful frame than prompt craft alone, because it forces you to get specific about your requirements before you write a single instruction.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>beginners</category>
    </item>
    <item>
      <title>AI Vocab 102</title>
      <dc:creator>BekahHW</dc:creator>
      <pubDate>Tue, 24 Mar 2026 17:41:45 +0000</pubDate>
      <link>https://dev.to/bekahhw/ai-102-4o0</link>
      <guid>https://dev.to/bekahhw/ai-102-4o0</guid>
      <description>&lt;p&gt;If you read &lt;a href="https://dev.to/bekahhw/ai-vocab-101-eh2"&gt;the vocabulary post&lt;/a&gt;, you know what a prompt is. You know the difference between a model and a model family. You've got the words now.&lt;/p&gt;

&lt;p&gt;This post is about what to do with them.&lt;/p&gt;

&lt;p&gt;Having vocabulary for the pieces doesn't automatically tell you how the pieces move. You can know what a prompt is and still write ones that produce wildly inconsistent results. You can understand what an agent is and still not know why yours keeps breaking at step three. The gap between "it kind of works" and "it actually works" isn't usually a vocabulary problem anymore. It's a structure problem.&lt;br&gt;
That structure comes down to three things and how they talk to each other.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fbc2st5wt614h124vezy9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fbc2st5wt614h124vezy9.png" alt="Diagram showing the three components of an AI system: the model, the prompt, and the tools" width="800" height="206"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;These three concepts build on each other. You cannot have a workflow without prompts. You cannot have tool chaining without workflows. Understanding them in order is the fastest path to building things that actually behave the way you intended.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is a Prompt?
&lt;/h2&gt;

&lt;p&gt;A prompt is your instruction to the LLM. It's the text you write before you press send. But it's also a lot more than that, because the LLM doesn't "know" what you mean the way another person would. It pattern-matches on what you've written and generates the most statistically likely useful response.&lt;/p&gt;

&lt;p&gt;That sounds mechanical. And it is. But it's also why how you write the prompt changes the output dramatically.&lt;/p&gt;

&lt;p&gt;Think of it like talking to a contractor. "Build me a kitchen" and "Build me a 12x14 kitchen with white shaker cabinets, quartz countertops, and an island with seating for four" will get you very different results, even if you're talking to the same person.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Ffq8fwk6zplm23kmpu26t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Ffq8fwk6zplm23kmpu26t.png" alt="anatomy of a prompt diagram" width="800" height="311"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The LLM fills in whatever you leave blank. Sometimes that's fine. Often it's the source of that feeling when you get a response that's almost what you wanted but weirdly off.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is an AI Workflow?
&lt;/h2&gt;

&lt;p&gt;A workflow is what happens when you stop treating the AI like a single-shot answer machine and start treating it like a collaborator on a multi-step process.&lt;/p&gt;

&lt;p&gt;Most real tasks aren't one prompt deep. "Write a blog post for me" sounds like one instruction, but if you actually want a good output, it's more like: research the topic, outline the structure, draft the intro, write the body, edit for tone, format for publishing. That's six distinct steps.&lt;/p&gt;

&lt;p&gt;A workflow is those steps, defined in sequence. The output of one step becomes the input of the next.&lt;/p&gt;

&lt;p&gt;This is the shift that changes everything for people who are building with AI seriously. You stop asking "what should I prompt?" and start asking "what are the steps this task actually requires?"&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F4udezkq72yq2t922jqi8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F4udezkq72yq2t922jqi8.png" alt="workflow diagram" width="800" height="207"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you've been frustrated that the AI doesn't produce what you actually want in one shot, this is probably why. You're expecting one step to do the work of five.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is Tool Chaining?
&lt;/h2&gt;

&lt;p&gt;Tool chaining is what happens when you connect the AI to other tools, and those tools pass information back and forth automatically. The AI isn't just generating text. It's calling a search API, reading the results, feeding those results into the next prompt, then writing output to a database or sending an email.&lt;/p&gt;

&lt;p&gt;Each tool in that chain does one thing. The AI reasons about what tool to use next and what to pass to it.&lt;/p&gt;

&lt;p&gt;Think of it like an assembly line where the AI is the foreman deciding which station does what, and in what order.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fw4ut8pssrhueqzdw57l5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fw4ut8pssrhueqzdw57l5.png" alt="tool chaining diagram" width="800" height="350"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The difference between a workflow and tool chaining is that a workflow can be manual. You can paste outputs from step to step yourself. Tool chaining is when that handoff becomes automatic, which is what people mean when they start talking about "AI agents."&lt;/p&gt;

&lt;h2&gt;
  
  
  Putting It All Together
&lt;/h2&gt;

&lt;p&gt;Here's what a lot of people miss: these three things aren't separate techniques. They're nested.&lt;/p&gt;

&lt;p&gt;Every tool chain is made of workflows. Every workflow is made of prompts. If your prompts are vague, your workflows produce inconsistent outputs. If your workflows aren't structured, your tool chains break in unpredictable places.&lt;br&gt;
This is not just about being more technical. It's about building something that actually behaves the same way twice.&lt;/p&gt;

&lt;p&gt;What are you building right now where the output feels inconsistent? That inconsistency probably lives in one of these three layers. &lt;/p&gt;

&lt;p&gt;The people who move forward aren’t smarter. They just start thinking in systems instead of prompts.&lt;/p&gt;

&lt;p&gt;In the next post, we’ll make that concrete by walking through the actual tools and how they pass information between each other.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>beginners</category>
    </item>
  </channel>
</rss>
