<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sam Novak</title>
    <description>The latest articles on DEV Community by Sam Novak (@sam_novak_574b07811e18495).</description>
    <link>https://dev.to/sam_novak_574b07811e18495</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3844731%2F3626a15e-2f42-4123-8176-443a2c43928c.png</url>
      <title>DEV Community: Sam Novak</title>
      <link>https://dev.to/sam_novak_574b07811e18495</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sam_novak_574b07811e18495"/>
    <language>en</language>
    <item>
      <title>Put your particle effects in CI: the silent-skip bug nobody catches by eye</title>
      <dc:creator>Sam Novak</dc:creator>
      <pubDate>Mon, 14 Sep 2026 11:23:20 +0000</pubDate>
      <link>https://dev.to/sam_novak_574b07811e18495/put-your-particle-effects-in-ci-the-silent-skip-bug-nobody-catches-by-eye-5cpa</link>
      <guid>https://dev.to/sam_novak_574b07811e18495/put-your-particle-effects-in-ci-the-silent-skip-bug-nobody-catches-by-eye-5cpa</guid>
      <description>&lt;p&gt;Most particle editors are GUI-first, which is correct: authoring a visual effect by typing numbers is a miserable way to work. But once the effect exists, it is a file, and files belong in the same machinery as the rest of your project.&lt;/p&gt;

&lt;p&gt;I spent an afternoon wiring our VFX into CI and it caught a class of bug we had been shipping for months.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bug class
&lt;/h2&gt;

&lt;p&gt;A portable effect format usually means one description consumed by multiple renderer backends. That is the whole point. What it does not mean is that every module is implemented on every backend.&lt;/p&gt;

&lt;p&gt;We moved an effect from the Three.js path to a 2D Pixi path. It rendered. It just rendered &lt;strong&gt;wrong&lt;/strong&gt;, in a way nobody could name, and it took an embarrassing amount of squinting to work out that a velocity-over-lifetime module had been silently dropped because that backend does not implement it.&lt;/p&gt;

&lt;p&gt;It was not an error. It was reported, correctly, in an unsupported-modules list on the stats object. Nobody was reading the stats object.&lt;/p&gt;

&lt;h2&gt;
  
  
  Making the machine read it
&lt;/h2&gt;

&lt;p&gt;The useful realisation is that this is a test, not a debugging technique. If the runtime will tell you which modules it skipped, you can assert on that in CI:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Load every effect file in the repo&lt;/li&gt;
&lt;li&gt;Instantiate it against each backend the game actually ships&lt;/li&gt;
&lt;li&gt;Fail the build if the unsupported-module list is non-empty for a backend that effect is used on&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is maybe thirty lines. It turns a silent visual regression into a red X on a pull request, which is the entire difference between a bug you fix in a minute and one you ship.&lt;/p&gt;

&lt;h2&gt;
  
  
  Validation and generation, while you are there
&lt;/h2&gt;

&lt;p&gt;Two more things fall out once effects are scriptable from the command line.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Validate on commit.&lt;/strong&gt; A malformed effect file should not reach a build. A validate step in a pre-commit hook or CI job is cheap and stops the "works on my branch" version of VFX bugs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Generate variants.&lt;/strong&gt; We had five near-identical pickup effects differing only in tint. Those are now generated from one source plus a colour table, which means a change to the shape of the effect does not require an artist to open five files and remember what the other four looked like.&lt;/p&gt;

&lt;p&gt;A caution on that last one: generation is great for mechanical variants and terrible for anything where an artist is making a judgement call. If someone would want to nudge it by eye, it should not be generated.&lt;/p&gt;

&lt;h2&gt;
  
  
  Getting started
&lt;/h2&gt;

&lt;p&gt;You need two things from your tool: a documented project-file schema, and a CLI that can create, validate and export without a display. If your tool has both, the CI work above is an afternoon. If it has neither, that is worth knowing before the format becomes load-bearing in your pipeline.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://nixiefx.com/cli-reference/" rel="noopener noreferrer"&gt;NixieFX CLI reference&lt;/a&gt; documents the project schema and the create, validate and export commands, which is what I used for the above. Disclosure: I work with that team. The CI idea is tool-agnostic, and the unsupported-module assertion is the part I would steal first.&lt;/p&gt;

</description>
      <category>javascript</category>
      <category>gamedev</category>
      <category>testing</category>
    </item>
    <item>
      <title>Rework is an outcome your tracker probably cannot express</title>
      <dc:creator>Sam Novak</dc:creator>
      <pubDate>Mon, 14 Sep 2026 11:20:46 +0000</pubDate>
      <link>https://dev.to/sam_novak_574b07811e18495/rework-is-an-outcome-your-tracker-probably-cannot-express-3lcc</link>
      <guid>https://dev.to/sam_novak_574b07811e18495/rework-is-an-outcome-your-tracker-probably-cannot-express-3lcc</guid>
      <description>&lt;p&gt;Your tracker has states for work that has not started, work in progress, and work that is finished. Ask it to express "this was delivered, reviewed, and sent back because it solved the wrong problem" and most of them quietly collapse that into moving the ticket back to In Progress.&lt;/p&gt;

&lt;p&gt;That collapse throws away the most useful signal you have.&lt;/p&gt;

&lt;h2&gt;
  
  
  What gets lost
&lt;/h2&gt;

&lt;p&gt;When a rejected delivery is represented as a state change rather than an event, you lose:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;That a delivery happened at all.&lt;/strong&gt; The cycle-time number now says the task took nine days, with no indication that it was done on day three and redone twice.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Why it came back.&lt;/strong&gt; "Requirements changed" and "the brief was ambiguous" and "the implementation was wrong" are three completely different problems with three different fixes, and they all look identical from the board.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;First-pass rate.&lt;/strong&gt; Which is the single number that tells you whether your briefs are getting better or worse.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This matters for human work. It matters much more once agents are doing some of the work, because the agent cannot tell you it was confused, and the rework reason is the only place that confusion becomes visible.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cheap version
&lt;/h2&gt;

&lt;p&gt;You do not need a new tool. You need rework to be an &lt;strong&gt;event with a reason&lt;/strong&gt;, not a state transition.&lt;/p&gt;

&lt;p&gt;Add one required field on send-back, with a short closed list. Ours is roughly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;brief was ambiguous&lt;/li&gt;
&lt;li&gt;brief was wrong&lt;/li&gt;
&lt;li&gt;requirements changed after the brief&lt;/li&gt;
&lt;li&gt;implementation did not match a clear brief&lt;/li&gt;
&lt;li&gt;out of scope work included&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Closed list, not free text. Free text gives you a thousand unique sentences and no aggregate. Five buckets give you a monthly number you can act on.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we found when we started counting
&lt;/h2&gt;

&lt;p&gt;The honest result: most of our send-backs were not the implementation being wrong. They were the brief being ambiguous in a way that was obvious in hindsight and invisible at writing time. We had been treating a specification problem as an execution problem, which is why none of the things we tried had been helping.&lt;/p&gt;

&lt;p&gt;That is not a profound insight. It is just one that was unavailable to us until rework had a reason code.&lt;/p&gt;

&lt;h2&gt;
  
  
  One design note
&lt;/h2&gt;

&lt;p&gt;If you are building or choosing tooling here, look at whether rework feedback is carried back into the task context that the next run of the work actually reads. A reason code that only lands in a report is an analytics feature. One that reaches whoever, or whatever, picks the task up next is a correctness feature. The &lt;a href="https://wagglet.com/blog/wagglet-workflow-request-draft-ticket-delivery" rel="noopener noreferrer"&gt;Wagglet workflow docs&lt;/a&gt; treat rework feedback as part of the task context for exactly this reason.&lt;/p&gt;

&lt;p&gt;Disclosure: I work with the team that builds Wagglet. The reason-code idea costs nothing and works in whatever you already have.&lt;/p&gt;

</description>
      <category>productivity</category>
      <category>ai</category>
      <category>management</category>
    </item>
    <item>
      <title>Attachment URLs are bearer links, and revoking the token does not revoke them</title>
      <dc:creator>Sam Novak</dc:creator>
      <pubDate>Mon, 14 Sep 2026 11:18:28 +0000</pubDate>
      <link>https://dev.to/sam_novak_574b07811e18495/attachment-urls-are-bearer-links-and-revoking-the-token-does-not-revoke-them-52oi</link>
      <guid>https://dev.to/sam_novak_574b07811e18495/attachment-urls-are-bearer-links-and-revoking-the-token-does-not-revoke-them-52oi</guid>
      <description>&lt;p&gt;When we started handing real tasks to coding agents, we did the obvious safety work first. The agent gets a read-only credential, scoped to one task, that stops working when the claim ends.&lt;/p&gt;

&lt;p&gt;Then a colleague asked a question I did not have a good answer to.&lt;/p&gt;

&lt;h2&gt;
  
  
  The question
&lt;/h2&gt;

&lt;p&gt;If the task context includes attachments, a spec PDF, a screenshot of the bug, a CSV of the failing rows, how does the agent actually fetch them?&lt;/p&gt;

&lt;p&gt;The usual answer is a signed URL. The context payload contains a link, the agent GETs it, done.&lt;/p&gt;

&lt;p&gt;Here is the part that took me a while to properly absorb: &lt;strong&gt;that URL is its own credential.&lt;/strong&gt; It is a bearer token wearing a URL costume. Which means:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Revoking the task credential does &lt;strong&gt;not&lt;/strong&gt; revoke it. The read token can be completely dead and the attachment link still works.&lt;/li&gt;
&lt;li&gt;It does not care who presents it. Anything that can read the URL can read the file.&lt;/li&gt;
&lt;li&gt;It leaks the way URLs leak, which is to say constantly. Logs, error reports, a screenshot of a terminal, a model provider request history, a traceback pasted into a group chat.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of that is exotic. It is the ordinary lifecycle of a string in a distributed system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why agents make this sharper
&lt;/h2&gt;

&lt;p&gt;A human opens an attachment in a browser and the URL dies quietly in their history. An agent puts it in a context window, which may be logged, retried, cached, or forwarded to a subagent. The number of places that string comes to rest goes up by an order of magnitude, and most of them are not places you thought about when you designed the expiry.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to actually do
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Treat attachment URLs as secrets in your logging policy.&lt;/strong&gt; If you redact tokens from logs, redact these too. They are the same class of thing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Give them a shorter life than the session.&lt;/strong&gt; The agent needs the file for one run, not for the seven days the claim might last.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Do not log whole context payloads at info level.&lt;/strong&gt; This is where most accidental exposure actually comes from, not from anything clever.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scope them to the resource, not the account.&lt;/strong&gt; A leaked link that exposes one PDF is an incident. One that exposes a bucket is a very different afternoon.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The documentation point
&lt;/h2&gt;

&lt;p&gt;The thing I have come to care about more than the mechanism is whether the vendor says it out loud. A system can handle this correctly and still leave you guessing about whether revocation covers attachment links, and you will guess wrong in the optimistic direction.&lt;/p&gt;

&lt;p&gt;The handoff skill docs for &lt;a href="https://wagglet.com/security" rel="noopener noreferrer"&gt;Wagglet&lt;/a&gt; state it flatly: attachment URLs returned in context are independent bearer links, keep them private, and revoking the task read credential does not revoke a URL already returned. That sentence saved me an experiment, which is what good documentation is for.&lt;/p&gt;

&lt;p&gt;Disclosure: I work with the team that builds Wagglet. The advice above applies to any signed-URL scheme you happen to be using.&lt;/p&gt;

&lt;p&gt;If you are designing this layer right now, the question worth writing on the wall is simple: when the credential dies, what is still alive?&lt;/p&gt;

</description>
      <category>security</category>
      <category>webdev</category>
      <category>api</category>
    </item>
    <item>
      <title>Seat-based AI tooling has an idle capacity problem nobody budgets for</title>
      <dc:creator>Sam Novak</dc:creator>
      <pubDate>Thu, 10 Sep 2026 10:00:18 +0000</pubDate>
      <link>https://dev.to/sam_novak_574b07811e18495/seat-based-ai-tooling-has-an-idle-capacity-problem-nobody-budgets-for-20kp</link>
      <guid>https://dev.to/sam_novak_574b07811e18495/seat-based-ai-tooling-has-an-idle-capacity-problem-nobody-budgets-for-20kp</guid>
      <description>&lt;p&gt;If your company buys per-seat AI coding subscriptions, you are almost certainly paying for a lot of capacity that expires unused every week, while some of your engineers hit their limit on Tuesday.&lt;/p&gt;

&lt;p&gt;Both things are true at once, and that combination is what makes it a distribution problem rather than a spending problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the distribution is so uneven
&lt;/h2&gt;

&lt;p&gt;Usage is bursty and it correlates with role in a way that seat allocation does not.&lt;/p&gt;

&lt;p&gt;Someone mid-migration will saturate a subscription for two weeks and then barely touch it for a month. Someone in meetings four days a week has a seat that is nearly untouched. A designer or a QA engineer who was given a seat during a rollout may never have opened it.&lt;/p&gt;

&lt;p&gt;Meanwhile the allowance does not roll over. Unused capacity is not banked, it just stops existing, and it stops existing quietly. There is no line item for it and no alert.&lt;/p&gt;

&lt;h2&gt;
  
  
  The wrong fix, which everyone tries first
&lt;/h2&gt;

&lt;p&gt;Share a login. One account, shared credentials, anyone can run anything.&lt;/p&gt;

&lt;p&gt;This is worth naming explicitly because it is the natural first idea and it is bad in four separate ways. It breaks attribution, so you cannot tell who ran what. It makes the rate limit collective, so one person's big job blocks everyone. It makes revocation all-or-nothing. And it generally violates the terms you agreed to when you bought the seats.&lt;/p&gt;

&lt;p&gt;The second wrong fix is buying more seats to relieve the pressure, which increases the amount of idle capacity you are paying for. You will feel like you solved it, because the person who was blocked stops complaining.&lt;/p&gt;

&lt;h2&gt;
  
  
  Move the task instead of the account
&lt;/h2&gt;

&lt;p&gt;The version that works is unglamorous: when someone is out of capacity and someone else has plenty, move the work item, not the credentials. The second person runs it on their own subscription, under their own identity, and the attribution stays intact.&lt;/p&gt;

&lt;p&gt;That requires the task to be portable, which is the actual cost. A task that only its author can run - because the brief lives in their head - cannot be moved no matter how much capacity is idle elsewhere. So the prerequisite for using spare capacity is writing tasks down properly, which is annoyingly the same prerequisite as everything else in this area.&lt;/p&gt;

&lt;h2&gt;
  
  
  Be honest about the two kinds of value
&lt;/h2&gt;

&lt;p&gt;This is where I would push back on most write-ups of the idea, including some enthusiastic internal ones.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Recovered output&lt;/strong&gt; is real: work got done that otherwise would have waited. Worth having.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cash saved&lt;/strong&gt; is usually not real. You only save money if you would otherwise have bought more seats. If you were not going to, then using idle capacity costs the same as not using it, and reporting a saving is fiction. It is a throughput gain, not a cost reduction, and the two get conflated constantly because one of them sounds better in a board update.&lt;/p&gt;

&lt;p&gt;There is a third thing, harder to measure and possibly the most valuable: someone who was blocked stops being blocked, which changes what they attempt next.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://wagglet.com/blog/use-unused-claude-code-codex-capacity" rel="noopener noreferrer"&gt;This write-up models the whole thing with the assumptions visible&lt;/a&gt;, including a ten-seat worked example and an explicit separation of recovered output from genuine cash savings. I would read the section on why an allowance percentage is not enough before quoting a utilisation number to anybody, because a percentage hides exactly the burstiness that makes the problem interesting.&lt;/p&gt;

&lt;p&gt;And if you do try it: run it as a pilot designed so it can come out negative. The failure mode of this idea is not that it does not work - it is that it is impossible to disprove once someone has announced it.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>management</category>
      <category>devops</category>
    </item>
    <item>
      <title>An AI task contains five decisions, and one person usually makes all of them</title>
      <dc:creator>Sam Novak</dc:creator>
      <pubDate>Thu, 10 Sep 2026 09:58:16 +0000</pubDate>
      <link>https://dev.to/sam_novak_574b07811e18495/an-ai-task-contains-five-decisions-and-one-person-usually-makes-all-of-them-j5k</link>
      <guid>https://dev.to/sam_novak_574b07811e18495/an-ai-task-contains-five-decisions-and-one-person-usually-makes-all-of-them-j5k</guid>
      <description>&lt;p&gt;Watch a piece of AI-assisted work go through a small team and you can pull out five distinct decisions:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Is this worth doing at all?&lt;/li&gt;
&lt;li&gt;Is it prepared well enough to hand to somebody else?&lt;/li&gt;
&lt;li&gt;Who runs it, and on whose subscription?&lt;/li&gt;
&lt;li&gt;Did the outcome actually meet the bar?&lt;/li&gt;
&lt;li&gt;Did it land?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;On most teams I have seen, one person makes all five. Usually the person who had the idea. That is the whole problem, and it does not look like a problem while it is happening - it looks like being efficient.&lt;/p&gt;

&lt;h2&gt;
  
  
  What each collapse costs
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1 and 2 collapsing&lt;/strong&gt; means an idea goes straight to a work item without anyone deciding it is worth preparing. You end up with a board that is a list of thoughts, and nobody can tell which entries are real.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2 and 3 collapsing&lt;/strong&gt; is the common one. The person who prepared the task is the person who runs it, so preparation never actually has to be legible to anyone else. The brief lives in their head, the task looks tiny in the tracker, and it cannot be reassigned. You discover the brief was never written the day that person is on holiday.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3 and 4 collapsing&lt;/strong&gt; is the expensive one. The person who ran the task judges whether the result is good. With a human doing the work that is merely optimistic. With an agent it is close to meaningless: the agent reports success, the runner reads the report, the report becomes the verdict. Nobody lied and nothing was checked.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4 and 5 collapsing&lt;/strong&gt; is subtler. Accepting the work and the work actually landing are different facts, and only one of them is verifiable by a machine. If "accepted" also means "merged", then either you are blocking acceptance on CI, or you are letting a human assertion stand in for git. Both are worse than keeping two fields.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one I would fix first
&lt;/h2&gt;

&lt;p&gt;Split 3 from 4. Whoever ran it does not judge it.&lt;/p&gt;

&lt;p&gt;This is cheap. It does not need tooling, a process document, or anyone's permission - it needs a second name on the item. And it converts an agent's self-report from a verdict into a claim, which is the single highest-leverage change available in this whole area.&lt;/p&gt;

&lt;p&gt;The objection is always throughput: we are two people, we do not have a spare reviewer. Fair. Then swap - you review mine, I review yours - and accept that the review is shallow. A shallow review by someone who did not do the work still catches the confident-wrong-answer failure mode, which is the one you actually have.&lt;/p&gt;

&lt;h2&gt;
  
  
  The account question is its own decision
&lt;/h2&gt;

&lt;p&gt;Decision 3 has a part people skip: whose subscription does the run consume?&lt;/p&gt;

&lt;p&gt;The tempting answer is to share a login so anyone can run anything. Do not. It destroys attribution, makes usage limits collective, and generally violates the terms you agreed to. The alternative that works is moving the task to whoever has capacity, and letting them run it on their own account under their own identity. Same flexibility, no shared credentials.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://wagglet.com/how-it-works" rel="noopener noreferrer"&gt;Wagglet's how-it-works page&lt;/a&gt; lays this out as five decisions with one owner each, which is where I got the framing, and it is blunt about the account rule - move the prepared task, keep every account with its owner. It also has a short section on work that should not be handed off at all, which I appreciated, because most tools in this space write as though everything should be.&lt;/p&gt;

&lt;p&gt;You do not need five people. You need the five decisions to be visibly separate, so that when one of them is being skipped, somebody notices.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>management</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Chat is a bad medium for handing off work</title>
      <dc:creator>Sam Novak</dc:creator>
      <pubDate>Thu, 10 Sep 2026 09:57:10 +0000</pubDate>
      <link>https://dev.to/sam_novak_574b07811e18495/chat-is-a-bad-medium-for-handing-off-work-22p9</link>
      <guid>https://dev.to/sam_novak_574b07811e18495/chat-is-a-bad-medium-for-handing-off-work-22p9</guid>
      <description>&lt;p&gt;Most work handed between people right now travels as a chat message. "Can you look at the flaky checkout test? I think it is the timezone thing, there is context in the thread from Tuesday."&lt;/p&gt;

&lt;p&gt;That message is fine as a nudge and terrible as a handoff, and the reasons are structural rather than a matter of writing it better.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four things chat cannot do
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;It is not addressable.&lt;/strong&gt; There is no id. You cannot refer to "that handoff" in a standup, a commit message, or a follow-up two weeks later. The best you get is a permalink to one message, which loses everything said after it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It does not accumulate.&lt;/strong&gt; The scope gets amended four messages later, a constraint arrives from someone else in a side thread, and the original message still says the original thing. There is no single current version - there is a history you have to replay, in order, correctly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It cannot be reviewed before it is sent.&lt;/strong&gt; A handoff is a thing you should be able to prepare, reread, and fix. Chat is send-then-clarify by design. You find out it was under-specified by watching someone work on the wrong thing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It has no completion state.&lt;/strong&gt; "Done" arrives as another message. Nothing distinguishes "I pushed something" from "someone checked it" from "it actually landed". Those are three separate facts and chat flattens all of them into a thumbs up.&lt;/p&gt;

&lt;p&gt;None of this is an argument against chat. It is an argument that chat is the transport, not the artifact.&lt;/p&gt;

&lt;h2&gt;
  
  
  It gets sharper when an agent is involved
&lt;/h2&gt;

&lt;p&gt;Everything above is an old problem. Handing work to a coding agent makes it acute, for one specific reason: the agent needs the brief to be complete at the moment it starts, because it will not come back and ask.&lt;/p&gt;

&lt;p&gt;A human receiving a vague handoff does the single most valuable thing in the whole exchange - they notice it is vague and ask. That check disappears. An agent given an ambiguous brief produces a confident, plausible, wrong answer, and the ambiguity is only discovered at review, if there is a review.&lt;/p&gt;

&lt;p&gt;So the artifact has to carry, at start time, all of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the actual goal, not the symptom that prompted it&lt;/li&gt;
&lt;li&gt;the acceptance check, written before the work&lt;/li&gt;
&lt;li&gt;the context that is not in the repo, which is the part that lives in someone's head or in Tuesday's thread&lt;/li&gt;
&lt;li&gt;the boundary of what the run is allowed to touch&lt;/li&gt;
&lt;li&gt;and separately, instructions for the human supervising it, which are not the same as instructions for the agent&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last one took me a while. The person running the task needs to know how to judge progress and when to stop. The agent needs the task. Writing one blob addressed to both produces something that reads as neither.&lt;/p&gt;

&lt;h2&gt;
  
  
  The credential is part of the artifact
&lt;/h2&gt;

&lt;p&gt;The piece most often left out: what the run is allowed to reach. If handing off a task means handing over your session, the handoff has no boundary and cannot be audited. A bounded, read-only, expiring credential scoped to the one task is not a security nicety bolted on afterwards - it is what makes the handoff a discrete thing with edges rather than a temporary transfer of your whole identity.&lt;/p&gt;

&lt;p&gt;And the delivery step wants to be separate from the working step, so that reporting an outcome is a deliberate act with evidence attached rather than the run trailing off.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://wagglet.com/docs/task-handoff" rel="noopener noreferrer"&gt;Wagglet's task handoff documentation&lt;/a&gt; is a compact description of that shape in practice: one claimed task, a bounded agent context, a named human supervising the run, a separate delivery prompt, and review by somebody who did not do the work. The part I would point at even if you build your own is the explicit statement of what the copied credential can and cannot reach.&lt;/p&gt;

&lt;p&gt;Keep using chat to say "have you got a minute". Just do not let it be the only record of what you agreed.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>devops</category>
    </item>
    <item>
      <title>tools/list is an API contract, and you are versioning it by accident</title>
      <dc:creator>Sam Novak</dc:creator>
      <pubDate>Thu, 10 Sep 2026 09:56:02 +0000</pubDate>
      <link>https://dev.to/sam_novak_574b07811e18495/toolslist-is-an-api-contract-and-you-are-versioning-it-by-accident-2ci7</link>
      <guid>https://dev.to/sam_novak_574b07811e18495/toolslist-is-an-api-contract-and-you-are-versioning-it-by-accident-2ci7</guid>
      <description>&lt;p&gt;If you run an MCP server, your tool catalog is a public API. I do not think most of us are treating it like one.&lt;/p&gt;

&lt;p&gt;Here is the shape of the problem. A client connects, calls &lt;code&gt;tools/list&lt;/code&gt;, gets your catalog, and starts using it. Then you rename a tool for clarity. Or you tighten a parameter schema because someone was passing nonsense. Or you split one tool into two better ones.&lt;/p&gt;

&lt;p&gt;Every one of those is a breaking change shipped with no version number, no deprecation window, and no changelog.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it is worse than a normal API break
&lt;/h2&gt;

&lt;p&gt;With an HTTP API, a break produces a 404 or a 400 and somebody's error tracker lights up. The failure is loud and it points at you.&lt;/p&gt;

&lt;p&gt;With a tool catalog, the consumer is a language model. It does not throw. It adapts. Give it a tool that no longer exists and it will try a plausible neighbour, or invent arguments that match the old schema, or quietly do something adjacent and report success. The failure surfaces as a weird outcome three steps later, in someone else's log, attributed to the model being unreliable.&lt;/p&gt;

&lt;p&gt;That is a genuinely bad debugging situation, and the person who has to debug it is not you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rules I would follow
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Tool names are permanent.&lt;/strong&gt; Pick badly and live with it. A slightly awkward name that has been stable for six months is worth more than a clean rename. If you must rename, keep the old name working and have it do the same thing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Additive changes only, on parameters.&lt;/strong&gt; New optional parameters are fine. Making an optional parameter required is a break. Narrowing an enum is a break. Renaming a field is a break. Tightening validation on a field that previously accepted sloppy input is a break, and it is the one people do not notice they are doing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A removed tool should fail loudly, not vanish.&lt;/strong&gt; If something has to go, leaving it in the catalog with a description that says it is removed and an error that says what to use instead is far kinder than deleting it. The model will read the error. It cannot read your absence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Say so if your catalog is dynamic.&lt;/strong&gt; This is the one I would most want documented. If the tools a client sees depend on its permissions, its connection, or the team it is pinned to, then no client can safely cache the catalog once and reuse it. That is a legitimate design - it is how you keep a connection from becoming a second, broader authority than the person who created it - but it has to be stated, because the default client assumption is that a catalog is stable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Make the catalog itself authoritative.&lt;/strong&gt; Documentation drifts. If your docs list tool names and your server also lists them, one of the two will be wrong within a quarter. Say plainly which one wins.&lt;/p&gt;

&lt;h2&gt;
  
  
  The read-before-write pattern
&lt;/h2&gt;

&lt;p&gt;The corollary on the client side, which I have come round to: for anything with an effect, read current state immediately before acting, in the same run. Not at the start of the session - immediately before.&lt;/p&gt;

&lt;p&gt;This is not paranoia about your server. It is that the gap between "agent decided to do X" and "agent does X" can contain a human doing something else entirely, and a stale read makes the agent confidently overwrite it. A revision or version check on write is the cheap version of the same protection.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://wagglet.com/docs/mcp" rel="noopener noreferrer"&gt;Wagglet's MCP documentation&lt;/a&gt; is a decent worked example of these constraints being written down rather than left implied - a connection that is named and revocable and pinned to one person and one team, an explicit statement that the connection never becomes a second authority, and reading current state before an explicit action. Whether or not you use it, the doc is a reasonable checklist of the things a server author should be deciding on purpose.&lt;/p&gt;

&lt;p&gt;The short version: your tool names are your API surface, your clients cannot see your git history, and your consumers do not raise exceptions. Version accordingly.&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>api</category>
      <category>ai</category>
    </item>
    <item>
      <title>Your tracker has no private draft state, and it shows</title>
      <dc:creator>Sam Novak</dc:creator>
      <pubDate>Thu, 10 Sep 2026 09:09:53 +0000</pubDate>
      <link>https://dev.to/sam_novak_574b07811e18495/your-tracker-has-no-private-draft-state-and-it-shows-377n</link>
      <guid>https://dev.to/sam_novak_574b07811e18495/your-tracker-has-no-private-draft-state-and-it-shows-377n</guid>
      <description>&lt;p&gt;Every issue tracker I have used has some version of Backlog, In Progress, Done. Not one of them ships a place to prepare a task in private before other people can see it.&lt;/p&gt;

&lt;p&gt;That missing state causes a specific, recognisable mess, and once you notice it you see it everywhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mess
&lt;/h2&gt;

&lt;p&gt;Somebody has an idea. The only way to record it is to create a ticket. Creating a ticket puts it on the board. Being on the board means it is now a thing the team can see, comment on, estimate, and start.&lt;/p&gt;

&lt;p&gt;So a half-formed thought becomes a work item in one step, and then the work item becomes the argument. You have all had the ticket with eleven comments where the first four are people asking what it means, the next four are scope negotiation, and the last three are somebody who already started building the wrong version.&lt;/p&gt;

&lt;p&gt;None of that is a communication problem. It is a missing lifecycle stage.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two jobs that are not the same job
&lt;/h2&gt;

&lt;p&gt;Capturing an idea and preparing a task are different activities with different requirements.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Capturing&lt;/strong&gt; wants to be as cheap as possible. Any friction here and people stop recording things, which is strictly worse. One line, no fields, no estimate, no owner.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Preparing&lt;/strong&gt; is the expensive part. It means deciding the scope, writing down what would make the result acceptable, attaching the context somebody would otherwise have to come and ask you for, and - most importantly - deciding whether the thing is worth doing at all.&lt;/p&gt;

&lt;p&gt;If those share one state, capturing inherits preparation's friction, or preparation gets skipped. Usually both.&lt;/p&gt;

&lt;h2&gt;
  
  
  Preparation is cheap in private and expensive in public
&lt;/h2&gt;

&lt;p&gt;This is the bit that convinced me. The same edit costs radically different amounts depending on whether anyone is watching.&lt;/p&gt;

&lt;p&gt;Rewriting the scope of a private draft costs nothing. Rewriting the scope of a published ticket means notifying whoever commented, correcting whoever started, and re-litigating whatever was already negotiated. Same keystrokes, ten times the cost.&lt;/p&gt;

&lt;p&gt;So the natural consequence of having no private stage is that tasks get published under-prepared and then never properly fixed, because fixing them in public is too expensive.&lt;/p&gt;

&lt;h2&gt;
  
  
  Publishing as an actual gate
&lt;/h2&gt;

&lt;p&gt;The other thing you get from a real draft state is that publishing becomes a signal.&lt;/p&gt;

&lt;p&gt;When someone deliberately moves a task from draft to the board, they are asserting something: this is ready to put in front of another person. That assertion is genuinely useful information. On a board where everything is published by default, no item carries it, and every item has to be independently assessed for whether it is real.&lt;/p&gt;

&lt;p&gt;I would go further: nothing should be able to leave the draft state without a written acceptance check. Not as a process ritual - as the definition of prepared.&lt;/p&gt;

&lt;h2&gt;
  
  
  If your tracker cannot do this
&lt;/h2&gt;

&lt;p&gt;Most cannot, so approximate it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Add a Draft or Not Ready label and treat items carrying it as invisible. No estimating, no starting, no commenting beyond the author.&lt;/li&gt;
&lt;li&gt;Default new items to it. The gate only works if passing through it is deliberate.&lt;/li&gt;
&lt;li&gt;Make the exit criterion explicit and boring: scope, acceptance check, context attached.&lt;/li&gt;
&lt;li&gt;Keep capture separate and frictionless. A one-line note that becomes a draft later is fine. A one-line note that becomes a ticket immediately is the problem.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The convention is weaker than a real state because nothing enforces it, but it recovers most of the value.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://wagglet.com/blog/wagglet-workflow-request-draft-ticket-delivery" rel="noopener noreferrer"&gt;Wagglet's walkthrough of its full workflow&lt;/a&gt; is the most explicit version of this I have seen implemented rather than described - a rough request, a private Draft used as a preparation room, then a deliberate publish onto the board, and separate stages after that for claiming, delivering evidence, and a reviewer choosing among genuinely different verdicts. It looks longer on paper. That is rather the point: the stages were always there, they were just happening implicitly inside one column and one comment thread.&lt;/p&gt;

</description>
      <category>productivity</category>
      <category>agile</category>
      <category>ai</category>
    </item>
    <item>
      <title>The delegation test is whether you can write the check first</title>
      <dc:creator>Sam Novak</dc:creator>
      <pubDate>Thu, 10 Sep 2026 09:08:05 +0000</pubDate>
      <link>https://dev.to/sam_novak_574b07811e18495/the-delegation-test-is-whether-you-can-write-the-check-first-2c08</link>
      <guid>https://dev.to/sam_novak_574b07811e18495/the-delegation-test-is-whether-you-can-write-the-check-first-2c08</guid>
      <description>&lt;p&gt;Every discussion I see about what to hand off to a coding agent sorts tasks by difficulty. Easy things to the agent, hard things to the human. It sounds obvious and it has been wrong in practice for me almost every time.&lt;/p&gt;

&lt;p&gt;The axis that actually predicts a good handoff is not difficulty. It is this: &lt;strong&gt;can you write down, before any work starts, what would make you accept the result?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If you can write the check, hand it off. If you cannot, the task is not ready to hand to anybody.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why difficulty is the wrong axis
&lt;/h2&gt;

&lt;p&gt;I have handed off genuinely hard work that went fine, because the definition of done was mechanical: migrate this to the new API, all existing tests keep passing, add coverage for the two error paths. Hard, bounded, checkable.&lt;/p&gt;

&lt;p&gt;I have also handed off something trivial that came back four times, because what I actually wanted was "make this error message less alarming", and I had no way to say what less alarming meant. That is not a hard task. It is an unspecified one, and the agent did nothing wrong - it produced four defensible answers to a question I had not asked properly.&lt;/p&gt;

&lt;p&gt;The failure was mine, and difficulty had nothing to do with it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Write the check, then write the prompt
&lt;/h2&gt;

&lt;p&gt;The useful discipline is to write the acceptance check first, as a separate artifact, before writing a single line of instruction.&lt;/p&gt;

&lt;p&gt;Checkable:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Cursor-based pagination on this endpoint. Existing tests pass. One new test for the empty-page case.&lt;/li&gt;
&lt;li&gt;Upgrade this dependency across the monorepo. Build is green. No new lint warnings.&lt;/li&gt;
&lt;li&gt;This function allocates in a hot loop. Same output for the existing fixtures, zero allocations in the steady state.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Not checkable, as written:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Make the onboarding feel less intimidating.&lt;/li&gt;
&lt;li&gt;Clean up this module.&lt;/li&gt;
&lt;li&gt;Improve performance.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The trick with the second list is that it is usually not one task. It is a decision wearing a task costume. "Make onboarding less intimidating" contains a judgement call - what counts as intimidating, and what are we willing to lose - and that part is yours. Once you have made it, what remains is checkable: cut the signup form to three fields, defer the rest to first use, keep the existing validation behaviour.&lt;/p&gt;

&lt;p&gt;So the sequence is decide, write the check, then delegate. Skipping to delegate is what produces four rounds of nearly-right work.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that surprised me
&lt;/h2&gt;

&lt;p&gt;Two things I did not expect from working this way.&lt;/p&gt;

&lt;p&gt;First, if you rigorously delegate everything checkable, what remains on your own plate is pure judgement. That is more tiring, not less. The mechanical tasks were also the ones you could do while half awake, and they let you warm up into a day. Losing all of them at once is a real cost and nobody warns you about it.&lt;/p&gt;

&lt;p&gt;Second, and more useful: &lt;strong&gt;the check is not agent infrastructure, it is people infrastructure.&lt;/strong&gt; A bounded task with a written acceptance criterion is something you can hand to a teammate who has never touched that part of the system. Without the check they need the context you have. With it, they need the check and a couple of hours.&lt;/p&gt;

&lt;p&gt;That reframing changed how I think about the whole exercise. I started writing acceptance criteria to make agents behave and ended up with a backlog that more of the team could pick up. &lt;a href="https://wagglet.com/blog/meaningful-work-with-human-ai-pairs" rel="noopener noreferrer"&gt;Wagglet has a good piece on that specific side effect&lt;/a&gt; - handoffs as a way for more people to contribute rather than purely a throughput trick, including the caveat that the same mechanism can hollow out a role if you only ever hand out the mechanical half.&lt;/p&gt;

&lt;p&gt;That caveat is real, and it is the reason I would not automate the decision step even when I could. The check is the part you write. It is also the part that is actually your job.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>programming</category>
    </item>
    <item>
      <title>Bounties for AI work have the same failure mode as game economies</title>
      <dc:creator>Sam Novak</dc:creator>
      <pubDate>Thu, 10 Sep 2026 09:05:18 +0000</pubDate>
      <link>https://dev.to/sam_novak_574b07811e18495/bounties-for-ai-work-have-the-same-failure-mode-as-game-economies-1opa</link>
      <guid>https://dev.to/sam_novak_574b07811e18495/bounties-for-ai-work-have-the-same-failure-mode-as-game-economies-1opa</guid>
      <description>&lt;p&gt;I spent a few years designing game economies before I spent any time on AI tooling, and the transition has been strange, because I keep watching people rediscover failure modes that mobile games documented years ago.&lt;/p&gt;

&lt;p&gt;The current one: bounties for AI-assisted work. Pay people for tasks completed by their agent. Put a leaderboard on it. Watch throughput go up.&lt;/p&gt;

&lt;p&gt;Throughput does go up. That is the problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rule from game economies
&lt;/h2&gt;

&lt;p&gt;Every reward loop has a currency, a source, and a sink. The loop stays healthy while the effort to satisfy the metric is higher than the effort to game it. The moment that inverts, the economy does not slow down - it accelerates in the wrong direction, and it looks like success on the dashboard the whole time.&lt;/p&gt;

&lt;p&gt;This is why a "1% drop rate" feels broken to players, why battle passes get abandoned mid-season, and why currency sinks are the least glamorous and most load-bearing part of an economy. Sources are easy. Sinks are the design.&lt;/p&gt;

&lt;h2&gt;
  
  
  Now apply it to agent work
&lt;/h2&gt;

&lt;p&gt;Ask what a submission actually costs.&lt;/p&gt;

&lt;p&gt;Before agents, submitting a completed task cost hours. The metric was expensive to satisfy and roughly impossible to game, so nobody bothered building anti-gaming controls. That cost is now close to zero. One prompt produces something task-shaped. It compiles. It has a plausible commit message.&lt;/p&gt;

&lt;p&gt;So if you pay per submitted task, you have built a source with no sink. You will get volume, immediately, and it will be indistinguishable from productivity in every chart you have.&lt;/p&gt;

&lt;p&gt;The three specific patterns I would expect, all of them rational:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Task splitting.&lt;/strong&gt; One real change submitted as four tasks, because the unit of payment is the task and nothing checks that a task is a meaningful unit of work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cheap-task farming.&lt;/strong&gt; People sort the backlog by effort ascending, which is exactly what you asked for, and the hard tickets rot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Confident non-delivery.&lt;/strong&gt; An agent reports success. Nobody checks. The reward pays out on the claim, not the outcome.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of that requires bad actors. Design for the employee who understands your metric perfectly and is not trying to cheat, because that person will find the cheapest honest path and take it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the sink looks like here
&lt;/h2&gt;

&lt;p&gt;In a game the sink removes currency. In a work system the sink is a decision that can say no, placed between the claim and the reward, and made by somebody who is not the claimant.&lt;/p&gt;

&lt;p&gt;Concretely, the parts I would not skip:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Pay on accepted, not submitted.&lt;/strong&gt; The reward event is a reviewer's verdict, not a state transition somebody performed on themselves.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Require evidence in the submission.&lt;/strong&gt; Not "done" - the actual output, the test result, the URL. An unverifiable claim should be structurally impossible to submit, not merely frowned upon.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Separate acceptance from merge truth.&lt;/strong&gt; A reviewer accepting the work and a branch actually landing are two different facts. Collapsing them means "done" quietly stops meaning anything.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Make the review load visible.&lt;/strong&gt; This is the one people miss. Reviewing is now the bottleneck and it is unpaid, so it will be done badly unless it is counted as work.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Point 4 is where I think most of these systems will actually fail. You can move the bottleneck from doing to reviewing and call it a productivity win, and the reviewers will absorb it silently for about two months.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest uncertainty
&lt;/h2&gt;

&lt;p&gt;I do not know what the right payout ratio is, and I am suspicious of anyone who says they do. Games get curve numbers by shipping and watching, and most of them get it wrong twice first.&lt;/p&gt;

&lt;p&gt;What I am fairly confident about is the shape: verified outcome, independent decision, evidence attached, review counted. &lt;a href="https://wagglet.com/blog/gamification-ai-work-bounties-rewards" rel="noopener noreferrer"&gt;Wagglet's write-up on bounties and rewards for AI work&lt;/a&gt; is the most careful version of that argument I have read, including the anti-gaming controls, which is the part usually left as an exercise for the reader.&lt;/p&gt;

&lt;p&gt;Run it on one team for a month before announcing it as a revolution. If your throughput triples in week one, that is not the good outcome - that is your economy telling you the metric is cheaper to game than to satisfy.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>gamedev</category>
    </item>
    <item>
      <title>Call whoami first: your agent should not trust a cached tool catalog</title>
      <dc:creator>Sam Novak</dc:creator>
      <pubDate>Thu, 27 Aug 2026 08:13:54 +0000</pubDate>
      <link>https://dev.to/sam_novak_574b07811e18495/call-whoami-first-your-agent-should-not-trust-a-cached-tool-catalog-1gkp</link>
      <guid>https://dev.to/sam_novak_574b07811e18495/call-whoami-first-your-agent-should-not-trust-a-cached-tool-catalog-1gkp</guid>
      <description>&lt;p&gt;There is a failure I have now watched three separate teams hit, and it always looks like a bug in the agent.&lt;/p&gt;

&lt;p&gt;An agent connects to a tool server, does useful work for a week, and then one day starts producing confidently wrong actions. It calls a tool that no longer exists. It tries to transition something it is not allowed to transition. It reports success on a mutation that silently did nothing. Everyone goes looking for the regression in the model, or in the prompt, and the actual cause is that the agent was working from a catalog of tools and permissions it had cached at some point in the past.&lt;/p&gt;

&lt;p&gt;The fix is one boring call at the start of every session, and it is worth understanding why it is not optional.&lt;/p&gt;

&lt;h2&gt;
  
  
  A token is a connection, not a role
&lt;/h2&gt;

&lt;p&gt;Here is the thing people get wrong about tokens on a permissioned tool server. The token identifies &lt;em&gt;which named connection&lt;/em&gt; is talking. It does not carry the caller's authority.&lt;/p&gt;

&lt;p&gt;That sounds like a distinction without a difference until something changes on the other side. The person the connection belongs to gets moved to a different team. Their role changes. An admin turns off tool access at the workspace level. A permission that used to be granted is revoked.&lt;/p&gt;

&lt;p&gt;In a system built correctly, the token keeps working as an identifier and the &lt;em&gt;authority&lt;/em&gt; is resolved fresh on every call from the person's current membership, roles, and permissions. Which means the answer to "can I do this?" can be yes on Monday and no on Tuesday with nothing about the token having changed. Wagglet's MCP server is explicit about this: a token identifies a named connection, and it never overrides the same person's current membership, roles, or permissions. Team-level access is off by default, and disabling it later rejects every connection request immediately without deleting the connections.&lt;/p&gt;

&lt;p&gt;If your agent cached a permission list at connection time, it is now confidently wrong about what it can do, and the only symptom you will see is a failed action it did not expect to fail.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tool schemas move under you
&lt;/h2&gt;

&lt;p&gt;The second half is the tool catalog. Any actively developed tool server adds tools, renames arguments, tightens validation, and deprecates things. That is normal and healthy.&lt;/p&gt;

&lt;p&gt;What is not healthy is an agent whose idea of the available tools came from documentation, a blog post, or a previous session's transcript. A copied old catalog is not a source of truth. The live &lt;code&gt;tools/list&lt;/code&gt; response is. This is the same discipline as not hardcoding an API response shape you saw once in a tutorial, except the failure is quieter, because a model will happily improvise a plausible call for a tool that no longer accepts that argument.&lt;/p&gt;

&lt;h2&gt;
  
  
  So: whoami first, then read, then act
&lt;/h2&gt;

&lt;p&gt;The sequence that holds up:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Establish identity and authority.&lt;/strong&gt; Call the server's &lt;code&gt;whoami&lt;/code&gt; (or equivalent) and treat its team, identity, roles, and permissions as the authority for this session. Not your prompt's belief about who you are. Not last week's.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Fetch the live tool schemas.&lt;/strong&gt; Use them, not a remembered list. If a tool you were planning to call is gone, that is information, not an error to work around.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Discover ids with bounded search, and page properly.&lt;/strong&gt; Cursors exist because the population can be bigger than one response. An agent that reads the first page and reports a total is producing a confident number that is wrong.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Read the complete authorized record before you change it.&lt;/strong&gt; Fetch the ticket or document, keep the returned record and its item revisions, and use those in the mutation. This is what makes concurrent edits safe instead of last-write-wins.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Use the matching lifecycle command, and give each intended mutation a fresh operation id.&lt;/strong&gt; Reuse an operation id only when you are deliberately retrying &lt;em&gt;the same&lt;/em&gt; intended change. This is the difference between a retry and a duplicate.&lt;/p&gt;

&lt;p&gt;The full sequence with the actual tool names is in the &lt;a href="https://wagglet.com/docs/mcp" rel="noopener noreferrer"&gt;Wagglet MCP workspace guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this is more than hygiene
&lt;/h2&gt;

&lt;p&gt;There is a design point buried in this that is worth pulling out.&lt;/p&gt;

&lt;p&gt;A permissioned tool interface is not a database with a chat wrapper. It keeps the product's own language and its own lifecycle rules, and it refuses actions that do not fit them. That refusal is the feature. An agent that could freely write any field would produce records that are structurally valid and semantically nonsense - a task marked delivered with no evidence, a review verdict with no reviewer.&lt;/p&gt;

&lt;p&gt;Which is why "read current state, then use the explicit action for what you actually mean" is not ceremony. It is the thing that keeps an agent's output reviewable by a human afterwards. If you want the reasoning behind treating delivery, acceptance, merge, and deploy as four separate facts rather than one status field, that is the argument in &lt;a href="https://wagglet.com/blog/wagglet-workflow-request-draft-ticket-delivery" rel="noopener noreferrer"&gt;the Wagglet workflow&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one-line version
&lt;/h2&gt;

&lt;p&gt;Do not let an agent start from what it remembers. Make it ask who it is, what it may do, and what tools exist - every session, out loud, before the first action. It costs one call. It saves the class of bug that looks like the model got worse and is actually your permissions changing underneath a cached answer.&lt;/p&gt;

&lt;p&gt;Background on the surrounding handoff model: &lt;a href="https://wagglet.com/how-it-works" rel="noopener noreferrer"&gt;how Wagglet works&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>devops</category>
      <category>programming</category>
    </item>
    <item>
      <title>Your coding agent needs a read-only, task-scoped, expiring credential</title>
      <dc:creator>Sam Novak</dc:creator>
      <pubDate>Thu, 27 Aug 2026 08:07:58 +0000</pubDate>
      <link>https://dev.to/sam_novak_574b07811e18495/your-coding-agent-needs-a-read-only-task-scoped-expiring-credential-2o2l</link>
      <guid>https://dev.to/sam_novak_574b07811e18495/your-coding-agent-needs-a-read-only-task-scoped-expiring-credential-2o2l</guid>
      <description>&lt;p&gt;Every team that starts handing work to coding agents hits the same wall in about week three. Someone prepares a task, someone else is going to run it, and the agent needs context. Not vague context - the ticket, the acceptance criteria, the linked spec, the three comments that changed the scope. So one of two things happens.&lt;/p&gt;

&lt;p&gt;Either the person who prepared the task pastes a wall of text into a chat window and hopes it is still accurate by the time the agent reads it, or somebody shares a credential. A personal access token. An API key with read/write on the whole workspace. In the worst version I have seen, a shared login to the tracker itself.&lt;/p&gt;

&lt;p&gt;Both are bad, and they are bad in ways that do not look like security problems at first. They look like reliability problems.&lt;/p&gt;

&lt;h2&gt;
  
  
  The snapshot goes stale, quietly
&lt;/h2&gt;

&lt;p&gt;A pasted brief is a snapshot. It was true when it was copied. Ten minutes later somebody narrows the scope in a comment, or attaches the design that answers the open question, and the agent has no way to know. It will complete the task it was given, confidently, and you will find out at review that it was the previous version of the task.&lt;/p&gt;

&lt;p&gt;The failure mode is not that the agent is wrong. It is that nothing in the system can tell you the agent was reading a stale copy. There is no version, no timestamp, no way to ask "is this still current?".&lt;/p&gt;

&lt;h2&gt;
  
  
  The credential is too big, always
&lt;/h2&gt;

&lt;p&gt;So the obvious fix is to give the agent live access. And the moment you do that with a normal token, you have handed a process that generates text the ability to change your workspace.&lt;/p&gt;

&lt;p&gt;Think about what a standard workspace token can usually do: read any ticket, comment anywhere, transition anything, sometimes delete. You needed exactly one of those powers - read this one task and the things it points at. You granted all of them, for as long as the token lives, which is usually forever.&lt;/p&gt;

&lt;p&gt;And you granted them to the wrong identity. A shared token is not "the agent" or "the runner", it is whoever created the token. Every action lands in the audit log under their name. Six weeks later, when you are trying to work out who moved the ticket, the log says a person who was on holiday.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a bounded credential looks like
&lt;/h2&gt;

&lt;p&gt;The version that actually works is narrower on four separate axes, and it is worth being explicit about all four, because most homegrown solutions get two right and forget the others.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Read-only.&lt;/strong&gt; Not "read-mostly". The capability the agent starts with should have no ability to deliver, transition, comment, or edit anything. If the agent's report is going to move the task forward, that should be a separate, later, deliberately authorized action - not something the same credential could do at any moment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rooted at one resource.&lt;/strong&gt; Scoped to this claim, on this task. Not to the project, not to the team. If the agent follows a link out of the task to something it was not granted, the answer should be a refusal, not a helpful response. This is the axis people skip, because project-wide scoping is so much easier to build.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Expiring.&lt;/strong&gt; Days, not months. The task will be done or abandoned long before then. A credential that outlives the work it was minted for is just a key under the doormat.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Revocable independently.&lt;/strong&gt; Turning off agent access for the team should reject connections immediately, without deleting them and without anybody having to hunt down individual tokens.&lt;/p&gt;

&lt;p&gt;Wagglet's handoff is built on exactly this shape: when a teammate claims a task, the copied prompt carries a read-only, claim-rooted context capability plus a fallback snapshot. The agent can refresh the permitted current context - so the staleness problem goes away - but that starting capability cannot deliver or change anything. The details are in the &lt;a href="https://wagglet.com/docs/mcp" rel="noopener noreferrer"&gt;MCP workspace guide&lt;/a&gt;, and the surrounding flow is on &lt;a href="https://wagglet.com/how-it-works" rel="noopener noreferrer"&gt;how it works&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Identity stays with the person
&lt;/h2&gt;

&lt;p&gt;The other half of this, which is easy to miss: the runner uses their own agent subscription. Their own Claude Code or Codex account, their own token balance. Nobody shares an agent login to move a task, which is the practice this whole pattern exists to kill.&lt;/p&gt;

&lt;p&gt;That matters for a boring reason and an interesting one. The boring reason is terms of service. The interesting one is that "who ran this" and "who prepared this" are genuinely different facts, and a shared login collapses them into one name. Once collapsed, you cannot answer basic questions about your own process.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cheap version, if you are not adopting anything
&lt;/h2&gt;

&lt;p&gt;You do not need a product to get most of this. If you are rolling your own handoff today:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Mint a token per task, not per person, and give it read scope only.&lt;/li&gt;
&lt;li&gt;Set an expiry you would be comfortable defending, then halve it.&lt;/li&gt;
&lt;li&gt;Put the resource id in the token's scope and enforce it server-side, not in the prompt. A prompt instruction to "only read ticket 412" is a suggestion.&lt;/li&gt;
&lt;li&gt;Log the human runner as the actor for anything the agent causes, and keep the agent's own report as evidence rather than as a state change.&lt;/li&gt;
&lt;li&gt;Make the delivery step a separate call with a separate authorization, so "the agent said it is done" and "the task is done" stay different facts.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The pattern is not complicated. It is just narrower than the credential you already have lying around, which is why almost everybody reaches for the wrong one first.&lt;/p&gt;

&lt;p&gt;More on the task-design side of this: &lt;a href="https://wagglet.com/blog/dual-prompt-human-agent-task-design" rel="noopener noreferrer"&gt;the dual prompt&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>devops</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
