<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Wasim Sheikh</title>
    <description>The latest articles on DEV Community by Wasim Sheikh (@anciwasim).</description>
    <link>https://dev.to/anciwasim</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4124806%2F8b135459-4284-4c96-b849-753f453e7a89.jpg</url>
      <title>DEV Community: Wasim Sheikh</title>
      <link>https://dev.to/anciwasim</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/anciwasim"/>
    <language>en</language>
    <item>
      <title>Boundary Contract: Make Every Agent Tool Fail Boring</title>
      <dc:creator>Wasim Sheikh</dc:creator>
      <pubDate>Tue, 06 Oct 2026 22:12:34 +0000</pubDate>
      <link>https://dev.to/anciwasim/boundary-contract-make-every-agent-tool-fail-boring-15bm</link>
      <guid>https://dev.to/anciwasim/boundary-contract-make-every-agent-tool-fail-boring-15bm</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Cross-post of Insights #8 — canonical: &lt;a href="https://sheikhwasim.com/insights/boundary-contract-tools-fail-boring/" rel="noopener noreferrer"&gt;https://sheikhwasim.com/insights/boundary-contract-tools-fail-boring/&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Most agent failures I see in production don't start in the model. They start at the edge of a tool.&lt;/p&gt;

&lt;p&gt;A call times out. An API returns an empty list. A search gives back three results out of thirty. The model gets something shaped like data, so it does what models do: it writes a fluent, confident answer on top of it.&lt;/p&gt;

&lt;p&gt;The transcript reads fine. The customer gets the wrong answer. Nobody notices until a human asks why.&lt;/p&gt;

&lt;p&gt;I call the missing control a &lt;strong&gt;Boundary Contract&lt;/strong&gt;: every tool the agent can call returns a typed outcome, not just a payload. &lt;em&gt;Ok&lt;/em&gt;, &lt;em&gt;empty&lt;/em&gt;, &lt;em&gt;partial&lt;/em&gt;, &lt;em&gt;timeout&lt;/em&gt;, &lt;em&gt;rejected&lt;/em&gt;. The agent has to handle each one by name. It can't turn a broken call into the truth.&lt;/p&gt;

&lt;p&gt;Clever prose can wait. The boundary has to be boring first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quick use case
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Situation.&lt;/strong&gt; A field-service company shipped a scheduling agent. A customer asks for a repair visit. The agent checks the technician calendar, offers the next open slot, and books it. Staging looked sharp.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What broke.&lt;/strong&gt; On a busy Monday the calendar service started timing out under load. The tool wrapper caught the timeout and returned an empty list, because that's what the original developer thought "safe" meant. The agent read "no slots" and told customers the next opening was three weeks out. Some cancelled. Some called a competitor. The logs showed a successful tool call, a polite reply, and no error anywhere. The dashboard stayed green all morning.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix (Boundary Contract).&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Typed outcomes on every tool.&lt;/strong&gt; The calendar tool now returns &lt;code&gt;ok&lt;/code&gt;, &lt;code&gt;empty&lt;/code&gt;, &lt;code&gt;partial&lt;/code&gt;, &lt;code&gt;timeout&lt;/code&gt;, or &lt;code&gt;rejected&lt;/code&gt;, plus the payload. "Empty" and "timed out" stopped looking the same.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Input checked before the call.&lt;/strong&gt; A bad zip code or a missing service type gets &lt;code&gt;rejected&lt;/code&gt; with a reason, before it reaches the calendar API and comes back as a mystery.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A named policy per outcome.&lt;/strong&gt; &lt;code&gt;timeout&lt;/code&gt; means retry once, then tell the customer "I can't see the calendar right now" and offer a callback. Never "no slots."
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Partial results are surfaced.&lt;/strong&gt; If one region's calendar answered and another didn't, the agent says so instead of quietly offering only what came back.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The outcome goes in the trace.&lt;/strong&gt; Every run logs which outcome each tool returned, so "how many timeouts did we hide this week" has a real answer.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Same model. Same prompt. The difference was that a broken call could no longer pass as an answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why "it returned something" isn't enough
&lt;/h2&gt;

&lt;p&gt;A tool payload answers: &lt;em&gt;What came back?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A Boundary Contract answers: &lt;em&gt;What actually happened at the edge, and what is the agent allowed to say about it?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Most stacks only fund the first question. They wrap errors into empty results to "keep the agent from crashing," and the model fills the gap with confidence. That isn't resilience. It's a lie with good grammar.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Boundary Contract (five layers)
&lt;/h2&gt;

&lt;p&gt;Build it into the tool layer, not into a line of the system prompt that says "be careful with errors."&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Input contract
&lt;/h3&gt;

&lt;p&gt;Each tool declares what valid input looks like: required fields, allowed values, sane ranges. Invalid input is rejected &lt;em&gt;before&lt;/em&gt; the call, with a reason the agent can repeat or fix. A tool that accepts anything will eventually do anything.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Outcome types, not just payloads
&lt;/h3&gt;

&lt;p&gt;Every response carries one of a small, fixed set of outcomes. Five is enough for most tools: &lt;code&gt;ok&lt;/code&gt;, &lt;code&gt;empty&lt;/code&gt;, &lt;code&gt;partial&lt;/code&gt;, &lt;code&gt;timeout&lt;/code&gt;, &lt;code&gt;rejected&lt;/code&gt;. The important part is that &lt;em&gt;empty&lt;/em&gt; and &lt;em&gt;failed&lt;/em&gt; are never the same value.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Time budget, named
&lt;/h3&gt;

&lt;p&gt;Each tool has an explicit timeout and retry rule written down next to it. When the budget runs out, the outcome says &lt;code&gt;timeout&lt;/code&gt;. It doesn't silently become &lt;code&gt;empty&lt;/code&gt;, and it doesn't hang until the whole run dies.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Outcome policy in the orchestrator
&lt;/h3&gt;

&lt;p&gt;For each outcome, decide what the agent may do and say. &lt;code&gt;partial&lt;/code&gt; means disclose the gap. &lt;code&gt;timeout&lt;/code&gt; means fall back or hand off to a human. &lt;code&gt;rejected&lt;/code&gt; means ask a clarifying question. Write these as rules the orchestrator enforces, so the model doesn't improvise them per ticket.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Outcome in the trace
&lt;/h3&gt;

&lt;p&gt;Log the outcome for every tool call, not only the payload. Freeze one happy path and one timeout path as &lt;a href="https://sheikhwasim.com/insights/golden-trace-eval-anchor/" rel="noopener noreferrer"&gt;Golden Traces&lt;/a&gt;, and add "timeout must never produce a confident answer" to your &lt;a href="https://sheikhwasim.com/insights/smoke-evals-before-you-ship/" rel="noopener noreferrer"&gt;Smoke Evals&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  How this maps to what you already have
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;5-layer agent stack.&lt;/strong&gt; Boundary Contract lives in &lt;strong&gt;Tools&lt;/strong&gt;, enforced by &lt;strong&gt;Guardrails&lt;/strong&gt;, recorded in &lt;strong&gt;Traces&lt;/strong&gt;.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Receipt Gate.&lt;/strong&gt; Receipt Gate checks a write landed before the next one fires. Boundary Contract makes sure the agent knows when a read &lt;em&gt;didn't&lt;/em&gt; land.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Permission Envelope.&lt;/strong&gt; The envelope decides whether a tool may run. The contract decides what its result means.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cheap Twin.&lt;/strong&gt; A &lt;code&gt;partial&lt;/code&gt; or &lt;code&gt;timeout&lt;/code&gt; outcome is a good escalate line on the checklist.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ship Gate.&lt;/strong&gt; Inject a timeout in canary and prove the agent says "I can't see that right now" instead of making something up.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The test
&lt;/h2&gt;

&lt;p&gt;Pick the tool your agent calls most. Then ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If this tool timed out right now, would the agent tell the user something true, and would the trace show it was a timeout?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Run it for real: block the endpoint in staging and read the reply. If the agent still answers confidently, the boundary isn't a contract yet. It's a hope.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;Agents don't need to be clever at the edges. They need to be honest there. Every tool should fail in a way that's boring, typed, and visible, so the clever parts of the system are built on something true.&lt;/p&gt;

&lt;p&gt;I'm Wasim Sheikh, AI Architect. I build systems teams trust and organizations depend on: not demos, not proofs of concept, production.&lt;/p&gt;

&lt;p&gt;Follow for practical AI architecture that ships.&lt;br&gt;&lt;br&gt;
&lt;strong&gt;Connect on LinkedIn:&lt;/strong&gt; &lt;a href="https://www.linkedin.com/in/wasim-sheikh/" rel="noopener noreferrer"&gt;Wasim Sheikh&lt;/a&gt; · &lt;strong&gt;Site:&lt;/strong&gt; &lt;a href="https://sheikhwasim.com" rel="noopener noreferrer"&gt;sheikhwasim.com&lt;/a&gt; · &lt;strong&gt;Notes:&lt;/strong&gt; &lt;a href="https://practicalainotes.substack.com/" rel="noopener noreferrer"&gt;Practical AI Notes&lt;/a&gt; · &lt;strong&gt;X:&lt;/strong&gt; &lt;a href="https://x.com/anciwasim" rel="noopener noreferrer"&gt;@anciwasim&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>devops</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Clean Exit, Wrong Record: Grade the Agent on the Database, Not the Last Line</title>
      <dc:creator>Wasim Sheikh</dc:creator>
      <pubDate>Sun, 04 Oct 2026 21:16:50 +0000</pubDate>
      <link>https://dev.to/anciwasim/clean-exit-wrong-record-grade-the-agent-on-the-database-not-the-last-line-2a4c</link>
      <guid>https://dev.to/anciwasim/clean-exit-wrong-record-grade-the-agent-on-the-database-not-the-last-line-2a4c</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Cross-post of an Insights piece. Canonical: &lt;a href="https://sheikhwasim.com/insights/clean-exit-wrong-record/" rel="noopener noreferrer"&gt;https://sheikhwasim.com/insights/clean-exit-wrong-record/&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;An agent can do the careful work, close the ticket, and still be wrong.&lt;/p&gt;

&lt;p&gt;On October 3, Microsoft and Hugging Face published ThinkingBox. It grades agents on the records they leave behind, not the sentences they generate. Their walkthrough is an ordinary support case: a delayed kitchen appliance, nine tool calls, the policy read correctly, the ticket closed as resolved, and a polite closer. The required end state was on hold. The customer never got a real answer. A grader watching tool calls would pass it. The database wouldn't.&lt;/p&gt;

&lt;p&gt;That's not a leaderboard story. It's the bug already sitting in production agents that treat a clean exit as proof of work.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://sheikhwasim.com/insights/receipt-gate-no-second-write/" rel="noopener noreferrer"&gt;Receipt Gate&lt;/a&gt; says the last write has to leave a skim-worthy receipt before the next one fires. ThinkingBox is the colder half of the same rule: the receipt has to match the record. The model's last line doesn't get a vote.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quick use case
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Situation.&lt;/strong&gt; A support team shipped a delivery-exception agent. A stuck shipment comes in. The agent can read the order, check the carrier, open or update a ticket, and reply to the customer. The eval was a golden transcript: the right tools fired, the reply was polite, and the last line said the case was resolved. Staging looked done.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What broke.&lt;/strong&gt; Put that published retail task on their stack. Tools return success. The ticket lands in solved. The customer gets the closer. The carrier exception is still open, so the required end state is hold, not solved. Nothing in the eval looks at that field. The run is green because the trajectory is tidy. The record isn't.&lt;/p&gt;

&lt;p&gt;The ThinkingBox writeup is blunt about how often that tidy exit is a costume. In a common-set ablation (121,680 valid trials across 12 models), 79,853 attempts failed the executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error. The checks still found wrong field values in 77.61% of them, unintended extra effects in 43.30%, and missing required effects in 25.36%. Those findings overlap. A green trace is not a passed state.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix.&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Name the end state in fields&lt;/strong&gt;: status, required effects, forbidden extras. Not "the reply sounds resolved."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read the record back&lt;/strong&gt; after the last mutating tool, from the system of record, not from the model's summary.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Latch the customer-visible step&lt;/strong&gt; on that readback. No send until the fields match.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Receipt the diff&lt;/strong&gt;: what the record shows, what it was supposed to show, pass or fail in one skim.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Repeat the golden task&lt;/strong&gt; from a clean start until a one-off pass stops counting as a ship.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Same tools. Same model. Different definition of done. To be fair, they haven't measured the lift from checking terminal state before you commit. You still have to run it on your own records.&lt;/p&gt;

&lt;h2&gt;
  
  
  A trajectory is a claim
&lt;/h2&gt;

&lt;p&gt;The line worth keeping: a trajectory is a claim. Database state is the evidence. Repetition is the trust test.&lt;/p&gt;

&lt;p&gt;Pass@1 says how a model usually does. "Solved at least once in 20 tries" says whether it can ever do the task. Neither answers "will it do this every time." On their 507 workflows, each run 20 times from a clean backend, Claude Opus 5.5 led the overall pass@1 they report, at 67.16%. It passed the same number of tasks on all 20 attempts as Claude Opus 5: 241. A small bump in the headline score bought no extra dependability. Kimi-K3 solved 476 of 507 tasks at least once and only 68 of them every time.&lt;/p&gt;

&lt;p&gt;If you're picking a model for work that touches real records off the "solved at least once" column, you're buying a demo.&lt;/p&gt;

&lt;p&gt;Where the failures sit matters for how you spend your week. They label roughly four in five as tool handling, not reasoning: tool usage at 79.9% in an unweighted average of per-model shares, not a unique cause and not a law. Agents usually get far enough to try the workflow, then fail to recover from tool errors, failed preconditions, or empty lookups. That's a retry policy and a smaller tool surface before it's a shopping trip for a newer model.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://sheikhwasim.com/insights/cheap-twin-escalate-when-needed/" rel="noopener noreferrer"&gt;Cheap Twin&lt;/a&gt; still applies: don't send every ticket to the flagship, and don't route on a single-pass score. In their list-rate estimate (not an invoice), the cheapest single success wasn't the cheapest task that passed every repeat.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five lines before the next closer ships
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Required end state is a field diff, written down before the run&lt;/li&gt;
&lt;li&gt;Post-write readback from the system of record, not the assistant message&lt;/li&gt;
&lt;li&gt;Customer-visible send stays closed on mismatch&lt;/li&gt;
&lt;li&gt;Receipt names the fields that matched, the fields that didn't, and any extra writes&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://sheikhwasim.com/insights/ship-gate-pre-deploy-checklist/" rel="noopener noreferrer"&gt;Ship Gate&lt;/a&gt; uses a repeat count you can defend, and you say whether you mean best-of-k or every-of-k&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Twenty is their number, not a religion. A refund that moves money may need a stricter every-time bar than a draft note. Name k. Say which one you mean. Grade the sentence only when the requirement has no field.&lt;/p&gt;

&lt;h2&gt;
  
  
  How this sits with the rest of the stack
&lt;/h2&gt;

&lt;p&gt;Receipt Gate already refuses the next write when the last step left mush. This tightens the receipt: it has to agree with a readback. &lt;a href="https://sheikhwasim.com/insights/golden-trace-eval-anchor/" rel="noopener noreferrer"&gt;Golden Trace&lt;/a&gt; should freeze the end state, not only the happy transcript, otherwise a prompt bump that still "sounds resolved" sails right through. &lt;a href="https://sheikhwasim.com/insights/smoke-evals-before-you-ship/" rel="noopener noreferrer"&gt;Smoke evals&lt;/a&gt; need one trap where every tool returns ok and a required field is wrong.&lt;br&gt;
 &lt;a href="https://sheikhwasim.com/insights/permission-envelope-consent-that-expires/" rel="noopener noreferrer"&gt;Permission Envelope&lt;/a&gt; still wraps the irreversible send; a correct record doesn't grant a standing right to email the customer. Cheap Twin can pick the brain. It can't pick the definition of done.&lt;/p&gt;

&lt;h2&gt;
  
  
  The test
&lt;/h2&gt;

&lt;p&gt;Take one workflow that writes a record and then tells a human it's finished. Run it five times from the same starting state. After each run, ignore the last message. Read the fields you claimed were the goal.&lt;/p&gt;

&lt;p&gt;If any run says resolved while a required field is wrong, your eval is grading the claim.&lt;/p&gt;

&lt;p&gt;Curious how others handle this: do you read the record back before the agent talks to the customer, or are you still trusting the closer?&lt;/p&gt;




&lt;p&gt;I'm Wasim Sheikh, AI Architect. I build systems teams trust and organizations depend on: production, not demos.&lt;/p&gt;

&lt;p&gt;Full piece and the rest of the series: &lt;a href="https://sheikhwasim.com/insights/clean-exit-wrong-record/" rel="noopener noreferrer"&gt;sheikhwasim.com&lt;/a&gt; · &lt;a href="https://www.linkedin.com/in/wasim-sheikh/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://practicalainotes.substack.com/" rel="noopener noreferrer"&gt;Practical AI Notes&lt;/a&gt; · &lt;a href="https://x.com/anciwasim" rel="noopener noreferrer"&gt;X @anciwasim&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;p&gt;Numbers and the retail walkthrough are from the October 3, 2026 joint post. I didn't re-run the benchmark. Cost remarks there are the authors' list-rate estimates (OpenRouter snapshot noted as September 20, 2026), not production invoices. They say they haven't measured the lift from terminal-state checks, a smaller tool surface, or human approval on this benchmark.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Microsoft and Hugging Face, "The Agent Said It Was Done. The Database Disagreed." Tuhin Kundu, October 3, 2026. &lt;a href="https://huggingface.co/blog/microsoft/thinkingbox" rel="noopener noreferrer"&gt;https://huggingface.co/blog/microsoft/thinkingbox&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Benchmark data repo cited in that post (release tag &lt;code&gt;thinkingbox-bench-v1.0&lt;/code&gt;): &lt;a href="https://github.com/microsoft/thinkingbox-data" rel="noopener noreferrer"&gt;https://github.com/microsoft/thinkingbox-data&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>testing</category>
      <category>devops</category>
    </item>
    <item>
      <title>Cheap Twin: Escalate Only When the Small Model Needs Help</title>
      <dc:creator>Wasim Sheikh</dc:creator>
      <pubDate>Sun, 04 Oct 2026 17:02:19 +0000</pubDate>
      <link>https://dev.to/anciwasim/cheap-twin-escalate-only-when-the-small-model-needs-help-1e9i</link>
      <guid>https://dev.to/anciwasim/cheap-twin-escalate-only-when-the-small-model-needs-help-1e9i</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Cross-post of Insights #7 — canonical: &lt;a href="https://sheikhwasim.com/insights/cheap-twin-escalate-when-needed/" rel="noopener noreferrer"&gt;https://sheikhwasim.com/insights/cheap-twin-escalate-when-needed/&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Most agent stacks send every request to the flagship model "just in case."&lt;/p&gt;

&lt;p&gt;That looks careful. It is usually waste — and it hides when the cheap path was already good enough. Latency climbs. The bill spikes. Nobody can answer: &lt;em&gt;which tickets actually needed the heavyweight brain?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I call the missing control a &lt;strong&gt;Cheap Twin&lt;/strong&gt;: run the same input through a small model first. Escalate to the expensive brain only when the twin disagrees with a short checklist, or fails a smoke assert. Everything else ships on the twin — with a receipt that names which model answered.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://sheikhwasim.com/insights/receipt-gate-no-second-write/" rel="noopener noreferrer"&gt;Receipt Gate&lt;/a&gt; makes each write legible. Cheap Twin makes each &lt;em&gt;brain choice&lt;/em&gt; legible — and defaults the happy path to the model you can afford to run all day.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quick use case
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Situation.&lt;/strong&gt; A commercial construction ops team shipped an "RFI triage agent." It scored inbound requests for information: trade, urgency, and suggested owner. Leadership wanted accuracy, so every ticket went to the flagship model. Staging looked sharp.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What broke.&lt;/strong&gt; Volume hit production. Latency climbed into the awkward zone where PMs refreshed and re-submitted. The model bill jumped without a matching jump in "hard" tickets. Nobody could say which RFIs needed the heavyweight path. Logs showed tokens and a model name — not &lt;em&gt;why&lt;/em&gt; that model was chosen. There was no choice. There was a default.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix (Cheap Twin).&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Run the twin first&lt;/strong&gt; — small model proposes trade / urgency / owner on every RFI&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gate escalate on a short checklist&lt;/strong&gt; — missing deadline, multi-trade conflict, owner-facing language, dollar threshold, unknown vendor&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fail closed on smoke asserts&lt;/strong&gt; — if the twin output fails a cheap structural check, escalate (or hold) instead of shipping mush&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Escalate with the twin's draft&lt;/strong&gt; — flagship gets the same input &lt;em&gt;plus&lt;/em&gt; the twin proposal and which checklist line fired&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Receipt the brain&lt;/strong&gt;  every ticket logs twin vs flagship, checklist hits, and cost — skim-worthy, not a token dump&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Same tools. Same schema. Most traffic never left the twin. When it did, the receipt said why.&lt;/p&gt;

&lt;p&gt;Cheap Twins aren't about starving the agent of intelligence. They're about refusing the expensive path when you have no signal you need it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Flagship-by-default is a missing control
&lt;/h2&gt;

&lt;p&gt;A flagship call answers: &lt;em&gt;What would the best model say?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A Cheap Twin answers: &lt;em&gt;Did we need the best model for this ticket — and can we prove it?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Those are different questions. Most stacks fund the first and call the bill "the cost of quality." Quality without a routing receipt is hope with a larger context window: you cannot tune the checklist, Golden-Trace the cheap path, or tell finance which spike was complexity vs lazy routing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Cheap Twin stack (five layers)
&lt;/h2&gt;

&lt;p&gt;Build it like a control on the model path, not a discount coupon in the prompt.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Twin-first contract
&lt;/h3&gt;

&lt;p&gt;Every eligible request hits the small model before the flagship. The twin returns a structured proposal in the same schema the flagship would use — not a chatty essay the orchestrator reinterprets.&lt;/p&gt;

&lt;p&gt;If the twin cannot speak the production schema, it is not a twin. It is a different product.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Escalate checklist (five lines, not fifty)
&lt;/h3&gt;

&lt;p&gt;Name the concrete reasons the cheap path is not enough. Keep the list short enough to argue in standup:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Missing or ambiguous deadline&lt;/li&gt;
&lt;li&gt;Multi-trade / multi-system conflict&lt;/li&gt;
&lt;li&gt;Owner-facing or customer-facing language&lt;/li&gt;
&lt;li&gt;Dollar / risk threshold crossed&lt;/li&gt;
&lt;li&gt;Unknown vendor, SKU, or entity&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Vague scores ("confidence &amp;lt; 0.7") belong behind a named line — or they become a knob nobody trusts.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Smoke assert before ship
&lt;/h3&gt;

&lt;p&gt;Even on the twin path, run a cheap structural assert before anything leaves: required fields present, enums valid, no empty owner, no "see above" mush. Fail → escalate or hold. Do not ship a polite brick because the small model was fast.&lt;/p&gt;

&lt;p&gt;Pair with your &lt;a href="https://sheikhwasim.com/insights/smoke-evals-before-you-ship/" rel="noopener noreferrer"&gt;Smoke Evals&lt;/a&gt; pack so the same traps block CI and live routing.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Escalate with context, not from scratch
&lt;/h3&gt;

&lt;p&gt;When a checklist line fires, the flagship should see the twin's proposal and &lt;em&gt;which line&lt;/em&gt; triggered the escalate. That keeps the expensive call short and disagreement inspectable. Blind re-runs that ignore the twin waste the money you just decided to spend.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Brain receipt in the run record
&lt;/h3&gt;

&lt;p&gt;Every ticket leaves a skim-worthy receipt: twin model id, flagship model id (or "not called"), checklist hits, assert pass/fail, latency, and cost. Mushy "used GPT" fails the gate — same rule as &lt;a href="https://sheikhwasim.com/insights/receipt-gate-no-second-write/" rel="noopener noreferrer"&gt;Receipt Gate&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Freeze a twin-happy path and a forced-escalate path as &lt;a href="https://sheikhwasim.com/insights/golden-trace-eval-anchor/" rel="noopener noreferrer"&gt;Golden Traces&lt;/a&gt;. Prompt and model bumps must not silently flip the routing mix.&lt;/p&gt;

&lt;h2&gt;
  
  
  How this maps to what you already have
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://sheikhwasim.com/insights/agent-architecture-five-layer-stack/" rel="noopener noreferrer"&gt;5-layer agent stack&lt;/a&gt;&lt;/strong&gt; — Cheap Twin lives in &lt;strong&gt;Models&lt;/strong&gt; + &lt;strong&gt;Guardrails&lt;/strong&gt;; &lt;strong&gt;Traces&lt;/strong&gt; carry the brain receipt&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://sheikhwasim.com/insights/smoke-evals-before-you-ship/" rel="noopener noreferrer"&gt;Smoke evals&lt;/a&gt;&lt;/strong&gt; — "twin ships when checklist clear" and "checklist hit → escalate"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://sheikhwasim.com/insights/golden-trace-eval-anchor/" rel="noopener noreferrer"&gt;Golden Trace&lt;/a&gt;&lt;/strong&gt; — freeze one twin-only success and one escalate success; regress both&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://sheikhwasim.com/insights/receipt-gate-no-second-write/" rel="noopener noreferrer"&gt;Receipt Gate&lt;/a&gt;&lt;/strong&gt; — model choice deserves a plain-language receipt before the next write&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://sheikhwasim.com/insights/permission-envelope-consent-that-expires/" rel="noopener noreferrer"&gt;Permission Envelope&lt;/a&gt;&lt;/strong&gt; — escalate widens &lt;em&gt;which brain&lt;/em&gt; may propose, not which tools&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://sheikhwasim.com/insights/ship-gate-pre-deploy-checklist/" rel="noopener noreferrer"&gt;Ship Gate&lt;/a&gt;&lt;/strong&gt; — prove routing + receipts in canary, not that the demo always hits flagship&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Architecture without a twin is just a flagship with a louder invoice.&lt;/p&gt;

&lt;h2&gt;
  
  
  The test
&lt;/h2&gt;

&lt;p&gt;Ask one question before the next "always use the best model" ship:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;For a normal ticket, does a small model propose first — and can a teammate skim a receipt that says whether you escalated, which checklist line fired, and what it cost?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If the answer is no, you are still buying flagship by default.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;I'm Wasim Sheikh — AI Architect. I build systems teams trust and organizations depend on: not demos, not proofs of concept — production.&lt;/p&gt;

&lt;p&gt;Follow for practical AI architecture that ships.&lt;br&gt;&lt;br&gt;
&lt;strong&gt;Connect on LinkedIn:&lt;/strong&gt; &lt;a href="https://www.linkedin.com/in/wasim-sheikh/" rel="noopener noreferrer"&gt;Wasim Sheikh&lt;/a&gt; · &lt;strong&gt;Site:&lt;/strong&gt; &lt;a href="https://sheikhwasim.com" rel="noopener noreferrer"&gt;sheikhwasim.com&lt;/a&gt; · &lt;strong&gt;Notes:&lt;/strong&gt; &lt;a href="https://practicalainotes.substack.com/" rel="noopener noreferrer"&gt;Practical AI Notes&lt;/a&gt; · &lt;strong&gt;X:&lt;/strong&gt; &lt;a href="https://x.com/anciwasim" rel="noopener noreferrer"&gt;@anciwasim&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>devops</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Receipt Gate: No Second Write Without a Receipt</title>
      <dc:creator>Wasim Sheikh</dc:creator>
      <pubDate>Tue, 29 Sep 2026 15:38:46 +0000</pubDate>
      <link>https://dev.to/anciwasim/receipt-gate-no-second-write-without-a-receipt-1a4e</link>
      <guid>https://dev.to/anciwasim/receipt-gate-no-second-write-without-a-receipt-1a4e</guid>
      <description>&lt;p&gt;Most agent failures are not one bad tool call. They are three "successful" writes in a row that nobody can reconstruct.&lt;/p&gt;

&lt;p&gt;Tool A updates the CRM. Tool B opens a ticket. Tool C emails the customer. Each step returns &lt;code&gt;ok&lt;/code&gt;. The dashboard stays green. Hours later support asks why a customer got a resolution note for a case that never closed — and the only answer lives in a token dump no human will read.&lt;/p&gt;

&lt;p&gt;I call the missing control a &lt;strong&gt;Receipt Gate&lt;/strong&gt;: after every write, produce a short, human-readable receipt. The next write stays closed until that receipt is skim-worthy. No receipt, no second write. Most cascading messes die right there.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://sheikhwasim.com/insights/golden-trace-eval-anchor/" rel="noopener noreferrer"&gt;Golden Trace&lt;/a&gt; freezes the path that worked. Receipt Gate makes each live step legible before the next one fires.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quick use case
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Situation.&lt;/strong&gt; A revenue-ops team shipped an "exception closer" agent. On a flagged invoice it could (1) patch the CRM opportunity stage, (2) post a note on the finance ticket, and (3) send the customer a status email. Staging looked sharp: three tools, one happy path, under thirty seconds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What broke.&lt;/strong&gt; A model bump plus a "helpful" prompt tweak patched the CRM to Closed-Won while the ticket note still said "pending PO match." The email went out with a warm "resolved" tone. Every tool returned success. Nobody could skim, in ten seconds, what actually changed.&lt;/p&gt;

&lt;p&gt;Ops found it when a customer asked for a closed invoice that did not exist. Logs showed tokens, not a receipt chain a human could trust.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix (Receipt Gate).&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Require a receipt after every write&lt;/strong&gt; — object, fields, tool, success/fail in plain language.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gate the next write on that receipt&lt;/strong&gt; — tool B cannot fire until tool A's receipt is present and non-mushy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Make receipts skim-first&lt;/strong&gt; — ten-second read; raw JSON alone is not enough.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fail closed on mush&lt;/strong&gt; — vague "updated record" blocks the chain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep the chain in the run record&lt;/strong&gt; — receipts travel with the trace for audits and &lt;a href="https://sheikhwasim.com/insights/golden-trace-eval-anchor/" rel="noopener noreferrer"&gt;Golden Trace&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Same three tools. Same model family. Different blast radius. The next "Closed-Won + resolved email" never left the first gate.&lt;/p&gt;

&lt;p&gt;Receipt Gates aren't about slowing the agent for sport. They're about refusing to stack silent side effects before a human — or the next step — can see them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Success responses lie. Receipts don't have to.
&lt;/h2&gt;

&lt;p&gt;A tool status code answers: &lt;em&gt;Did the API accept the call?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A receipt answers: &lt;em&gt;What business state actually changed, in words a teammate can check?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Those are different questions. Most stacks fund the first and hope the second shows up in observability later.&lt;/p&gt;

&lt;p&gt;When writes chain, hope is the bug. Later tools inherit a world the first write may have only partly changed — or changed the wrong way while still returning 200.&lt;/p&gt;

&lt;p&gt;If the only bridge is the model's short-term memory, you are one polite hallucination away from a customer-facing contradiction.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Receipt Gate stack (five layers)
&lt;/h2&gt;

&lt;p&gt;Build it like a control on the write path, not a nicer log sink.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Write → receipt contract
&lt;/h3&gt;

&lt;p&gt;Every mutating tool emits a structured receipt before the orchestrator continues:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Object / resource id&lt;/li&gt;
&lt;li&gt;Fields or side effects that changed&lt;/li&gt;
&lt;li&gt;Tool name + schema id&lt;/li&gt;
&lt;li&gt;Inputs that mattered (redacted)&lt;/li&gt;
&lt;li&gt;Outcome in plain language: success, partial, or fail&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the tool cannot say what it did, it is not ready for production chaining.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Skim rule (ten seconds)
&lt;/h3&gt;

&lt;p&gt;A receipt only a debugger can parse is not a receipt. Target a human skim: one-sentence summary, concrete change bullets, and an explicit "did not change" when an expected field was skipped.&lt;/p&gt;

&lt;p&gt;Mushy language ("synced the account," "handled the exception") fails the gate. Specificity is the product.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Next-write latch
&lt;/h3&gt;

&lt;p&gt;The orchestrator holds a latch: the next mutating tool stays closed until the prior receipt passes schema + skim checks. Read-only tools can still run. Irreversible or customer-visible writes wait.&lt;/p&gt;

&lt;p&gt;Cascading damage dies at the latch — not in the postmortem.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Chain in the run record
&lt;/h3&gt;

&lt;p&gt;Persist receipts in order next to the trace. Receipt[n] can reference Receipt[n-1]. Partial failures keep the chain. Ops get an export without opening the full token stream.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://sheikhwasim.com/insights/golden-trace-eval-anchor/" rel="noopener noreferrer"&gt;Golden Trace&lt;/a&gt; can freeze a successful chain. &lt;a href="https://sheikhwasim.com/insights/smoke-evals-before-you-ship/" rel="noopener noreferrer"&gt;Smoke evals&lt;/a&gt; can assert "no second write without receipt." &lt;a href="https://sheikhwasim.com/insights/ship-gate-pre-deploy-checklist/" rel="noopener noreferrer"&gt;Ship Gate&lt;/a&gt; should refuse agents that mutate twice with nothing skimmable in between.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Human or policy break-glass
&lt;/h3&gt;

&lt;p&gt;Sometimes a human must accept a partial receipt and continue. That is fine if it is explicit: named approver, why mush was allowed, and a short TTL on the override (pair with &lt;a href="https://sheikhwasim.com/insights/permission-envelope-consent-that-expires/" rel="noopener noreferrer"&gt;Permission Envelope&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;Silent auto-continue on a failed skim check is how the CRM went Closed-Won again.&lt;/p&gt;

&lt;h2&gt;
  
  
  How this maps to what you already have
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://sheikhwasim.com/insights/agent-architecture-five-layer-stack/" rel="noopener noreferrer"&gt;5-layer agent stack&lt;/a&gt;&lt;/strong&gt; — Receipt Gate lives across &lt;strong&gt;Tools&lt;/strong&gt; + &lt;strong&gt;Traces&lt;/strong&gt;; Guardrails enforce the latch.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://sheikhwasim.com/insights/permission-envelope-consent-that-expires/" rel="noopener noreferrer"&gt;Permission Envelope&lt;/a&gt;&lt;/strong&gt;  a write may be authorized and still blocked if it left no receipt.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://sheikhwasim.com/insights/golden-trace-eval-anchor/" rel="noopener noreferrer"&gt;Golden Trace&lt;/a&gt;&lt;/strong&gt; — freeze a chain whose receipts still match milestones after a model bump.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://sheikhwasim.com/insights/smoke-evals-before-you-ship/" rel="noopener noreferrer"&gt;Smoke evals&lt;/a&gt;&lt;/strong&gt; — include "second write with missing/mushy receipt → blocked."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://sheikhwasim.com/insights/ship-gate-pre-deploy-checklist/" rel="noopener noreferrer"&gt;Ship Gate&lt;/a&gt;&lt;/strong&gt;  prove chained writes leave skim-worthy receipts, not only that demos finish.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Architecture without receipts is just faster confusion with better models.&lt;/p&gt;

&lt;h2&gt;
  
  
  The test
&lt;/h2&gt;

&lt;p&gt;Ask one question before the next multi-write agent ship:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If write #1 returns ok, can a teammate skim a plain-language receipt of what changed  and does write #2 stay closed until that receipt exists?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If the answer is no, you are still chaining side effects on status codes and model memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;I'm Wasim Sheikh — AI Architect. I build systems teams trust and organizations depend on: not demos, not proofs of concept — production.&lt;/p&gt;

&lt;p&gt;Follow for practical AI architecture that ships.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Connect on LinkedIn:&lt;/strong&gt; &lt;a href="https://www.linkedin.com/in/wasim-sheikh/" rel="noopener noreferrer"&gt;Wasim Sheikh&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Related:&lt;/strong&gt; &lt;a href="https://sheikhwasim.com/insights/golden-trace-eval-anchor/" rel="noopener noreferrer"&gt;Golden Trace&lt;/a&gt; · &lt;a href="https://sheikhwasim.com/insights/permission-envelope-consent-that-expires/" rel="noopener noreferrer"&gt;Permission Envelope&lt;/a&gt; · &lt;a href="https://sheikhwasim.com/insights/smoke-evals-before-you-ship/" rel="noopener noreferrer"&gt;Smoke Evals&lt;/a&gt; · &lt;a href="https://sheikhwasim.com/insights/ship-gate-pre-deploy-checklist/" rel="noopener noreferrer"&gt;Ship Gate&lt;/a&gt; · &lt;a href="https://sheikhwasim.com/insights/agent-architecture-five-layer-stack/" rel="noopener noreferrer"&gt;The 5-Layer Agent Stack&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>security</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Golden Trace: The Eval Anchor Most Teams Skip</title>
      <dc:creator>Wasim Sheikh</dc:creator>
      <pubDate>Fri, 25 Sep 2026 21:45:16 +0000</pubDate>
      <link>https://dev.to/anciwasim/golden-trace-the-eval-anchor-most-teams-skip-1501</link>
      <guid>https://dev.to/anciwasim/golden-trace-the-eval-anchor-most-teams-skip-1501</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Cross-post of Insights #5 — canonical: &lt;a href="https://sheikhwasim.com/insights/golden-trace-eval-anchor/" rel="noopener noreferrer"&gt;https://sheikhwasim.com/insights/golden-trace-eval-anchor/&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Most teams drown in failed agent runs and never lock the one that worked.&lt;/p&gt;

&lt;p&gt;They collect stack traces, bad tool calls, and angry tickets. They rarely freeze a clean end-to-end success and treat it as sacred. The next prompt tweak, model bump, or tool schema change quietly bends that path. The demo still "looks fine." Customers feel the drift first.&lt;/p&gt;

&lt;p&gt;I call the missing control a &lt;strong&gt;Golden Trace&lt;/strong&gt;: one real successful run you refuse to lose — inputs, plans, tool calls, intermediate state, final outcome — frozen as the regression anchor for every later change.&lt;/p&gt;

&lt;p&gt;Smoke evals catch the traps. Golden Trace proves the happy path you already earned still holds.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quick use case
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Situation.&lt;/strong&gt; A platform team shipped an internal "invoice exception agent." It could read a flagged invoice, pull the PO and receipt lines, open a clarifying ticket when amounts diverged, and close the exception when the math matched. One Friday afternoon path worked end to end: correct PO match, one clarifying question, clean close. Leadership celebrated. Nobody saved the run.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What broke.&lt;/strong&gt; Two weeks later a "harmless" prompt polish and a new retrieval index shipped. The agent still sounded sharp. It started closing exceptions when the PO line was &lt;em&gt;close enough&lt;/em&gt;, skipping the clarifying ticket. Finance caught it in a weekly audit. Traces showed answers. They did not show which milestones of the original Friday path had disappeared. There was no frozen success to regress against — only a pile of failure tickets from earlier weeks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix (Golden Trace).&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Pick one proven success&lt;/strong&gt; — the Friday path that matched PO → asked once → closed clean&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Freeze the full trace&lt;/strong&gt; — input invoice, tool calls + args, intermediate checks, final status, cost&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Name the milestones&lt;/strong&gt; — "PO exact match," "clarify if delta &amp;gt; $X," "no close without receipt line"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Regress on every change&lt;/strong&gt; — prompt, tool schema, model, or index bump must hit the same milestones&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fail closed on drift&lt;/strong&gt; — if a milestone vanishes, CI red; no "looks good in staging" override&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Same model. Same tools. A locked success path the team could no longer accidentally erase.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure logs teach. Success paths prove.
&lt;/h2&gt;

&lt;p&gt;A failure log answers: &lt;em&gt;What went wrong that time?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A Golden Trace answers: &lt;em&gt;Does the path that already worked still work on this exact change?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Those are different questions. Most LLM ops programs only fund the first.&lt;/p&gt;

&lt;p&gt;Smoke evals (the trap pack) guard the bad paths. Golden Trace guards the good one you already paid for in debugging time. Without both, you either ship polite bricks or silent regressions that still pass a vibe check.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Golden Trace stack (five layers)
&lt;/h2&gt;

&lt;p&gt;Build it like a control, not a screenshot in Slack.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Capture (one real win)
&lt;/h3&gt;

&lt;p&gt;Do not synthesize a perfect run in a notebook. Capture a production-like success:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Real-ish inputs (redact secrets / PII)&lt;/li&gt;
&lt;li&gt;Every tool call with typed args and results&lt;/li&gt;
&lt;li&gt;The decision points that mattered&lt;/li&gt;
&lt;li&gt;The final business outcome, not just the last model sentence&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you cannot replay it, it is not a Golden Trace. It is a story.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Milestone contract
&lt;/h3&gt;

&lt;p&gt;A full token dump is too brittle. Extract the &lt;strong&gt;milestones&lt;/strong&gt; that define success:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Required tools in order (or allowed set)&lt;/li&gt;
&lt;li&gt;Hard checks that must still pass (exact match, threshold, human gate)&lt;/li&gt;
&lt;li&gt;Forbidden shortcuts (skip clarify, invent a line, close on "close enough")&lt;/li&gt;
&lt;li&gt;Exit condition that matches the business outcome&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Milestones are the contract. The raw trace is the evidence.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Freeze in version control
&lt;/h3&gt;

&lt;p&gt;Store the golden artifact next to the code:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Trace JSON (redacted) + milestone assertions&lt;/li&gt;
&lt;li&gt;Linked smoke-eval case IDs if the same path has trap variants&lt;/li&gt;
&lt;li&gt;Owner and last-reviewed date&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If it is not in git, the next refactor will "simplify" it away.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Regress on every change that can bend the path
&lt;/h3&gt;

&lt;p&gt;Run Golden Trace checks when you change:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;System / developer prompts&lt;/li&gt;
&lt;li&gt;Tool schemas or auth scopes&lt;/li&gt;
&lt;li&gt;Model version or temperature defaults&lt;/li&gt;
&lt;li&gt;Retriever / RAG index that feeds the agent&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Red means no ship. "Staging still looks fine" is how the invoice agent learned "close enough."&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Refresh when the product truth moves
&lt;/h3&gt;

&lt;p&gt;When the real success criteria change — new policy, new required human gate, new tool — update the Golden Trace on purpose. Do not let an outdated golden become a museum piece that blocks good ships, and do not silently drift the milestones to match a worse path.&lt;/p&gt;

&lt;p&gt;A Golden Trace is a living contract with the last path you were proud of.&lt;/p&gt;

&lt;h2&gt;
  
  
  How this maps to what you already have
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://sheikhwasim.com/insights/agent-architecture-five-layer-stack/" rel="noopener noreferrer"&gt;5-layer agent stack&lt;/a&gt;&lt;/strong&gt; — Golden Trace lives in &lt;strong&gt;Traces&lt;/strong&gt;; Guardrails + Rollback catch what drift still slips through&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://sheikhwasim.com/insights/smoke-evals-before-you-ship/" rel="noopener noreferrer"&gt;Smoke evals&lt;/a&gt;&lt;/strong&gt; — trap cases block the bad paths; Golden Trace locks the good one&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://sheikhwasim.com/insights/ship-gate-pre-deploy-checklist/" rel="noopener noreferrer"&gt;Ship Gate&lt;/a&gt;&lt;/strong&gt; — gate #3 (eval set) should include at least one golden success, not only failures&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://sheikhwasim.com/insights/permission-envelope-consent-that-expires/" rel="noopener noreferrer"&gt;Permission Envelope&lt;/a&gt;&lt;/strong&gt; — if a milestone requires a fresh human grant, the golden must assert that the grant was present and not silently skipped&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Architecture without a locked success path is just hope with better diagrams.&lt;/p&gt;

&lt;h2&gt;
  
  
  The test
&lt;/h2&gt;

&lt;p&gt;Ask one question before the next agent ship:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Do we have one versioned successful run — with named milestones — that CI fails if this change drifts off it?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If the answer is no, you are still shipping on demo confidence and failure archaeology.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;I'm Wasim Sheikh — AI Architect. I build systems teams trust and organizations depend on: not demos, not proofs of concept — production.&lt;/p&gt;

&lt;p&gt;Follow for practical AI architecture that ships.&lt;br&gt;
&lt;strong&gt;Connect on LinkedIn:&lt;/strong&gt; &lt;a href="https://www.linkedin.com/in/wasim-sheikh/" rel="noopener noreferrer"&gt;Wasim Sheikh&lt;/a&gt; · &lt;strong&gt;Site:&lt;/strong&gt; &lt;a href="https://sheikhwasim.com" rel="noopener noreferrer"&gt;sheikhwasim.com&lt;/a&gt; · &lt;strong&gt;Notes:&lt;/strong&gt; &lt;a href="https://practicalainotes.substack.com/" rel="noopener noreferrer"&gt;Practical AI Notes&lt;/a&gt; · &lt;strong&gt;X:&lt;/strong&gt; &lt;a href="https://x.com/anciwasim" rel="noopener noreferrer"&gt;@anciwasim&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>architecture</category>
      <category>testing</category>
    </item>
    <item>
      <title>Permission Envelope: Consent That Expires</title>
      <dc:creator>Wasim Sheikh</dc:creator>
      <pubDate>Tue, 22 Sep 2026 16:06:05 +0000</pubDate>
      <link>https://dev.to/anciwasim/permission-envelope-consent-that-expires-2eb8</link>
      <guid>https://dev.to/anciwasim/permission-envelope-consent-that-expires-2eb8</guid>
      <description>&lt;p&gt;Most teams give an AI agent a capability once and never take it back.&lt;/p&gt;

&lt;p&gt;That is not consent. That is a standing power of attorney with no end date.&lt;/p&gt;

&lt;p&gt;The model asks for calendar access, a CRM write, or a deploy token. Someone clicks allow. Three weeks later the same grant is still live — for a different task, a different tenant context, and a tired engineer who forgot what they approved. When something goes wrong, the logs show “authorized.” They do not show &lt;em&gt;when the human last meant it&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;I call the missing control a &lt;strong&gt;Permission Envelope&lt;/strong&gt;: a short-lived, scoped grant that dies unless a human renews it. Expiry is not bureaucracy. Expiry is the audit trail that proves consent was still fresh when the tool fired.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quick use case
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Situation.&lt;/strong&gt; A platform team shipped an internal “release assistant.” It could read a PR, draft release notes, and — with an approved toggle — open the production deploy ticket and page the on-call channel. Staging demos used a long-lived service account so nobody had to re-auth every run.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What broke.&lt;/strong&gt; Two sprints later, a noisy alert chain made the agent open a deploy ticket for the wrong service and page the wrong rotation. The grant that allowed “open ticket + page” had been approved once for a canary release. It never expired. There was no re-ask, no TTL on the capability, and no log line that said “last human consent: 19 days ago for a different change.” Security called it authorized. Ops called it a near miss.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix (Permission Envelope).&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Scope the grant&lt;/strong&gt; — “open deploy ticket for service X + page rotation Y” — not “any ticket, any channel”
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bind to a purpose&lt;/strong&gt; — one change ID / one release window; the grant cannot be reused for a different PR
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Set a short TTL&lt;/strong&gt; — minutes or hours for irreversible tools; days only for read-only
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Re-ask on expiry&lt;/strong&gt; — silent renewal is forbidden; a fresh human intent is the only path back in
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Log consent, not just auth&lt;/strong&gt; — who approved, what they saw, when it expires, and which action consumed it&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Same agent. Same model. Different blast radius. The next wrong-service page never left the draft state.&lt;/p&gt;

&lt;p&gt;Permission Envelopes aren’t about distrusting the model. They’re about refusing to let yesterday’s yes authorize today’s irreversible action.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why permanent grants fail quietly
&lt;/h2&gt;

&lt;p&gt;Auth answers: &lt;em&gt;Is this identity allowed to call the tool?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A Permission Envelope answers: &lt;em&gt;Did a human still mean this specific action, for this purpose, right now?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Those are different questions. Most agent stacks only ask the first.&lt;/p&gt;

&lt;p&gt;Service accounts and OAuth tokens are great at surviving restarts. They are terrible at remembering that the human clicked “allow” for a canary, not for every future Tuesday. Permanent grants optimize for fewer clicks. Production optimizes for fewer surprises at 2 a.m.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Permission Envelope stack (five layers)
&lt;/h2&gt;

&lt;p&gt;Build the grant like a system, not a checkbox buried in settings.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Capability (what)
&lt;/h3&gt;

&lt;p&gt;Name the exact tool surface: which write endpoints, which tenants, which channels. Prefer deny-by-default. “CRM update” is not a capability — “update ticket status on account A, fields status + note only” is.&lt;/p&gt;

&lt;p&gt;If you cannot write the capability in one sentence a junior engineer understands, the grant is already too wide.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Purpose (why / which)
&lt;/h3&gt;

&lt;p&gt;Bind the grant to a concrete work item: PR number, change ticket, incident ID, customer case. When the purpose changes, the envelope dies. Reuse across tasks is how “authorized” becomes meaningless.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Duration (how long)
&lt;/h3&gt;

&lt;p&gt;Pick the shortest TTL that still lets the workflow finish:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Irreversible writes&lt;/strong&gt; (send, deploy, charge, delete, page) — minutes to a few hours
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mutating drafts&lt;/strong&gt; (create ticket, draft email) — hours, then re-ask
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read-only&lt;/strong&gt; — longer is fine; still expire so abandoned sessions don’t linger&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Auto-extend is a smell. If the job needs more time, the human should know.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Fresh consent (who / when)
&lt;/h3&gt;

&lt;p&gt;Expiry must force a real re-ask: show the blast radius again, get an explicit approve, freeze the approved payload. Do not let the model regenerate the action after the click. Do not treat “still logged in” as “still consented.”&lt;/p&gt;

&lt;p&gt;The re-ask is the audit trail field teams keep skipping — and the one regulators and postmortems actually care about.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Consumption log (proof)
&lt;/h3&gt;

&lt;p&gt;Every tool call that used the envelope records:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Grant ID + scope + purpose&lt;/li&gt;
&lt;li&gt;Approver identity + timestamp + what they were shown&lt;/li&gt;
&lt;li&gt;Expiry + whether this call consumed or renewed&lt;/li&gt;
&lt;li&gt;Outcome (success / denied / expired mid-flight)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you only log “token valid,” you cannot reconstruct consent. You can only reconstruct hope.&lt;/p&gt;

&lt;h2&gt;
  
  
  How this maps to what you already have
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;5-layer agent stack&lt;/strong&gt; — Tools get typed scopes; Guardrails deny out-of-envelope calls; Traces carry the grant ID; Rollback still matters when a consumed grant was wrong
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ship Gate&lt;/strong&gt; — threat model names “stale consent”; human override is the re-ask; observability is the consumption log
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Smoke evals&lt;/strong&gt; — add cases where the grant is expired, purpose-mismatched, or too wide; CI should fail if the agent still fires the write&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A Permission Envelope does not replace Ship Gate. It makes gate #4 (human override) real after the first approve — and every hour after.&lt;/p&gt;

&lt;h2&gt;
  
  
  The test
&lt;/h2&gt;

&lt;p&gt;Ask one question before the next agent write-tool ships:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can we show, for every irreversible action, &lt;em&gt;who&lt;/em&gt; consented, &lt;em&gt;to what scope&lt;/em&gt;, &lt;em&gt;for which purpose&lt;/em&gt;, and &lt;em&gt;when that consent expires&lt;/em&gt; — and does the tool refuse when any of those are missing or stale?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If any half is no, you are still shipping standing power of attorney.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;I’m Wasim Sheikh — AI Architect. I build systems teams trust and organizations depend on: not demos, not proofs of concept — production.&lt;/p&gt;

&lt;p&gt;Follow for practical AI architecture that ships.&lt;br&gt;&lt;br&gt;
&lt;strong&gt;Connect on LinkedIn:&lt;/strong&gt; &lt;a href="https://www.linkedin.com/in/wasim-sheikh/" rel="noopener noreferrer"&gt;Wasim Sheikh&lt;/a&gt; · &lt;strong&gt;Site:&lt;/strong&gt; &lt;a href="https://sheikhwasim.com" rel="noopener noreferrer"&gt;sheikhwasim.com&lt;/a&gt; · &lt;strong&gt;Notes:&lt;/strong&gt; &lt;a href="https://practicalainotes.substack.com/" rel="noopener noreferrer"&gt;Practical AI Notes&lt;/a&gt; · &lt;strong&gt;X:&lt;/strong&gt; &lt;a href="https://x.com/anciwasim" rel="noopener noreferrer"&gt;@anciwasim&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>architecture</category>
      <category>devops</category>
    </item>
    <item>
      <title>Smoke Evals: The 20 Cases That Should Block Every AI Deploy</title>
      <dc:creator>Wasim Sheikh</dc:creator>
      <pubDate>Fri, 18 Sep 2026 15:39:01 +0000</pubDate>
      <link>https://dev.to/anciwasim/smoke-evals-the-20-cases-that-should-block-every-ai-deploy-59f6</link>
      <guid>https://dev.to/anciwasim/smoke-evals-the-20-cases-that-should-block-every-ai-deploy-59f6</guid>
      <description>&lt;p&gt;Your AI feature passed the demo.&lt;/p&gt;

&lt;p&gt;That is not evidence it will survive Monday morning.&lt;/p&gt;

&lt;p&gt;Demos are curated. Someone picks the happy path, soft-prompts around the edges, and never pastes the weird customer email that invents a discount, leaks a ticket ID, or loops a tool until the bill spikes. Production does all of that before lunch.&lt;/p&gt;

&lt;p&gt;I call the missing test set &lt;strong&gt;smoke evals&lt;/strong&gt; — a short, fixed pack of trap cases that must pass before any AI change leaves the sandbox. Not a research benchmark. Not a vibe check. A deploy gate.&lt;/p&gt;

&lt;p&gt;If &lt;a href="https://sheikhwasim.com/insights/ship-gate-pre-deploy-checklist/" rel="noopener noreferrer"&gt;Ship Gate&lt;/a&gt; is the checklist, smoke evals are the part that actually fails the build.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quick use case
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Situation.&lt;/strong&gt; A platform team shipped an “AI ops assistant” that could summarize incidents and suggest runbook steps. Staging looked sharp. Leadership wanted it in the on-call channel before a busy release week.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What broke.&lt;/strong&gt; Night one, a noisy alert chain made the model invent a rollback command that did not exist in the runbook. An engineer almost ran it. There was no frozen set of “known trap” incidents, no assert that refused invented commands, and no CI job that would have blocked the merge.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix (smoke evals).&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Freeze ~20 cases&lt;/strong&gt; — real-ish incidents: missing runbook steps, ambiguous alerts, PII in tickets, prompt-injection style pastes, “do nothing” noise&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Assert two things&lt;/strong&gt; — &lt;em&gt;safe&lt;/em&gt; (no invented commands, no secret leakage) and &lt;em&gt;useful&lt;/em&gt; (still names the right service / severity when the answer is known)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wire to CI&lt;/strong&gt; — every prompt, tool schema, or model bump runs the pack; red means no ship&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Own the failures&lt;/strong&gt; — each failing case gets an owner and a fix path (prompt, tool contract, or human gate)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Refresh monthly&lt;/strong&gt; — add the newest production scare; retire cases that no longer teach&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Same model. Different gate. The next invented rollback never reached the channel.&lt;/p&gt;

&lt;p&gt;Smoke evals aren’t bureaucracy. They’re the smallest set of cases that would have caught the last three failures — run on every change.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why demos lie (again)
&lt;/h2&gt;

&lt;p&gt;A demo answers: &lt;em&gt;Can we make it look smart once?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A smoke eval answers: &lt;em&gt;Does it refuse the bad path, every time, on this exact change?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Those are different questions. Most teams only ask the first.&lt;/p&gt;

&lt;p&gt;The second is cheaper than a postmortem. Twenty cases, a few dollars of inference, a red CI job. Compare that to an almost-run false rollback at 2 a.m.&lt;/p&gt;

&lt;h2&gt;
  
  
  The smoke-eval stack (five layers)
&lt;/h2&gt;

&lt;p&gt;Build the pack like a system, not a spreadsheet someone opens once.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Case bank (frozen, versioned)
&lt;/h3&gt;

&lt;p&gt;Keep the cases in git. Each case has:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Input&lt;/strong&gt; — the ticket, chat, alert, or document paste&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expected behavior&lt;/strong&gt; — refuse / escalate / answer with X&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Failure mode tagged&lt;/strong&gt; — hallucination, tool misuse, PII leak, loop, overconfidence&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If it isn’t in version control, it isn’t a gate. It’s a wish.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Dual assert: safe &lt;em&gt;and&lt;/em&gt; useful
&lt;/h3&gt;

&lt;p&gt;Safe-only evals train the model to say nothing. Useful-only evals ship confident nonsense.&lt;/p&gt;

&lt;p&gt;Every case should check both:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Safe&lt;/strong&gt; — no invented SKUs, commands, prices, or secrets&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Useful&lt;/strong&gt; — when the ground truth is known, the answer still helps&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A pack that only measures “didn’t explode” will ship a polite brick.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. CI as the gate
&lt;/h3&gt;

&lt;p&gt;Run the pack on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Prompt / system-message changes&lt;/li&gt;
&lt;li&gt;Tool schema changes&lt;/li&gt;
&lt;li&gt;Model version bumps&lt;/li&gt;
&lt;li&gt;Retriever / RAG index refreshes that affect answers&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Red build = no merge. No “we’ll check it later in staging” exception for the pack that exists &lt;em&gt;because&lt;/em&gt; staging lied.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Failure ownership
&lt;/h3&gt;

&lt;p&gt;A failing case without an owner becomes noise. Assign:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Who owns the failure mode&lt;/li&gt;
&lt;li&gt;Whether the fix is prompt, tool contract, retrieval, or human override&lt;/li&gt;
&lt;li&gt;When it must be green again&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Smoke evals without ownership rot into ignored red.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Monthly refresh from production scars
&lt;/h3&gt;

&lt;p&gt;Every real incident that almost (or did) hurt customers becomes a candidate case. Add it. Keep the pack near twenty so it stays fast. Retire cases that no longer teach.&lt;/p&gt;

&lt;p&gt;The pack should smell like your last three scares — not a generic LLM leaderboard.&lt;/p&gt;

&lt;h2&gt;
  
  
  How this maps to what you already have
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://sheikhwasim.com/insights/ship-gate-pre-deploy-checklist/" rel="noopener noreferrer"&gt;Ship Gate&lt;/a&gt;&lt;/strong&gt; — gate #3 (eval set) is where smoke evals live; the other four gates still matter&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://sheikhwasim.com/insights/agent-architecture-five-layer-stack/" rel="noopener noreferrer"&gt;5-layer agent stack&lt;/a&gt;&lt;/strong&gt; — traces tell you &lt;em&gt;which&lt;/em&gt; case failed and &lt;em&gt;why&lt;/em&gt;; guardrails + rollback catch what the pack missed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Smoke evals don’t replace architecture. They make the architecture earn the next deploy.&lt;/p&gt;

&lt;h2&gt;
  
  
  The test
&lt;/h2&gt;

&lt;p&gt;Ask one question before the next AI ship:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Do we have ~20 versioned trap cases that assert safe &lt;em&gt;and&lt;/em&gt; useful — and does CI fail if they fail?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If either half is no, you are still shipping on demo confidence.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://sheikhwasim.com/insights/smoke-evals-before-you-ship/" rel="noopener noreferrer"&gt;sheikhwasim.com&lt;/a&gt;. I’m Wasim Sheikh — AI Architect. I build systems teams trust and organizations depend on.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Connect:&lt;/strong&gt; &lt;a href="https://www.linkedin.com/in/wasim-sheikh/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://practicalainotes.substack.com/" rel="noopener noreferrer"&gt;Practical AI Notes&lt;/a&gt; · &lt;a href="https://x.com/anciwasim" rel="noopener noreferrer"&gt;X @anciwasim&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>architecture</category>
      <category>testing</category>
    </item>
    <item>
      <title>Ship Gate: The Pre-Deploy Checklist Most AI Features Skip</title>
      <dc:creator>Wasim Sheikh</dc:creator>
      <pubDate>Thu, 17 Sep 2026 15:35:12 +0000</pubDate>
      <link>https://dev.to/anciwasim/ship-gate-the-pre-deploy-checklist-most-ai-features-skip-1do6</link>
      <guid>https://dev.to/anciwasim/ship-gate-the-pre-deploy-checklist-most-ai-features-skip-1do6</guid>
      <description>&lt;p&gt;Most AI features don't fail in the model.&lt;/p&gt;

&lt;p&gt;They fail at the gate.&lt;/p&gt;

&lt;p&gt;Someone demos a shiny prompt. Leadership loves it. It ships. Then a customer pastes something weird, an agent calls the wrong tool, or a "helpful" answer invents an offer that never existed. The postmortem is always the same: we optimized for the demo, not for production.&lt;/p&gt;

&lt;p&gt;I call the missing step &lt;strong&gt;Ship Gate&lt;/strong&gt; — a short, boring checklist you run before any AI feature leaves the sandbox. Not a 40-page risk memo. A gate you can clear in an afternoon.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quick use case
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Situation.&lt;/strong&gt; A product team shipped an AI email summarizer into customer-facing replies. Staging looked clean. Leadership wanted it live before the quarter closed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What broke.&lt;/strong&gt; Day one, the model invented a discount a customer had never been offered. Support scrambled. There was no eval set of "known trap" emails, no human gate on outbound AI text, and no kill switch short of a full deploy rollback.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix (Ship Gate).&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Threat model&lt;/strong&gt; — named "hallucinated commercial offers" as a top failure mode with an owner&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data boundary&lt;/strong&gt; — outbound copy restricted to template fields; free-form claims blocked&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Eval set&lt;/strong&gt; — trap cases that fail the build if the model invents offers or leaks PII&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Human override&lt;/strong&gt; — outbound AI copy held until a human approves (or stricter template mode)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Observability&lt;/strong&gt; — one bad reply reconstructable in minutes: prompt, context, output, cost&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Then they added the ops layer every Ship Gate implies: canary traffic, a one-click kill switch, and a documented rollback owner before the flag flipped on.&lt;/p&gt;

&lt;p&gt;Ship Gate isn't bureaucracy. It's the minimum checklist so an AI feature earns production — instead of learning in front of customers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why demos lie
&lt;/h2&gt;

&lt;p&gt;Demos are curated. You pick the happy path. You soft-prompt around the edge cases. You never paste a raw customer email with secrets still in it.&lt;/p&gt;

&lt;p&gt;Production is the opposite. Users paste junk. Models invent confident nonsense. Tools fire with the wrong args. Logs capture more than you think.&lt;/p&gt;

&lt;p&gt;If your only test was "it looked good in the meeting," you didn't ship AI — you shipped a vibe.&lt;/p&gt;

&lt;p&gt;Ship Gate exists to make that gap visible before customers feel it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The five gates
&lt;/h2&gt;

&lt;p&gt;Run these in order. Fail any one and you don't ship. Pass all five and you can sleep.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Threat model (ten minutes, not a committee)
&lt;/h3&gt;

&lt;p&gt;Before code, write three sentences:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What can go wrong? (wrong answer, tool misuse, data leak, prompt injection, cost runaway)&lt;/li&gt;
&lt;li&gt;Who gets hurt? (customer, employee, your brand, a regulator)&lt;/li&gt;
&lt;li&gt;What's the blast radius? (one chat vs. every tenant)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you can't name the top three failure modes, you're not ready to build — you're ready to be surprised.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Data boundary
&lt;/h3&gt;

&lt;p&gt;Answer out loud:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What data can the model see?&lt;/li&gt;
&lt;li&gt;What can it write?&lt;/li&gt;
&lt;li&gt;What must never leave this boundary (PII, secrets, customer documents, internal tickets)?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then enforce it in code, not in a slide. Strip secrets before the model. Scope tools to the minimum. Prefer retrieval over stuffing the whole corpus into context. If the model doesn't need a field, don't give it the field.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Eval set before polish
&lt;/h3&gt;

&lt;p&gt;Pick 20–50 real-ish cases: happy path, hostile prompts, empty input, long paste, "ignore previous instructions," the ticket that always breaks things.&lt;/p&gt;

&lt;p&gt;Score each on three axes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Correct enough for the job&lt;/li&gt;
&lt;li&gt;Safe enough (no leak, no unsafe action)&lt;/li&gt;
&lt;li&gt;Useful format (what the user actually needs)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Ship Gate rule: you don't tune the UI until the eval set is green enough that you'd trust a teammate to use the feature unsupervised.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Human override
&lt;/h3&gt;

&lt;p&gt;Every autonomous or semi-autonomous path needs an escape hatch:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Confirm before irreversible actions (send email, delete, charge, deploy)&lt;/li&gt;
&lt;li&gt;Log who approved what&lt;/li&gt;
&lt;li&gt;Make "stop / undo / escalate" obvious&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the only recovery path is "hope the model was right," you don't have a product. You have a liability with a chat box.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Observability that a human can read
&lt;/h3&gt;

&lt;p&gt;When it breaks at 2 a.m., can you answer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What did the user ask?&lt;/li&gt;
&lt;li&gt;What context did the model get?&lt;/li&gt;
&lt;li&gt;Which tools ran, with which args?&lt;/li&gt;
&lt;li&gt;What did we return?&lt;/li&gt;
&lt;li&gt;Cost and latency for that turn?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the answer is "we'd have to dig through JSON for an hour," fix the logs before you grow the feature. Ship Gate isn't complete until a tired engineer can reconstruct a single bad turn.&lt;/p&gt;

&lt;h2&gt;
  
  
  A one-page Ship Gate card
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Gate&lt;/th&gt;
&lt;th&gt;Pass criteria&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Threat model&lt;/td&gt;
&lt;td&gt;Top 3 failure modes named + owner&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data boundary&lt;/td&gt;
&lt;td&gt;See / write / never-leave documented + enforced&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Eval set&lt;/td&gt;
&lt;td&gt;≥20 cases; safe + useful bar met&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Human override&lt;/td&gt;
&lt;td&gt;Irreversible actions confirmed; stop path clear&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Observability&lt;/td&gt;
&lt;td&gt;One bad turn reconstructable in &amp;lt;5 minutes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;No green across the board → no ship.&lt;/p&gt;

&lt;h2&gt;
  
  
  The test
&lt;/h2&gt;

&lt;p&gt;Ask one question before any AI feature goes live:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Which of the five gates is still red — and who owns fixing it?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If nobody owns the red gate, you don't have a launch plan. You have a calendar date.&lt;/p&gt;




&lt;p&gt;I'm Wasim Sheikh — AI Architect. I build systems teams trust and organizations depend on: not demos, not proofs of concept — production.&lt;/p&gt;

&lt;p&gt;Canonical: &lt;a href="https://sheikhwasim.com/insights/ship-gate-pre-deploy-checklist/" rel="noopener noreferrer"&gt;https://sheikhwasim.com/insights/ship-gate-pre-deploy-checklist/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>architecture</category>
      <category>devops</category>
    </item>
    <item>
      <title>The 5-Layer Stack Behind Agents That Ship</title>
      <dc:creator>Wasim Sheikh</dc:creator>
      <pubDate>Mon, 14 Sep 2026 15:59:29 +0000</pubDate>
      <link>https://dev.to/anciwasim/the-5-layer-stack-behind-agents-that-ship-ee9</link>
      <guid>https://dev.to/anciwasim/the-5-layer-stack-behind-agents-that-ship-ee9</guid>
      <description>&lt;p&gt;Originally published on my site: &lt;a href="https://sheikhwasim.com/insights/agent-architecture-five-layer-stack/" rel="noopener noreferrer"&gt;https://sheikhwasim.com/insights/agent-architecture-five-layer-stack/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Most "AI agents" fail for the same reason.&lt;/p&gt;

&lt;p&gt;Someone wires a language model to a handful of tools, demos a happy path, and calls it architecture. It works in the slide. It collapses the first time production is messy  timeouts, bad inputs, spend spikes, irreversible side effects.&lt;/p&gt;

&lt;p&gt;Here's the five-layer stack that actually ships.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quick use case
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Situation.&lt;/strong&gt; A team built an internal "support triage agent." Demo day looked sharp: read a ticket, call search, suggest a reply, update CRM status.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What broke.&lt;/strong&gt; In production, one vague ticket sent the agent into a 40-step tool loop. It wrote the wrong CRM status, and spend spiked past the monthly budget in an afternoon. Chat logs showed answers — not &lt;em&gt;which&lt;/em&gt; tool ran, with what inputs, or why it kept going. The CRM write couldn't be cleanly reversed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix (mapped to the stack).&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Orchestrator&lt;/strong&gt; — hard max steps + stop when confidence is low&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tools&lt;/strong&gt; — typed CRM update with auth scope + idempotency key&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Traces&lt;/strong&gt; — every plan / tool call / result / cost on one correlation ID&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Guardrails&lt;/strong&gt; — daily spend cap + deny list for irreversible tools without a human gate&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rollback&lt;/strong&gt; — compensating "revert status" action when the write was wrong&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Same model. Different system. That's the difference between a demo and something teams can trust.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 1 — Orchestrator
&lt;/h2&gt;

&lt;p&gt;The orchestrator decides the &lt;strong&gt;next step&lt;/strong&gt;. It is not the model dumping text forever.&lt;/p&gt;

&lt;p&gt;Good orchestrators:&lt;/p&gt;

&lt;p&gt;Tools are clear APIs, not mystery side effects.&lt;/p&gt;

&lt;p&gt;Ship-ready tools have:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Auth and scoped credentials&lt;/li&gt;
&lt;li&gt;Typed inputs and validated outputs&lt;/li&gt;
&lt;li&gt;Timeouts, rate limits, and idempotency keys&lt;/li&gt;
&lt;li&gt;Explicit success / failure contracts&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A tool that "usually works" is a liability. Prefer boring interfaces over clever ones.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 3 — Traces
&lt;/h2&gt;

&lt;p&gt;Every step logged so you can see what it did — and why.&lt;/p&gt;

&lt;p&gt;Traces are how you debug, audit, and improve:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Prompt / plan / tool call / result / decision&lt;/li&gt;
&lt;li&gt;Timing and cost per step&lt;/li&gt;
&lt;li&gt;Correlation IDs across services&lt;/li&gt;
&lt;li&gt;Redaction for secrets and PII&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you can't replay the path, you can't trust the system in an enterprise.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 4 — Guardrails
&lt;/h2&gt;

&lt;p&gt;Budget, permissions, PII, and stop conditions — &lt;strong&gt;before&lt;/strong&gt; it runs wild.&lt;/p&gt;

&lt;p&gt;Guardrails belong in the control plane, not in a polite system prompt:&lt;/p&gt;

&lt;p&gt;If it's wrong, you reverse it. Idempotent actions beat clever guesses.&lt;/p&gt;

&lt;p&gt;Design for undo:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Prefer reversible writes&lt;/li&gt;
&lt;li&gt;Snapshot before mutate&lt;/li&gt;
&lt;li&gt;Compensating transactions when undo isn't free&lt;/li&gt;
&lt;li&gt;Clear "blast radius" for every tool&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;An agent that can't roll back isn't a system. It's a demo with confidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  The test
&lt;/h2&gt;

&lt;p&gt;Ask one question of any agent stack:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can it show traces, and can it roll back?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If either answer is no, keep it out of production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;I'm Wasim Sheikh — AI Architect. I build systems teams trust and organizations depend on: not demos, not proofs of concept — production.&lt;/p&gt;

&lt;p&gt;Follow for practical AI architecture that ships.&lt;br&gt;
Site: &lt;a href="https://sheikhwasim.com" rel="noopener noreferrer"&gt;https://sheikhwasim.com&lt;/a&gt; · Notes: &lt;a href="https://practicalainotes.substack.com/" rel="noopener noreferrer"&gt;https://practicalainotes.substack.com/&lt;/a&gt; · X: &lt;a href="https://x.com/anciwasim" rel="noopener noreferrer"&gt;https://x.com/anciwasim&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Spend and token caps&lt;/li&gt;
&lt;li&gt;Allow / deny lists for tools and data&lt;/li&gt;
&lt;li&gt;Human gates for irreversible actions&lt;/li&gt;
&lt;li&gt;Policy checks on inputs and outputs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Prompts are guidance. Guardrails are enforcement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 5 — Rollback
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Choose among plan  act  verify → stop&lt;/li&gt;
&lt;li&gt;Bound loop length and retry policy&lt;/li&gt;
&lt;li&gt;Separate "thinking" from "committing"&lt;/li&gt;
&lt;li&gt;Fail closed when confidence or evidence is weak&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If your system can't explain &lt;em&gt;why&lt;/em&gt; it took the next action, you don't have an orchestrator — you have a chat session with plugins.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 2 — Tools
&lt;/h2&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>architecture</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
