<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Anthony Garces</title>
    <description>The latest articles on DEV Community by Anthony Garces (@anthonygarces).</description>
    <link>https://dev.to/anthonygarces</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4077859%2F606ef6b9-dec7-4434-b4cb-87ad820c0744.png</url>
      <title>DEV Community: Anthony Garces</title>
      <link>https://dev.to/anthonygarces</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/anthonygarces"/>
    <language>en</language>
    <item>
      <title>Inside the Append-Only, Hash-Chained Journal</title>
      <dc:creator>Anthony Garces</dc:creator>
      <pubDate>Mon, 28 Sep 2026 09:07:53 +0000</pubDate>
      <link>https://dev.to/anthonygarces/inside-the-append-only-hash-chained-journal-pf4</link>
      <guid>https://dev.to/anthonygarces/inside-the-append-only-hash-chained-journal-pf4</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Ranex keeps verdict records in an append-only, hash-chained SQLite journal so ordinary rewrites are blocked and edited rows are detectable, while rollback and truncation remain an explicit open limit. &lt;a href="https://ranex.dev/blog/how-the-kernel-works" rel="noopener noreferrer"&gt;Read how the kernel works.&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Someone asks why a change was allowed out the door. You open the audit record and find a clean story. What stops that story from being cleaned up after the fact?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A record that can be quietly rewritten is not evidence of a decision.&lt;/strong&gt; Ranex records verdict activity in an append-only hash-chained journal. The design does not turn a local database into an untouchable historical archive. It does give you a concrete way to detect a row edit and a concrete command to check the chain.&lt;/p&gt;

&lt;h2 id="in-this-note"&gt;In this note&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The record must resist a quiet rewrite&lt;/li&gt;
&lt;li&gt;Two controls cover different paths&lt;/li&gt;
&lt;li&gt;Check the journal you depend on&lt;/li&gt;
&lt;li&gt;The limit is rollback and truncation&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id="the-record-must-resist-a-quiet-rewrite"&gt;The record must resist a quiet rewrite&lt;/h2&gt;

&lt;p&gt;An append-only journal preserves the sequence of admitted events rather than a polished final summary. That matters when your question is not “what does the database say now?” but “what evidence and verdict existed when this decision happened?”&lt;/p&gt;

&lt;p&gt;Ranex’s journal is a SQLite record. Ordinary updates and deletes are prohibited by SQLite triggers. Each journal row also carries a link into a hash chain, so a row has a relationship to the history before it. The operator command &lt;code&gt;ranex journal verify&lt;/code&gt; recomputes that chain. Those facts provide a useful control: the normal application path cannot edit or delete a row, and an edit made around that normal path is not supposed to blend into the sequence unnoticed.&lt;/p&gt;

&lt;p&gt;Do not reduce this to “there is a database.” A database can retain the latest story while losing the important question. When did the story change, and what did it replace? Append-only changes the shape of what the journal is for. The journal is a record to replay alongside the deterministic verdict path, not a mutable dashboard state that happens to contain audit-looking fields.&lt;/p&gt;

&lt;p&gt;This is particularly useful when a worker produces a result you did not watch arrive. Ranex does not take the worker’s summary as proof; it reads the diff on disk and runs checks. The evidence, verdict, and journal give the operator something stronger than a completion message. They give the operator materials to inspect after the excitement has passed.&lt;/p&gt;

&lt;h2 id="two-controls-cover-different-paths"&gt;Two controls cover different paths&lt;/h2&gt;

&lt;p&gt;SQLite triggers and a hash chain are not duplicates. The triggers block ordinary rewrite attempts; the chain makes an out-of-band row edit visible when verification recomputes the links.&lt;/p&gt;

&lt;p&gt;Imagine a program, a stray maintenance script, or an agent with access to the database trying to issue an ordinary update or delete. The SQLite triggers prohibit that normal operation. That is a direct guard where application behavior meets the journal. It helps keep “fix the record” from becoming a casual recovery move.&lt;/p&gt;

&lt;p&gt;Now take a different route. An attacker edits the database file outside the normal path. A trigger cannot fire for a change that bypasses the SQL operation it protects. The hash chain is for that case. A changed row breaks the relationship with following rows, and &lt;code&gt;ranex journal verify&lt;/code&gt; recomputes the chain for the operator. Verification is the important verb here. A chain nobody checks is only a claim about how rows were written.&lt;/p&gt;

&lt;p&gt;There is no magic in the word hash. The value of the design is that it gives you a defined test against a defined class of edits. You can run the verifier instead of accepting an application’s self-description. You can ask whether the record still links as it should. If it does not, the journal has evidence of a problem rather than a quietly amended history.&lt;/p&gt;

&lt;p&gt;The repository also says concurrent appenders are serialised before reading the previous link. That keeps append operations from racing into incompatible predecessor links. It is a consistency property for the chain. It is not a claim that the journal has solved every preservation problem, which brings us to the line that needs to stay visible.&lt;/p&gt;

&lt;h2 id="check-the-journal-you-depend-on"&gt;Check the journal you depend on&lt;/h2&gt;

&lt;p&gt;You can apply the lesson before you use Ranex. Find the record your release process calls an audit trail, then ask whether it can tell an edit from an original event.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Does the normal application path prohibit updates and deletes of decision records?&lt;/li&gt;
&lt;li&gt;Can an operator independently verify the relationship between successive records?&lt;/li&gt;
&lt;li&gt;Does the record name the evidence, subject, and verdict it is preserving?&lt;/li&gt;
&lt;li&gt;Would an edit outside the application leave a detectable break?&lt;/li&gt;
&lt;li&gt;Who can copy, replace, or restore the record store?&lt;/li&gt;
&lt;li&gt;Can a reviewer distinguish a current status from the sequence that produced it?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These questions are the reason the journal exists. Generated work can leave you with more activity than a person can remember. A future reviewer needs to connect a requirement, evidence, the exact artifact evaluated, and the verdict that allowed a next step. The journal is one part of retaining that connection. It does not decide whether the requirement was good; the approved flow graph is intended as the root of trust in the broader designed loop. It retains what the kernel saw and decided.&lt;/p&gt;

&lt;p&gt;There is a natural boundary with &lt;a href="https://ranex.dev/blog/ed25519-evidence-signing" rel="noopener noreferrer"&gt;signed evidence&lt;/a&gt;. Signatures establish something about who held a registered private key. The journal records the admitted sequence. Neither control replaces the other, and neither should be stretched into a claim it does not make.&lt;/p&gt;

&lt;h2 id="the-limit-is-rollback-and-truncation"&gt;The limit is rollback and truncation&lt;/h2&gt;

&lt;p&gt;The journal does not detect rollback or truncation. An internally consistent earlier prefix still verifies after later rows are removed.&lt;/p&gt;

&lt;p&gt;That is the boundary named in the README’s known gaps. If later rows disappear and the remaining database ends at an older valid point, recomputing the retained chain succeeds. The verifier detects an edited row in the retained history; it does not know, by itself, that history was made shorter. Serialising appenders does not close this gap. Trigger protection does not close it either.&lt;/p&gt;

&lt;p&gt;Keep the distinction sharp. “Tamper-evident for out-of-band row edits” is a useful claim. “Permanent history that detects every deletion” is not a claim Ranex makes. The &lt;a href="https://github.com/anthonykewl20/ranex#status" rel="noopener noreferrer"&gt;repository’s Status section&lt;/a&gt; lists the append-only hash-chained journal as working today and names rollback or truncation as a known gap. The project remains pre-release.&lt;/p&gt;

&lt;p&gt;That clarity is better for your incident response. If verification fails, you have a reason to investigate an altered chain. If it passes, you know the retained rows are internally consistent, not that no later history was removed. You can decide what additional retention or external anchoring your own risk requires without pretending the local journal already supplied it.&lt;/p&gt;

&lt;h2 id="questions-people-actually-ask"&gt;Questions people actually ask&lt;/h2&gt;

&lt;h3&gt;How does a hash-chained journal detect tampering?&lt;/h3&gt;

&lt;p&gt;The Ranex journal links each row hash to the previous row, and journal verify recomputes the chain to expose an out-of-band row edit.&lt;/p&gt;

&lt;h3&gt;Can a hash-chained journal detect deleted entries?&lt;/h3&gt;

&lt;p&gt;The Ranex journal cannot detect rollback or truncation because an internally consistent earlier prefix still verifies after later rows are removed.&lt;/p&gt;

&lt;h3&gt;What stops ordinary journal updates and deletes?&lt;/h3&gt;

&lt;p&gt;Ranex uses SQLite triggers to prohibit ordinary updates and deletes, while the hash chain detects edits outside that normal path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Try it. Break it. Tell me what broke.&lt;/strong&gt; Read the &lt;a href="https://github.com/anthonykewl20/ranex" rel="noopener noreferrer"&gt;Ranex repository&lt;/a&gt;, then inspect whether your own audit record can show a rewrite.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records — the same fact gate the product enforces on code. It ships only after Anthony’s own review.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>evidence</category>
      <category>hashing</category>
    </item>
    <item>
      <title>The Watchdog That Would Have Passed Every Test and Hung Anyway</title>
      <dc:creator>Anthony Garces</dc:creator>
      <pubDate>Fri, 25 Sep 2026 09:00:16 +0000</pubDate>
      <link>https://dev.to/anthonygarces/the-watchdog-that-would-have-passed-every-test-and-hung-anyway-1ekc</link>
      <guid>https://dev.to/anthonygarces/the-watchdog-that-would-have-passed-every-test-and-hung-anyway-1ekc</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; A watchdog fixes a stall only when timeout errors are terminal and non-retryable; idle and absolute budgets cover different waits. That is the shape to hunt for after reading &lt;a href="https://ranex.dev/blog/building-ranex-slice-log" rel="noopener noreferrer"&gt;the Field Notes slice log&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;You have a timeout. The test turns red when the provider goes silent. The build turns green when you add the watchdog. Then a real stalled turn sits there, your session stays busy, and somebody has to intervene by hand. &lt;strong&gt;A retryable timeout would have let the watchdog fire, let the criteria pass, and restarted the exact stall that caused the timeout.&lt;/strong&gt;&lt;/p&gt;

&lt;h2 id="in-this-note"&gt;In this note&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The timeout that fixes nothing&lt;/li&gt;
&lt;li&gt;Why one timeout cannot do two jobs&lt;/li&gt;
&lt;li&gt;What to inspect in your own pipeline&lt;/li&gt;
&lt;li&gt;The second hang hiding behind the first&lt;/li&gt;
&lt;li&gt;Your next failure drill&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That was the defect behind SLICE-012 in the Ranex harness fork. A stalled provider stream had no timeout, so its coordinator never settled and the active session stayed busy. The shipped change makes that stream reach a terminal state on its own. Good. The more useful part is the hole the tests could have missed.&lt;/p&gt;

&lt;p&gt;The session would still hang anyway.&lt;/p&gt;

&lt;h2 id="the-timeout-that-fixes-nothing"&gt;The timeout that fixes nothing&lt;/h2&gt;

&lt;p&gt;A timeout is only a fix when its failure path reaches a terminal state. If your retry policy recreates the condition that timed out, retrying it is a loop, not recovery.&lt;/p&gt;

&lt;p&gt;Picture the order of events. A provider stream stops producing. The watchdog fires. The runner sees an error it is allowed to retry. It starts the stream again. The provider stalls again. The watchdog fires again. Your test can prove the timer fired. It can prove the retry happened. It can still miss the fact that the turn never ends.&lt;/p&gt;

&lt;p&gt;SLICE-012 made the classification explicit: both watchdog failures are typed, &lt;a href="https://github.com/anthonykewl20/ranex/blob/main/docs/slices/done/SLICE-012-provider-watchdog.md" rel="noopener noreferrer"&gt;&lt;strong&gt;non-retryable&lt;/strong&gt; errors&lt;/a&gt;. That detail is load-bearing. The earlier prototype had avoided the retryability problem by accident. A pre-implementation review found that no requirement had said the timeout must be non-retryable, so the completed slice made the decision and required a test that proves it is not retried.&lt;/p&gt;

&lt;p&gt;That is a small wording change with a large operational consequence. “Times out” is not enough. “Fails terminally, without retry” says what the system is allowed to do next.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If retry recreates the failure condition, retry is a hang with better logging.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If retry would only restart the stall, the failure class needs to say so. Your operator needs a terminal state, not a more energetic version of stuck.&lt;/p&gt;

&lt;h2 id="why-one-timeout-cannot-do-two-jobs"&gt;Why one timeout cannot do two jobs&lt;/h2&gt;

&lt;p&gt;You need two thresholds when the waiting patterns differ by orders of magnitude. An idle deadline catches silence between chunks; an absolute budget limits the whole turn.&lt;/p&gt;

&lt;p&gt;The slice wraps the stream with two separate controls. The idle deadline resets every time a chunk arrives. It answers: “Did the provider go quiet in the middle of a response?” The absolute budget races against the whole consumer. It answers: “Has this turn taken too long, even though activity continues?” Each one has to be proven without the other.&lt;/p&gt;

&lt;p&gt;That independence matters. Feed chunks often enough and idle will never fire; only the absolute budget can stop an overlong turn. Disable idle and stall the stream; absolute must still cut it. One silent-stream fixture cannot prove both. It only proves whichever control fires first.&lt;/p&gt;

&lt;p&gt;The first pull is deliberately untimed by idle. The records explain why: inter-chunk gaps run around 10–100ms, while a reasoning model can take minutes to its first token. One number cannot serve both distributions. An idle threshold tight enough for the gaps would cut the first response on every call.&lt;/p&gt;

&lt;p&gt;There is a cost to that choice, and it is written down instead of tucked behind a pleasant name. A provider that accepts the connection and never sends a chunk is bounded only by the absolute budget. At the default, that can be up to 30 minutes. The record does not claim that setting is ideal. It names the limitation: a separate first-chunk budget, identified as the proper fix and left out of this slice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do not let one comforting timeout setting pretend it covers three different waits.&lt;/strong&gt; First response, inter-chunk silence, and whole-turn duration need their own evidence. If you cannot show that evidence, say what remains unbounded.&lt;/p&gt;

&lt;h2 id="what-to-inspect-in-your-own-pipeline"&gt;What to inspect in your own pipeline&lt;/h2&gt;

&lt;p&gt;You can find this class of failure without adopting Ranex. Start where a streaming call, a retry policy, and a session state meet. The job is to make the broken path fail before you trust the repaired one.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Find every timeout and write down its failure class: retryable, non-retryable, interrupt, or something else.&lt;/li&gt;
&lt;li&gt;For each retry, ask whether the next attempt changes the condition that failed. If it does not, prove the system settles instead of cycling.&lt;/li&gt;
&lt;li&gt;Separate time to first token, gaps between chunks, and total work time. Do not reuse a threshold merely because it is nearby.&lt;/li&gt;
&lt;li&gt;Make a stream send one chunk and then stall. Confirm the session reaches a terminal state without a person stopping it.&lt;/li&gt;
&lt;li&gt;Make a healthy slow stream complete below the idle threshold and below the whole-turn budget. A watchdog that cuts legitimate work is not a watchdog you can trust.&lt;/li&gt;
&lt;li&gt;Run the timeout while tool work is in flight. Check the actual tool state and settlement, not a log line that says cleanup happened.&lt;/li&gt;
&lt;li&gt;Use a non-default timeout and prove behavior changes. Reading a configuration value back is not proof that it controls the running system.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There is no glamorous trick here. You are looking for the exit from failure, not the detection of failure. That distinction saves you from a green test whose only achievement is proving a timer owns a clock.&lt;/p&gt;

&lt;h2 id="the-second-hang-hiding-behind-the-first"&gt;The second hang hiding behind the first&lt;/h2&gt;

&lt;p&gt;A watchdog failure is an error, not an interrupt. That difference left tool fibers dispatched during streaming uncleared, and the later settlement wait held the turn open forever.&lt;/p&gt;

&lt;p&gt;This is the defect that showed up while building the fix. Cleanup had been tied to interrupts. The watchdog produced a typed error instead, so the cleanup did not run. The runner then waited for tool fibers that had not been cleared. The visible timeout existed; the terminal state did not.&lt;/p&gt;

&lt;p&gt;The completed slice required a specific observable for this path: the tool fiber terminates and the tool is recorded as interrupted. That is stronger than a feeling that a tool was not stranded. It makes the question checkable after the fact.&lt;/p&gt;

&lt;p&gt;That is also why a green test suite deserves a hard question. &lt;a href="https://ranex.dev/blog/concurrency-test-that-proved-nothing" rel="noopener noreferrer"&gt;What did it actually exercise?&lt;/a&gt; In the slice record, the unsafe baseline had to hang a real provider stream inside the runner. Testing a timeout helper in isolation would have been decoration. The runner, the failure class, the retry behavior, the tool cleanup, and the terminal session state all had to meet in the same proof.&lt;/p&gt;

&lt;p&gt;SLICE-012 closed on 2026-08-07 with all nine criteria met, landing as commit &lt;code&gt;23d6a5b4ee&lt;/code&gt; in &lt;code&gt;anthonykewl20/ranex-harness&lt;/code&gt;. This is harness work, not a claim that Ranex is ready for use. Ranex is pre-release; much of the broader system is designed, not built. The watchdog is one shipped durability claim in the harness fork.&lt;/p&gt;

&lt;p&gt;If you want the surrounding model for why the verdict must sit outside an agent’s own loop, read &lt;a href="https://ranex.dev/blog/how-the-kernel-works" rel="noopener noreferrer"&gt;how the kernel works&lt;/a&gt;. The slice record itself is in the &lt;a href="https://github.com/anthonykewl20/ranex" rel="noopener noreferrer"&gt;Ranex repository&lt;/a&gt; under &lt;code&gt;docs/slices/done/&lt;/code&gt;.&lt;/p&gt;

&lt;h2 id="questions-people-actually-ask"&gt;Questions people actually ask&lt;/h2&gt;

&lt;p&gt;These answers explain timeout failures, terminal states, and the waits each budget covers.&lt;/p&gt;

&lt;h3&gt;Why can a watchdog timeout still leave a session hanging?&lt;/h3&gt;

&lt;p&gt;A retryable timeout restarts the same stall when the condition that timed out will recur after retry. The watchdog can fire while the turn still never reaches a terminal state.&lt;/p&gt;

&lt;h3&gt;Why use both idle and absolute timeouts?&lt;/h3&gt;

&lt;p&gt;An idle deadline detects a quiet mid-stream provider, while an absolute budget bounds the whole turn even when chunks keep arriving.&lt;/p&gt;

&lt;h3&gt;Does an idle timeout cover time to first token?&lt;/h3&gt;

&lt;p&gt;No. The first pull is deliberately untimed because inter-chunk gaps and reasoning-model time to first token have very different latency distributions. A connection that never sends is bounded by the absolute budget.&lt;/p&gt;

&lt;h3&gt;What happened to tool work when the watchdog failed?&lt;/h3&gt;

&lt;p&gt;The watchdog produced an error rather than an interrupt, so streaming tool fibers were not cleared and settlement could keep the turn open. The slice added coverage for that path.&lt;/p&gt;

&lt;h2 id="your-next-failure-drill"&gt;Your next failure drill&lt;/h2&gt;

&lt;p&gt;Take one timeout in your pipeline this week. Do not start by making the timer shorter. Stall the real operation after it begins. Then follow the failure all the way through: retry decision, cleanup, settlement, and the state your operator sees.&lt;/p&gt;

&lt;p&gt;Write the failure class down. Split thresholds that are serving different distributions. Name what remains unbounded.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Try it. Break it. Tell me what broke.&lt;/strong&gt; If this record helps you find a test that passes while its failure loops forever, give the &lt;a href="https://github.com/anthonykewl20/ranex" rel="noopener noreferrer"&gt;repository a GitHub star&lt;/a&gt; and send an honest critique. The critique is more useful than applause.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records — the same fact gate the product enforces on code. It ships only after Anthony’s own review.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>concurrency</category>
      <category>flakytests</category>
      <category>testdesign</category>
    </item>
    <item>
      <title>Ranex vs. GitHub Rulesets and Copilot Hooks: What’s Actually Different</title>
      <dc:creator>Anthony Garces</dc:creator>
      <pubDate>Wed, 23 Sep 2026 09:05:05 +0000</pubDate>
      <link>https://dev.to/anthonygarces/ranex-vs-github-rulesets-and-copilot-hooks-whats-actually-different-32om</link>
      <guid>https://dev.to/anthonygarces/ranex-vs-github-rulesets-and-copilot-hooks-whats-actually-different-32om</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Rulesets, editor hooks, and CI gates can control where work moves. Ranex asks what the evidence for that work must establish. Those boundaries can sit together. &lt;a href="https://ranex.dev/blog/ai-accountability-apparatus" rel="noopener noreferrer"&gt;The accountability apparatus&lt;/a&gt; needs both a route and a record.&lt;/p&gt;

&lt;p&gt;You already have rules around a push, a merge, or an editor session. Good. The question is not whether you should throw them away. The question is whether a green result carries the information you need about the artifact that reached that channel.&lt;/p&gt;

&lt;h2 id="in-this-note"&gt;In this note&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;What channel controls are for&lt;/li&gt;
&lt;li&gt;What an evidence control asks instead&lt;/li&gt;
&lt;li&gt;How to use the boundaries together&lt;/li&gt;
&lt;li&gt;Why pre-release status changes the choice&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id="channel-controls-decide-where-work-can-move"&gt;Channel controls decide where work can move&lt;/h2&gt;

&lt;p&gt;GitHub rulesets sit around repository actions such as push and merge. Editor hooks sit in an editor workflow. CI pipeline gates sit in the pipeline. Each can enforce the policy you configure at its own boundary.&lt;/p&gt;

&lt;p&gt;That is worth having. A branch should not advance when required conditions for that branch are absent. An editor action can be stopped before it runs. A pipeline can stop before it publishes its next result. Those controls make authority visible at the channel where the action happens.&lt;/p&gt;

&lt;p&gt;Keep the comparison boring because it is boring. These are different places to put a boundary. Your existing controls may be the right answer for your repository, your team, and the action you need to constrain.&lt;/p&gt;

&lt;p&gt;But a channel rule does not answer every evidence question by itself. A pipeline can report a result. A merge policy can require that result. An editor hook can permit or refuse an action. You still need to ask what command ran, what claim it was meant to support, and what exact artifact it observed.&lt;/p&gt;

&lt;p&gt;That is not a complaint about those controls. It is the normal limit of a boundary. A lock on a door decides who can enter; it does not describe what happened in the room. Your workflow needs the channel decision and a record that can answer the evidence question later.&lt;/p&gt;

&lt;h2 id="evidence-controls-decide-what-a-claim-must-show"&gt;Evidence controls decide what a claim must show&lt;/h2&gt;

&lt;p&gt;Ranex puts its boundary around the claim. Its kernel evaluates gate, evidence, subject, and approver as a pure function. The evidence is bound to a subject digest. A required claim without satisfying evidence is FAIL rather than a default or a skip.&lt;/p&gt;

&lt;p&gt;This is a narrower question than “can this merge happen?” It is: can this record support this claim about this code? The README’s &lt;a href="https://github.com/anthonykewl20/ranex" rel="noopener noreferrer"&gt;“Status” section&lt;/a&gt; says the same command run against a different commit proves nothing about this one. That is the evidence boundary in plain language.&lt;/p&gt;

&lt;p&gt;Suppose a test job produces green output. The channel control can decide whether that output is required before merge. The evidence question remains open until you can state what the job ran, what it measured, and whether its record is bound to the code under judgment. A result with no stated proposition is hard to audit later.&lt;/p&gt;

&lt;p&gt;The evidence boundary also separates production from approval. Ranex lists no self-approval as working behavior: whoever produced the evidence cannot approve it. That does not make a channel control less useful. It supplies a different question for a workflow where an agent can produce both the change and the report about the change.&lt;/p&gt;

&lt;h2 id="use-the-boundaries-together"&gt;Use the boundaries together&lt;/h2&gt;

&lt;p&gt;You can adopt the evidence question without adopting Ranex. Start with the controls you already use, then make each required result answerable.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which channel is being controlled: push, merge, editor action, or pipeline stage?&lt;/li&gt;
&lt;li&gt;What exact claim does the required check support?&lt;/li&gt;
&lt;li&gt;What command is authorized to support that claim?&lt;/li&gt;
&lt;li&gt;Which code digest did the command observe?&lt;/li&gt;
&lt;li&gt;Who produced the evidence, and who is allowed to approve it?&lt;/li&gt;
&lt;li&gt;What happens when the evidence is missing or belongs to older code?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That checklist does not ask you to abandon a familiar stack. It asks you to make the green light legible. Your ruleset can still protect the merge channel. Your hook can still protect the editor channel. Your pipeline can still coordinate work. The evidence record needs to carry its own meaning.&lt;/p&gt;

&lt;p&gt;There is no trophy for replacing working controls. Composition is the point. Let a channel control stop an unauthorized transition. Let an evidence control refuse a claim that lacks the right proof. The first answers where an action may go; the second answers what the action established.&lt;/p&gt;

&lt;h2 id="the-status-is-not-a-replacement-claim"&gt;The status is not a replacement claim&lt;/h2&gt;

&lt;p&gt;Ranex is pre-release. The README says it is a kernel with a working verdict path and very little else. Subject-bound evidence, absence blocks, no self-approval, the journal, and the run-to-evaluate path are listed as working today. The full flow graph and scenario compilation are designed, not built.&lt;/p&gt;

&lt;p&gt;Its limits are part of the comparison. Approver identity is unauthenticated: &lt;code&gt;--approver&lt;/code&gt; is a plain string. Same-UID key theft remains open. The journal does not detect rollback or truncation. Those gaps mean you should read the status before treating any design goal as present capability.&lt;/p&gt;

&lt;p&gt;So do not read this as a replacement pitch. Use the channel controls you trust. Ask the evidence question wherever agent-generated work raises the stakes. If Ranex earns a place later, it earns it by making that question executable, not by declaring your current setup inadequate.&lt;/p&gt;

&lt;h2 id="questions-people-actually-ask"&gt;Questions people actually ask&lt;/h2&gt;

&lt;h3&gt;How does Ranex differ from GitHub rulesets?&lt;/h3&gt;

&lt;p&gt;GitHub rulesets govern repository channels such as push and merge, while Ranex evaluates evidence for a claim against a subject digest.&lt;/p&gt;

&lt;h3&gt;Can Ranex work with editor hooks and CI gates?&lt;/h3&gt;

&lt;p&gt;Ranex can sit beside editor hooks and CI gates because their channel controls and its evidence question address different boundaries.&lt;/p&gt;

&lt;h3&gt;Is Ranex ready to replace an existing governance stack?&lt;/h3&gt;

&lt;p&gt;Ranex is pre-release, so its README describes a working verdict path and known gaps rather than a replacement claim.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Try it. Break it. Tell me what broke.&lt;/strong&gt; Read the &lt;a href="https://github.com/anthonykewl20/ranex" rel="noopener noreferrer"&gt;MIT-licensed repository&lt;/a&gt;, then write down what one required check in your stack actually proves.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records — the same fact gate the product enforces on code. It ships only after Anthony’s own review.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>evidence</category>
      <category>gateverdicts</category>
    </item>
    <item>
      <title>The First Real Delegation: An Agent, a Worktree, and a Judge With No Key</title>
      <dc:creator>Anthony Garces</dc:creator>
      <pubDate>Mon, 21 Sep 2026 09:13:17 +0000</pubDate>
      <link>https://dev.to/anthonygarces/the-first-real-delegation-an-agent-a-worktree-and-a-judge-with-no-key-42bg</link>
      <guid>https://dev.to/anthonygarces/the-first-real-delegation-an-agent-a-worktree-and-a-judge-with-no-key-42bg</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Run agent-written code without a signing key, measure it in a separate keyless process, and keep approval with another authority. This is the next entry after &lt;a href="https://ranex.dev/blog/building-ranex-slice-log" rel="noopener noreferrer"&gt;the Ranex slice log&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;You delegated work to an agent, then watched it run code it had just written. &lt;strong&gt;Would you let that same process hold the key that signs the result?&lt;/strong&gt; A real delegation is built around the answer “no.”&lt;/p&gt;

&lt;h2 id="in-this-note"&gt;In this note&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The dangerous moment after the agent edits the tree&lt;/li&gt;
&lt;li&gt;The lesson: separate execution, measurement, and judgement&lt;/li&gt;
&lt;li&gt;A checklist for your own delegated loop&lt;/li&gt;
&lt;li&gt;What this delegation proved — and what it did not&lt;/li&gt;
&lt;li&gt;Your next move&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id="the-dangerous-moment-after-the-agent-edits-the-tree"&gt;The dangerous moment after the agent edits the tree&lt;/h2&gt;

&lt;p&gt;The risk is not only that an agent writes the wrong code. The risk is that code it wrote gets to reach the authority that declares the work trustworthy.&lt;/p&gt;

&lt;p&gt;A headless agent in a worktree feels contained until you follow the process tree. It can edit the repository. It can trigger the suite. The suite is code it may have changed minutes earlier. If the execute process carries a signing key, the boundary exists only in the story you tell about it.&lt;/p&gt;

&lt;p&gt;SLICE-008 began with a blunt defect: there was no front door. A person had to connect dispatch, harness, emission file, and judgement by hand. The earlier gear-mesh proof used a model that did no work and hand-wrote its own evidence. That proved parts could turn. It did not prove that the kernel could judge work an agent actually produced.&lt;/p&gt;

&lt;p&gt;There was a sharper failure underneath. The existing &lt;code&gt;ranex run&lt;/code&gt; path read the signing key before the bound command spawned. The slice records that &lt;a href="https://github.com/anthonykewl20/ranex/blob/main/docs/slices/done/SLICE-008-first-delegation.md" rel="noopener noreferrer"&gt;&lt;code&gt;RISK-06&lt;/code&gt; reproduced that command taking the key from the process environment&lt;/a&gt;. If an agent-controlled test hook can reach that key, it can produce a signature while the rest of the system looks tidy.&lt;/p&gt;

&lt;p&gt;Do not file that concern away. Design it out of the execution phase.&lt;/p&gt;

&lt;h2 id="the-lesson-separate-execution-measurement-and-judgement"&gt;The lesson: separate execution, measurement, and judgement&lt;/h2&gt;

&lt;p&gt;The practical split is direct: run the agent without a signing key, measure its committed work in a separate keyless invocation, then attest later. The judge produces a candidate with named missing claims, never a self-issued PASS.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;ranex task delegate&lt;/code&gt; dispatches a worktree and starts the harness headless. Its environment is built from empty: a pinned &lt;code&gt;PATH&lt;/code&gt;, a scratch &lt;code&gt;HOME&lt;/code&gt; outside the worktree, bridge variables, one model-credential variable, an auto-approve flag, and a wall-clock bound. When the signing-key variable is present, delegation refuses before it spawns anything.&lt;/p&gt;

&lt;p&gt;“Built from empty” means one concrete thing: the execution environment is deliberately assembled instead of inheriting whatever trust happened to be in the parent shell. That is a much better place to start asking what the worker can reach.&lt;/p&gt;

&lt;p&gt;When the agent finishes, the kernel parses its emission and matches the worktree and commit against its own dispatch record. A missing emission blocks. A mismatch blocks. An empty delegation where the commit equals the base is refused rather than judged.&lt;/p&gt;

&lt;p&gt;Next, the kernel runs the frozen suite against the emitted commit. The commit is materialised, sealed, and offline for that measurement. The invocation holds no key.&lt;/p&gt;

&lt;p&gt;Only after that does a separate, later invocation sign the result. Untrusted code is not running at that time.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The worker can make a change. The measurement can describe the change. Neither action grants approval.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The &lt;code&gt;task judge&lt;/code&gt; result is a journalled &lt;code&gt;CANDIDATE&lt;/code&gt; naming its missing claims. The stamp stays a human’s and out of band. That distinction is the whole design, not a label applied at the end.&lt;/p&gt;

&lt;p&gt;For the wider reason a separate check port matters, read &lt;a href="https://ranex.dev/blog/ai-accountability-apparatus" rel="noopener noreferrer"&gt;the accountability apparatus&lt;/a&gt;. An accountability system needs an inspector outside the actor it inspects.&lt;/p&gt;

&lt;h2 id="a-checklist-for-your-own-delegated-loop"&gt;A checklist for your own delegated loop&lt;/h2&gt;

&lt;p&gt;You can audit this without adopting a kernel. Start with the process that executes agent-produced code, then trace every secret, timeout, artifact, and verdict out to its owner.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Start from an empty environment.&lt;/strong&gt; Write down the exact variables and paths the worker receives. Treat inherited shell state as unreviewed input.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Refuse execution when a signing key is present.&lt;/strong&gt; Do this before spawning the agent or its process tree, not after it has begun.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Inspect the live process tree.&lt;/strong&gt; Do not accept a source-level assertion that a key was removed. Check what the child and parent can actually reach.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bind a task to its worktree before work begins.&lt;/strong&gt; Make the judge compare the emitted worktree and commit with a dispatch record it owns.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure a frozen suite against the emitted commit.&lt;/strong&gt; Do not let a worker summary stand in for an independently observed result.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bound the whole process group.&lt;/strong&gt; A wall-clock limit must kill the group on a stall, not leave child processes behind.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use a state short of approval.&lt;/strong&gt; Candidate is an honest output when claims remain missing. A worker should not award itself PASS.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One dad-joke-sized warning: a timeout that kills only the parent is not a timeout. It is a process group hug with the interesting parts still running.&lt;/p&gt;

&lt;h2 id="what-this-delegation-proved-and-what-it-did-not"&gt;What this delegation proved — and what it did not&lt;/h2&gt;

&lt;p&gt;SLICE-008 proved an end-to-end delegated run against a real free model. It ended in a journalled &lt;code&gt;CANDIDATE&lt;/code&gt; naming missing claims, with no PASS, and a reviewable diff.&lt;/p&gt;

&lt;p&gt;It also tested the shapes that tend to become decoration. The execute phase refuses a signing key before spawn. The live process test checks that the delegated command cannot reach the key through environment, file path, or parent. A planted &lt;code&gt;conftest.py&lt;/code&gt; may run and fail the suite, but it produces no signed record. Forged and missing emissions block before materialisation.&lt;/p&gt;

&lt;p&gt;The wall-clock bound terminates a stalled run’s whole process group, records the timeout, and journals no candidate. The fork’s operator-facing command presents as &lt;code&gt;ranex&lt;/code&gt;, while opencode’s MIT attribution remains in the tree.&lt;/p&gt;

&lt;p&gt;Those are useful facts. They are not a claim that all delegation is safe.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;RISK-06&lt;/code&gt; stayed open when this slice closed.&lt;/strong&gt; The model credential sat in a network-open loop, where it could be posted elsewhere. The record said to use a scoped, spend-limited key, and that the existing &lt;code&gt;ranex run&lt;/code&gt; path still read the key before spawning. Recorded is not mitigated — until it is: SLICE-046 later closed &lt;code&gt;RISK-06&lt;/code&gt; by binding &lt;code&gt;ranex run&lt;/code&gt;’s command inside the qualified confinement session, so the worker can no longer take the signing key. The controller that runs that session is still same-uid trusted infrastructure — see &lt;a href="https://ranex.dev/gaps" rel="noopener noreferrer"&gt;the gap list&lt;/a&gt; for what that still leaves open.&lt;/p&gt;

&lt;p&gt;The slice also does not close merge collisions, verifiable separation, or gate quality. A weak gate can still accept plausible code. You own the target and the quality of the gate. The kernel cannot make either judgement disappear.&lt;/p&gt;

&lt;p&gt;The source record lives in &lt;a href="https://github.com/anthonykewl20/ranex" rel="noopener noreferrer"&gt;&lt;code&gt;docs/slices/done/SLICE-008-first-delegation.md&lt;/code&gt; in the Ranex repository&lt;/a&gt;. Read it if you are designing this boundary. The failures are part of the useful material.&lt;/p&gt;

&lt;h2 id="questions-people-actually-ask"&gt;Questions people actually ask&lt;/h2&gt;

&lt;p&gt;These answers cover keyless measurement and the remaining model-credential risk.&lt;/p&gt;

&lt;h3&gt;Can a delegated agent expose its model API key?&lt;/h3&gt;

&lt;p&gt;The model credential remains in a network-open loop — use a scoped, spend-limited key. &lt;code&gt;RISK-06&lt;/code&gt;, the risk that the worker could take the signing key, was open when SLICE-008 closed; SLICE-046 closed it later by binding &lt;code&gt;ranex run&lt;/code&gt; inside the confinement session.&lt;/p&gt;

&lt;h3&gt;What does an environment built from empty protect against?&lt;/h3&gt;

&lt;p&gt;The delegated harness receives a deliberate environment rather than inherited trust, and delegation refuses to start when the signing-key variable is present.&lt;/p&gt;

&lt;h3&gt;Did SLICE-008 solve model credential exposure?&lt;/h3&gt;

&lt;p&gt;No. The credential still sits in a network-open loop, so use a scoped, spend-limited key. SLICE-008 left &lt;code&gt;RISK-06&lt;/code&gt; open for &lt;code&gt;ranex run&lt;/code&gt;; SLICE-046 closed it later by binding the command inside the confinement session.&lt;/p&gt;

&lt;h2 id="your-next-move"&gt;Your next move&lt;/h2&gt;

&lt;p&gt;Take one agent job that runs tests or scripts. Identify the secret that can authorise its result. Then inspect the live child and parent processes while the job runs. If that secret is reachable, split execution from measurement before you trust another green status.&lt;/p&gt;

&lt;p&gt;Write down what remains open too. “Recorded, not mitigated” is not a failure of the notes. It is how the next person avoids mistaking a boundary for a finished fortress.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Try it. Break it. Tell me what broke.&lt;/strong&gt; If this was useful, &lt;a href="https://github.com/anthonykewl20/ranex" rel="noopener noreferrer"&gt;star the Ranex repository&lt;/a&gt; and leave an honest critique. A hard question is more valuable than a polite nod.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records — the same fact gate the product enforces on code. It ships only after Anthony’s own review.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>agentsafety</category>
      <category>separationofduties</category>
      <category>signing</category>
    </item>
    <item>
      <title>Absence Blocks: Why No Evidence Is a Fail, Never a Skip</title>
      <dc:creator>Anthony Garces</dc:creator>
      <pubDate>Fri, 18 Sep 2026 09:04:21 +0000</pubDate>
      <link>https://dev.to/anthonygarces/absence-blocks-why-no-evidence-is-a-fail-never-a-skip-1p86</link>
      <guid>https://dev.to/anthonygarces/absence-blocks-why-no-evidence-is-a-fail-never-a-skip-1p86</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; A required claim without satisfying evidence must fail, because a skipped check is not proof; a fresh Ranex clone starts red for exactly that reason. &lt;a href="https://ranex.dev/blog/how-the-kernel-works" rel="noopener noreferrer"&gt;Read how the kernel works.&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Your dashboard is green, yet a job did not run. Do you know which one? If the answer lives in a log nobody reads, your green result is carrying more confidence than the evidence earned.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ranex treats absence as a blocker.&lt;/strong&gt; When a gate requires a claim and no satisfying evidence exists, the verdict is FAIL — not a default, not a skip, not a warning waiting for a tired person to scroll past.&lt;/p&gt;

&lt;h2 id="in-this-note"&gt;In this note&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;A silent skip is not a pass&lt;/li&gt;
&lt;li&gt;A fresh clone starts red&lt;/li&gt;
&lt;li&gt;Make absence visible in your pipeline&lt;/li&gt;
&lt;li&gt;Failure names the missing proof&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id="a-silent-skip-is-not-a-pass"&gt;A silent skip is not a pass&lt;/h2&gt;

&lt;p&gt;A check that did not run has established nothing. Treating that empty space as success turns missing work into a green light.&lt;/p&gt;

&lt;p&gt;You know the shapes. A condition bypasses a job. A suite collects no tests and returns cleanly. A scanner stops early and leaves a warning below the fold. The summary says green because the system recorded an exit path, not because it recorded the evidence you needed. That is a failure mode, not an inconvenient edge case.&lt;/p&gt;

&lt;p&gt;Ranex makes the missing piece explicit. A gate carries required claims. Evidence must satisfy those claims for the subject under judgment. If the required evidence is absent, evaluation returns FAIL. The rule is an invariant of the kernel, alongside subject-bound evidence and no self-approval — though approver identity is unauthenticated today, so that check compares unverified strings. A gate that cannot block is refused at construction, because a gate that cannot stop anything is decoration.&lt;/p&gt;

&lt;p&gt;This is why absence and determinism belong together. &lt;a href="https://ranex.dev/blog/verdict-pure-function" rel="noopener noreferrer"&gt;A pure-function verdict&lt;/a&gt; gives the same answer for the same inputs. Absence blocks answers FAIL when an input cannot satisfy the rule. Without both, “same inputs” can still mean “missing evidence became permission.”&lt;/p&gt;

&lt;p&gt;There is a cost. A fail-closed system will make you produce evidence before it lets you call the work done. That can feel fussy when you know the command would have passed. But the entire point is that the system does not accept what you know without a record it can evaluate.&lt;/p&gt;

&lt;h2 id="a-fresh-clone-starts-honestly"&gt;A fresh clone starts red&lt;/h2&gt;

&lt;p&gt;A fresh Ranex clone fails gate evaluation because it has no evidence yet. That failure is correct: no record exists for the required &lt;code&gt;tests-executed&lt;/code&gt; claim.&lt;/p&gt;

&lt;p&gt;The README gives the operator path. Run the test suite with frozen dependencies, then ask the gate to evaluate the current subject. On a fresh clone, evaluation returns FAIL with a nonzero exit and names the missing claim. It does not create a cheerful initial state. It does not infer success from an empty evidence file. It tells you that nothing has been proven.&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;PYTHONPATH=src uv run --frozen python -m ranex.cli.main gate evaluate HEAD \
    --approver reviewer_alice&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;To produce evidence, the operator needs a signing identity whose private key lives outside the repository. The committed keyring contains public halves, not private keys. After deliberate dependency provisioning, &lt;code&gt;ranex run&lt;/code&gt; records the observed command result and subject digest; evaluation then judges that evidence against the gate. The steps matter because the verdict is not an applause button for a command. It is a judgment about whether the bound evidence satisfies the required claim for this subject.&lt;/p&gt;

&lt;p&gt;Fresh-clone failure is a fast test of your own mental model. If a system can announce PASS before any relevant evidence exists, what did PASS mean? It meant the system had another rule, even if nobody wrote it down. Ranex writes its rule down in behavior: no satisfying evidence, no pass.&lt;/p&gt;

&lt;h2 id="make-absence-visible-in-your-pipeline"&gt;Make absence visible in your pipeline&lt;/h2&gt;

&lt;p&gt;You do not need to adopt a kernel to look for this fault line. Pick a recent green deployment and inspect the evidence path rather than the badge.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;List every claim the release depends on, not only every job name.&lt;/li&gt;
&lt;li&gt;For each claim, identify the record that satisfies it and the exact subject it covers.&lt;/li&gt;
&lt;li&gt;Make a missing result fail rather than disappear into a default branch.&lt;/li&gt;
&lt;li&gt;Make skipped, errored, and missing test outcomes visible as distinct states.&lt;/li&gt;
&lt;li&gt;Change the tree and confirm earlier evidence no longer satisfies the new subject.&lt;/li&gt;
&lt;li&gt;Ask whether an operator can tell the difference between unfinished work and a rejected record.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last distinction matters. A record that fails verification is not reported as “no evidence.” Ranex reports it as refused and includes a reason. Missing evidence means no satisfying record exists. Refused evidence means a record arrived but did not clear verification. An unfinished task and an attack are different events; hiding them under the same empty label destroys useful information.&lt;/p&gt;

&lt;p&gt;Staleness is absence in another form. Evidence is bound to a subject digest, so the same command on a different commit proves nothing about the current one. Once the tree moves past the digest that evidence covered, it stops satisfying the claim. That is not the system being difficult. It is the system refusing to let yesterday’s proof stand in for today’s code.&lt;/p&gt;

&lt;h2 id="failure-names-the-missing-proof"&gt;Failure names the missing proof&lt;/h2&gt;

&lt;p&gt;A useful FAIL tells you which required proof is missing or refused. It does not pretend that all failures are identical, and it does not claim a guarantee it lacks.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://github.com/anthonykewl20/ranex#status" rel="noopener noreferrer"&gt;repository’s Status section&lt;/a&gt; lists absence blocks as working today and says the project is pre-release. The full flow-graph and scenario-compilation picture is designed rather than built. Keep that boundary in view. The working claim here is smaller: the verdict path fails when a required claim has no satisfying evidence.&lt;/p&gt;

&lt;p&gt;That small claim changes the conversation during a release. Instead of “the pipeline was green,” you can ask “which evidence satisfied this claim for this commit?” If there is no answer, you do not need a debate about whether the omission feels safe. The gate already has the right response.&lt;/p&gt;

&lt;p&gt;Go to the last build you trusted because it was quiet. Find one claim it was supposed to establish. Then remove its evidence in a disposable copy of the process. If the release still looks green, you found the work to do.&lt;/p&gt;

&lt;h2 id="questions-people-actually-ask"&gt;Questions people actually ask&lt;/h2&gt;

&lt;h3&gt;What does absence blocks mean in CI?&lt;/h3&gt;

&lt;p&gt;Ranex treats a required claim without satisfying evidence as FAIL, never as a default or a skip.&lt;/p&gt;

&lt;h3&gt;Why does a fresh Ranex clone fail evaluation?&lt;/h3&gt;

&lt;p&gt;A fresh Ranex clone has no evidence for the required tests-executed claim, so gate evaluation correctly returns FAIL with a nonzero exit.&lt;/p&gt;

&lt;h3&gt;Is refused evidence the same as missing evidence?&lt;/h3&gt;

&lt;p&gt;Ranex reports a record that fails verification as refused with a reason, while missing evidence means no satisfying record exists.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Try it. Break it. Tell me what broke.&lt;/strong&gt; Read the &lt;a href="https://github.com/anthonykewl20/ranex" rel="noopener noreferrer"&gt;Ranex repository&lt;/a&gt;, then make one required check absent in a safe copy of your own pipeline.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records — the same fact gate the product enforces on code. It ships only after Anthony’s own review.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>evidence</category>
      <category>gateverdicts</category>
    </item>
    <item>
      <title>I Forked My Agent Harness So It Couldn’t Grade Its Own Homework</title>
      <dc:creator>Anthony Garces</dc:creator>
      <pubDate>Wed, 16 Sep 2026 09:02:16 +0000</pubDate>
      <link>https://dev.to/anthonygarces/i-forked-my-agent-harness-so-it-couldnt-grade-its-own-homework-551d</link>
      <guid>https://dev.to/anthonygarces/i-forked-my-agent-harness-so-it-couldnt-grade-its-own-homework-551d</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; An agent harness cannot independently approve work when it controls both the artifact and the success path; move “done” outside the loop. This is one record from &lt;a href="https://ranex.dev/blog/building-ranex-slice-log" rel="noopener noreferrer"&gt;the Ranex slice log&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Your agent finished the task. The harness says it is done. &lt;strong&gt;Who, exactly, was allowed to make that call?&lt;/strong&gt; That question is the next layer under every green check you got this week: the moment the harness had to stop grading its own homework.&lt;/p&gt;

&lt;h2 id="in-this-note"&gt;In this note&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The problem hiding inside a successful run&lt;/li&gt;
&lt;li&gt;The lesson: move “done” outside the loop&lt;/li&gt;
&lt;li&gt;A checklist for finding self-grading in your pipeline&lt;/li&gt;
&lt;li&gt;What was actually proven in this slice&lt;/li&gt;
&lt;li&gt;Your next move&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id="the-problem-hiding-inside-a-successful-run"&gt;The problem hiding inside a successful run&lt;/h2&gt;

&lt;p&gt;A check bolted around an agent is not independent if the agent’s harness can still author the route to success. You need a different actor, with its own record, to decide what the emitted work means.&lt;/p&gt;

&lt;p&gt;You know the shape. A task starts. An agent edits files. A wrapper collects a summary, a commit, a test result, or all three. Then that same wrapper turns around and announces success.&lt;/p&gt;

&lt;p&gt;That is not a tiny trust gap — it is the whole gap.&lt;/p&gt;

&lt;p&gt;The agent does not need bad intent for this to fail. The harness can be wrong, incomplete, configured too broadly, or simply allowed to load something you did not expect. If its output is both the evidence and the verdict, there is no independent place to ask the rude question: &lt;em&gt;Did this task produce what the dispatch actually asked for?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Ranex is pre-release. It is a kernel with a working verdict path and very little else. But this boundary is built because a useful system needs to make that rude question executable, not ceremonial.&lt;/p&gt;

&lt;p&gt;I did not get to skip this by calling the surrounding code a harness. Names do not create separation. &lt;strong&gt;Authority does.&lt;/strong&gt;&lt;/p&gt;

&lt;h2 id="the-lesson-move-done-outside-the-loop"&gt;The lesson: move “done” outside the loop&lt;/h2&gt;

&lt;p&gt;The fix is structural: let the harness produce work and references, then let the kernel cross-check them against a record it owns. The harness can emit evidence; it cannot emit a gate, a merge, a stamp, or an approver.&lt;/p&gt;

&lt;p&gt;For SLICE-007, &lt;a href="https://github.com/anthonykewl20/ranex/blob/main/docs/slices/done/SLICE-007-trimmed-fork-to-the-kernel.md" rel="noopener noreferrer"&gt;Ranex forked &lt;code&gt;opencode&lt;/code&gt; at &lt;code&gt;v1.18.11&lt;/code&gt;, commit &lt;code&gt;012c2f57&lt;/code&gt;&lt;/a&gt;. The fork was trimmed to a defined keep-set. Its plugin surface was locked to compiled-in built-ins. It refuses to start when the required bridge is missing or unbound.&lt;/p&gt;

&lt;p&gt;Those details sound operational. They are the point.&lt;/p&gt;

&lt;p&gt;A harness that can take configuration or npm plugin paths can change what is running around the work. A harness that starts without its bridge can run unobserved. You do not repair either condition by asking the agent to be more careful.&lt;/p&gt;

&lt;p&gt;Here is the deal: &lt;strong&gt;the actor that performs the work should not own the meaning of its own evidence.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The kernel gained &lt;code&gt;task dispatch&lt;/code&gt; and &lt;code&gt;task judge&lt;/code&gt;. At dispatch, it records the task-to-worktree relationship in an append-only, hash-chained journal. The harness works in that worktree, commits, and emits references. Then the kernel reads the committed worktree itself, compares the emitted references with its own dispatch record, materialises the commit, and evaluates the evidence.&lt;/p&gt;

&lt;p&gt;The resulting journal entry is &lt;code&gt;CANDIDATE&lt;/code&gt;. Never &lt;code&gt;PASS&lt;/code&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A candidate is work ready to be judged. It is not permission granted by the worker that made it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The approval stamp stays human and out of band. That can feel slower than letting an agent complete the sentence with its own praise. Good. A gate exists to make an important transition harder to fake.&lt;/p&gt;

&lt;p&gt;If you want the broader shape of the boundary, &lt;a href="https://ranex.dev/blog/how-the-kernel-works" rel="noopener noreferrer"&gt;the kernel topology&lt;/a&gt; explains why workers return a diff while the check port produces the only verdict that counts.&lt;/p&gt;

&lt;h2 id="a-checklist-for-finding-self-grading-in-your-pipeline"&gt;A checklist for finding self-grading in your pipeline&lt;/h2&gt;

&lt;p&gt;You can find the dangerous shapes without adopting Ranex. Follow a task from dispatch to the word “done,” and mark every place one process supplies both the artifact and its interpretation.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;One wrapper owns task creation and completion.&lt;/strong&gt; Hunt for a task identifier created by the same process that later marks it complete.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The worker reports the commit it wants judged.&lt;/strong&gt; Check whether an independent process reads the worktree or commit itself and compares it with a dispatch record.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plugins can arrive through configuration or a package path.&lt;/strong&gt; List what can load into your harness at runtime. If the list is open, treat that as authority entering through a side door.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The harness can run without its observer.&lt;/strong&gt; Remove or unbind the bridge in a disposable environment. Does startup refuse, or does work continue without the boundary?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A summary becomes a verdict.&lt;/strong&gt; Separate “the worker says tests ran” from an independently measured result against the committed subject.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The word PASS appears before a person or separate authority acts.&lt;/strong&gt; Replace it with a state that says what it is: candidate, pending, refused, or failed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Do not settle for a diagram where these responsibilities have different labels. Ask which process can write each record and which process reads the artifact from disk. The answers matter more than the labels.&lt;/p&gt;

&lt;h2 id="what-was-actually-proven-in-this-slice"&gt;What was actually proven in this slice&lt;/h2&gt;

&lt;p&gt;SLICE-007 proved the path end to end: dispatch, worker loop, hooks, kernel judgement, evidence, and a journalled &lt;code&gt;CANDIDATE&lt;/code&gt;. The gear-mesh end-to-end test ran that loop using a deterministic in-fork model with zero credentials.&lt;/p&gt;

&lt;p&gt;That proof is narrow on purpose. It does not claim &lt;a href="https://ranex.dev/blog/keyless-agent-delegation" rel="noopener noreferrer"&gt;delegation&lt;/a&gt;, clean-room orchestration, confinement, authenticated approval, or a finished product. The slice record says those boundaries remain outside its closure.&lt;/p&gt;

&lt;p&gt;It also does not turn an emitted candidate into a universal claim of correctness. Ranex judges work by evidence and executable checks, and it cannot decide whether the target itself was the right target. The owner still owns that judgement.&lt;/p&gt;

&lt;p&gt;The close is recorded in &lt;code&gt;docs/adr/ADR-008-fork-opencode-and-bridge-to-the-kernel.md&lt;/code&gt;. The slice record is in &lt;a href="https://github.com/anthonykewl20/ranex" rel="noopener noreferrer"&gt;&lt;code&gt;docs/slices/done/&lt;/code&gt; in the Ranex repository&lt;/a&gt;. The fork retained opencode’s MIT attribution. That is one small line with a large habit behind it: say what you used, say what you changed, and do not borrow credit with the code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Attribution compounds.&lt;/strong&gt; Especially in a system built to distinguish a claim from its proof.&lt;/p&gt;

&lt;h2 id="questions-people-actually-ask"&gt;Questions people actually ask&lt;/h2&gt;

&lt;p&gt;These answers explain how a harness can be kept from approving its own work.&lt;/p&gt;

&lt;h3&gt;Why is an agent harness grading itself a problem?&lt;/h3&gt;

&lt;p&gt;When the same actor can produce work, shape the success path, and declare completion, the check does not independently establish the claim.&lt;/p&gt;

&lt;h3&gt;How can an AI agent harness be stopped from approving its own work?&lt;/h3&gt;

&lt;p&gt;The harness can be stopped by letting it produce work and references while a separate kernel cross-checks them against a record it owns.&lt;/p&gt;

&lt;h3&gt;Does the kernel issue PASS for harness work?&lt;/h3&gt;

&lt;p&gt;No. The kernel cross-checks the dispatch record, materialises the committed work, and journals &lt;code&gt;CANDIDATE&lt;/code&gt;. A human stamp remains out of band.&lt;/p&gt;

&lt;h2 id="your-next-move"&gt;Your next move&lt;/h2&gt;

&lt;p&gt;Pick one agent task from this week. Trace its task ID, its worktree or branch, its commit, its evidence, and its final status. Then ask whether the actor that changed the files also had a path to write the final status.&lt;/p&gt;

&lt;p&gt;If the answer is yes, do not call the task done yet. Make a separate authority read the committed artifact and compare it with a record it created before the work began.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Try it. Break it. Tell me what broke.&lt;/strong&gt; If this helps, &lt;a href="https://github.com/anthonykewl20/ranex" rel="noopener noreferrer"&gt;star the Ranex repository&lt;/a&gt; — then leave an honest critique. The critique is more useful than applause.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records — the same fact gate the product enforces on code. It ships only after Anthony’s own review.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>agentsafety</category>
      <category>separationofduties</category>
      <category>testdesign</category>
    </item>
    <item>
      <title>A Report Is Not the Artifact</title>
      <dc:creator>Anthony Garces</dc:creator>
      <pubDate>Mon, 14 Sep 2026 09:35:49 +0000</pubDate>
      <link>https://dev.to/anthonygarces/a-report-is-not-the-artifact-4pk4</link>
      <guid>https://dev.to/anthonygarces/a-report-is-not-the-artifact-4pk4</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; A report tells you what a session says happened. The artifact on disk and a fresh rerun tell you what you can check. &lt;a href="https://ranex.dev/blog/building-ranex-slice-log" rel="noopener noreferrer"&gt;The parent slice log&lt;/a&gt; is where the distinction was recorded.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Your agent says the gate passed. Before you approve it, can you point to the artifact and rerun the command without the agent in the room?&lt;/strong&gt; If not, you have a description of evidence, not evidence you can independently inspect.&lt;/p&gt;

&lt;h2 id="in-this-note"&gt;In this note&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;A report is self-description&lt;/li&gt;
&lt;li&gt;The artifact is where the claim meets reality&lt;/li&gt;
&lt;li&gt;Approval does not create an observation&lt;/li&gt;
&lt;li&gt;Make your pipeline answerable&lt;/li&gt;
&lt;li&gt;Rerun before you rely on it&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id="a-report-is-self-description"&gt;A report is self-description&lt;/h2&gt;

&lt;p&gt;A session report is useful for navigation. It is not the same thing as the files, command, and output it describes.&lt;/p&gt;

&lt;p&gt;That sounds obvious until the report is green and the review queue is long. A worker says it changed a fixture, ran a suite, and saw a pass. A reviewer reads that sentence and gives it a second approval. The words become more official, but nobody has created another observation of the work.&lt;/p&gt;

&lt;p&gt;Ranex is built around a smaller rule: read the diff on disk; discard the worker summary; run the checks. The rule is not distrust for its own sake. It gives a later reviewer something that exists outside the session which made the claim. A summary can be wrong, stale, incomplete, or written before the final artifact existed. The artifact gives you a place to ask again.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/anthonykewl20/ranex/blob/main/docs/slices/done/SLICE-011-durable-execution-prototype.md" rel="noopener noreferrer"&gt;SLICE-011 says the supervisor reran every gate against the worktree on disk&lt;/a&gt; and calls the session summary discarded self-report. That wording is worth borrowing for your own pipeline. A report is a claim about a run. It does not get promoted to evidence just because it is formatted cleanly.&lt;/p&gt;

&lt;h2 id="the-artifact-is-where-the-claim-meets-reality"&gt;The artifact is where the claim meets reality&lt;/h2&gt;

&lt;p&gt;The artifact lets you test the claim again, using the command that is supposed to support it.&lt;/p&gt;

&lt;p&gt;That is the transferable rule. Check out the exact worktree or commit. Read the actual diff. Run the declared command. Preserve its output where another person can retrieve it. When the behavior has a failure mode, force that path too. Each step turns a statement into something a later person can challenge.&lt;/p&gt;

&lt;p&gt;A report can still be retained. It can name the command, commit, environment, artifact digest, and outcome so you know what to reproduce. Its proper job is a map, not a substitute for the terrain.&lt;/p&gt;

&lt;p&gt;This matters most for concurrency, timing, process cleanup, or shared state. Those failures can disappear in a clean session and return when the same artifact is run again. It also matters for ordinary build work: a report that describes a test without preserving the exact subject has asked you to trust memory at the moment you need measurement.&lt;/p&gt;

&lt;h2 id="approval-does-not-create-an-observation"&gt;Approval does not create an observation&lt;/h2&gt;

&lt;p&gt;Two people approving one report have read one observation twice. They have not independently verified the artifact.&lt;/p&gt;

&lt;p&gt;That is not a criticism of review. Review can catch a bad instruction, a mistaken assumption, or a missing case. But no amount of careful reading can make an unrerun report answer a command it never ran.&lt;/p&gt;

&lt;p&gt;The incident narrative belongs in &lt;a href="https://ranex.dev/blog/flaky-suite-approved-twice" rel="noopener noreferrer"&gt;How a 12% Flaky Test Suite Got Approved Twice&lt;/a&gt;: the suite was reported stable, approved by two reviewers, and then a rerun against disk found two failures in sixteen full-suite runs. This note keeps the principle: approvals are readers; the artifact is the thing that can be observed again.&lt;/p&gt;

&lt;p&gt;Use that distinction when someone proposes “another sign-off” as the fix for uncertain evidence. Ask what new observation the sign-off adds. If the answer is none, schedule a rerun or a fresh measurement instead.&lt;/p&gt;

&lt;h2 id="make-your-pipeline-answerable"&gt;Make your pipeline answerable&lt;/h2&gt;

&lt;p&gt;You do not need a governance kernel to make reports earn their authority. Give every important result a route back to its artifact.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Record the immutable subject: a commit, build digest, or other exact artifact identity.&lt;/li&gt;
&lt;li&gt;Record the exact command and the inputs it used, not a paraphrase such as “tests passed.”&lt;/li&gt;
&lt;li&gt;Make the artifact available to the person who approves the result.&lt;/li&gt;
&lt;li&gt;Rerun blocking checks outside the producing session before using them to merge, deploy, or certify a control.&lt;/li&gt;
&lt;li&gt;Force a relevant refusal or negative path and confirm the check can distinguish it from success.&lt;/li&gt;
&lt;li&gt;Keep stderr and other diagnostics that explain a failure instead of retaining only a green summary.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are not ceremonies to add after the real work. They are how you discover whether the work was real before your pipeline gives it teeth.&lt;/p&gt;

&lt;h2 id="questions-people-actually-ask"&gt;Questions people actually ask&lt;/h2&gt;

&lt;p&gt;These questions keep a report in its proper role: a map back to a checkable artifact.&lt;/p&gt;

&lt;h3&gt;Why is a session report not evidence by itself?&lt;/h3&gt;

&lt;p&gt;A session report is self-description from the run that produced it. Ranex treats the worktree and a rerun of the recorded command as evidence that can contradict that description.&lt;/p&gt;

&lt;h3&gt;What should a reviewer verify instead of trusting a report?&lt;/h3&gt;

&lt;p&gt;A reviewer should inspect the artifact on disk, rerun the exact gate command, force the relevant failure path, and preserve the result outside the producing session.&lt;/p&gt;

&lt;h3&gt;Can two approvals make one report independently verified?&lt;/h3&gt;

&lt;p&gt;No. Ranex learned that two reviewers can approve the same self-description; approval adds readers, not a new observation of the artifact.&lt;/p&gt;

&lt;h3&gt;What did SLICE-011 prove about reports and artifacts?&lt;/h3&gt;

&lt;p&gt;SLICE-011 proved that rerunning every gate against the worktree on disk found two failures in sixteen full-suite runs after the suite had been reported stable and approved.&lt;/p&gt;

&lt;h2 id="rerun-before-you-rely-on-it"&gt;Rerun before you rely on it&lt;/h2&gt;

&lt;p&gt;Pick one report that can block a release or approve an agent change. Locate the artifact it claims to describe. Rerun its exact command against that artifact. Then make the command fail on purpose and verify that the report would not hide it.&lt;/p&gt;

&lt;p&gt;Ranex is pre-release. The lesson does not depend on waiting for it: a report is a useful map, and the artifact is still the ground.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Try it. Break it. Tell me what broke.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records — the same fact gate the product enforces on code. It ships only after Anthony’s own review.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>agentsafety</category>
      <category>evidence</category>
      <category>gateverdicts</category>
    </item>
    <item>
      <title>Nine Defects, Zero Unit Tests: Drive Your Tool Like a Stranger</title>
      <dc:creator>Anthony Garces</dc:creator>
      <pubDate>Fri, 11 Sep 2026 09:06:14 +0000</pubDate>
      <link>https://dev.to/anthonygarces/nine-defects-zero-unit-tests-drive-your-tool-like-a-stranger-1e4n</link>
      <guid>https://dev.to/anthonygarces/nine-defects-zero-unit-tests-drive-your-tool-like-a-stranger-1e4n</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Unit tests missed nine CLI journey defects; run the tool from zero state, using only its published instructions. &lt;a href="https://ranex.dev/blog/building-ranex-slice-log" rel="noopener noreferrer"&gt;The parent slice log is a reminder to test the artifact people actually touch&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Your command works on your machine, your unit tests are green, and a new operator still cannot get through the first run.&lt;/strong&gt; Can Ranex gate Ranex for real?&lt;/p&gt;

&lt;h2 id="in-this-note"&gt;In this note&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The CLI is where your system becomes real&lt;/li&gt;
&lt;li&gt;Nine defects appeared only in the journey&lt;/li&gt;
&lt;li&gt;Drive these shapes like a new operator&lt;/li&gt;
&lt;li&gt;Approval reduces hidden change. It cannot make code truthful.&lt;/li&gt;
&lt;li&gt;Run your own first-day test&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Your tool lives at its CLI surface. That is where paths, locks, machines, empty state, instructions, and human assumptions collide.&lt;/p&gt;

&lt;p&gt;SLICE-006 found nine defects by driving the Ranex CLI as a person does. None came from a unit test. That is not an argument against unit tests. It is an argument against asking them to prove a journey they never take.&lt;/p&gt;

&lt;h2 id="the-cli-is-where-your-system-becomes-real"&gt;The CLI is where your system becomes real&lt;/h2&gt;

&lt;p&gt;Your tool’s real failure modes live where a person invokes it. A unit test can prove a component; it cannot automatically prove that setup, documentation, state, and the command path join into a usable and truthful run.&lt;/p&gt;

&lt;p&gt;Ranex had a specific problem. It materialises committed blobs to observe the subject tree. That is the right subject, but ignored directories such as &lt;code&gt;.venv&lt;/code&gt; and &lt;code&gt;node_modules&lt;/code&gt; are not in that tree. The bound test command needed dependencies and a resolver that the observed party could not choose from an ambient writable path.&lt;/p&gt;

&lt;p&gt;So the question became practical: can Ranex gate Ranex for real?&lt;/p&gt;

&lt;p&gt;The answer built in SLICE-006 has four deliberate steps. &lt;code&gt;deps fetch&lt;/code&gt; derives the lock clean under pinned inputs and byte-compares it with the committed lock. Only SHA-256-addressed wheels enter the store. &lt;code&gt;deps approve&lt;/code&gt; records the named package delta an approver accepted. Before anything runs, a fresh environment is assembled from the verified store entries and made read-only. Then &lt;code&gt;run&lt;/code&gt; executes sealed and offline.&lt;/p&gt;

&lt;p&gt;The gate command remains the catalog-bound &lt;code&gt;uv run pytest -q&lt;/code&gt;. The environment is assembled from verified store entries, made read-only before spawn, and set so the command cannot sync, build, or fetch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That is a dependency process, not a truth machine.&lt;/strong&gt; A hash tells you which bytes arrived. Approval tells you which package change a person accepted. Neither tells you that dependency code will behave honestly once it runs.&lt;/p&gt;

&lt;h2 id="nine-defects-appeared-only-in-the-journey"&gt;Nine defects appeared only in the journey&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/anthonykewl20/ranex/blob/main/docs/slices/done/SLICE-006-gating-a-real-test-suite.md" rel="noopener noreferrer"&gt;Nine defects were found&lt;/a&gt; by driving the CLI as a person does, and none by a unit test. They appeared at the joins: real locks, cold-start state, operator instructions, CI, and the fresh repository used for observation.&lt;/p&gt;

&lt;p&gt;The first defect is a familiar one. Plain &lt;code&gt;uv run pytest -q&lt;/code&gt; re-locked and rewrote &lt;code&gt;uv.lock&lt;/code&gt;. In this system, that lock is a trust root. After a clean re-lock, the rewrite silently dropped the resolution epoch block. Then &lt;code&gt;deps fetch&lt;/code&gt; refused the committed lock against its own clean derivation.&lt;/p&gt;

&lt;p&gt;The fix had two parts. The gated run sets &lt;code&gt;UV_FROZEN=1&lt;/code&gt;, while the repository’s own commands use &lt;code&gt;uv run --frozen&lt;/code&gt;. The difference is not cosmetic. The gated argv must stay exactly as the catalog binds it.&lt;/p&gt;

&lt;p&gt;Other journey failures had the same practical flavor. A real lock held one package at several versions, and the parser rejected it as corruption. A check for a fabricated hash was passing for the wrong reason because the resolution epoch was omitted. CI was also invoking &lt;code&gt;uv&lt;/code&gt; unfrozen, silently mutating the same trust root where it would be hardest to notice.&lt;/p&gt;

&lt;p&gt;Then the cold-start journey caught product-facing defects. &lt;code&gt;keygen&lt;/code&gt; told a first-time operator to create an invalid keyring. The README walkthrough had rotted: it omitted the fetch and approval steps, failed to name installation of the pinned resolver, and described a requirement that had been removed. Another journey re-entered itself inside the materialised sample because its recursion guard lived in an environment that the sample deliberately built from empty.&lt;/p&gt;

&lt;p&gt;None of those defects is a tiny detail to the person blocked by it. The tool either works from zero state or it does not.&lt;/p&gt;

&lt;h2 id="drive-these-shapes-like-a-new-operator"&gt;Drive these shapes like a new operator&lt;/h2&gt;

&lt;p&gt;Test the command surface as a stranger with no warm cache, no hidden setup, and no memory of why the tool works. The checklist is not “add more end-to-end tests.” It is a list of states your unit suite is built to avoid.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Start from zero state: no generated keyring, no provisioned store, no prior approval, no warm cache.&lt;/li&gt;
&lt;li&gt;Follow the README command by command. Confirm every command it requires is present, ordered, and still names current behavior.&lt;/li&gt;
&lt;li&gt;Run with the real committed lock and manifest, then deliberately alter a package, graph edge, URL, or hash and require byte comparison to refuse.&lt;/li&gt;
&lt;li&gt;Use the pinned resolver path and prove a user-writable resolver is rejected.&lt;/li&gt;
&lt;li&gt;Observe the dependency run offline. A network attempt must deny and produce no evidence.&lt;/li&gt;
&lt;li&gt;Exercise the tool in CI, where a separate command spelling can rewrite a trust root without a developer noticing.&lt;/li&gt;
&lt;li&gt;Run inside every clean-room or materialised repository shape the tool creates for itself.&lt;/li&gt;
&lt;li&gt;Force an ordinary success path too. A refusal-only test can pass because the command never ran.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The last item is non-negotiable. SLICE-006 records an actual dependency-bearing execution that provisions once, reuses the store without downloads, runs offline, signs evidence, and evaluates PASS. A system that only proves its refusals is an outage wearing responsible clothes.&lt;/p&gt;

&lt;p&gt;You do not need a large product to borrow this method. Take one onboarding command. Use a fresh environment. Follow your own docs without filling gaps from memory. Every place you have to “just know” something is a candidate defect.&lt;/p&gt;

&lt;h2 id="approval-reduces-hidden-change-it-cannot-make-code-truthful"&gt;Approval reduces hidden change. It cannot make code truthful.&lt;/h2&gt;

&lt;p&gt;An approved, hash-correct wheel can still choose its own exit code. SLICE-006 demonstrates this with &lt;code&gt;tests/security/test_slice006_approved_wheel_can_lie.py&lt;/code&gt;, where an approved wheel forces a passing verdict.&lt;/p&gt;

&lt;p&gt;That test is labelled not caught. It should be. Calling integrity proof would be a lie.&lt;/p&gt;

&lt;p&gt;Python makes the route concrete through installed &lt;code&gt;pytest11&lt;/code&gt; entry points, but the boundary is larger than one plugin mechanism. Direct imports also execute dependency code. The dependency is part of the trusted computing base for that run.&lt;/p&gt;

&lt;p&gt;This is why &lt;a href="https://ranex.dev/blog/ai-accountability-apparatus" rel="noopener noreferrer"&gt;an accountability apparatus&lt;/a&gt; needs its limits stated plainly. You can reduce hidden change with a clean derivation, pinned inputs, SHA-256-addressed wheels, a reviewable package delta, and sealed offline execution. You cannot derive truthful behavior from a package hash.&lt;/p&gt;

&lt;p&gt;Ranex is pre-release. This slice closes a runnable self-gate for its Python dependency path. Other ecosystems need their own manifest, lock, and artifact policies. No generic abstraction is claimed here.&lt;/p&gt;

&lt;p&gt;Also, the lock story is worth keeping in your operating memory. A lockfile is not boring generated clutter when it decides what code enters a measured run. Plain &lt;code&gt;uv run&lt;/code&gt; rewrote &lt;code&gt;uv.lock&lt;/code&gt; and silently dropped the resolution epoch once. The later refusal was the control doing its job, not the control being inconvenient.&lt;/p&gt;

&lt;h2 id="questions-people-actually-ask"&gt;Questions people actually ask&lt;/h2&gt;

&lt;p&gt;These answers cover CLI testing from zero state and the limits of unit tests.&lt;/p&gt;

&lt;h3&gt;What does Ranex gates Ranex mean?&lt;/h3&gt;

&lt;p&gt;Ranex provisions and runs its own committed suite sealed and offline, then evaluates the signed evidence against the gate.&lt;/p&gt;

&lt;h3&gt;What does end-to-end CLI testing catch that unit tests miss?&lt;/h3&gt;

&lt;p&gt;End-to-end CLI testing catches defects in real operator journeys across command setup, lock handling, documentation, cold-start state, and the materialised repository boundary.&lt;/p&gt;

&lt;h3&gt;Can an approved hash-correct dependency force a passing verdict?&lt;/h3&gt;

&lt;p&gt;Yes. The SLICE-006 security test demonstrates an approved, hash-correct wheel forcing a passing verdict. Approval reduces hidden change; it cannot make third-party code truthful.&lt;/p&gt;

&lt;h2 id="run-your-own-first-day-test"&gt;Run your own first-day test&lt;/h2&gt;

&lt;p&gt;Pick the most consequential command your users run. Clear the state it is meant to create. Open the docs. Execute only what they say. Watch every file it reads, every tool it resolves, every network call it attempts, and every trust root it can rewrite.&lt;/p&gt;

&lt;p&gt;Then preserve the journey. Put the defect reproduction beside the code, not only in a ticket. The source records for this work are in &lt;a href="https://github.com/anthonykewl20/ranex" rel="noopener noreferrer"&gt;the Ranex repository&lt;/a&gt;, under &lt;code&gt;docs/slices/done/&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Try it. Break it. Tell me what broke.&lt;/strong&gt; If the journey catches something your unit suite missed, star the repository and send an honest critique. That is the result worth keeping.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records — the same fact gate the product enforces on code. It ships only after Anthony’s own review.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>evidence</category>
      <category>testdesign</category>
    </item>
    <item>
      <title>Uptime Kuma: Why HTTP 200 Can Hide a Broken App</title>
      <dc:creator>Anthony Garces</dc:creator>
      <pubDate>Thu, 10 Sep 2026 10:14:02 +0000</pubDate>
      <link>https://dev.to/anthonygarces/uptime-kuma-why-http-200-can-hide-a-broken-app-2hkj</link>
      <guid>https://dev.to/anthonygarces/uptime-kuma-why-http-200-can-hide-a-broken-app-2hkj</guid>
      <description>&lt;p&gt;&lt;strong&gt;Uptime Kuma can show UP while your endpoint reports a failed dependency.&lt;/strong&gt; If the endpoint still answers HTTP 200, add a check for the specific response field that matters. A broad keyword such as &lt;code&gt;ok&lt;/code&gt; can miss the failure too.&lt;/p&gt;

&lt;p&gt;The example below uses four small test responses and records what 12 real Uptime Kuma monitors reported. Use it before trusting the green badge on a self-hosted app or a service you just shipped.&lt;/p&gt;

&lt;p&gt;&lt;small&gt;Source checks and local experiment: September 9, 2026. Tested with Uptime Kuma 2.5.3. The responses are synthetic; no production database was disconnected.&lt;/small&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The quick fix:&lt;/strong&gt; For the example response below, select &lt;strong&gt;HTTP(s) – JSON Query&lt;/strong&gt;, set the query to &lt;code&gt;checks.database&lt;/code&gt;, the condition to &lt;code&gt;==&lt;/code&gt;, and the expected value to &lt;code&gt;ok&lt;/code&gt;. Then deliberately return &lt;code&gt;down&lt;/code&gt; and confirm the monitor changes state. Test delivery to your notification channel separately.&lt;/p&gt;

&lt;p&gt;See the recorded results · Configure the monitor · Run the lab · Check the alert path · FAQ&lt;/p&gt;

&lt;h2 id="results"&gt;HTTP 200 and a broken dependency can coexist&lt;/h2&gt;

&lt;p&gt;An accepted HTTP status tells you that the response passed your status-code rule. It does not establish that the database, queue, or feature behind that endpoint works.&lt;/p&gt;

&lt;p&gt;Here is the deliberately broken response used in the lab. It arrives with HTTP 200:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;{
  "status": "ok",
  "checks": {
    "database": "down"
  }
}&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;A status-only check accepts it. A keyword check for &lt;code&gt;ok&lt;/code&gt; also accepts it because that text still appears in the top-level field. The database-specific query rejects it.&lt;/p&gt;

&lt;p&gt;These are the observed states from Uptime Kuma 2.5.3. Every fixture in this table returns HTTP 200. Each monitor used zero retries and accepted only status 200.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Response fixture&lt;/th&gt;
&lt;th&gt;HTTP status check&lt;/th&gt;
&lt;th&gt;Keyword: ok&lt;/th&gt;
&lt;th&gt;JSON: checks.database == ok&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Healthy database field&lt;/td&gt;
&lt;td&gt;UP&lt;/td&gt;
&lt;td&gt;UP&lt;/td&gt;
&lt;td&gt;UP&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Database field says down&lt;/td&gt;
&lt;td&gt;UP&lt;/td&gt;
&lt;td&gt;UP&lt;/td&gt;
&lt;td&gt;DOWN&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Database field is missing&lt;/td&gt;
&lt;td&gt;UP&lt;/td&gt;
&lt;td&gt;UP&lt;/td&gt;
&lt;td&gt;DOWN&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HTML sign-in page&lt;/td&gt;
&lt;td&gt;UP&lt;/td&gt;
&lt;td&gt;DOWN&lt;/td&gt;
&lt;td&gt;DOWN&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://cms.ranex.dev/wp-content/uploads/2026/09/uptime-kuma-http-200-lab.zip" rel="noopener noreferrer"&gt;Download the lab and recorded JSON results&lt;/a&gt; to inspect the exact response bodies, timestamps, configuration, and monitor messages. The HTML fixture directly returns a sign-in page; this experiment does not simulate a redirect or an expired session.&lt;/p&gt;

&lt;p&gt;Each monitor followed its configured rule. The broad rules simply allowed this broken response to pass.&lt;/p&gt;

&lt;h2 id="tool"&gt;Why use Uptime Kuma for this check?&lt;/h2&gt;

&lt;p&gt;Uptime Kuma is a self-hosted monitoring app whose documented monitor types include HTTP, keyword matching, and JSON queries. The pinned release uses the MIT license. You can inspect both the &lt;a href="https://github.com/louislam/uptime-kuma/blob/1f0755fb044fe08e99fccde6722062fb2bf6c8f4/README.md" rel="noopener noreferrer"&gt;project README&lt;/a&gt; and &lt;a href="https://github.com/louislam/uptime-kuma/blob/1f0755fb044fe08e99fccde6722062fb2bf6c8f4/LICENSE" rel="noopener noreferrer"&gt;license&lt;/a&gt; on GitHub.&lt;/p&gt;

&lt;p&gt;You can point all three monitor types at the same endpoint and compare the results. Start with your existing instance or a disposable local installation. This guide focuses on choosing the check; use the &lt;a href="https://github.com/louislam/uptime-kuma/wiki/%F0%9F%94%A7-How-to-Install" rel="noopener noreferrer"&gt;official installation guide&lt;/a&gt; for deployment instructions.&lt;/p&gt;

&lt;p&gt;If you are still deciding whether to operate another service, use our &lt;a href="/blog/oss-signal-assess-new-github-repo"&gt;six checks before adopting a repository&lt;/a&gt;. Hosting the monitor also means owning its availability and maintenance.&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fcms.ranex.dev%2Fwp-content%2Fuploads%2F2026%2F09%2Fuptime-kuma-http-200-explainer.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fcms.ranex.dev%2Fwp-content%2Fuploads%2F2026%2F09%2Fuptime-kuma-http-200-explainer.webp" alt="A server indicator, a magnifying glass inspecting a response, and an alert bell represent three separate checks." width="800" height="450"&gt;&lt;/a&gt;Conceptual illustration created with AI: a responding server, a useful response, and a delivered alert are separate things to verify. This is not an Uptime Kuma screenshot.&lt;p&gt;&lt;/p&gt;

&lt;h2 id="setup"&gt;Configure a JSON query for the field you need&lt;/h2&gt;

&lt;p&gt;For this test response, check whether &lt;code&gt;checks.database&lt;/code&gt; equals &lt;code&gt;ok&lt;/code&gt;. Set it through the JSON query monitor rather than searching the whole response for the word &lt;code&gt;ok&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Start the fixture from the download with &lt;code&gt;python3 fixture.py&lt;/code&gt;. On the Linux host-network setup described in its README, create this monitor:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Setting&lt;/th&gt;
&lt;th&gt;Lab value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Monitor type&lt;/td&gt;
&lt;td&gt;HTTP(s) – JSON Query&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;URL&lt;/td&gt;
&lt;td&gt;&lt;code&gt;http://127.0.0.1:8765/healthy&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Method / accepted status&lt;/td&gt;
&lt;td&gt;GET / 200&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;JSON query&lt;/td&gt;
&lt;td&gt;&lt;code&gt;checks.database&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Condition&lt;/td&gt;
&lt;td&gt;&lt;code&gt;==&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Expected value&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;ok&lt;/code&gt; (no quotation marks)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Interval / retry interval&lt;/td&gt;
&lt;td&gt;20 seconds / 20 seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Retries / timeout / maximum redirects&lt;/td&gt;
&lt;td&gt;0 / 5 seconds / 0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The &lt;a href="https://github.com/louislam/uptime-kuma/blob/1f0755fb044fe08e99fccde6722062fb2bf6c8f4/src/pages/EditMonitor.vue#L925-L980" rel="noopener noreferrer"&gt;versioned editor source&lt;/a&gt; shows the query, condition, and expected-value inputs. Labels can change across versions. These are demonstration settings, not universal production defaults.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Save the monitor and wait for a heartbeat against &lt;code&gt;/healthy&lt;/code&gt;. It should report UP.&lt;/li&gt;
&lt;li&gt;Change only the URL path to &lt;code&gt;/broken&lt;/code&gt;. Wait for a new heartbeat. The recorded result is DOWN.&lt;/li&gt;
&lt;li&gt;Repeat with &lt;code&gt;/missing&lt;/code&gt;, then &lt;code&gt;/login&lt;/code&gt;. Both were DOWN in the recorded run.&lt;/li&gt;
&lt;li&gt;Restore &lt;code&gt;/healthy&lt;/code&gt; and confirm recovery in your instance. Record the observation in the included worksheet.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If Kuma runs in an ordinary Docker bridge network, its &lt;code&gt;127.0.0.1&lt;/code&gt; is the container itself. Follow the download’s Linux host-network example or choose an address reachable from your Kuma deployment. A connection failure caused by the wrong address is a different test.&lt;/p&gt;

&lt;h2 id="keyword"&gt;When keyword monitoring helps, and when it misses&lt;/h2&gt;

&lt;p&gt;A keyword check helps when the response has a distinctive success marker. It is too broad when the same marker also appears in an error response.&lt;/p&gt;

&lt;p&gt;In the pinned implementation, keyword matching checks whether the response contains the configured text. The invert option reverses that condition. A JSON query evaluates the selected expression and comparison instead. See the &lt;a href="https://github.com/louislam/uptime-kuma/blob/1f0755fb044fe08e99fccde6722062fb2bf6c8f4/server/model/monitor.js#L683-L724" rel="noopener noreferrer"&gt;monitor implementation&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;For an HTML page, choose text tied to the expected result and test an error page that keeps the usual navigation and footer. For a structured API response, select the field whose value changes when the dependency fails. Avoid treating a generic page title or the word &lt;code&gt;ok&lt;/code&gt; anywhere in the body as proof of the whole app.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The field still needs to mean something.&lt;/strong&gt; A hard-coded &lt;code&gt;"database": "ok"&lt;/code&gt; is just another green badge. In your own service, define what the endpoint must actually observe before it returns that value, and make that observation fail in a controlled test.&lt;/p&gt;

&lt;h2 id="lab"&gt;Run the same failure lab yourself&lt;/h2&gt;

&lt;p&gt;The download contains the fixture server, a dependency-free Python check, a script that creates the real Kuma monitors, both recorded result files, and an acceptance worksheet.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://cms.ranex.dev/wp-content/uploads/2026/09/uptime-kuma-http-200-lab.zip" rel="noopener noreferrer"&gt;&lt;strong&gt;Download the Uptime Kuma HTTP 200 lab ZIP&lt;/strong&gt;&lt;/a&gt;. Extract it and open the &lt;code&gt;uptime-kuma-lab&lt;/code&gt; directory in a terminal:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;python3 check-fixture.py&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;That command exercises the HTTP responses with independent Python checks. To reproduce the actual Uptime Kuma results, follow the README’s disposable Docker setup and run &lt;code&gt;run-kuma.cjs&lt;/code&gt; inside that container. The download distinguishes the two runs explicitly.&lt;/p&gt;

&lt;p&gt;The Kuma run used &lt;a href="https://github.com/louislam/uptime-kuma/releases/tag/2.5.3" rel="noopener noreferrer"&gt;release 2.5.3&lt;/a&gt;, Node 22.22.3, SQLite storage, Linux host networking, and the pinned Docker image digest recorded in the README. It created 12 monitors and checked their first observed heartbeat states. It did not measure detection latency, exercise a real database, or send external notifications.&lt;/p&gt;

&lt;p&gt;Use the worksheet to record recovery and alert delivery in your deployment. Those cells are blank because this experiment did not test them.&lt;/p&gt;

&lt;h2 id="alerts"&gt;Verify the alert reaches someone&lt;/h2&gt;

&lt;p&gt;A DOWN state is only one step in the operational test. Attach your intended notification channel to the monitor and check that a controlled failure reaches the person who must respond.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Use a disposable endpoint or a planned test window. Record the healthy response and the expected failure before changing anything.&lt;/li&gt;
&lt;li&gt;Trigger the failure. Confirm that the monitor reaches DOWN after your configured retries, then confirm receipt in the attached channel.&lt;/li&gt;
&lt;li&gt;Restore the response. Confirm the monitor recovers and inspect the recovery notification if your channel is configured to send it.&lt;/li&gt;
&lt;li&gt;Record who received the alert and when. Separately decide how you will notice if the monitor’s own host becomes unavailable.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For production, choose intervals, retries, and timeouts around the service’s response time and the interruption your team can tolerate. This lab’s zero-retry setting makes the example easy to inspect; it is not a paging policy.&lt;/p&gt;

&lt;p&gt;Checking one response field does not prove that a user can sign in, save a document, or complete a purchase. Where that workflow matters, add a separate test that performs it with safe test data. Define the expected outcome before running the test. The &lt;a href="/painting-the-bullseye"&gt;painted bullseye failure mode&lt;/a&gt; has the same problem: the success rule is too easy to satisfy.&lt;/p&gt;

&lt;h2 id="faq"&gt;Questions about Uptime Kuma HTTP and JSON checks&lt;/h2&gt;

&lt;h3&gt;Why does Uptime Kuma show UP when my app is broken?&lt;/h3&gt;

&lt;p&gt;Your monitor may be checking an HTTP status that the broken response still satisfies. In this lab, every fixture returned 200. Add a condition tied to the response field or user outcome you need, then verify it with a controlled failure.&lt;/p&gt;

&lt;h3&gt;What JSON query should I use for a database health field?&lt;/h3&gt;

&lt;p&gt;For this guide’s response shape, use checks.database, condition ==, and expected value ok without quotation marks. Adapt the field and value to your actual response, and test healthy, failed, and missing-field cases.&lt;/p&gt;

&lt;h3&gt;Does a keyword check for ok prove that an API is healthy?&lt;/h3&gt;

&lt;p&gt;No. In the recorded lab, the broken response kept a top-level status of ok while its database field said down. The keyword monitor stayed UP. Use a more specific marker or a query for the relevant field.&lt;/p&gt;

&lt;h3&gt;Did this lab test notifications or a real database outage?&lt;/h3&gt;

&lt;p&gt;No. It tested 12 real Uptime Kuma monitors against four synthetic HTTP responses. External notification delivery, recovery, real dependency failures, and complete user workflows require separate checks.&lt;/p&gt;

&lt;p&gt;Start with one monitor you already rely on. Write down what its green state should prove, then give it a response that violates that promise. Keep the response and the observed result so the next person can repeat the check.&lt;/p&gt;

&lt;p&gt;&lt;small&gt;Disclosure: this article and its illustrations were prepared with AI assistance. The local experiment was executed during preparation; the exact scripts and outputs are in the download. Ranex publishes this guide for developers who want verifiable checks. No affiliation with the Uptime Kuma project is claimed.&lt;/small&gt;&lt;/p&gt;

</description>
      <category>tutorial</category>
    </item>
    <item>
      <title>Your SQLite backup says OK. Two rows are missing.</title>
      <dc:creator>Anthony Garces</dc:creator>
      <pubDate>Wed, 09 Sep 2026 10:14:56 +0000</pubDate>
      <link>https://dev.to/anthonygarces/your-sqlite-backup-says-ok-two-rows-are-missing-46p0</link>
      <guid>https://dev.to/anthonygarces/your-sqlite-backup-says-ok-two-rows-are-missing-46p0</guid>
      <description>&lt;p&gt;A SQLite backup can pass &lt;code&gt;PRAGMA integrity_check&lt;/code&gt; and still be missing records you committed.&lt;/p&gt;

&lt;p&gt;The attached lab produces exactly that result. The live database has three rows. A copy of its main file has one row, and SQLite reports &lt;code&gt;ok&lt;/code&gt; for both.&lt;/p&gt;

&lt;p&gt;If your backup job copies an &lt;code&gt;app.db&lt;/code&gt; file while the application is running, check which journal mode it uses. With write-ahead logging, committed changes can still be in the WAL file.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://cms.ranex.dev/wp-content/uploads/2026/09/sqlite-backup-lab.zip" rel="noopener noreferrer"&gt;Download the SQLite backup and restore kit (ZIP)&lt;/a&gt;. It includes a repeatable failure test, a backup utility, and the recorded results. Python 3.11 or newer with its standard &lt;code&gt;sqlite3&lt;/code&gt; module is enough. The lab runs offline with Python’s standard library.&lt;/p&gt;

&lt;h2&gt;The file was readable. The recent records were absent.&lt;/h2&gt;

&lt;p&gt;The fixture first wrote one row to the main database. It then enabled WAL mode, disabled automatic checkpointing for the test, and committed two more rows while keeping the connection open. No writes occurred during either copy.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Database checked&lt;/th&gt;
&lt;th&gt;Integrity result&lt;/th&gt;
&lt;th&gt;Rows found&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Live source&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ok&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Main-file-only copy&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ok&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Restored copy made with the backup utility&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ok&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;/p&gt;&lt;a href="https://cms.ranex.dev/wp-content/uploads/2026/09/daily-2026-09-08-sqlite-explainer.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcms.ranex.dev%2Fwp-content%2Fuploads%2F2026%2F09%2Fdaily-2026-09-08-sqlite-explainer.png" alt="SQLite WAL test: copying the main file restores one row; the backup API restores three. Both pass integrity_check." width="800" height="450"&gt;&lt;/a&gt;Illustration of the recorded SQLite test. The source kept all three rows; the main-file copy omitted the two committed through WAL. Select the image to view it full size.&lt;p&gt;&lt;/p&gt;

&lt;p&gt;The run used Python 3.14.7 and SQLite 3.53.1 on Linux on September 8, 2026. The kit’s &lt;code&gt;result.json&lt;/code&gt; includes the actual row IDs and contents, not just the counts.&lt;/p&gt;

&lt;p&gt;Nothing in this example required a damaged file. The copy was structurally sound and old. An integrity check cannot tell you that an application record should have been there.&lt;/p&gt;

&lt;h2&gt;Where the two rows went&lt;/h2&gt;

&lt;p&gt;SQLite’s &lt;a href="https://www.sqlite.org/wal.html" rel="noopener noreferrer"&gt;WAL documentation&lt;/a&gt; explains that writes can be committed to a separate log before a checkpoint moves them into the main database. The main file alone can therefore leave out committed work.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://www.sqlite.org/backup.html" rel="noopener noreferrer"&gt;online backup API&lt;/a&gt; reads through SQLite to produce a database snapshot. The included utility calls that API through Python’s &lt;a href="https://docs.python.org/3/library/sqlite3.html#sqlite3.Connection.backup" rel="noopener noreferrer"&gt;Connection.backup method&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;That is also why the test keeps the source connection open. Closing the last connection can change the WAL state. A demo that closes everything before copying can accidentally remove the condition it was supposed to test.&lt;/p&gt;

&lt;h2&gt;Run the failure test first&lt;/h2&gt;

&lt;p&gt;Extract the kit and read the two Python files. From that directory, run:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;python3 run-lab.py&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;The script creates disposable databases. It makes the incomplete file copy, runs the backup utility, restores that backup to another path, and opens the restored database through a separate connection. It compares the restored rows with the source records.&lt;/p&gt;

&lt;p&gt;It also checks that the utility refuses an existing destination, rejects a missing source, rejects an invalid database, and cleans up its temporary files. Your application’s database is not used by this lab.&lt;/p&gt;

&lt;h2&gt;Make a backup you can inspect&lt;/h2&gt;

&lt;p&gt;For a database you choose, the separate utility takes a source path and a new destination:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;mkdir -p backups
python3 sqlite-backup.py app.db backups/app-snapshot.db&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;The destination directory must exist. The script refuses to replace an existing file, so use a new name for the next backup.&lt;/p&gt;

&lt;p&gt;It opens the source through SQLite in read-only mode, writes to a temporary file, runs an integrity check, and publishes the destination after those steps succeed. The output contains its size and a SHA-256 checksum. That checksum can help you check a later transfer; it does not establish that every application record is present.&lt;/p&gt;

&lt;p&gt;The final publication step uses a same-directory hard link. Use a filesystem that supports that operation. If it does not, the script fails without an overwrite fallback. The default backup progress limit is 30 seconds; &lt;code&gt;--timeout 120&lt;/code&gt; allows a longer backup. The later integrity scan is outside that progress limit.&lt;/p&gt;

&lt;h2&gt;Check something your application cares about&lt;/h2&gt;

&lt;p&gt;After restoring to a separate location, find a recent record whose ID and content you know. Check a relationship or attachment that the application needs. For the tiny lab, the expected rows are fixed before either copy runs. In your system, decide what must survive before judging the restore.&lt;/p&gt;

&lt;p&gt;The utility copies the main SQLite database only. Separately attached databases, uploaded files, configuration, and keys need their own handling. If your application documents a backup procedure, follow it. A generic database copy cannot replace application-specific recovery steps.&lt;/p&gt;

&lt;p&gt;This kit does not provide retention, off-site storage, encryption, or a scheduled backup service. The test also does not simulate writers changing the database during backup, disk failure, or a power cut.&lt;/p&gt;

&lt;p&gt;Take one backup your current job has already produced. Restore it away from production and check a recent record. If the file opens but the record is missing, your backup check has found work worth doing.&lt;/p&gt;

&lt;p&gt;The &lt;a href="/blog/oss-signal-release-changes-2026-09-03"&gt;Gitea upgrade guide&lt;/a&gt; has a worksheet for recording recovery checks before a rollout. More runnable resources are collected at &lt;a href="/open-source"&gt;Open Source&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;small&gt;Research, code, and editing used AI assistance. This article reports the attached controlled test. The original kit code is available under the MIT license.&lt;/small&gt;&lt;/p&gt;

&lt;h2&gt;FAQ&lt;/h2&gt;

&lt;h3&gt;Can a SQLite backup pass integrity_check and still miss data?&lt;/h3&gt;

&lt;p&gt;Yes. In this controlled WAL test, a main-file-only copy passed integrity_check but contained one of three committed rows. Structural integrity did not establish that the copy included the recent writes.&lt;/p&gt;

&lt;h3&gt;What does the included backup utility copy?&lt;/h3&gt;

&lt;p&gt;It uses the SQLite backup API to copy the main database into a new file. It does not copy attached databases, application uploads, configuration, or encryption keys.&lt;/p&gt;

&lt;h3&gt;Will the backup script overwrite an existing backup?&lt;/h3&gt;

&lt;p&gt;No. It refuses an existing destination. The lab checks that a second run fails without changing the first backup.&lt;/p&gt;

&lt;h3&gt;Does the lab prove my application can recover?&lt;/h3&gt;

&lt;p&gt;No. It checks a small SQLite fixture. Restore your own backup separately, verify recent application records, and follow the application's supported recovery procedure.&lt;/p&gt;

</description>
      <category>backup</category>
      <category>opensource</category>
      <category>selfhosting</category>
      <category>sqlite</category>
    </item>
    <item>
      <title>No Self-Approval: The One Rule That Survives Perfect AI</title>
      <dc:creator>Anthony Garces</dc:creator>
      <pubDate>Wed, 09 Sep 2026 09:19:01 +0000</pubDate>
      <link>https://dev.to/anthonygarces/no-self-approval-the-one-rule-that-survives-perfect-ai-m4e</link>
      <guid>https://dev.to/anthonygarces/no-self-approval-the-one-rule-that-survives-perfect-ai-m4e</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; If the same actor produces the evidence and approves it, you do not have verification. You have agreement. &lt;a href="https://ranex.dev/blog/ai-accountability-apparatus" rel="noopener noreferrer"&gt;The accountability apparatus&lt;/a&gt; starts outside the worker.&lt;/p&gt;

&lt;p&gt;Your agent hands over a diff, a test it wrote, and a green summary. You are tired, the deadline is loud, and every line looks plausible. Who is left to say the proof counts?&lt;/p&gt;

&lt;h2 id="in-this-note"&gt;In this note&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Why production and approval must split&lt;/li&gt;
&lt;li&gt;How a test stops becoming a moving target&lt;/li&gt;
&lt;li&gt;A separation-of-duties checklist&lt;/li&gt;
&lt;li&gt;What the current limit is&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id="the-producer-cannot-approve-its-own-proof"&gt;The producer cannot approve its own proof&lt;/h2&gt;

&lt;p&gt;That is the rule. Whoever produced evidence cannot approve it. The README calls it one of the rules that makes a verdict trustworthy, and the reason is not subtle: one actor can write code, write the test around that code, and declare success. The target then follows the dart.&lt;/p&gt;

&lt;p&gt;Picture the merge you need to ship. An agent chooses the behavior, changes the implementation, changes the assertion, runs the command, and reports green. Nothing in that chain required a lie. The agent can be sincere. The result can still establish nothing beyond internal consistency.&lt;/p&gt;

&lt;p&gt;That is the third failure described in the README’s &lt;a href="https://github.com/anthonykewl20/ranex" rel="noopener noreferrer"&gt;“The problem” section&lt;/a&gt;: the person making the change gets to paint the bullseye afterward. A green check is only useful when the thing it measured was not arranged by the party being measured.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Competence does not dissolve a conflict of interest.&lt;/strong&gt; A stronger model changes the quality of the work. It does not turn self-approval into independent approval. Give the producer perfect recall, flawless code generation, and a complete map of your repository — the boundary still matters because the roles did not change.&lt;/p&gt;

&lt;p&gt;This is why the rule survives the fantasy of perfect AI. The question is not whether a producer is clever enough to check itself. The question is whether your release process asks it to be the witness, the judge, and the beneficiary of the same claim. Would you accept that shape from a person?&lt;/p&gt;

&lt;h2 id="freeze-the-target-before-the-work-starts"&gt;Freeze the target before the work starts&lt;/h2&gt;

&lt;p&gt;Separation cannot begin at the final approval button. It has to reach back to the test. Ranex describes tests frozen before BUILD, generated in COMPILE, digested, and read-only to implementers. A test file changed by the implementation diff fails the gate.&lt;/p&gt;

&lt;p&gt;That gives you a sequence worth copying even if you never run Ranex:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;State the behavior before the implementation begins.&lt;/li&gt;
&lt;li&gt;Turn that behavior into an executable test before the producer can shape it.&lt;/li&gt;
&lt;li&gt;Show that the test fails against the pre-implementation tree.&lt;/li&gt;
&lt;li&gt;Keep the implementer from authoring or judging that test.&lt;/li&gt;
&lt;li&gt;Let a separate approval step decide whether the resulting evidence counts.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The red-then-green step matters. A test that already passes before the behavior exists is not a target. It is an alibi. The no-self-approval rule matters beside it. A producer who can rewrite the test until it passes has moved the target even if the final command is real.&lt;/p&gt;

&lt;p&gt;The intended build loop draws the boundary again. A worker returns a diff. The kernel reads the diff on disk and runs checks; the worker summary is discarded. Workers do not merge. The README is explicit that models can propose, criticize, or translate, but cannot pass a gate.&lt;/p&gt;

&lt;p&gt;That full governed loop is designed, not built. Ranex is pre-release. What works today includes the narrower kernel rule: evaluation is a pure function of gate, evidence, subject, and approver, and no self-approval is listed as working behavior. Keep that distinction intact. A design is not evidence just because it is a good design.&lt;/p&gt;

&lt;h2 id="check-the-separation-before-you-trust-green"&gt;Check the separation before you trust green&lt;/h2&gt;

&lt;p&gt;You do not need to replace your stack to ask this question. Take a recent agent-assisted change and trace the roles, not the tools.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Who produced the evidence for the claim?&lt;/li&gt;
&lt;li&gt;Who approved that evidence?&lt;/li&gt;
&lt;li&gt;Could the implementer author or edit the test that judged the change?&lt;/li&gt;
&lt;li&gt;Did the test fail before the implementation existed?&lt;/li&gt;
&lt;li&gt;Did someone inspect the artifact on disk rather than accept the worker summary?&lt;/li&gt;
&lt;li&gt;Can the producer publish the result without a separate decision?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If one actor owns every answer, your process has a green light but no independent basis for it. That is not an accusation. It is a missing control. You can fix it with role separation, frozen tests, and a reviewer who has authority to refuse the claim.&lt;/p&gt;

&lt;p&gt;Do this with agent sessions too. “Different prompt” is not a different approver when the same session can revise the evidence until it likes the result. Name the producer. Name the approver. Preserve the boundary.&lt;/p&gt;

&lt;h2 id="the-current-limit-is-identity-not-the-rule"&gt;The current limit is identity, not the rule&lt;/h2&gt;

&lt;p&gt;Here is the part you should not skip. Ranex compares producer and approver values, but the approver identity is unauthenticated today. &lt;code&gt;--approver&lt;/code&gt; is a plain string. A producer can name someone else as the approver.&lt;/p&gt;

&lt;p&gt;Evidence signing establishes that the holder of a registered private key signed an evidence record. It does not establish who approved it. The no-self-approval comparison therefore works on unauthenticated strings. The README lists that as a known gap, plainly.&lt;/p&gt;

&lt;p&gt;So do not turn this post into a larger claim than the repository supports. The working rule blocks matching producer and approver values. It does not solve identity. The stamp stays a human’s, out of band, in the intended loop — and authenticating that human remains unfinished work.&lt;/p&gt;

&lt;p&gt;That stated limit is useful to you. It tells you where to inspect your own process next: a separation rule without trustworthy identities has a hole at the label. Keep the rule. Do not pretend the label closes the hole.&lt;/p&gt;

&lt;p&gt;The separate approver protects a bounded claim, not a promise that everything is right. &lt;a href="https://ranex.dev/blog/what-a-gate-verdict-asserts" rel="noopener noreferrer"&gt;A gate verdict has its own narrow meaning&lt;/a&gt; — and that is why the approval boundary matters.&lt;/p&gt;

&lt;h2 id="questions-people-actually-ask"&gt;Questions people actually ask&lt;/h2&gt;

&lt;h3&gt;What is the no-self-approval rule?&lt;/h3&gt;

&lt;p&gt;Ranex no-self-approval means whoever produced evidence cannot approve it.&lt;/p&gt;

&lt;h3&gt;Why does no self-approval still matter with better AI?&lt;/h3&gt;

&lt;p&gt;Separation of duties still separates production from approval when the producer is more capable.&lt;/p&gt;

&lt;h3&gt;Is approver identity authenticated in Ranex today?&lt;/h3&gt;

&lt;p&gt;Ranex does not authenticate approver identity today because &lt;code&gt;--approver&lt;/code&gt; is a plain string.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Try it. Break it. Tell me what broke.&lt;/strong&gt; Read the &lt;a href="https://github.com/anthonykewl20/ranex" rel="noopener noreferrer"&gt;MIT-licensed repository&lt;/a&gt;, then trace one change through your own approval path.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records — the same fact gate the product enforces on code. It ships only after Anthony’s own review.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiaccountability</category>
      <category>separationofduties</category>
    </item>
    <item>
      <title>uv locked vs frozen: the missing dependency test</title>
      <dc:creator>Anthony Garces</dc:creator>
      <pubDate>Tue, 08 Sep 2026 12:58:32 +0000</pubDate>
      <link>https://dev.to/anthonygarces/uv-locked-vs-frozen-the-missing-dependency-test-478c</link>
      <guid>https://dev.to/anthonygarces/uv-locked-vs-frozen-the-missing-dependency-test-478c</guid>
      <description>&lt;p&gt;You add a dependency to &lt;code&gt;pyproject.toml&lt;/code&gt;. CI runs &lt;code&gt;uv sync --frozen&lt;/code&gt;, exits successfully, and the package is still missing.&lt;/p&gt;

&lt;p&gt;There is a small, reproducible reason for that. &lt;code&gt;--frozen&lt;/code&gt; uses the lockfile you already have. It skips the check that would tell you the manifest has changed.&lt;/p&gt;

&lt;p&gt;If your CI requirement is “the checked-in lockfile must match this project,” use &lt;code&gt;uv sync --locked&lt;/code&gt;. In the test below, that command caught the mismatch.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://cms.ranex.dev/wp-content/uploads/2026/09/uv-lock-lab.zip" rel="noopener noreferrer"&gt;Download the uv lockfile lab (ZIP)&lt;/a&gt;. It includes the script, recorded output, and an MIT license. You need Python 3.11 or newer, uv, and access to PyPI. You do not need Ranex.&lt;/p&gt;

&lt;h2&gt;The failure, in five commands&lt;/h2&gt;

&lt;p&gt;The lab started with a project that had no dependencies and generated its lockfile. It then added &lt;code&gt;idna==3.10&lt;/code&gt; to the manifest without updating that lockfile.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Command&lt;/th&gt;
&lt;th&gt;Exit code&lt;/th&gt;
&lt;th&gt;What the lab observed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;uv lock&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;Created the original lockfile before the manifest edit.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;uv sync --locked&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Rejected the stale lockfile. Its bytes did not change.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;uv sync --frozen&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;Accepted the old lockfile. The new dependency was absent.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;uv sync&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;Updated the lockfile and installed idna 3.10.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;uv lock --check&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;Accepted the refreshed lockfile.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The run used uv 0.11.26 and Python 3.14.7 on Linux on September 8, 2026. These are the versions tested. The download keeps the exact command arguments, exit codes, and lockfile hashes in &lt;code&gt;result.json&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;Run it without touching your repository&lt;/h2&gt;

&lt;p&gt;Extract the kit, read &lt;code&gt;run-lab.py&lt;/code&gt;, and run it from the extracted directory:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;python3 run-lab.py&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;The script makes its own temporary project and cache. It uses your current Python interpreter, downloads idna from PyPI into a temporary virtual environment, and removes the lab files when it exits. It does not inherit your private index settings or &lt;code&gt;UV_*&lt;/code&gt; environment overrides.&lt;/p&gt;

&lt;p&gt;Look for these two fields in the output:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;"idna_present_after_frozen": false,
"idna_version_after_sync": "3.10"&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;The script checks the installed package separately from uv’s exit status. That is how it catches the successful sync with a missing dependency.&lt;/p&gt;

&lt;p&gt;There is one detail worth keeping if you adapt the test: inspect the virtual environment’s Python directly. An ordinary &lt;code&gt;uv run&lt;/code&gt; can update the lockfile and environment before running your check, which would change the condition you meant to inspect.&lt;/p&gt;

&lt;h2&gt;Why both flags exist&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://docs.astral.sh/uv/concepts/projects/sync/" rel="noopener noreferrer"&gt;official uv documentation&lt;/a&gt; distinguishes checking a lockfile from using one. &lt;code&gt;--locked&lt;/code&gt; raises an error when the lockfile needs an update. &lt;code&gt;--frozen&lt;/code&gt; skips that freshness check. Neither command promises to do the other’s job.&lt;/p&gt;

&lt;p&gt;A frozen install can be intentional when a workflow is meant to consume an already approved lockfile. The trouble starts when you treat its success as evidence that the current manifest and lockfile agree. That was never the check you asked it to run.&lt;/p&gt;

&lt;h2&gt;Fix the mismatch before changing the CI flag&lt;/h2&gt;

&lt;p&gt;For a project where the dependency edit is intentional, update the lockfile locally and review the change:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;uv lock
git diff -- pyproject.toml uv.lock
uv sync --locked&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Commit the manifest and lockfile together. Once CI has checked out that commit and installed your chosen Python and uv versions, make &lt;code&gt;uv sync --locked&lt;/code&gt; a step that must succeed before the application tests run.&lt;/p&gt;

&lt;p&gt;If you only need to validate the lockfile, &lt;code&gt;uv lock --check&lt;/code&gt; is the smaller command. It does not replace running the application tests or checking what ends up in a built container.&lt;/p&gt;

&lt;p&gt;Do not replace a failing locked sync with a frozen sync just to clear the failure. First check whether the dependency edit was intended and whether the corresponding lockfile change was committed.&lt;/p&gt;

&lt;h2&gt;What this test covers&lt;/h2&gt;

&lt;p&gt;This is one project with one newly declared dependency. It does not test every uv release, workspaces, optional groups, private indexes, or your build image. The timings in the raw output are incidental; this is not a speed comparison.&lt;/p&gt;

&lt;p&gt;The useful check to take back to your repository is simple: add a dependency without refreshing the lockfile on a throwaway branch. Does your required CI job reject the mismatch? Check that before assuming a green install tested it.&lt;/p&gt;

&lt;p&gt;For a broader dependency decision, use &lt;a href="/blog/oss-signal-assess-new-github-repo"&gt;the repository evaluation guide&lt;/a&gt;. More practical resources live at &lt;a href="/open-source"&gt;Open Source&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;small&gt;Research, code, and editing used AI assistance. Results come from the attached local run. The kit is free to use under its MIT license.&lt;/small&gt;&lt;/p&gt;

&lt;h2&gt;FAQ&lt;/h2&gt;

&lt;h3&gt;What is the difference between uv locked and frozen?&lt;/h3&gt;

&lt;p&gt;In this lab, &lt;code&gt;--locked&lt;/code&gt; rejected the stale lockfile. &lt;code&gt;--frozen&lt;/code&gt; used the existing lockfile without checking whether it matched the edited project manifest.&lt;/p&gt;

&lt;h3&gt;How do I fix a stale uv lockfile?&lt;/h3&gt;

&lt;p&gt;Review the dependency change, run uv lock, inspect the diff, and commit pyproject.toml and uv.lock together. Then rerun uv sync &lt;code&gt;--locked&lt;/code&gt; and your application tests.&lt;/p&gt;

&lt;h3&gt;Does the lab change my project?&lt;/h3&gt;

&lt;p&gt;No. It creates a temporary project, virtual environment, and cache, then removes them on exit. It requires Python 3.11 or newer, uv, and network access to PyPI.&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>python</category>
      <category>troubleshooting</category>
      <category>uv</category>
    </item>
  </channel>
</rss>
