<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: 藍凰</title>
    <description>The latest articles on DEV Community by 藍凰 (@empenguin).</description>
    <link>https://dev.to/empenguin</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4154998%2Fb017fddf-c5aa-4c67-b33d-b70fd9f64df6.png</url>
      <title>DEV Community: 藍凰</title>
      <link>https://dev.to/empenguin</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/empenguin"/>
    <language>en</language>
    <item>
      <title>Should Probabilistic AI Judges Be Treated as Sensors Rather Than Verifiers?</title>
      <dc:creator>藍凰</dc:creator>
      <pubDate>Fri, 02 Oct 2026 10:57:09 +0000</pubDate>
      <link>https://dev.to/empenguin/should-probabilistic-ai-judges-be-treated-as-sensors-rather-than-verifiers-18m6</link>
      <guid>https://dev.to/empenguin/should-probabilistic-ai-judges-be-treated-as-sensors-rather-than-verifiers-18m6</guid>
      <description>&lt;p&gt;I’ve been thinking about a class of AI systems that do not primarily generate text.&lt;/p&gt;

&lt;p&gt;Instead, they take some state or evidence and return bounded judgments such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is_risky: 0.91

category:
  safe: 0.03
  suspicious: 0.79
  unknown: 0.18

review_priority: 2.6 / 3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Recent systems such as Jev make this pattern particularly visible: instead of asking a language model to generate an explanation or JSON object and then parsing it back into a decision, the model directly produces typed probabilistic judgments.&lt;/p&gt;

&lt;p&gt;That seems extremely useful for agent runtimes.&lt;/p&gt;

&lt;p&gt;But it raises a question I’m not sure how these systems should be placed inside the control plane.&lt;/p&gt;

&lt;h2&gt;
  
  
  A tempting architecture
&lt;/h2&gt;

&lt;p&gt;Suppose an agent wants to use a tool.&lt;/p&gt;

&lt;p&gt;A probabilistic model inspects the request and returns:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;dangerous_operation = 0.94
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The obvious implementation is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;if dangerous_operation &amp;gt; 0.8:
    reject()
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is simple, fast, and probably useful in many applications.&lt;/p&gt;

&lt;p&gt;But I’m increasingly uncomfortable with calling the probabilistic model a &lt;strong&gt;verifier&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The model did not prove that the operation was unsafe.&lt;/p&gt;

&lt;p&gt;It produced an observation about the operation.&lt;/p&gt;

&lt;p&gt;That seems closer to a sensor.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sensor versus verifier
&lt;/h2&gt;

&lt;p&gt;The architecture I am experimenting with separates four roles:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Artifact / Runtime State
          ↓
   Finding Producer
          ↓
 Finding + confidence
          ↓
      Policy
          ↓
      Decision
          ↓
      Verifier
          ↓
      Enforcer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A probabilistic AI model would live at the first stage:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AI judge
   ↓
FindingProducer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Finding:
  type: SEMANTIC_RISK
  probability: 0.94
  model_version: X
  evidence_ref: Y
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A separate policy layer could then decide what that finding means:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;if SEMANTIC_RISK &amp;gt;= 0.90:
    require_human_review
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The verifier would check whether the policy was applied consistently, whether the correct policy version was used, and whether required preconditions were satisfied.&lt;/p&gt;

&lt;p&gt;The enforcer would finally allow, block, quarantine, or escalate the action.&lt;/p&gt;

&lt;p&gt;So:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;probabilistic judgment
        ≠
policy decision
        ≠
verification
        ≠
enforcement
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Why bother separating them?
&lt;/h2&gt;

&lt;p&gt;Imagine the model changes.&lt;/p&gt;

&lt;p&gt;Version A says:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;risk = 0.92
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Version B evaluates the exact same evidence and says:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;risk = 0.41
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Was the original decision wrong?&lt;/p&gt;

&lt;p&gt;Maybe.&lt;/p&gt;

&lt;p&gt;But another interpretation is that the historical decision was valid under:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;model A
+
policy version 3
+
evidence snapshot E
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;while a new evaluation would produce a different result under a new model.&lt;/p&gt;

&lt;p&gt;That suggests model outputs might need provenance similar to other evidence:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;model identity
model version
question / rubric version
input evidence version
timestamp / epoch
confidence
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Otherwise, replaying or auditing an agent decision later becomes ambiguous.&lt;/p&gt;

&lt;h2&gt;
  
  
  Another problem: calibrated does not mean correct
&lt;/h2&gt;

&lt;p&gt;Even if a probabilistic model is well calibrated, a result such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;0.93
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;does not mean:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;this statement is objectively 93% true
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It is still a model judgment.&lt;/p&gt;

&lt;p&gt;And even a perfectly type-safe system can confidently choose the wrong member of the allowed output space.&lt;/p&gt;

&lt;p&gt;So I currently prefer this interpretation:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A probabilistic AI judge is a semantic observation mechanism, not an authority source.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Its result can influence a decision, but it should not silently become the decision itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  But perhaps this separation is too strict
&lt;/h2&gt;

&lt;p&gt;There are obvious counterexamples.&lt;/p&gt;

&lt;p&gt;Spam filters routinely make automated decisions.&lt;/p&gt;

&lt;p&gt;Fraud systems block transactions.&lt;/p&gt;

&lt;p&gt;Content moderation models directly suppress material.&lt;/p&gt;

&lt;p&gt;Anomaly detectors can trigger circuit breakers.&lt;/p&gt;

&lt;p&gt;At some point, a probabilistic judgment clearly does become operational authority.&lt;/p&gt;

&lt;p&gt;Maybe the right distinction is not:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AI model may never decide
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;but instead:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;model produces judgment

policy explicitly grants that judgment
a bounded amount of authority
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;confidence &amp;gt;= 0.99
AND
effect is reversible
AND
blast radius is low
→ automatic action allowed

otherwise
→ review / escalation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In that architecture, the model still does not define its own authority.&lt;/p&gt;

&lt;p&gt;The runtime does.&lt;/p&gt;

&lt;h2&gt;
  
  
  The question
&lt;/h2&gt;

&lt;p&gt;So I’m curious how others model this boundary.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Should probabilistic AI judges be treated primarily as sensors / finding producers rather than verifiers?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;More specifically:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Should model output ever directly authorize an irreversible action?&lt;/li&gt;
&lt;li&gt;Should model version and rubric version be part of the evidence provenance?&lt;/li&gt;
&lt;li&gt;If a newer model disagrees with an older model, should historical decisions be re-evaluated?&lt;/li&gt;
&lt;li&gt;Where should confidence thresholds live: inside the model interface, inside policy, or inside the application?&lt;/li&gt;
&lt;li&gt;Is there established terminology for separating probabilistic semantic judgment from deterministic policy enforcement?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I’m especially interested in related patterns from:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;runtime assurance&lt;/li&gt;
&lt;li&gt;policy engines&lt;/li&gt;
&lt;li&gt;autonomous systems&lt;/li&gt;
&lt;li&gt;fraud / risk systems&lt;/li&gt;
&lt;li&gt;capability security&lt;/li&gt;
&lt;li&gt;agent runtimes&lt;/li&gt;
&lt;li&gt;formal methods&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The architecture that led me to this question is an experimental agent-runtime project called HANDS, but I’m deliberately trying to phrase the problem independently of that implementation.&lt;/p&gt;

&lt;p&gt;I would be particularly interested in counterexamples where treating the model as “just a sensor” is the wrong abstraction.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Disclosure: I used ChatGPT to help structure and edit this post. The underlying architecture question and design are from my own ongoing project, and I reviewed the content before publishing.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>discuss</category>
      <category>ai</category>
      <category>architecture</category>
      <category>security</category>
    </item>
    <item>
      <title>Can a Self-Improving Agent Runtime Safely Modify Its Own Verifier?</title>
      <dc:creator>藍凰</dc:creator>
      <pubDate>Thu, 01 Oct 2026 14:00:50 +0000</pubDate>
      <link>https://dev.to/empenguin/gptcan-a-self-improving-agent-runtime-safely-modify-its-own-verifier-5ih</link>
      <guid>https://dev.to/empenguin/gptcan-a-self-improving-agent-runtime-safely-modify-its-own-verifier-5ih</guid>
      <description>&lt;p&gt;I’m experimenting with an agent-runtime architecture where execution results can become evidence for changing the runtime itself.&lt;/p&gt;

&lt;p&gt;The simplified loop looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;execution
    ↓
evidence
    ↓
modification candidate
    ↓
verification
    ↓
activation
    ↓
new runtime state
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This creates a problem I’m trying to reason about.&lt;/p&gt;

&lt;p&gt;Suppose runtime epoch &lt;code&gt;e&lt;/code&gt; uses verifier &lt;code&gt;V_e&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;An agent observes its execution history and proposes a modification &lt;code&gt;C&lt;/code&gt;. That modification includes an update to the verifier itself, producing &lt;code&gt;V_(e+1)&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;My current rule is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A modification must not change the verifier, authorization policy, or evidence-acceptance rules used to authorize that same modification.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Accept(C) = V_e(C, Evidence_e)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;may authorize:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;V_e → V_(e+1)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;but this should not be allowed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Accept(C) = V_(e+1)(C, Evidence_e)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;because the candidate would effectively participate in defining the rules by which it is accepted.&lt;/p&gt;

&lt;p&gt;I’ve been calling this a &lt;strong&gt;no same-epoch self-authorization&lt;/strong&gt; rule.&lt;/p&gt;

&lt;p&gt;But I’m not convinced that this is sufficient.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What if the candidate leaves the verifier unchanged but changes the evidence-selection mechanism?&lt;/li&gt;
&lt;li&gt;Is validation of &lt;code&gt;V_(e+1)&lt;/code&gt; by &lt;code&gt;V_e&lt;/code&gt; enough, or does that merely create a chain of inherited trust?&lt;/li&gt;
&lt;li&gt;If the verifier itself contains a bug, what mechanism should be allowed to replace it?&lt;/li&gt;
&lt;li&gt;Does safe self-modification ultimately require a small non-self-modifying root of trust?&lt;/li&gt;
&lt;li&gt;Or can verifier evolution itself be safely modeled as a layered adaptation process?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I’m especially interested in existing work or implementation experience from:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;self-adaptive systems&lt;/li&gt;
&lt;li&gt;runtime assurance&lt;/li&gt;
&lt;li&gt;capability security&lt;/li&gt;
&lt;li&gt;formal methods&lt;/li&gt;
&lt;li&gt;reflective systems&lt;/li&gt;
&lt;li&gt;agent runtimes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I’m not claiming this is a new problem. Quite the opposite: I suspect there are already better terms and established models for parts of it.&lt;/p&gt;

&lt;p&gt;If you know of a relevant concept, paper, system, or counterexample to the rule above, I’d like to know what I should be comparing this design against.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Disclosure: I used ChatGPT to help structure and edit this post. The underlying architecture question and design are from my own ongoing project, and I reviewed the content before publishing.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>discuss</category>
      <category>architecture</category>
      <category>security</category>
    </item>
  </channel>
</rss>
