<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: NAIF Gravity</title>
    <description>The latest articles on DEV Community by NAIF Gravity (@naifgravity).</description>
    <link>https://dev.to/naifgravity</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4167573%2Fdb63b42a-73e1-464f-8a4b-4532c655f954.png</url>
      <title>DEV Community: NAIF Gravity</title>
      <link>https://dev.to/naifgravity</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/naifgravity"/>
    <language>en</language>
    <item>
      <title>Code review challenge: which AI-generated edit changes behavior?</title>
      <dc:creator>NAIF Gravity</dc:creator>
      <pubDate>Wed, 07 Oct 2026 05:55:10 +0000</pubDate>
      <link>https://dev.to/naifgravity/code-review-challenge-which-ai-generated-edit-changes-behavior-m6i</link>
      <guid>https://dev.to/naifgravity/code-review-challenge-which-ai-generated-edit-changes-behavior-m6i</guid>
      <description>&lt;p&gt;Can you review these three proposed edits before looking at the answers?&lt;/p&gt;

&lt;p&gt;Assume a &lt;strong&gt;preserve-behavior&lt;/strong&gt; policy. The input domain is integers from &lt;strong&gt;-64 through 64, plus None&lt;/strong&gt;. Evaluate each candidate independently against this original function:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Candidate A
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Candidate B
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Candidate C
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;operator&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;operator&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;mul&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Your review
&lt;/h2&gt;

&lt;p&gt;For each candidate, choose:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ALLOW candidate:&lt;/strong&gt; appears to preserve behavior within the stated domain, subject to the actual evaluator admitting and checking it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;QUARANTINE:&lt;/strong&gt; has a behavior difference that needs review.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;UNSUPPORTED:&lt;/strong&gt; exceeds the documented evaluator scope, even if ordinary Python reasoning suggests equivalence.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Reply in the comments with &lt;code&gt;A: ..., B: ..., C: ...&lt;/code&gt; and one sentence of reasoning. For a changed result, include a counterexample input. No signup or purchase is needed to solve the exercise.&lt;/p&gt;

&lt;h2&gt;
  
  
  Answers — read after deciding
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A: behavior-preserving; an ALLOW candidate, not an issued AMF decision.&lt;/strong&gt; For integers, adding a value to itself gives the same result as multiplying it by two. Both versions return zero for None.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;B: QUARANTINE under preserve-behavior.&lt;/strong&gt; At &lt;code&gt;value = 0&lt;/code&gt;, the original returns &lt;code&gt;0&lt;/code&gt;; the proposal returns &lt;code&gt;2&lt;/code&gt;. It happens to match at &lt;code&gt;value = 2&lt;/code&gt;, so a single passing example would miss the regression. Intentional changes still need a different acceptance policy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;C: UNSUPPORTED in the documented NAIF AMF Pilot Kit subset.&lt;/strong&gt; Imports, attributes and arbitrary calls are excluded. Ordinary Python equivalence does not make an out-of-scope program an admissible ALLOW.&lt;/p&gt;

&lt;h2&gt;
  
  
  What was actually verified?
&lt;/h2&gt;

&lt;p&gt;We checked A and B with a simple local Python comparison across all &lt;strong&gt;130 domain inputs&lt;/strong&gt;: None plus 129 integers. A had zero differences; B had 128. This is a small language-level check, &lt;strong&gt;not an AMF execution, Mutation Passport, hosted PR Check or product benchmark&lt;/strong&gt;. C was not run; its scope classification follows the public documentation.&lt;/p&gt;

&lt;p&gt;A real review record should also identify the policy, admitted scope, analyzed base and proposed revision. No observed difference in a bounded domain is not a general software-safety guarantee.&lt;/p&gt;

&lt;p&gt;The exercise illustrates the separation between behavioral analysis and scope admission behind &lt;a href="https://dev.to/naifgravity/naif-agent-mutation-firewall-allow-quarantine-and-unsupported-explained-39en"&gt;NAIF Agent Mutation Firewall&lt;/a&gt;. The &lt;a href="https://github.com/naief9961-tech/naif-amf-pilot-sandbox" rel="noopener noreferrer"&gt;public Pilot Kit documentation&lt;/a&gt; describes the actual integration limits.&lt;/p&gt;

&lt;h2&gt;
  
  
  Institutional enquiries
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Enterprise Access — Coming Soon.&lt;/strong&gt; Teams interested in discussing a bounded review workflow can email &lt;strong&gt;&lt;a href="mailto:contact@naifgravity.com"&gt;contact@naifgravity.com&lt;/a&gt;&lt;/strong&gt; with the subject &lt;strong&gt;NAIF AMF — Code Review Challenge&lt;/strong&gt;. Share your use case and a synthetic or redacted example, without credentials or customer data. Availability will be announced separately.&lt;/p&gt;

&lt;p&gt;Disclosure: Published by NAIF Gravity. AI authored this exercise and ran the small Python comparison described above.&lt;/p&gt;

</description>
      <category>python</category>
    </item>
    <item>
      <title>NAIF AMF Pilot Kit: binding AI-code review evidence to the exact GitHub PR</title>
      <dc:creator>NAIF Gravity</dc:creator>
      <pubDate>Wed, 07 Oct 2026 05:43:48 +0000</pubDate>
      <link>https://dev.to/naifgravity/naif-amf-pilot-kit-binding-ai-code-review-evidence-to-the-exact-github-pr-2kge</link>
      <guid>https://dev.to/naifgravity/naif-amf-pilot-kit-binding-ai-code-review-evidence-to-the-exact-github-pr-2kge</guid>
      <description>&lt;p&gt;A pull-request workflow can finish successfully while the proposed change still needs review. That distinction matters when AI coding agents produce frequent edits: collecting evidence successfully is not the same as allowing the edit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;NAIF AMF Pilot Kit v1&lt;/strong&gt;, the third component of NAIF Enterprise Assurance, is a documented read-only/check-only adapter around frozen AMF v0.4 and the engine hashes referenced by RDR V2.5.7. Its role is to connect a bounded review decision to a GitHub PR and package inspectable evidence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Enterprise Access — Coming Soon.&lt;/strong&gt; Institutional enquiries are welcome for requirements and scope discussion. This is a technical explanation, not a production-readiness or certification announcement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Follow the decision Check, not just the job
&lt;/h2&gt;

&lt;p&gt;The public integration documents a Check named &lt;strong&gt;NAIF AMF / Decision&lt;/strong&gt; on the exact PR head SHA:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Review outcome&lt;/th&gt;
&lt;th&gt;Decision Check&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ALLOW&lt;/td&gt;
&lt;td&gt;Successful&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;QUARANTINE&lt;/td&gt;
&lt;td&gt;Failure&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;UNSUPPORTED&lt;/td&gt;
&lt;td&gt;Failure&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The orchestration job can be green after a rejected change because it successfully published a failing Decision Check. Reviewers should inspect the named Check and the revision it covers.&lt;/p&gt;

&lt;p&gt;For example, if an agent pushes revision B after revision A was checked, evidence for A is not evidence for B. The receipt also binds the analyzed base. Base-only updates do not automatically revalidate every open PR; base divergence requires an update/rebase and a fresh run.&lt;/p&gt;

&lt;p&gt;The kit does not automatically merge PRs or change branch protection. Any future owner decision to require the Check is separate from this article.&lt;/p&gt;

&lt;h2&gt;
  
  
  Collect proposed code without running the PR project
&lt;/h2&gt;

&lt;p&gt;The documented workflow uses code from the trusted base commit and fetches PR contents as inert Git blobs. It checks blob identity, file count and regular-file mode before applying a restrictive scalar AST grammar.&lt;/p&gt;

&lt;p&gt;It does not check out, install, build or import the proposed PR project. Admitted functions are evaluated by the frozen interpreters. CPU, memory and time limits add containment, but do not turn the adapter into a general Python sandbox.&lt;/p&gt;

&lt;p&gt;Collection and publishing use the built-in ephemeral GitHub token with contents/read, pull-requests/read and checks/write permissions. The token is not passed to the analysis child. Fork approval and source-access policies still apply; denied collection cannot produce ALLOW.&lt;/p&gt;

&lt;h2&gt;
  
  
  Know the admitted subset before evaluating fit
&lt;/h2&gt;

&lt;p&gt;The documented default supports up to eight modified Python files, one unannotated pure positional-argument function in each, and integers [-64,64] plus None. Imports, loops, classes, arbitrary calls, strings/floats, new/deleted/renamed files and arbitrary multi-language projects are unsupported.&lt;/p&gt;

&lt;p&gt;ALLOW means no observed behavior change within that scope and domain. An intentional behavior-changing repair may be quarantined under preserve-behavior. Trusted policy changes require review on the base; a PR cannot silently configure its own approval policy.&lt;/p&gt;

&lt;h2&gt;
  
  
  What an evidence package contains
&lt;/h2&gt;

&lt;p&gt;The documented package includes raw engine receipts, wrapper passports, hashes, logs, timings, an offline HTML dashboard and a minimal-in-domain counterexample when found.&lt;/p&gt;

&lt;p&gt;Raw UUID and time fields vary between runs; semantic hashes are intended to reproduce. Receipts are unsigned: hashes provide integrity linkage, not issuer authentication.&lt;/p&gt;

&lt;p&gt;The public repository describes 40 synthetic scenarios. Those are distinct from actual hosted PR executions. Readiness requires the preregistered criteria, hosted artifact verification and live PR Check testing; the README alone does not establish readiness. No new run was performed for this article.&lt;/p&gt;

&lt;p&gt;A useful institutional evaluation would start with a supported change class, explicit acceptance criteria and evidence tied to the exact revisions being reviewed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Institutional enquiries
&lt;/h2&gt;

&lt;p&gt;Email &lt;strong&gt;&lt;a href="mailto:contact@naifgravity.com"&gt;contact@naifgravity.com&lt;/a&gt;&lt;/strong&gt; with the subject &lt;strong&gt;NAIF AMF Pilot Kit — Institutional Enquiry&lt;/strong&gt;. Describe your GitHub review workflow, code languages, desired evidence and approval constraints. Use synthetic or redacted examples; omit tokens and customer data.&lt;/p&gt;

&lt;p&gt;Availability will be announced separately after scope review.&lt;/p&gt;

&lt;p&gt;Technical documentation: &lt;a href="https://github.com/naief9961-tech/naif-amf-pilot-sandbox" rel="noopener noreferrer"&gt;NAIF AMF Pilot Kit repository&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Related: &lt;a href="https://dev.to/naifgravity/naif-agent-mutation-firewall-allow-quarantine-and-unsupported-explained-39en"&gt;AMF decision semantics&lt;/a&gt; · &lt;a href="https://dev.to/naifgravity/reviewing-ai-generated-code-changes-with-bounded-evidence-naif-enterprise-assurance-3k3c"&gt;Enterprise overview&lt;/a&gt; · &lt;a href="https://naifgravity.com" rel="noopener noreferrer"&gt;NAIF Gravity&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Disclosure: Published by NAIF Gravity. AI drafted this explanation from the public documentation; no new benchmark or production validation was performed for this article.&lt;/p&gt;

</description>
      <category>devops</category>
    </item>
    <item>
      <title>NAIF Agent Mutation Firewall: ALLOW, QUARANTINE and UNSUPPORTED explained</title>
      <dc:creator>NAIF Gravity</dc:creator>
      <pubDate>Wed, 07 Oct 2026 05:41:49 +0000</pubDate>
      <link>https://dev.to/naifgravity/naif-agent-mutation-firewall-allow-quarantine-and-unsupported-explained-39en</link>
      <guid>https://dev.to/naifgravity/naif-agent-mutation-firewall-allow-quarantine-and-unsupported-explained-39en</guid>
      <description>&lt;p&gt;An AI coding agent proposes a small refactor. The tests pass. Before approving it, a reviewer still needs to know which behavior was compared, which inputs were covered, and what happened when the change exceeded the evaluator's scope.&lt;/p&gt;

&lt;p&gt;This is the review problem behind &lt;strong&gt;NAIF Agent Mutation Firewall (AMF)&lt;/strong&gt;, the second component of NAIF Enterprise Assurance. RDR provides the analysis foundation; AMF provides the mutation review gate; the Pilot Kit packages the GitHub integration and evidence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Enterprise Access — Coming Soon.&lt;/strong&gt; This article explains the documented approach for institutional enquiries; it does not announce general production availability.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three outcomes, with different meanings
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Decision&lt;/th&gt;
&lt;th&gt;Review meaning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ALLOW&lt;/td&gt;
&lt;td&gt;No observed behavior change within the admitted scope and configured input domain.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;QUARANTINE&lt;/td&gt;
&lt;td&gt;Review the change and any reported counterexample before proceeding. An intentional behavior-changing repair can also be quarantined under a preserve-behavior policy.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;UNSUPPORTED&lt;/td&gt;
&lt;td&gt;The evaluator cannot assess the change within its supported scope. Route it to another review method.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;An unsupported result must never become ALLOW just because the agent produced plausible code or the surrounding workflow finished successfully.&lt;/p&gt;

&lt;h2&gt;
  
  
  A boundary change worth investigating
&lt;/h2&gt;

&lt;p&gt;Consider this illustrative edit:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Before
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;classify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;

&lt;span class="c1"&gt;# Proposed change
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;classify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For an input of zero, the two versions take different branches. A preserve-behavior review should investigate that difference rather than treating it as a cosmetic refactor. Whether the new behavior is desirable is a separate product decision.&lt;/p&gt;

&lt;p&gt;This example explains the review question; it is not a new execution result or benchmark claim.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the documented integration admits
&lt;/h2&gt;

&lt;p&gt;The public Pilot Kit uses frozen AMF v0.4 and the engine hashes referenced by RDR V2.5.7. Its default admitted scope is narrow: up to eight modified Python files, one unannotated pure positional-argument function per file, and integers from -64 to 64 plus None.&lt;/p&gt;

&lt;p&gt;Imports, attributes, loops, classes, arbitrary calls, strings, floats and general multi-language projects fall outside that scope. Both engine observations must agree before ALLOW in the documented adapter. Syntax-derived invariant candidates are not proven invariants.&lt;/p&gt;

&lt;p&gt;This makes the result assessable: a reviewer can see the domain and exclusions instead of reading an unqualified “safe” label.&lt;/p&gt;

&lt;h2&gt;
  
  
  Evidence to keep with the decision
&lt;/h2&gt;

&lt;p&gt;The documented artifacts include raw engine receipts, wrapper passports, hashes, logs, timings and a minimal-in-domain counterexample when one is found. The receipt must stay associated with the analyzed base and proposed revision.&lt;/p&gt;

&lt;p&gt;Receipts are unsigned. SHA-256 links support integrity checking; they do not authenticate the issuer. The gate does not certify arbitrary software, automatically merge a PR, or replace review of authentication, network behavior and unsupported code.&lt;/p&gt;

&lt;p&gt;A practical institutional evaluation starts by choosing a small supported change class and defining the behavior that must remain unchanged.&lt;/p&gt;

&lt;h2&gt;
  
  
  Institutional enquiries
&lt;/h2&gt;

&lt;p&gt;Interested teams can email &lt;strong&gt;&lt;a href="mailto:contact@naifgravity.com"&gt;contact@naifgravity.com&lt;/a&gt;&lt;/strong&gt; with the subject &lt;strong&gt;NAIF Agent Mutation Firewall — Institutional Enquiry&lt;/strong&gt;. Include your review use case, languages, current approval workflow and a synthetic or redacted example. Do not include credentials or customer data.&lt;/p&gt;

&lt;p&gt;Enquiries are for requirements and scope discussion. Availability will be announced separately.&lt;/p&gt;

&lt;p&gt;Technical reference: &lt;a href="https://github.com/naief9961-tech/naif-amf-pilot-sandbox" rel="noopener noreferrer"&gt;NAIF AMF Pilot Kit documentation&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://naifgravity.com" rel="noopener noreferrer"&gt;NAIF Gravity&lt;/a&gt; · &lt;a href="https://dev.to/naifgravity/reviewing-ai-generated-code-changes-with-bounded-evidence-naif-enterprise-assurance-3k3c"&gt;Enterprise overview&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Disclosure: Published by NAIF Gravity. AI drafted this explanation from the public pilot documentation; no new benchmark or production validation was performed for this article.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
    </item>
    <item>
      <title>Reviewing AI-generated code changes with bounded evidence: NAIF Enterprise Assurance</title>
      <dc:creator>NAIF Gravity</dc:creator>
      <pubDate>Wed, 07 Oct 2026 05:31:05 +0000</pubDate>
      <link>https://dev.to/naifgravity/reviewing-ai-generated-code-changes-with-bounded-evidence-naif-enterprise-assurance-3k3c</link>
      <guid>https://dev.to/naifgravity/reviewing-ai-generated-code-changes-with-bounded-evidence-naif-enterprise-assurance-3k3c</guid>
      <description>&lt;p&gt;A green test suite does not always explain whether an AI-generated code change preserves the behavior a team cares about. For an enterprise reviewer, the useful question is narrower: what was evaluated, under which policy, and what evidence supports the decision?&lt;/p&gt;

&lt;p&gt;NAIF Enterprise Assurance is the NAIF Gravity initiative for that review workflow. It brings together RDR, NAIF Agent Mutation Firewall (AMF), and the AMF Pilot Kit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Enterprise Access — Coming Soon.&lt;/strong&gt; We welcome institutional enquiries to discuss requirements and scope. This post is a technical overview, not a claim of general production readiness or enterprise certification.&lt;/p&gt;

&lt;h2&gt;
  
  
  From a code change to reviewable evidence
&lt;/h2&gt;

&lt;p&gt;The documented pilot combines an analysis foundation referenced by RDR with the AMF mutation review gate. The Pilot Kit collects an admitted code change, runs the frozen engines under a defined policy, and packages observations and receipts for review.&lt;/p&gt;

&lt;p&gt;The decision states matter:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ALLOW:&lt;/strong&gt; no behavior change was observed within the admitted scope and evaluated domain. This is not proof that arbitrary software is safe.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;QUARANTINE:&lt;/strong&gt; the change requires review under the policy; an intentional behavior-changing repair may also be quarantined by a preserve-behavior policy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;UNSUPPORTED:&lt;/strong&gt; the change lies outside the supported analysis scope. It must not silently become an approval.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The GitHub pilot creates a Decision Check for the exact PR head SHA. A successfully completed orchestration job is not the same as an ALLOW decision. The owner must inspect the Decision Check; the kit does not automatically merge code or change branch protection.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the scope must be explicit
&lt;/h2&gt;

&lt;p&gt;Consider a small pure Python function:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Before
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;calculate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;fee&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;amount&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;fee&lt;/span&gt;

&lt;span class="c1"&gt;# Proposed change
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;calculate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;fee&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;amount&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;fee&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A useful evaluation asks whether an admitted input exposes the behavior difference, records the relevant witness if found, and binds the result to the exact code and policy. This example illustrates the review question; it is not a published benchmark result.&lt;/p&gt;

&lt;p&gt;The current documented kit has a deliberately narrow scope: up to eight modified Python files, one unannotated pure positional-argument function per file, and integers from -64 to 64 plus None by default. Imports, loops, classes, arbitrary calls, strings/floats and general multi-language projects are outside that scope.&lt;/p&gt;

&lt;p&gt;It does not train a model, execute the submitted PR as a project, write client code, deploy a repair, or automatically merge changes. Resource limits add containment but do not make it a general Python sandbox.&lt;/p&gt;

&lt;h2&gt;
  
  
  What evidence should reviewers expect?
&lt;/h2&gt;

&lt;p&gt;The documented artifact package includes engine receipts, wrapper passports, hashes, logs, timings, an in-domain counterexample when found, and an offline review dashboard. Receipts are unsigned: hash linkage supports integrity checks, not issuer authentication.&lt;/p&gt;

&lt;p&gt;A meaningful institutional discussion starts by agreeing on the supported code subset, policy, evaluation domain and acceptance criteria. Results outside that boundary should remain unsupported or inconclusive rather than being promoted into an assurance claim.&lt;/p&gt;

&lt;p&gt;Public technical reference: &lt;a href="https://github.com/naief9961-tech/naif-amf-pilot-sandbox" rel="noopener noreferrer"&gt;NAIF AMF Pilot Kit documentation&lt;/a&gt;. Readiness must be established from verified evidence, not inferred from an overview or a successful job badge.&lt;/p&gt;

&lt;h2&gt;
  
  
  Institutional enquiries
&lt;/h2&gt;

&lt;p&gt;If your organization is interested in NAIF Enterprise Assurance, email &lt;strong&gt;&lt;a href="mailto:contact@naifgravity.com"&gt;contact@naifgravity.com&lt;/a&gt;&lt;/strong&gt; with the subject &lt;strong&gt;NAIF Enterprise Assurance — Institutional Enquiry&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Include your team’s use case, languages, PR workflow, the behavior you need to preserve, and a small synthetic or redacted example. Please do not send tokens, private keys, production credentials or customer data. We can discuss fit, the evaluation boundary and availability by email; access is subject to scope review and a separate availability announcement.&lt;/p&gt;

&lt;p&gt;Project: &lt;a href="https://naifgravity.com" rel="noopener noreferrer"&gt;NAIF Gravity&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Disclosure: this overview is published by NAIF Gravity. It was drafted by an autonomous AI assistant against the public pilot documentation. No new benchmark execution or production validation was performed for this post.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>python</category>
    </item>
    <item>
      <title>HTTP 200 is not enough: checking MCP discovery without executing tools</title>
      <dc:creator>NAIF Gravity</dc:creator>
      <pubDate>Wed, 07 Oct 2026 01:32:58 +0000</pubDate>
      <link>https://dev.to/naifgravity/http-200-is-not-enough-checking-mcp-discovery-without-executing-tools-2c94</link>
      <guid>https://dev.to/naifgravity/http-200-is-not-enough-checking-mcp-discovery-without-executing-tools-2c94</guid>
      <description>&lt;p&gt;An MCP endpoint can return HTTP 200 while the integration still fails. A successful HTTP request does not establish that the response matches the JSON-RPC request, that initialization negotiated a supported revision, or that subsequent requests carry the session headers.&lt;/p&gt;

&lt;p&gt;NAIF Gravity MCP Diagnostics is a small Python standard-library probe for that narrower question: can a client complete supported Streamable HTTP discovery?&lt;/p&gt;

&lt;p&gt;Repository: &lt;a href="https://github.com/naief9961-tech/naif-gravity-mcp-diagnostics" rel="noopener noreferrer"&gt;https://github.com/naief9961-tech/naif-gravity-mcp-diagnostics&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The sequence matters
&lt;/h2&gt;

&lt;p&gt;For the handshake-era revisions supported by this probe, discovery follows this sequence:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Send &lt;code&gt;initialize&lt;/code&gt; and validate the response ID, result shape and negotiated revision.&lt;/li&gt;
&lt;li&gt;Forward the negotiated &lt;code&gt;MCP-Protocol-Version&lt;/code&gt; and any returned &lt;code&gt;Mcp-Session-Id&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Send &lt;code&gt;notifications/initialized&lt;/code&gt; without a request ID.&lt;/li&gt;
&lt;li&gt;If the server advertises tools, request &lt;code&gt;tools/list&lt;/code&gt; and follow bounded pagination.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A JSON-RPC error inside an HTTP 200 response is still a failure. A mismatched response ID is also a failure: it cannot establish that the reply belongs to this request.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it on an endpoint you may test
&lt;/h2&gt;

&lt;p&gt;Requires Python 3.10 or later; no third-party packages.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/naief9961-tech/naif-gravity-mcp-diagnostics.git
&lt;span class="nb"&gt;cd &lt;/span&gt;naif-gravity-mcp-diagnostics
python3 tools/mcp_health_check.py https://your-mcp.example/mcp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Replace the example endpoint with one you own or are authorized to test. For bearer authentication, provision a scoped test token securely in your environment and pass only its variable name:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 tools/mcp_health_check.py https://your-mcp.example/mcp &lt;span class="nt"&gt;--token-env&lt;/span&gt; MCP_TEST_TOKEN
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This does not perform OAuth login or refresh. Bearer authentication requires HTTPS except on loopback, and redirects are refused.&lt;/p&gt;

&lt;h2&gt;
  
  
  JSON and SSE are both valid response shapes
&lt;/h2&gt;

&lt;p&gt;A POST response may contain JSON or an SSE stream. The probe reads SSE comments and multiline data, ignores unrelated notifications, and stops when it receives the matching response. It does not require the stream to close first.&lt;/p&gt;

&lt;p&gt;Output includes HTTP statuses, protocol revision, whether a session exists, and tool count. It omits response bodies, tool names, session values and bearer tokens. Failures return a nonzero exit status.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproduce behavior without production access
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 &lt;span class="nt"&gt;-m&lt;/span&gt; unittest discover &lt;span class="nt"&gt;-s&lt;/span&gt; tests &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="s1"&gt;'test_*.py'&lt;/span&gt; &lt;span class="nt"&gt;-v&lt;/span&gt;
python3 examples/webhook_signature_fixture.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The current checkout passes 28 test methods covering loopback discovery, JSON reports, the fault lab, issue-draft field selection and the offline regression guard. Discovery cases cover sessions, authentication, JSON/SSE, pagination, redirect refusal and malformed responses. The offline webhook fixture shows a separate integration pitfall: raw-body HMAC verification can fail after JSON is parsed and reserialized. Its key and payload are fictional.&lt;/p&gt;

&lt;h2&gt;
  
  
  A seven-case local fault lab
&lt;/h2&gt;

&lt;p&gt;The repository now includes a loopback-only lab for healthy discovery, missing authentication, mismatched response IDs, invalid JSON, unsupported revisions, pagination cursor loops and invalid tool schemas. Each case explains a next check.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 examples/fault_lab.py
python3 examples/fault_lab.py &lt;span class="nt"&gt;--case&lt;/span&gt; wrong-id &lt;span class="nt"&gt;--report&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; probe-report.json
&lt;span class="c"&gt;# Exit 1 is expected for the intentional fault. Continue with:&lt;/span&gt;
python3 tools/issue_report.py probe-report.json &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; issue-draft.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The lab itself exits 0 when the expected outcomes are reproduced. Single-case &lt;code&gt;--report&lt;/code&gt; preserves the probe exit code. The issue draft generator copies only selected diagnostic fields and never submits an issue; add a synthetic reproduction and environment details, then review it before sharing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Detect discovery regressions after an update
&lt;/h2&gt;

&lt;p&gt;Capture a healthy &lt;code&gt;before.json&lt;/code&gt; and a current &lt;code&gt;after.json&lt;/code&gt; using the same authorized endpoint, requested revision, probe commit and authorization scope:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 tools/mcp_health_check.py https://your-authorized-mcp.example/mcp &lt;span class="nt"&gt;--json&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; before.json
&lt;span class="c"&gt;# Repeat after your authorized update, writing after.json.&lt;/span&gt;
python3 tools/integration_guard.py before.json after.json &lt;span class="nt"&gt;--json&lt;/span&gt;
python3 examples/integration_guard_demo.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The offline guard exits 1 for failed discovery, lost tool capability or decreased tool count. Healthy changes are reported for review; &lt;code&gt;--strict-changes&lt;/code&gt; blocks those too. Invalid or incomparable input exits 2. It sends no network requests and performs no repair or deployment.&lt;/p&gt;

&lt;p&gt;This catches bounded discovery changes, not every integration regression: reports omit tool names and schemas, so equal-count tool substitutions are invisible. Endpoint identity is also omitted; keeping both captures comparable is the operator’s responsibility.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/naief9961-tech/naif-gravity-mcp-diagnostics/blob/main/docs/INTEGRATION-GUARD.md" rel="noopener noreferrer"&gt;Guard policy and CI usage&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Know what PASS means
&lt;/h2&gt;

&lt;p&gt;Supported revisions are 2025-03-26, 2025-06-18 and 2025-11-25. The probe does not implement newer stateless lifecycles, stdio, legacy separate SSE transport, SSE reconnection or server-initiated RPC requests.&lt;/p&gt;

&lt;p&gt;Requests are bounded to 1 MiB each and tools discovery to 10 pages. The socket timeout is an inactivity timeout, not a strict overall runtime budget.&lt;/p&gt;

&lt;p&gt;PASS means supported discovery completed. It does not certify conformance, production health or successful execution. The probe never calls a server tool.&lt;/p&gt;

&lt;p&gt;The repository includes a &lt;a href="https://github.com/naief9961-tech/naif-gravity-mcp-diagnostics/blob/main/docs/DIAGNOSTIC-QUICKSTART.md" rel="noopener noreferrer"&gt;quick start with limits&lt;/a&gt;. Synthetic bug reproductions and compatibility corrections are welcome.&lt;/p&gt;

&lt;p&gt;Disclosure: this utility belongs to the NAIF Gravity project. This article was drafted by an autonomous AI assistant, which checked the claims against the implementation and ran the local tests.&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>python</category>
      <category>showdev</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
