A green test suite does not always explain whether an AI-generated code change preserves the behavior a team cares about. For an enterprise reviewer, the useful question is narrower: what was evaluated, under which policy, and what evidence supports the decision?
NAIF Enterprise Assurance is the NAIF Gravity initiative for that review workflow. It brings together RDR, NAIF Agent Mutation Firewall (AMF), and the AMF Pilot Kit.
Enterprise Access — Coming Soon. We welcome institutional enquiries to discuss requirements and scope. This post is a technical overview, not a claim of general production readiness or enterprise certification.
From a code change to reviewable evidence
The documented pilot combines an analysis foundation referenced by RDR with the AMF mutation review gate. The Pilot Kit collects an admitted code change, runs the frozen engines under a defined policy, and packages observations and receipts for review.
The decision states matter:
- ALLOW: no behavior change was observed within the admitted scope and evaluated domain. This is not proof that arbitrary software is safe.
- QUARANTINE: the change requires review under the policy; an intentional behavior-changing repair may also be quarantined by a preserve-behavior policy.
- UNSUPPORTED: the change lies outside the supported analysis scope. It must not silently become an approval.
The GitHub pilot creates a Decision Check for the exact PR head SHA. A successfully completed orchestration job is not the same as an ALLOW decision. The owner must inspect the Decision Check; the kit does not automatically merge code or change branch protection.
Why the scope must be explicit
Consider a small pure Python function:
# Before
def calculate(amount, fee):
return amount + fee
# Proposed change
def calculate(amount, fee):
return amount - fee
A useful evaluation asks whether an admitted input exposes the behavior difference, records the relevant witness if found, and binds the result to the exact code and policy. This example illustrates the review question; it is not a published benchmark result.
The current documented kit has a deliberately narrow scope: up to eight modified Python files, one unannotated pure positional-argument function per file, and integers from -64 to 64 plus None by default. Imports, loops, classes, arbitrary calls, strings/floats and general multi-language projects are outside that scope.
It does not train a model, execute the submitted PR as a project, write client code, deploy a repair, or automatically merge changes. Resource limits add containment but do not make it a general Python sandbox.
What evidence should reviewers expect?
The documented artifact package includes engine receipts, wrapper passports, hashes, logs, timings, an in-domain counterexample when found, and an offline review dashboard. Receipts are unsigned: hash linkage supports integrity checks, not issuer authentication.
A meaningful institutional discussion starts by agreeing on the supported code subset, policy, evaluation domain and acceptance criteria. Results outside that boundary should remain unsupported or inconclusive rather than being promoted into an assurance claim.
Public technical reference: NAIF AMF Pilot Kit documentation. Readiness must be established from verified evidence, not inferred from an overview or a successful job badge.
Institutional enquiries
If your organization is interested in NAIF Enterprise Assurance, email contact@naifgravity.com with the subject NAIF Enterprise Assurance — Institutional Enquiry.
Include your team’s use case, languages, PR workflow, the behavior you need to preserve, and a small synthetic or redacted example. Please do not send tokens, private keys, production credentials or customer data. We can discuss fit, the evaluation boundary and availability by email; access is subject to scope review and a separate availability announcement.
Project: NAIF Gravity.
Disclosure: this overview is published by NAIF Gravity. It was drafted by an autonomous AI assistant against the public pilot documentation. No new benchmark execution or production validation was performed for this post.
Top comments (0)