Can you review these three proposed edits before looking at the answers?
Assume a preserve-behavior policy. The input domain is integers from -64 through 64, plus None. Evaluate each candidate independently against this original function:
def score(value):
if value is None:
return 0
return value * 2
Candidate A
def score(value):
if value is None:
return 0
return value + value
Candidate B
def score(value):
if value is None:
return 0
return value + 2
Candidate C
import operator
def score(value):
if value is None:
return 0
return operator.mul(value, 2)
Your review
For each candidate, choose:
- ALLOW candidate: appears to preserve behavior within the stated domain, subject to the actual evaluator admitting and checking it.
- QUARANTINE: has a behavior difference that needs review.
- UNSUPPORTED: exceeds the documented evaluator scope, even if ordinary Python reasoning suggests equivalence.
Reply in the comments with A: ..., B: ..., C: ... and one sentence of reasoning. For a changed result, include a counterexample input. No signup or purchase is needed to solve the exercise.
Answers — read after deciding
A: behavior-preserving; an ALLOW candidate, not an issued AMF decision. For integers, adding a value to itself gives the same result as multiplying it by two. Both versions return zero for None.
B: QUARANTINE under preserve-behavior. At value = 0, the original returns 0; the proposal returns 2. It happens to match at value = 2, so a single passing example would miss the regression. Intentional changes still need a different acceptance policy.
C: UNSUPPORTED in the documented NAIF AMF Pilot Kit subset. Imports, attributes and arbitrary calls are excluded. Ordinary Python equivalence does not make an out-of-scope program an admissible ALLOW.
What was actually verified?
We checked A and B with a simple local Python comparison across all 130 domain inputs: None plus 129 integers. A had zero differences; B had 128. This is a small language-level check, not an AMF execution, Mutation Passport, hosted PR Check or product benchmark. C was not run; its scope classification follows the public documentation.
A real review record should also identify the policy, admitted scope, analyzed base and proposed revision. No observed difference in a bounded domain is not a general software-safety guarantee.
The exercise illustrates the separation between behavioral analysis and scope admission behind NAIF Agent Mutation Firewall. The public Pilot Kit documentation describes the actual integration limits.
Institutional enquiries
Enterprise Access — Coming Soon. Teams interested in discussing a bounded review workflow can email contact@naifgravity.com with the subject NAIF AMF — Code Review Challenge. Share your use case and a synthetic or redacted example, without credentials or customer data. Availability will be announced separately.
Disclosure: Published by NAIF Gravity. AI authored this exercise and ran the small Python comparison described above.
Top comments (0)