DEV Community

Cover image for The Prompt Changed. The Acceptance Evidence Didn't.
Alex Agafonov
Alex Agafonov

Posted on

The Prompt Changed. The Acceptance Evidence Didn't.

Imagine a form that a company uses to collect requests. AI helps suggest the right department from the request text, while an employee reviews ambiguous cases.

The team tests the scenario. A valid request reaches the right department. An invalid email address is rejected, including when the request goes directly to the API. An ambiguous request stays in the manual review queue.

Before rollout, someone improves the prompt. Another team member changes the model configuration. A third decides that confident answers can proceed without employee review.

The folder still contains a report marked "tested."

That report now describes a different system, even if its name and interface have not changed.

I would bind acceptance evidence to a specific state of the tool. Here is a small evidence record and a local check that can reveal when that record has become stale before rollout.

Define exactly what you are releasing

For the first rollout, I would limit this example to preparing a recommendation. The system receives a request and suggests a department. An employee sees the original text and makes the decision. Sending messages to customers and modifying customer records are outside this stage.

That boundary gives acceptance testing a clear scope: one part of the process and its handoff to a person.

People define the criteria. The business owner is responsible for the desired outcome. The scenario owner defines the routing rules. The implementation team is responsible for the selected tools and constraints. A specialist makes decisions in the cases reserved for human review.

Each criterion needs an observable result. An invalid email should cause the API to reject the request. An ambiguous message should enter the manual review queue without automatic forwarding.

Test the path that bypasses the interface

Testing a form in a browser is useful, but client-side validation does not cover a direct API request.

An illustrative negative case could look like this:

{
  "id": "invalid-email-direct-api",
  "request": {
    "method": "POST",
    "path": "/requests",
    "body": {
      "email": "not-an-email",
      "message": "Please route my request"
    }
  },
  "expected": {
    "status": 400,
    "requestCreated": false,
    "routingStarted": false
  }
}
Enter fullscreen mode Exit fullscreen mode

Here, 400 is the chosen contract of the hypothetical API. In your own system, specify the actual contract and check the side effects. An error status alone does not prove that the request was never stored or forwarded.

The set also needs a valid request, an ambiguous message, and a failed model call. Define both the acceptable result and the prohibited actions for each case before running it.

Model behavior may require repeated runs and a separate analysis of failures. One successful response does not establish that the scenario is reliable. The team should define the acceptance rules and acceptable error rates before testing.

Bind the report to the tested inputs

Suppose four files describe the behavior of this illustrative tool:

  • prompt.txt: the instruction for classifying a request.
  • model.json: the selected model and call parameters, without secrets.
  • actions.json: allowed actions and the conditions for manual review.
  • cases.json: test cases and expected results.

Their contents can be bound to one SHA-256 fingerprint. Changing any of these files makes the previous evidence record stale.

A real product needs a broader inventory. Include the code revision, sources, schemas, and other dependencies that affect behavior. These four files keep the example small. Anything outside the list is invisible to the script.

An evidence record next to the report could look like this:

{
  "fingerprint": "<SHA-256 of the tested configuration>",
  "tests": "passed",
  "report": "reports/form-routing-001.json"
}
Enter fullscreen mode Exit fullscreen mode

First capture the files' state, then test that state. Keep it unchanged during the run and record the fingerprint with the result. If the files change, run the tests again. Do not rewrite the fingerprint in an old report to make it match the current configuration.

A small freshness check

This script runs locally. It reads four files from the current directory and compares their fingerprint with the record in evidence.json. It uses Node.js's hashing API and synchronous file reads.

import { createHash } from "node:crypto";
import { readFileSync } from "node:fs";
import { resolve } from "node:path";
import { pathToFileURL } from "node:url";

export const paths = [
  "prompt.txt", "model.json", "actions.json", "cases.json"
];

export function fingerprint(read) {
  const entries = paths.map(path => [
    path, createHash("sha256").update(read(path)).digest("hex")
  ]);
  return createHash("sha256")
    .update(JSON.stringify(entries)).digest("hex");
}

export function checkEvidence(record, current) {
  if (record?.fingerprint !== current) return "stale-or-missing";
  if (record.tests !== "passed") return "tests-not-passed";
  if (typeof record.report !== "string" || !record.report.trim()) {
    return "report-missing";
  }
  return "evidence-current";
}

if (process.argv[1] && import.meta.url ===
    pathToFileURL(resolve(process.argv[1])).href) {
  try {
    const current = fingerprint(path => readFileSync(path));
    if (process.argv[2] === "--fingerprint") {
      console.log(current);
    } else {
      const record = JSON.parse(readFileSync("evidence.json", "utf8"));
      const result = checkEvidence(record, current);
      console.log(result);
      if (result !== "evidence-current") process.exitCode = 1;
    }
  } catch (error) {
    console.error(`evidence-unavailable: ${error.message}`);
    process.exitCode = 1;
  }
}
Enter fullscreen mode Exit fullscreen mode

Save the code as check-evidence.mjs. In the directory containing the four input files, run:

node check-evidence.mjs --fingerprint
Enter fullscreen mode Exit fullscreen mode

This prints the current fingerprint. After testing that exact state, save the evidence record in evidence.json. Then run:

node check-evidence.mjs
Enter fullscreen mode Exit fullscreen mode

The command compares the record with the current files. It exits with a nonzero status when the inputs have changed, the evidence is missing, tests are not marked as passed, the report reference is missing, or a file cannot be read.

An evidence-current result means only that the fingerprints match, the test field says passed, and the report reference is nonempty. The script does not open the report, verify its authenticity, or run the tests. In a working process, a trusted testing process should create the record, and the report must be available to the reviewer.

Unchanged local files also do not guarantee unchanged behavior from an external model. Provider-side versions and changes require separate controls.

The rollout decision is still a separate step

Before a limited rollout, a specialist opens the report, checks the results, and approves that specific scenario and version. Preserve the decision together with the person, rollout scope, and tested fingerprint.

If the system was accepted only in recommendation mode, that decision cannot authorize automatic sending. The action boundary has changed, so it needs a separate test and decision.

A specialist's approval does not remove the team's responsibility for correct rules and constraints. Acceptance cannot repair the absence of protection against a prohibited action.

I would start with a small part of the process that can be tested and stopped. Expand automation only after defining the next scope of acceptance.

We began with a folder marked "tested." We should finish with a specific connection: what was tested, which inputs were used, what happened, and who authorized the next step.

Then a changed prompt can no longer quietly inherit the approval given to another version.

References

Top comments (0)