DEV Community

Cover image for Demonstrable, Not Documented: What Regulators Are Actually Asking For, and What It Takes to Answer
Jb
Jb

Posted on

Demonstrable, Not Documented: What Regulators Are Actually Asking For, and What It Takes to Answer

For years, AI governance has been treated primarily as a documentation problem.

Write the policy.

Record the model owner.

Complete the risk assessment.

Document human oversight.

Keep the logs.

Show the auditor the dashboard.

All of those things matter.

But across regulation, standards, accountability regimes and litigation, a different expectation is becoming visible:

It is no longer enough to document what was supposed to happen. Organisations increasingly need to demonstrate what actually happened.

Those are not the same thing.

A policy can show that human oversight was required.

It cannot, by itself, show what a human actually saw during one specific decision, whether they intervened, what the AI produced before that intervention, or whether the record being reviewed six months later is the same record that existed when the decision was made.

A governance dashboard can show that a control exists.

It cannot necessarily demonstrate that the control operated during the execution now being challenged.

An internal log can show what a system recorded.

It does not automatically give somebody outside the organisation a reason to trust that the record has remained unchanged.

This distinction between documented compliance and demonstrable compliance is becoming increasingly important.

What makes it particularly interesting is that the pressure is coming from several directions at once.

Different regimes, similar direction

There is no global regulation saying that every AI execution must produce a cryptographically verifiable record.

That would be an overstatement.

The EU AI Act does not prescribe a particular cryptographic evidence architecture.

NIST does not require Certified Execution Records.

ISO/IEC 42001 does not tell organisations which hashing algorithm to use.

The UK's Senior Managers and Certification Regime does not require banks to cryptographically certify AI decisions.

US law does not impose one common AI evidence protocol.

These frameworks were created for different purposes.

But when you look at the practical questions they create for organisations, a pattern appears.

The language is different.

The legal mechanisms are different.

The sectors are different.

The underlying question keeps moving in the same direction:

Can you demonstrate what happened?

The EU AI Act: record the system, then make oversight real

Article 12 of the EU AI Act requires high-risk AI systems to technically allow for the automatic recording of events over the lifetime of the system.

The regulation explicitly connects those logging capabilities with traceability, post-market monitoring and monitoring the operation of high-risk systems.

That already moves the conversation beyond simply knowing which model an organisation deployed.

There must be a record of relevant events during operation.

Article 14 approaches the problem from another direction.

High-risk AI systems must be designed so that natural persons can effectively oversee them.

Depending on the context, those people must be able to understand relevant capabilities and limitations, monitor operation, detect anomalies, interpret outputs, avoid over-reliance on AI, disregard or override outputs, and intervene where necessary.

Article 14 does not say every human intervention needs a cryptographic certificate.

Article 12 does not prescribe one particular evidence format.

But put the two together and a practical question appears:

If you say meaningful human oversight existed, what evidence shows what the human was actually overseeing?

What did the system produce?

What information was available at that point?

Was the recommendation accepted?

Was it overridden?

What system or model version was involved?

What happened immediately before and after the decision?

And is the record being presented today the same record that existed then?

That is where the difference between documenting a control and demonstrating its operation becomes important.

The application timetable has also changed.

Following Regulation (EU) 2026/1744, the amended calendar now sets the relevant high-risk requirements for systems classified under Article 6(2) and Annex III from 2 December 2027, while the corresponding requirements for high-risk systems embedded in regulated products under Article 6(1) apply from 2 August 2028.

That gives organisations more time.

It does not remove the evidence problem.

NIST: traceability has to survive beyond the engineering team

The NIST AI Risk Management Framework is voluntary rather than law.

But the AI RMF Playbook points in a remarkably similar direction.

NIST asks organisations to consider mechanisms that facilitate AI system auditability, including traceability of the development process, data sourcing, and logging system processes and outcomes.

It also explicitly asks whether an AI system can be audited by independent third parties.

That distinction matters.

Operational teams already collect enormous amounts of information.

Logs.

Traces.

Metrics.

Alerts.

Model-performance data.

Those are important for operating AI systems safely and reliably.

But a system being observable to its operator is not the same as its evidence being independently checkable by another party.

NIST does not require cryptographic execution evidence.

That is not the claim.

The important point is that auditability has to survive beyond the engineering team that operates the system.

At some point, somebody outside that trust boundary may need to examine what happened.

ISO/IEC 42001: evidence that a management system is operating

ISO/IEC 42001 approaches AI from the management-system side.

It specifies requirements for establishing, implementing, maintaining and continually improving an Artificial Intelligence Management System.

That includes areas such as responsibilities, policies, AI risk management, data governance, lifecycle controls, monitoring and performance evaluation.

The standard is broader than execution evidence.

It should remain broader.

But management-system assurance has never been only about having the correct policy document.

ISO itself explains that an organisation claiming conformance with a management-system standard needs evidence that it is meeting the relevant requirements, typically gathered through audits. Where an organisation chooses ISO/IEC 42001 certification, that certification is performed by an independent certification body, not by ISO itself.

Now imagine a financial institution has a policy stating:

AI-assisted credit decisions below a defined confidence threshold require human review.

The policy establishes the rule.

The interesting question during an audit is whether the organisation can show what happened when that rule should have applied.

Which execution crossed the threshold?

What recommendation did the system produce?

What information was available?

Was the case routed correctly?

Did somebody review it?

Was the recommendation accepted or overridden?

Which policy version applied at the time?

The policy tells you what should happen.

Evidence tells you what happened when the policy met a real execution.

That distinction becomes more important as AI performs a larger part of the processes those management systems are supposed to govern.

SM&CR: accountability still has a name attached to it

UK financial services introduces another dimension.

The Senior Managers and Certification Regime is not an AI regulation.

It does not require cryptographic AI evidence.

Its relevance here is accountability.

Senior Managers have Statements of Responsibilities identifying what they are responsible and accountable for.

The FCA also places a Duty of Responsibility on Senior Managers and requires them to take reasonable steps in the areas for which they are responsible.

Now introduce AI into a material workflow.

Credit.

Fraud.

Insurance.

Financial crime.

Customer treatment.

Market surveillance.

Trading controls.

The accountable person does not disappear because software performs part of the decision-making.

A different problem emerges.

The person retains responsibility while becoming further removed from individual executions inside the process.

A Senior Manager may be able to explain why a model was approved.

They may be able to show the governance framework.

They may be able to describe the controls.

But if one particular decision is challenged months later, can the organisation establish what the system actually recorded at that decision point?

To be clear, this is an inference from the accountability structure of SM&CR.

It is not a claim that SM&CR requires cryptographically verifiable execution records.

But the accountability logic is difficult to ignore.

Named responsibility becomes harder to discharge when the person who owns the process cannot establish what happened inside the execution they nominally own.

The United States: existing law can create the same evidence problem

The US picture is different again.

Rather than one horizontal federal regime equivalent to the EU AI Act, organisations face a mixture of existing statutes, sector rules, agency enforcement, state requirements and litigation.

That can create the same evidentiary pressure through a different route.

The question may arrive not as:

“Are you compliant with the AI Act?”

but as:

“This decision affected me. Explain why it happened.”

Credit provides a useful example.

Under current Regulation B, 12 CFR §1002.9, a creditor taking adverse action must provide the principal reasons for that decision. The official interpretation states that those specific reasons must relate to and accurately describe the factors actually considered or scored by the creditor.

That obligation matters whether the decision process is simple or technically complex.

There is an important historical clarification here.

The CFPB previously issued Circular 2022-03 specifically addressing adverse-action notices and complex algorithms.

That circular was withdrawn on 12 May 2025 as part of a wider withdrawal of CFPB guidance.

The circular should therefore not be cited as current guidance.

But withdrawing the circular did not remove the underlying Regulation B requirement to provide specific principal reasons for adverse action.

That distinction is useful because it illustrates the broader point.

AI-specific policy can change.

The underlying need to account for consequential decisions may remain.

Employment provides another example of how existing law can reach automated decision systems.

In the iTutorGroup case, the US Equal Employment Opportunity Commission alleged that the company's tutor application software was programmed to automatically reject female applicants aged 55 or older and male applicants aged 60 or older. The case settled for $365,000 plus other relief.

It is more accurate to describe this as an automated hiring software case than as proof of some general AI-specific enforcement regime.

But that is precisely why it is relevant.

Existing legal duties do not necessarily disappear because software performs the decision.

When the dispute concerns one particular outcome, explaining how a system generally works is different from establishing what happened in that specific case.

The structural gap in today's evidence

Across all of these regimes, organisations already produce enormous quantities of documentation.

Risk registers.

Model cards.

Policies.

Audit logs.

Observability traces.

Approval workflows.

Governance dashboards.

Incident reports.

Model validation documents.

None of these are useless.

Most are necessary.

The problem is more structural.

Almost all of the evidence is produced, stored and presented by the same organisation whose system is being examined.

That does not mean the evidence is false.

It means the evidence has a trust boundary.

Imagine a bank presents an internal record showing that an AI-assisted recommendation was produced at 14:32 on 6 February.

Who produced the record?

The bank's system.

Who stored it?

The bank.

Who controls the database?

The bank.

Who generated the export?

The bank.

Who is now asking the external reviewer to rely on it?

The bank.

Again, that does not make the record dishonest.

And some organisations already use strong append-only, WORM or externally anchored logging systems that materially improve this position.

But the fundamental question remains:

What does the examining party have to trust?

The organisation cannot become an independent third party to itself.

And this distinction becomes even more important when cryptography is introduced.

A cryptographic signature produced entirely inside the system being questioned can authenticate that system's assertion. It does not create independent corroboration.

That is subtle, but important.

A digital signature can show that a particular private key signed a particular payload.

A hash can show whether protected content has changed relative to a previously established digest.

Neither automatically proves that the original record was complete, truthful, or faithfully captured from the underlying event.

A useful way to think about it is this:

Operator-controlled evidence is testimony from the system. Independently verifiable evidence gives somebody else a way to check the testimony.

While everyone trusts everyone, that distinction may not matter very much.

When a customer disputes a decision, an insurer challenges an outcome, a regulator investigates an incident, or litigation begins, the trust assumption changes.

The evidence is now being examined by somebody who may explicitly not trust the system that produced it.

That is the point where many AI audit trails become weaker than organisations expect.

Agents make reconstruction harder

The problem becomes more difficult as AI becomes agentic.

A conventional model invocation may look roughly like this:

input → model → output

An agent workflow can look more like this:

input → model → retrieval → tool call → another model → external API → policy check → another agent → action

Some steps are deterministic.

Some are probabilistic.

Some depend on external systems whose state changes.

Some happen inside infrastructure operated by different vendors.

Some may never be directly visible to the human ultimately accountable for the outcome.

The final decision may be straightforward.

The path that produced it may not be.

This weakens one of the traditional fallbacks in governance.

Historically, when the record was incomplete, organisations could often ask the person who made the decision.

With increasingly autonomous systems, that person may no longer exist in quite the same way.

A human may approve the framework.

A human may supervise exceptions.

A human may own the business process.

But software may perform dozens of intermediate actions before anything reaches them.

So months later, asking the accountable person what happened may not answer very much.

The evidentiary question shifts from:

Who remembers what happened?

to:

What record survived what happened?

That is a significant change.

What would a record need to survive challenge?

This is where the discussion should move away from products.

Before asking what software to buy, ask what properties the evidence itself needs.

If an AI execution record is expected to survive external scrutiny, there are at least five properties worth considering.

1. Tamper-evidence over a defined protected set

Saying that a record is “immutable” is not enough.

What exactly is protected?

The input?

The output?

The model identifier?

Parameters?

Tool responses?

Human review?

The execution sequence?

The policy version?

A credible evidence format should explicitly define its protected set.

If a protected field changes, verification should detect it.

Just as importantly, fields outside that integrity boundary should be clearly identified.

Otherwise two people can examine the same artifact while assuming different things were actually protected.

Tamper-evidence is not a vague property of “the record.”

It has a scope.

That scope needs to be explicit.

2. Independent verifiability

This is the property many internal audit trails are missing.

A third party should be able to verify the relevant integrity claims without relying solely on the organisation whose system produced the record.

That does not mean eliminating all trust.

No practical evidence architecture eliminates every trust assumption.

It means making those assumptions visible.

A verifier should be able to answer questions such as:

  • Does the integrity proof match the protected artifact?
  • Is the signature valid?
  • Which key signed it?
  • What fields were protected?
  • Has the execution chain changed?
  • What verification material was used?
  • If verification fails, why did it fail?

That last point matters.

A verification system that returns only a green badge is difficult to audit itself.

Independent verification should be repeatable enough that two conforming verifiers examining the same artifact under the same rules reach the same result.

A backend saying VALID is an assertion.

A verifier showing why the integrity checks pass is something different.

3. An independent assertion of time

A timestamp written by the application itself is weak evidence of when the record actually existed.

The application controls the value.

That becomes important when timing is part of the dispute.

Was the record created when the execution happened?

Or after the complaint arrived?

Those are very different evidentiary positions.

For higher-assurance records, an external time assertion can reduce dependence on the operator.

Trusted timestamping, including mechanisms based on RFC 3161, is one established approach.

The point is not that every AI execution requires an external timestamp.

The point is that when time itself matters, the evidence should not rely exclusively on the party whose chronology is being questioned.

4. Portability

Evidence should not disappear because a vendor relationship ends.

A record expected to survive challenge should be capable of leaving the system that created it.

That means portable artifacts.

Published verification rules.

Stable identifiers.

Exportable trust material.

Historical public keys where signatures are involved.

A reviewer should not need an active subscription to the original SaaS platform simply to determine whether an old record remains intact.

This becomes increasingly important as organisations change vendors, merge systems, replace models and retire infrastructure.

Long-term accountability and vendor lock-in are uncomfortable partners.

5. Privacy

There is one final requirement that is easy to overlook.

The evidence cannot create a second data problem.

Many AI workflows where evidence matters most involve highly sensitive information:

financial data.

medical records.

customer communications.

legal documents.

identity information.

internal investigations.

commercially sensitive material.

An evidence architecture that requires all of that information to be disclosed to every external verifier has solved one trust problem by creating another.

Privacy therefore has to be part of the evidence model.

Possible mechanisms include:

  • minimisation
  • redaction
  • hashing
  • cryptographic commitments
  • selective disclosure
  • protected evidence packs
  • controlled reviewer access

The exact mechanism can vary.

The requirement is broader:

Producing evidence that a record exists and remains intact should not automatically require disclosing all of the sensitive data behind it.

Why a common specification matters

None of these properties is individually new.

Hashing exists.

Digital signatures exist.

Trusted timestamping exists.

Canonical serialization exists.

Redaction exists.

Evidence packaging exists.

The problem appears when every organisation implements them differently.

One vendor hashes the entire JSON object.

Another hashes selected fields.

One changes serialization rules between software versions.

Another signs a digest but does not retain historical verification keys.

One stores an application timestamp in its own database and describes that as independent time.

Another provides a downloadable artifact that only its own SaaS product can interpret.

Another removes fields after sealing them without clearly defining what happens to the integrity boundary.

Every one of these systems may call the result an “AI audit trail.”

The artifacts remain incompatible.

The verification semantics remain inconsistent.

And the person examining the evidence has to understand a different trust model for every provider.

That is why a common specification matters.

AIEF: define the evidence requirements before defining the vendor

AIEF, the AI Execution Framework, is an open, implementation-agnostic framework for verifiable AI execution evidence.

Its current public specification is v0.2.5, Public Comment.

That is also the current canonical public name used by aief.dev and its suggested citation.

AIEF does not define how AI should make decisions.

It does not determine whether a model output was correct.

It does not define fairness.

It does not replace an AI governance programme.

And it does not require one vendor's technical architecture.

It asks a narrower question:

What properties should an execution artifact have if another party may need to verify it later?

The AIEF specification defines concepts including:

  • execution artifacts
  • declared protected sets
  • stable serialization
  • tamper-evidence
  • deterministic verification
  • standardised PASS and FAIL outcomes with reason codes
  • independent validation
  • external dependency evidence
  • chain integrity for multi-step workflows
  • retention and portability
  • privacy and redaction controls

For probabilistic AI systems, AIEF makes another important distinction.

Verification is not:

“Rerun the model and hope it produces the same answer.”

AIEF uses a cache-as-truth model.

The recorded output is the historical object being audited.

Verification checks the integrity of that recorded artifact rather than pretending a probabilistic model, an external API, or a changing retrieval system will necessarily reproduce the same historical result.

AIEF also defines its boundary carefully.

Successful verification can establish that protected fields remain consistent with the declared integrity proof.

It does not prove:

  • that the decision was correct
  • that the model was fair
  • that the output was truthful
  • that every relevant event was captured
  • that the business decision was appropriate
  • that the originating system acted in good faith

That last limitation is especially important.

AIEF explicitly distinguishes artifact integrity from artifact authenticity.

A perfectly intact artifact can still have been falsely issued in the first place.

Higher-assurance situations may therefore require additional provenance, attestation, runtime trust, external timestamping or other mechanisms.

A standard becomes more credible when it is explicit about what it does not prove.

AIEF is published openly and invites comment from builders, auditors, regulators, academics and practitioners.

Its implementation registry is separate from the specification itself.

That separation matters.

AIEF defines the evidence model. Implementations are free to implement it differently.

One implementation: NexArt

NexArt is one implementation listed in AIEF's implementation registry.

The registry describes entries as factual and self-declared, and explicitly states that inclusion is not an endorsement.

That is the correct relationship.

AIEF is the open framework. NexArt is an AI Evidence Layer and Verifiable Execution Infrastructure that implements this class of evidence in a production system.

NexArt represents individual executions as Certified Execution Records, or CERs.

Multi-step executions can be assembled into Project Bundles.

The evidence flow is conceptually straightforward:

execution data → protected artifact → integrity anchor → attestation material → verification

The distinction between those stages matters.

The application or integration remains responsible for capture.

NexArt does not independently prove that every relevant event was captured, or that application-supplied facts were truthful.

What the evidence layer can establish is that the protected record has not subsequently changed without detection and, where attestation is present, that the corresponding signing key signed the relevant record.

The cryptography is ordinary on purpose

NexArt CERs use SHA-256 certificate hashes over a defined protected projection.

Canonicalization is protocol-bound.

Current protocol 1.3.1 uses RFC 8785 JSON Canonicalization Scheme, while older supported records can use their earlier canonicalization profile.

When a CER is node-certified, the NexArt attestation node adds an Ed25519-signed receipt.

The node publishes its verification key material publicly through a well-known endpoint so that signature checking does not require asking the originating application whether the signature is valid.

As of this writing, the production node publicly identifies its operator as NexArt.

That is worth saying explicitly.

Independent verification is not the same claim as independent node ownership.

The value is that a third party holding the evidence and authentic public-key material can perform the cryptographic checks independently of the organisation whose AI workflow produced the record.

That is a narrower claim.

It is also a more defensible one.

External time evidence

NexArt also supports external RFC 3161 timestamp evidence as an additional verification layer.

This should not be confused with the node's own attestedAt time.

A node-issued timestamp comes from the signing system itself.

An RFC 3161 timestamp can provide a separate time-of-existence assertion from an external Time Stamp Authority.

Timestamp state is record-specific, so a verifier should inspect whether external timestamp evidence is actually present rather than assume it from the existence of a CER.

That is exactly the kind of distinction an evidence system should make visible.

Confidential execution without putting raw input and output into the proof record

Privacy creates another challenge.

A financial institution may need to prove that a record has not changed while having no legal or commercial reason to disclose the customer's underlying data to the verifier.

NexArt protocol 1.3.1 addresses this using HMAC-SHA256 commitment envelopes for raw input and output.

The implementation boundary is important.

For 1.3.1, the client transmits the raw input and output to the NexArt attestation node over TLS.

The node processes those values transiently and replaces them with cryptographic commitments before building the certified record.

According to the current protocol documentation, the raw input and output are excluded from the certified snapshot and from the persistent proof record, and the certification flow is designed not to persist them.

The resulting CER contains commitments rather than the plaintext input and output.

This does not mean every field in the record is confidential.

Fields such as the prompt and model parameters can remain plaintext, so integrations still have to decide carefully what information belongs in each part of the evidence record.

And NexArt does not claim that the raw values never cross the customer boundary.

They do cross to the node over TLS in the documented 1.3.1 flow.

That precision matters.

The goal is not to turn “confidential execution” into a vague privacy slogan.

It is to define exactly where plaintext exists, where it does not, and what remains independently verifiable afterward.

Check the evidence yourself

The easiest way to understand independently verifiable evidence is to stop describing it and give somebody something they can inspect.

A public NexArt record is identified by its certificate hash.

One public sample used in a July 2026 public verification audit was:

sha256:2e9cfd2bdaa5125578722ed12508f5f5091cc7108bc4f87c94d835c841719082

Open the record in the public verifier.

The July audit recorded:

  • SHA-256 integrity: PASS
  • Ed25519 signature: PASS
  • key source: public well-known key
  • external timestamp authority: DigiCert
  • RFC 3161 timestamp status: OK

The public verification URL continues to resolve at the time of writing.

Public record lookup requires no NexArt account or API key.

You can also query the public resolver directly:

CERT="sha256:2e9cfd2bdaa5125578722ed12508f5f5091cc7108bc4f87c94d835c841719082"

curl -fsS \
  "https://node.nexart.io/v1/cer/public?certificate_hash=${CERT}"
Enter fullscreen mode Exit fullscreen mode

The important point is not that a NexArt webpage displays the word Verified.

The useful property is that the evidence exposes the material needed to check its cryptographic claims against published verification rules and public keys.

And the scope of that result needs to remain precise.

Successful verification can establish that the protected record remains intact and that the relevant signing key signed it.

Where valid external RFC 3161 evidence is present, it can also establish an external time-of-existence assertion.

It does not prove the AI's answer was correct.

It does not prove the underlying model call was independently observed.

It does not prove the application's capture boundary was complete.

It does not prove the decision was lawful, fair or appropriate.

That is not a weakness in the definition.

That is what makes the definition useful.

Documented compliance and demonstrable compliance are different things

The AI governance industry has spent a lot of time helping organisations document what they intend to do.

That work is necessary.

Policies matter.

Governance matters.

Model validation matters.

Human accountability matters.

Logs matter.

Observability matters.

But as AI systems begin making and influencing more consequential decisions, another layer becomes increasingly important.

Evidence of what actually happened inside the defined capture boundary.

The regulatory regimes discussed here do not use identical language.

They do not impose identical legal requirements.

And none of them mandates AIEF, NexArt, Certified Execution Records or any particular cryptographic architecture.

That is not the argument.

The argument is simpler.

Across record-keeping, human oversight, auditability, management-system assurance, named accountability and legal challenge, organisations are being pushed toward a stronger evidentiary position.

From:

“Here is our policy.”

to:

“Here is what our records show happened.”

And eventually from:

“Here is what our system says happened.”

to:

“Here is the evidence. Check it yourself.”

That is the difference between documented compliance and demonstrable compliance.

The gap between the two is where a growing amount of AI risk now sits.


Further reading and primary sources

The open framework

NexArt implementation

Regulatory and standards sources

Top comments (0)