DEV Community

Cover image for Semantic Versioning Is Not Enough for AI Agent Libraries
Raju Dandigam
Raju Dandigam

Posted on

Semantic Versioning Is Not Enough for AI Agent Libraries

A library release can preserve every exported TypeScript type and still change what users observe.

An adapter may capture different fields. A trace check may interpret repeated tools differently. A CLI command may return a new exit code. A redaction profile may remove more data. A persisted trace may remain parseable but reconstruct into a different execution tree.

Semantic Versioning remains necessary. For AI-agent tooling, it is not the entire compatibility story.

I maintain AgentInspect, a TypeScript-first local evidence debugger and trajectory-testing toolkit. Maintaining it has made me think about compatibility as several surfaces rather than one package API.

The examples here refer to agent-inspect@6.19.0, persisted schema 1.0, and Node.js 20 or later.

Map every compatibility surface

For an agent library, review at least these contracts:

Surface Example breaking behavior
Source API Renamed export or changed option type
Runtime behavior Same option captures different data
Persisted schema Reader cannot open an older trace
Interpretation Same trace produces different findings
CLI Changed flag, output shape, or exit status
Privacy Capture or redaction boundary changes
Adapter Framework events map differently
Evidence Bundle manifest or integrity semantics change

SemVer describes how a published package version communicates compatibility. It does not decide which of these behaviors your users depend on. Maintainers must make that inventory explicit.

“The type still compiles” is a weak test

Consider an adapter option:

type CaptureMode = "metadata-only" | "preview";
Enter fullscreen mode Exit fullscreen mode

Suppose two adapters accept preview, but one silently records metadata only. A release later makes preview behavior consistent and adds bounded, redacted previews. The public union did not change. The meaning of a configuration did.

That kind of improvement is desirable, but it still deserves:

  • release notes describing the behavioral change;
  • explicit privacy guidance;
  • adapter-specific fixtures before and after;
  • tests proving metadata-only remains the default;
  • tests proving no unexpected network transmission occurs.

AgentInspect's 6.18.0 work on bounded preview parity is a real example of why runtime semantics need their own change record beside API signatures.

Persisted artifacts need reader contracts

A local trace can outlive the code that produced it. A reader should either:

  1. open the schema correctly;
  2. migrate it explicitly; or
  3. reject it with an actionable error.

Do not “best effort” unknown fields into a convincing but false tree.

Golden fixtures make this testable:

fixtures/
  v0.1/minimal-success.jsonl
  v1.0/parallel-tools.jsonl
  v1.0/multi-run-session.jsonl
  malformed/missing-parent.jsonl
Enter fullscreen mode Exit fullscreen mode

Each release should run current readers, checks, reports, and exporters against supported historical fixtures. For custom ingestion, fixture contracts should cover the mapping from foreign events into the canonical read model.

Version 6.19.0 added custom TraceReader authoring and richer failure-role interoperability. Those capabilities increase the number of integrations; they also increase the importance of testing architectural intent, not only whether a file parsed.

CLI exit codes are an API

Humans read CLI prose. CI reads exit status and JSON.

A harmless wording change can be fine. Changing a failed contract from exit 1 to exit 0, renaming a JSON key, or printing diagnostics to stdout instead of stderr can break automation without changing a TypeScript declaration.

Maintain fixtures for:

  • successful and failed checks;
  • malformed input;
  • unknown schema versions;
  • JSON output;
  • redaction warnings;
  • integrity verification failures.

Document which output is stable for machines and which is presentation for humans.

Privacy changes deserve a higher bar

Capture defaults and redaction semantics are compatibility concerns with a security impact. A minor-looking change that writes prompt previews where only metadata existed before can violate a user's data boundary.

For every release, test negative promises:

metadata-only capture contains no prompt or response preview
redaction produces a separate artifact
no adapter uploads by default
integrity verification does not imply safety certification
Enter fullscreen mode Exit fullscreen mode

These assertions are as important as “the command succeeds.”

Publish a behavioral compatibility note

For each change, I now find four questions more useful than “major or minor?”

  1. Which surface changed?
  2. What will an existing user observe?
  3. Which old artifacts and configurations were tested?
  4. What migration or opt-in is required?

Turn the answers into a release artifact that machines can compare:

{
  "packageVersion": "6.19.0",
  "node": ">=20",
  "persistedSchemasRead": ["0.1", "0.2", "1.0"],
  "defaultWriterSchema": "1.0",
  "stableCliContracts": ["check", "gate", "report"],
  "privacyDefaults": { "capture": "metadata-only", "upload": false }
}
Enter fullscreen mode Exit fullscreen mode

The fields above illustrate the shape of a compatibility manifest; they should be generated from the exact release contract rather than copied by hand. Store it beside golden traces and CLI snapshots so an upgrade can reveal both intended and accidental movement.

SemVer remains the label on the release. The behavioral note explains the engineering reality behind it.

AI-agent libraries sit between nondeterministic models and deterministic software systems. Their job is often to turn behavior into evidence, policy, or control. Users therefore depend on more than function signatures. They depend on what gets captured, how it is interpreted, what CI sees, and which data stays out.

That broader contract deserves versioning too.

References

Top comments (0)