This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend.
It's Monday morning, CI goes red, and an automated test fails. You don't panic or blame compiler magicβyou open git diff and check package-lock.json or Cargo.lock. You know down to the exact commit hash what moved, what broke, and why.
So why do we accept total chaos the moment we start building AI agents?
Over the past year, we leaped straight from writing prompt files (SKILL.md) to running automated CI evaluations (like NVIDIA's AEVAL). Yet every time an evaluation test fails, we sigh, throw our hands up, and blame "unpredictable AI flukes" or "model drift."
But is it really a model fluke? Or did an unpinned sub-dependency shift under your feet? Without a lockfile, you literally can't tell.
The Illusion of Lockfiles in Current Tooling
You might think: "Wait, don't skill lockfiles already exist?"
Yes, some experimental agent harnesses now generate a skills-lock.json or parse SKILL.md frontmatter. But existing tools treat lockfiles as passive inventory lists on disk rather than active dependency owners. They don't track transitive sub-dependencies, don't validate environment capability contracts, and don't bridge runtime prompt context names back to content-addressed package hashes.
Because current tools don't own the supply chain, agent debugging remains a guessing game:
- Did a bare skill name in prompt context resolve to a different file version than last week?
- Was a capability contract missing from the execution environment?
- Did an unrecorded sub-dependency change?
We want evaluations to give us control over privileged agents. But real control requires two things:
- Reproducibility β running the exact same bits across every CI node and machine (requires a Package Manager & Lockfile).
- Attributable Failures β knowing precisely which versioned input broke the test (requires a Dependency Graph).
Two Graphs, One Bridge
The core architectural confusion comes from mixing up two completely different jobs:
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Dependency Graph (Tooling) β
β Resolves package refs, content hashes & pinned addresses β
ββββββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββ
β
Bridge (Name β Address)
β
ββββββββββββββββββββββββββββββββΌβββββββββββββββββββββββββββββββ
β Context Graph (LLM) β
β Loads bare skill names into model prompt context β
ββββββββββββββββββββββββββββββββ΄βββββββββββββββββββββββββββββββ
-
Dependency Graph (Tooling Layer): Resolved offline by developer CLI tooling before execution. It handles version constraints, content hashes, and package manifests, pinning everything into
skills-lock.json. The model never sees this graph. It deals in immutable package addresses. - Context Graph (LLM Layer): Resolved dynamically by the LLM at runtime. The agent reads skill metadata and loads instructions into its prompt context window. It deals in bare skill names.
The Bridge connects them. When a skill named allergen-screening is loaded into context, the bridge verifies that this bare name maps directly back to an immutable content hash pinned in skills-lock.json.
The Security & Attribution Bonus
Connecting prompt names to content hashes does double duty:
- Attribution & Blast Radius: When a test fails, the bridge translates loaded context names back to pinned hashes and runs reverse reachability over the dependency graph. Instead of a vague "AI fluke," you get an exact node-precise blast radius showing every module and dependent skill affected.
-
Prompt Injection Defense: If an indirect prompt injection attempts to trick the agent into invoking an unverified skill name or spoofing an existing one, the bridge acts as a hard boundary: if the requested skill name isn't pinned in
skills-lock.jsonor its content hash on disk has been tampered with, validation fails immediately.
Model-Agnostic Control
Models driftβwhether through unannounced API updates in closed models or variations across open-weight fine-tunes and quantizations. If your agent's safety depends entirely on prompt obedience inside model weights, every run is a gamble.
By moving governance into an explicit, versioned dependency graph backed by a content-addressed lockfile, safety becomes a model-independent invariant.
Whether you route requests to a commercial API endpoint, a hosted provider, or a local open-weight model, the guarantee stays anchored in the graph infrastructure. The LLM is treated as what it should be: an interchangeable reasoning engine.
What We Built & What's Next
We built a working prototype as an open-source research sandbox, patching the skills CLI to bring package manager semantics to agent evaluation pipelines, wrapped with full observability and execution capabilities:
- Repository: TheLazzziest/research
-
Try the Tooling:
-
bun run skills:lockβ Generates a content-addressableskills-lock.json. -
bun run skills:verifyβ Verifies local vendored skills against the lockfile. -
bun run skills:graphβ Renders the attributable skill dependency graph. -
bun run skills:lintβ Validates frontmatter, manifests, and bridge constraints.
-
-
Full Stack & Observability: Instrumented agent decorators with Sentry
gen_aispan/latency tracing, ElevenLabs voice narration (TtsPortadapter), Gemma open-weight model presets, and a Render deployment service blueprint (render.yaml).
An Honest Note on Implementation: Today, our tooling enforces this bridge offline for CI and evaluation pipelinesβmaking sure test sessions are bit-for-bit reproducible and failure attribution works before code hits production. Extending this bridge into an active runtime guardrail at the LLM dispatch boundary is the next frontier, but bringing package management to evals gives us deterministic control today.
Stop blaming model flukes for unpinned dependencies. Let me know what you think in the comments below!
References & Related Reading
- AEVAL: NVIDIA Research, AEVAL: Change-Triggered Deterministic CI for Agent Skills, arXiv:2607.16345 (July 2026).
- Agent Skill Supply Chains: Skills Are Not Islands: Measuring Dependency and Risk in Agent Skill Supply Chains, arXiv:2607.01136 (July 2026).
- Formal Skill Analysis: Formal Analysis and Supply Chain Security for Agentic AI Skills, arXiv:2603.00195 (March 2026).
- Harness Security: Scanning the Harness: Empirical Study of Supply-Chain Defects in AI Coding-Agent Configurations, arXiv:2609.07360 (September 2026).
- Operational SKILL.md: Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry, arXiv:2605.11418 (May 2026).
Prize Categories
- Best Use of Backboard (Model routing across open-weight providers via SDK adapter)
- Best Use of GitHub Copilot (CLI patch series and agent harness development)
- Best Use of Gemma (Open-weight model presets including Gemma 3 27B)
-
Best Use of Sentry Agent Tracing (Instrumented agent decorators with
gen_aispans and latency metrics) -
Best Use of ElevenLabs (Voice narration with
TtsPortadapter and/api/ttsendpoint) -
Best Use of Render (Deployment service blueprint via
render.yaml)
Top comments (0)