DEV Community

Kalislav Smirnov
Kalislav Smirnov

Posted on

An SBOM is a tool output, not a fact about your software

Two papers appeared on arXiv in September that describe, from opposite ends, a pipeline control I have shipped many times. The first, from Inria and ANSSI, ran three SBOM generators over the same 3,326 repositories and found that two of them miss more than half of the dependencies in JavaScript projects. The second ran 1,920 trials of AI coding assistants installing software and found that they opened an SBOM, signature or attestation before installing in 0.5% of trials and verified one in none.

My position: an SBOM generated in CI is evidence about the tool you ran, at the version you ran it, under the flags you set. It is weak evidence about the software. That still leaves it useful, but the thing to put under version control and defend in an audit is the generation recipe. The component count on its own means little.

This is commentary on those two papers, linked where used; I have not reproduced them.

What the Inria paper shows

Prado, Zendra, Boinot and Barais (17 September 2026, accepted at SCORED 2026) took 2,050 JavaScript and 1,276 Rust repositories with over 1,000 GitHub stars, regenerated a fresh lockfile for each, and used that lockfile as ground truth. They ran Syft 1.38.2, Trivy 0.68.2 and cdxgen 12.0.0, all producing CycloneDX 1.6, and compared.

With all lockfiles exposed, Syft reported 43.78% of the lockfile packages on JavaScript and Trivy 32.71%. On Rust, Syft reached 100% and Trivy 83.85%. cdxgen was at 98.72% and 99.34%. With no lockfile at all, Syft and Trivy produced an empty SBOM, and the paper is specific that this happens "with no error or warning shown to the operator."

The useful part is why. Most of the JavaScript gap is devDependencies. Syft excludes them by design, but only for some package managers: in one npm project it reported 2249 of 2249 production packages and 0 of 4930 dev packages, while in a Yarn and a pnpm project it reported both at 100%. Trivy excludes them by default too, but the exclusion stops at the first level, so a transitive dependency of a dev package can show up attached to the root as if it were a top-level production dependency. The authors call Syft's behaviour a design choice applied inconsistently and Trivy's an implementation error. cdxgen's small residual gap was a hard-coded list that skips any directory named examples, along with docs, tests and .github.

The same package also comes out with a different identifier depending on the tool. An npm alias is reported under the alias by cdxgen and Trivy and under the real scoped name by Syft. pnpm peer-dependency suffixes end up in the version string for cdxgen and Syft. One git-pinned dependency came back as a commit hash from one tool, a tarball URL from the second and a semantic version from the third. Vulnerability matching works on these identifiers, so the same code scanned with two tools can produce two different vulnerability lists.

The Supplier element, one of NTIA's minimum elements since 2021, is populated by none of the three tools on either ecosystem. Licenses: 75.8% for cdxgen, 60.2% for Syft, 0% for Trivy on JavaScript, and 0% for all three on Rust. Hashes come almost only from cdxgen, CPEs almost only from Syft.

What it does not show

The sample is popular GitHub repositories, and the authors say they did not measure how the 1,000-star threshold affects coverage. The results are tied to three specific versions; the authors note that a 2026 study using Trivy 0.66.0 and Syft 1.33.0 got empty SBOMs for package-lock.json, and a few releases later the same tools produce partial ones. Trivy 0.75.0 and Syft 1.54.0 both shipped on 1 October, so some of these numbers may already be stale. The field-completeness measurement checks whether a field is present, not whether its content is right. And the paper covers JavaScript and Rust only; I would not assume Java or Python look better.

What the second paper adds

Pengyin Shan's pre-registered audit (7 September 2026) asks the other question: when the thing installing software is an AI assistant, does it read any of this? Six research software projects, nine variants each (no signal, a valid SBOM, a valid signature, a valid attestation, two with wrong-issuer material, one with every signal, one with contradictory metadata), three models, two harnesses, every trial in a network-less container with file access and commands logged, protocol deposited with a DOI before the first trial.

In 9 of 1,920 trials the assistant opened a provenance file before installing. In 0 of 2,114 trials, including a frontier-model supplement, did any assistant run a verification command. Signal presence had no measurable effect (p = 0.50). The model that cost $1.00 per trial verified nothing in 50 trials; the only one that opened signals more than once cost about $0.10 per trial and did it in 9 of 360. The author's conclusion is that verification "has to be built into the program that runs the assistant."

The limits are real: six projects, research software, a container with no package index so most installs failed at dependency resolution, and material signed by the author's own key rather than the upstream's. It measures the decision to install, not a successful install. Still, zero verification commands across six models and two harnesses is hard to explain away with sample size.

Where this goes

The CRA's main obligations apply from 11 December 2027, with reporting obligations from 11 September 2026, per the Commission's own page. Most manufacturers will discharge that with one of these three tools, and the Inria authors put it bluntly: tool choice "is no longer just an engineering decision, it is a compliance-exposure question."

I think SBOMs matter in two years as an inventory you can defend, and I doubt they will matter as a trust signal anyone checks at install time. The second paper shows the install-time consumer is already often an agent, and the agent does not look. The first shows that a human who looks cannot tell from the document whether it is complete. The fix for the first is standards work: the authors ask CycloneDX and SPDX to define which dependency scopes are in, how provenance is recorded, and what the root of the graph is, which will take years. The fix for the second is a harness that runs cosign verify or slsa-verifier before pip install.

Monday

I have run security checks (Gitleaks, Trivy, SCA) in every pipeline of a secure development platform through shared GitLab CI templates, and I have delivered signed packages into an air-gapped bank, where the receiving side cannot regenerate anything and works with what is in the bundle. This is what I would change this week.

Pin the generator and write its version into the SBOM metadata. The paper's results changed between minor releases of the same tool. If an auditor asks why a component is missing, "Trivy 0.68.2 with default flags on a Yarn project" is an answer. "We run Trivy" is not.

Fail the job on an empty SBOM. Syft and Trivy return zero components and exit 0 when there is no lockfile. A jq '(.components // []) | length' with a threshold is five lines in a shared template and catches the case the paper calls more severe than under-reporting.

Decide the dev-dependency question on purpose. Trivy's --include-dev-deps exists; the default excludes them, and the exclusion leaks. For a deployment SBOM you may want them out. For a view of your build's supply chain you want them in, since the paper notes development dependencies have themselves been targets of supply-chain attacks. Write the choice down.

Commit the lockfile, or generate the SBOM from the built artifact instead of the source tree. Without a resolved lockfile the tools are either guessing or silent.

Run two generators once, on your five most important repositories, and diff the component sets. An afternoon of that tells you which scope and identifier rules bite your stack, before you promise anyone a complete inventory.

For air-gapped delivery, put the SBOM inside the signed bundle together with the generator name, version and command line, and have the receiving side verify the bundle signature. The SBOM records what you believe is in there; the signature records who is accountable if that belief is wrong.

If an AI agent sits anywhere in your dependency flow, put the verification step in the pipeline that runs it, because the model will not do it on its own.

What would change my mind

A benchmark that pins these edge cases, which the Inria authors say they are building, with Syft and Trivy above 95% on JavaScript. A CycloneDX or SPDX release that defines dependency scope as a standard attribute, so "complete" means one thing across tools. Or a re-run of the agent study in which a production harness verifies attestations by default and leaves a log.

Sources

Top comments (0)