DEV Community

Cover image for Beyond the package name: why following the call is the hard part of supply-chain security
jj1423
jj1423

Posted on

Beyond the package name: why following the call is the hard part of supply-chain security

Most supply-chain security checks the package: does it exist, is it new, is there a CVE against this version. Those checks answer is this safe to install?

They can't answer is the way my code uses it safe? A clean library can still be where your service breaks: a request URL passed to an HTTP client, an uploaded archive handed to a parser, a user path joined into a file operation. Nothing is wrong with the library, so no advisory will ever be filed, and every package-level check will keep saying "safe".

The obvious answer is to follow the data: from where input enters your code, through the call into the dependency, down to wherever it lands. We have been trying to do that for Python. It is much harder than it sounds. These are the problems that actually bit.

1. Who controls this argument?

Everything depends on knowing whether a value comes from outside. Statically, you mostly can't. The safe assumption, every function parameter might be attacker-controlled, turns your own configuration into "input": we had a storage client built from environment variables flagged as nine dangerous paths, twice.

It is worse for libraries, where every public parameter is outside input by definition. On ten popular open-source projects, more than half the paths we reported rested on that assumption rather than on a traced flow.

2. The direction of the danger matters

httpx validates URLs with regular expressions it wrote itself. A naive rule ("user input reaches a regex") flagged every httpx call we make, including one where the argument was a timeout in seconds. In our first run on our own service, 25 of 39 reported paths were this one mistake.

A regex is dangerous when the attacker writes the pattern, not when their string is matched against yours. Every sink needs rules like that: which argument, which direction, which mode. ZipFile(upload) reading is a decompression risk; ZipFile(target, "w") writing is not.

3. Right place, wrong reason

Our one real finding was a spreadsheet upload that could expand to two million rows in memory. The analysis pointed at the correct line and described it as a file-path operation, a regex and global state. None of those was the problem. The real category, decompression amplification, didn't exist in our sink list until we went looking for it.

Pointing at the right line is not the same as understanding it, and a wrong explanation sends a reviewer the wrong way.

4. Summarising every package, at every version

To follow an argument into a dependency you need to know where that package sends it, for the exact version in the lockfile. Python makes that hard: calls on objects, dynamic dispatch, thin wrapper modules, native code that can't be read at all. Every change to the summariser means rebuilding summaries for every version of every package, and results shift when it does.

5. The standard library is not a dependency

subprocess.run(cmd), eval, pickle.loads are the most dangerous calls in Python, and they aren't in your lockfile. An analysis built on the dependency graph skips them unless it handles them separately. Then the noise problem from #1 returns: on 39 real releases, path and regex calls produced 140 of 283 standard-library findings, for the lowest-severity class.

6. Paths explode

A tainted value can reach a sink along many routes through a dependency tree. Without a budget the analysis never finishes; with one, it stops early on exactly the large projects where it matters. A fixed budget cut eight of twenty real cases short before they reached the answer.

7. What counts as "right"?

Our first benchmark, calls we wrote from library docs, scored 21 of 21. That measured nothing. The useful test is real advisories, where an application really misused a dependency and someone filed a report: would the analysis have pointed at that line, for that reason? Against twenty of those, it does about half the time. Two of our own benchmark cases turned out to have the wrong answer key.

Where that leaves it

Following the call is where supply-chain security has to go, because it is the only way to see risk that no advisory will ever describe. But today it is a way to decide where to look, not a verdict. The hard parts are the ones above: knowing who controls a value, knowing which direction is dangerous, and measuring against real misuse rather than examples written to pass.

If you have worked on any of these, especially telling outside input from configuration statically, I would like to hear how you approached it.

Top comments (0)