Ask a language model which package to use and, now and then, it names one that does not exist. If someone registers that name first, the next developer — or the next coding agent — who follows the same suggestion installs whatever they put there. That is slopsquatting.
We collect those names and score each one. A name that returns 404 is easy: it is a target. The hard case is a name that does resolve. Is it the project the model meant, or someone who registered the name after models started suggesting it?
Our answer starts with age:
age_days = (now - info.created_at).days
if age_days < 14:
return "high"
if age_days < 60:
return "medium"
return "low"
Crude, but honest. A package that appeared last week and matches a name models keep inventing deserves suspicion.
The trouble was everything we put above that ladder. Over the past few weeks we had added exceptions — rules that let a young package off — each one to fix a specific false positive. This week a reader went through them in a comment, and by the end of the next day three were gone.
1. "It has a signed build statement"
npm provenance and PyPI's PEP 740 attestations record that a package was built by a named workflow in a named repository, through a Trusted Publisher flow. The registry produces the statement at publish time, so the package cannot just claim it in its manifest. That felt like strong evidence, so we added:
if prov is not None and prov.attested and prov.repository:
return "low", f"built from {prov.repository} with a signed build statement"
Read what the statement actually says. It binds the artifact to a repository. It says nothing about whose repository. Anyone can create a GitHub repo, configure a Trusted Publisher for it, and publish a squatted name with a perfectly valid attestation. It takes minutes.
So the rule did not lower the risk of an attack. It cleared exactly the package an attacker would publish, as long as they spent five extra minutes on CI.
The reader put it more generally, from a different system they work on: declarations pool where they are cheap. In a log they scan, 23 of 31 signing keys had zero accepted work and zero deliveries — and emitted 91% of all capability declarations. A signed statement that costs nothing to obtain tells you that someone obtained it.
Now: provenance is still recorded and published next to each finding, for a reader to weigh. The classifier does not take it as an input at all — not ignored, removed from the signature — so nobody can wire it back in by accident.
2. "It was registered before anyone suggested it"
This one came from a real false positive. cdx-rs is a cd replacement, published on 30 August. On 4 September a model offered cdx-rs as a library for parsing CycloneDX SBOMs, which it is not. Five days old: high risk. But the package predated the suggestion, so it could not be a response to it. We added:
if first_hallucinated and info.created_at < first_hallucinated:
return "low", "registered before the name was ever suggested"
first_hallucinated is the earliest time we saw a model produce the name. That is not when the name was first suggested. It is when our sampling first happened to catch it. Models had been producing these names long before we started collecting; we only know about the ones that came up after.
So near the day our collection began, the rule read: registered before we started looking. Someone who registered a name early and waited would have cleared through it. The comparison only means something when both dates are pinned by something outside the person making the claim — and one of ours was pinned by us.
The reader also spotted something worse about the shape of it. If first_hallucinated was missing, the check was skipped and the package fell through to the age ladder, where it could score high. If it was present and early, the package scored low immediately. Two kinds of not-knowing, two opposite verdicts, and no branch anywhere that said "we don't know".
Is there an independent date? The best I found is the model's training cutoff. The vendor sets it, and our sampling cannot move it. A package registered after every recommending model's cutoff could not have been learned from — the name was invented first and registered later, whenever we happened to look. But every model we sample has a cutoff older than our 60-day window, so a package that predates one is already old enough to clear by age. The rule had nothing honest left to do.
Now: age alone decides. That has a cost: cdx-rs is flagged again. We cannot tell a coincidence from a squat by its registration date, and a list of attack targets should fail toward caution.
3. "We don't know when it was registered"
This one the reader did not mention. Their framing found it. Once you start looking for safe verdicts that rest on absence of evidence, you look at every branch that returns low:
if info.created_at is None:
return "low", "registered"
A package that exists but has no creation date: fine. When does that happen? For every Go module, it turned out. The Go module proxy answers one question — does this module exist — and our probe asked it nothing else. Every Go module a model recommended, including one pushed yesterday, was cleared without a date ever being checked.
Now: an undated package scores medium. Go modules are dated from deps.dev, which lists every version the proxy has served with its time.
That fix comes with a caveat. For Go, that time is the commit's time, and a commit's date is whatever its author wrote. npm and PyPI record when the registry received an upload, which the publisher cannot choose. A backdated Go module still looks old. It is weaker evidence than the other ecosystems — and better than none, which is what we had.
The audit we couldn't run
The reader suggested a cheap check: take every name rule 2 cleared, plot its registration date against the day our collection began, and see whether they cluster at the boundary. If they do, the rule is reading its own start date.
We couldn't. A finding is stored. A name that scored low was simply dropped. The rule had cleared names and left nothing behind to audit.
Now every low verdict is recorded with the rule that produced it, the registration date and our first sighting. It is never published; it is there so the next rule we get wrong can be caught by looking rather than by a stranger reading our source. We'll run the reader's check once a few weeks of data have built up.
Findings that had been withdrawn under rules 1 and 2 were re-scored under the new rules. Two came back onto the list.
What I'd ask of any risk score
Every rule we've had to remove was one that cleared something. The rules that flag things are easy to defend: a 404, a package two days old. The exceptions are where the reasoning goes soft, because each was written to stop one false positive, by someone looking at that false positive and not at an attacker.
So, for each branch in your scoring that returns "safe":
- Can an attacker produce this input? A signed build, a README, a repository field, a star count.
- Does it describe the world, or only you? Your first sighting, your crawl date, your cache.
- What happens when it is missing? If "unknown" can come out as "safe", it will, at scale, for a whole ecosystem.
The rule that says "this one is fine" is the one an attacker reads.
Thanks to ANP2 Network for the comment that started this. If you can find a fourth, the comments are open.
Top comments (0)