A Hallucinated Package in Your Notebook Is Not a Code Problem
Slopsquatting guidance assumes a repo, a lockfile, and a scanner. Your notebooks have none of those, and the cluster running them holds the catalog credential.
Ask a coding agent for a PySpark transform and it may open with an install line for a package that sounds exactly right and does not exist. Not a typo, not a deprecated library. An invented name, produced with complete confidence, that a data engineer pastes into a notebook and runs.
The attack built on that behavior is called slopsquatting. Attackers collect the package names language models invent, register those names on PyPI or npm, and wait. The Cloud Security Alliance's AI Safety Initiative documented live cases in an April 2026 research note, including a fabricated 'huggingface-cli' package that drew more than 30,000 downloads in three months.
The underlying numbers come from a USENIX Security 2025 study by researchers at UT San Antonio, the University of Oklahoma, and Virginia Tech. Across 2.23 million generated code samples, 19.7% referenced a package that did not exist. What makes that an attack rather than a nuisance is repeatability: when the same prompt was rerun, 43% of hallucinated names came back every single time. These are stable, predictable, and therefore registrable.
The models improved, and the target got sharper
A 2026 re-evaluation published on arXiv tested five frontier models released between October 2025 and March 2026 and found hallucination rates of 4.62% to 6.10%, down from the 5.2% to 21.7% spread measured across sixteen older models. But the spread collapsed from 16.5 points to 1.48, and 127 package names were hallucinated by all five models, 53 of them still available to register.
A name that one model occasionally invents is a weak trap. A name that every frontier model invents is a reliable one. The rate fell and the targeting improved, and for an attacker the second matters more.
Why the published advice does not fit a data platform
Almost everything written about slopsquatting assumes a software supply chain: a repository, a pull request where a new dependency gets a second pair of eyes, a lockfile that pins what actually shipped, and a scanner in the pipeline.
A data platform has close to none of that. It also has something application teams do not, which is why this matters more here than there. The compute holds the credential to the data.
Notebooks are often not in source control. No repository means no pull request, so nothing puts a human in front of a new dependency before it executes.
Installs resolve at runtime, against the public index. A '%pip install' line is not a reviewed build step. It reaches the open internet at the moment the cell runs, from a cluster that already has a catalog credential attached.
Cluster libraries outlive the notebook that added them. A package installed once during exploration is then present for every job scheduled on that cluster, including production jobs nobody connects to the exploration.
Unpinned job dependencies re-resolve on every run. This is the one that should worry you. A requirements file naming a package that did not exist last month will install happily this month, once somebody registers the name. The hallucinated name is not a broken reference. It is a reserved slot sitting in your job definition, waiting for an attacker to fill it.
Init scripts get edited without a diff. Workspace admin is a wider group than most teams think, and startup scripts rarely go through review.
What to do about it
None of this needs a new category of tool.
Point clusters at a private package mirror and block the public index at the network layer. Advising engineers not to install from PyPI is not a control. Making it impossible is.
Pin and lock every production dependency. Jobs should resolve from a lockfile with hashes, never from a bare name. This alone closes the reserved-slot problem.
Put notebooks under git integration, through Databricks Repos or the Microsoft Fabric equivalent, so a new dependency produces a diff somebody can see.
Separate exploration from production identity. The cluster where people try things should not carry the service principal that holds production catalog grants. Most of the damage comes from those two being the same cluster.
Allowlist at the cluster policy level, not in guidance documents, so the control travels with the compute rather than depending on who is driving it.
Treat a new dependency as a change, because it is one. A package entering a production job deserves the scrutiny a schema change gets, and usually gets less.
What it costs to get this wrong
Two things make the data-platform version worse than the application version.
The first is blast radius. On an application server, a malicious package compromises the application. On a cluster, it inherits whatever that cluster can reach, which is frequently the warehouse, the lake, and every table the service principal was granted.
The second is detectability. A package that opens a reverse shell is loud. A package that reads a table and quietly posts it somewhere else is not, because reading that table is exactly what the job is supposed to do. Nothing in your monitoring flags a job for doing its job, so the exfiltration can run for months.
Most teams find out which of these applies to them during an incident review. The cheaper version is an afternoon spent listing three things: which clusters can reach the public internet, which job definitions pin their dependencies, and which service principal is attached to the notebook somebody is using to explore an idea right now.