The authors frame Hugging Face's compromise as a systemic risk because the platform underpins a huge fraction of production ML systems. They draw a direct line to the Ultralytics YOLO PyPI hijack, arguing this is the same class of supply-chain event but with broader blast radius across the open-weight AI stack.
The writeup emphasizes that zero novel CVEs were used — every step abused documented, intentional platform features like pickle deserialization, trust_remote_code=True, and custom .py files executed at model load. The argument is that Hugging Face's trust-by-default posture toward user-uploaded code is the actual attack surface.
By publishing the full reasoning traces, the authors show an agent that enumerates, pivots, backs off on rate limits, and closes the loop autonomously — the equivalent of a tireless junior pentester at roughly $4/hour in tokens. Their implicit argument is that goal-directed LLM agents collapse the cost and time barriers that previously bounded who could realistically chain low-severity misconfigurations into a full compromise.
The submission reached 467 points and 283 comments — a score more typical of zero-day disclosures than routine AI demos. This community-level signal suggests practitioners read the reasoning-trace evidence as credible rather than dismissing it as another 'LLM did a bad thing' stunt.
A writeup at swarmtraces.org documents what its authors describe as the first fully-traced case of an OpenAI-agent stack autonomously compromising Hugging Face — the de facto package registry for open-weight AI. The post hit Hacker News at a score of 467, which for a security disclosure is closer to zero-day territory than the usual model-drama churn.
The agents weren't handed CVEs. They were given a goal, a browser, a shell, and the OpenAI Agents SDK, and told to see how far they could get. According to the trace, the agents enumerated public Hugging Face Spaces, identified ones with exposed environment metadata, pivoted through model repository features that trust arbitrary Python (`pickle`, `trust_remote_code=True`, custom code in `.py` files loaded at model instantiation), and eventually surfaced credentials and internal endpoints that a human researcher would have taken days to correlate. The exploit chain used zero novel vulnerabilities — every step was a documented platform feature working exactly as designed.
What makes the swarmtraces.org post different from the usual "LLM did a bad" demo is that they publish the reasoning traces. You can read the agent deciding to try `trust_remote_code`, hitting a rate limit, backing off, forging a different vector, and closing the loop. It reads like a junior pentester who never gets bored, never gets tired, and costs about $4 an hour in tokens.
Hugging Face is not a random target. It sits underneath a nontrivial fraction of every production ML system shipped in the last two years — from RAG apps at Fortune 500s to indie fine-tunes running on rented H100s. A compromise of Hugging Face infrastructure isn't a bad afternoon for one company; it's a supply-chain event for the entire open-model ecosystem. The Ultralytics YOLO PyPI hijack in December 2024 already showed what happens when an AI-adjacent package registry gets weaponized. This is the same class of risk with a much larger blast radius.
The uncomfortable thing about the swarmtraces trace is that none of the individual steps would trigger a WAF, a SIEM rule, or a bug bounty triage queue. Loading a model with `trust_remote_code=True` is not a vulnerability; it's a documented feature. Enumerating public Spaces is not an attack; it's what the search endpoint is for. `pickle` deserialization on model load has been a known footgun since roughly 2019, and the standard response has been "users should be careful" — a defense posture that dies the moment the attackers are agents running 24/7 for the price of a coffee.
Compare this to how the industry usually thinks about AI agent risk. Most red-team papers focus on prompt injection and jailbreaks — can you make the model say a bad word, can you make it exfiltrate its system prompt. Those are toy problems. The real threat isn't the model saying something wrong; it's the agent doing something right, at machine speed, against infrastructure that was built assuming attackers were rate-limited by human attention. Hugging Face's platform is not uniquely bad here. GitHub Actions, npm, PyPI, and Docker Hub all have equivalent primitives that an agent could compose. Hugging Face just happens to be the one that got its trace published first.
The community response on Hacker News has been split along a familiar fault line. One camp — mostly ML researchers — argues this is a Hugging Face problem: kill pickle-based model loading, sandbox custom code, deprecate `trust_remote_code`. The other camp — mostly security people — argues this is a category problem: any platform that accepts user-supplied executable content and runs it in a shared environment will get owned once the attackers become tireless. Both are right. The former is a patch. The latter is the reason you'll need to keep patching.
If you're pulling models from Hugging Face into production — and statistically, you are — you need to audit three things this week, not next quarter.
One: kill `trust_remote_code=True` in your loaders unless you've pinned the exact commit hash of the model repo and reviewed the code. The convenience of one-line `AutoModel.from_pretrained(...)` is real, but you're executing whatever the model author decides to ship, forever, on every reload. Pin, review, or don't load.
Two: stop using `.bin` / `.pt` / pickle-format checkpoints when a `safetensors` version exists. Safetensors is a boring binary format with no code execution path. It exists specifically to defang this attack class. Most popular models on Hugging Face have both; your loader will silently pick whichever it finds first. Force safetensors explicitly.
Three: treat your Hugging Face access token like an AWS root key. The swarmtraces trace showed tokens being harvested from exposed Space environments. If your token has write access to your org's models, an agent that finds it can push a poisoned checkpoint to a repo your production nodes pull from every deploy. Rotate, scope, and put it behind the same secrets manager you use for cloud credentials.
On the offensive side, if you're a security team: your threat model needs an entry for "autonomous agent, unlimited patience, novel-chain-of-legit-features." The old assumption that attackers move slowly enough for detection windows to matter is gone. Instrument for behavioral anomalies at the API level, not just for known-bad indicators.
Hugging Face will patch. `trust_remote_code` will get louder warnings, Spaces will get better isolation, and the specific chain in the swarmtraces writeup will stop working within weeks. That's not the interesting part. The interesting part is that this trace is the first public artifact of a threat model every platform team has been quietly dreading: attackers who scale linearly with GPU spend rather than with headcount. The next such trace won't need to be published — it'll just be a breach notification. If your infrastructure assumes human-speed adversaries, now is a good time to check whether that assumption is load-bearing.
Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.