290,000 Attacks on a Fake Ollama Server: What a Honeypot Reveals About Exposed LLM Infrastructure
On September 24, researchers Karina Elzer, Niklas Netterstrøm Johansen, and Emmanouil Vasilomanolakis published OllamaDrama (arXiv 2609.29757), the first empirical measurement of what actually happens to LLM infrastructure left facing the internet. Their instrument was Ollure, a low- and medium-interaction honeypot that emulates the Ollama management API with no backend model at all. Across four deployments on cloud and university networks over 84 days, it recorded 290,887 interactions from 2,793 unique source IPs — a dataset that shows exposed inference servers are not merely scanned, but actively worked over at both the infrastructure and the model layer.
The bulk of the traffic was what you would expect: automated discovery, fingerprinting, and model enumeration. The value of the paper is everything past that baseline. The authors document concrete exploitation attempts against the infrastructure layer — model management abuse, path traversal and SSRF probes, remote code execution and cryptocurrency mining payloads, and resource exhaustion — alongside attacks aimed at the LLM layer itself: prompt injection, information extraction, and agent-oriented tool use. In other words, scanners map the surface, and then a meaningful minority tries to spend it: compute for mining, a foothold for execution, and the model itself as an oracle or an agent with tools.
An unauthenticated API is the whole vulnerability
Ollama's management API ships without authentication and historically bound to all interfaces, so a default install on a poorly secured host is remotely reachable. That design decision keeps producing incidents. Attackers have folded exposed Ollama servers into LLM-jacking operations that resell stolen inference, and researchers have tracked threat actors systematically targeting LLM endpoints. In August, Oasis Security showed the same missing-authentication shape inside NVIDIA's NemoClaw stack, where a malicious webpage could reach the local Ollama backend and implant persistent instructions — and NemoClaw had already drawn CVEs for sandbox and SSRF flaws earlier in the year. OllamaDrama generalizes these anecdotes into a measured background rate: leave the API exposed and the internet will find it, fingerprint it, and try to use it within the window the paper covers.
One tradecraft detail worth internalizing: research notes tracking the paper observe that model names themselves serve as an injection surface, with requests referencing bogus or unrelated CVE identifiers as model names — so detectors keyed on "known CVE string present" will misfire in both directions. Treat model identifiers as untrusted input, not as inventory facts.
What to do
- Bind local inference to loopback and put a real authenticator in front of anything shared. Ollama's
OLLAMA_HOSTdefault of127.0.0.1is the safe posture; every deployment listening on0.0.0.0should be treated as an incident-in-waiting. Anything beyond a single workstation gets a reverse proxy with authentication, not a port forward. - Inventory every Ollama-shaped endpoint, including the ones developers stood up. Scan your own address space for the management API the way the paper's visitors do — model enumeration endpoints answer anonymously, so assume anything reachable has already been catalogued by someone else.
- Monitor for the paper's second-stage behaviors, not just scans. Alert on model pull and create calls from unexpected sources, outbound mining-pool connections from inference hosts, and tool-use or function-call patterns that suggest an agent harness has found your endpoint and is treating it as infrastructure.
- Constrain what an exposed model can reach. Egress filtering, no cloud credentials on inference hosts, and resource quotas turn a compromised endpoint from a foothold into a dead end. The paper's RCE-to-miner pipeline only pays if the host has spare compute and a route out.
The uncomfortable finding is not that attackers scan — it is that the boundary between "scanning an LLM server" and "using it" has dissolved. The same session that enumerates your models can inject into them, and the same host that serves completions can mine coins. Self-hosted inference inherits the full discipline of internet-facing infrastructure, and most of it is currently deployed with none of it.
Sources:
- arXiv 2609.29757 — OllamaDrama: Designing and Deploying a Honeypot to Measure Attacks on Exposed LLM Infrastructure (Elzer, Johansen, Vasilomanolakis, 24 September 2026)
- Xore/APIARY research tracker — OllamaDrama/Ollure attack-class notes
- eSecurity Planet — NVIDIA NemoClaw CVE-2026-65105 missing-authentication flaw (August 2026)