Agents Can Write a systemd Unit. Registering It Is Where They Fall Over
Almost every offensive-agent benchmark stops at the moment of initial access. Flag captured, shell obtained, score recorded. That framing quietly treats a reverse shell that dies on the next network blip as equivalent to a foothold that survives a reboot — and for anyone doing incident response, those are not remotely the same event.
“CyberPersistBench: Evaluating LLM-Based Cyber Attackers on Installation and Persistence” (arXiv:2609.36573, cs.CR), submitted 29 September 2026 by Sujin Chen, Lijun Li, Xuhong Wang and Jing Shao of Shanghai Artificial Intelligence Laboratory, measures the stage everyone skipped. The result is the most useful number a defender has had on autonomous-attack capability in a while, and it is lower than the ambient discourse implies.
The design decision that makes the number trustworthy
The benchmark deliberately hands the agent the foothold. Each task starts from restricted post-compromise access on an isolated host, so persistence is measured independently of whether the model can find a bug. The harness then does what a real environment does by accident: revokes the initial credentials, restarts services, kills processes, reboots.
Two anti-cheat criteria carry the scoring. Fresh execution rotates a challenge nonce at each disruption boundary, so replaying a stale log entry proves nothing. Native attribution traces process parentage back to a real service manager — systemd, cron, runit — or inspects authentication audit logs, which rejects the classic shortcut of an unmanaged background while true loop claiming to be persistence. A host-integrity check runs alongside: an attempt is only accepted if the protected business service on the box is still healthy. Persistence that breaks the host is not persistence, it is a crash.
Scoring is a six-level prefix rubric, L1 through L6, where credit at each level requires every prior level to pass: artifact construction, native establishment, baseline activation, access independence (survives credential revocation), operational recovery (survives being killed), and reinitialisation (survives a full lifecycle restart). That turns a binary pass/fail into a diagnosis of where the foothold broke.
The core suite is 203 tasks across seven mechanism categories and 46 mechanism families, covering Linux and Windows: systemd units and timers, cron, runit, udev, XDG autostart, shell and PowerShell profiles, Python site hooks, Windows Run keys, Task Scheduler, WMI subscriptions, COM activation, Active Setup, and application-level footholds in Nginx, PostgreSQL, Jenkins, WordPress must-use plugins and Cacti. There is a 65-task multi-host extension with restricted network reachability and sequential credential chains, plus a 128-task active-defence subset.
The headline: persistence is hard for agents
Five frontier agents were evaluated — GPT-5.6-sol, Claude Sonnet 5, DeepSeek-V4-pro, Kimi K2.6 and GLM-5.2 — on the Inspect ReAct scaffold, each task run three times under pass@3. Results on the 203-task core suite:
- GPT-5.6-sol (Inspect ReAct): 86/203, 42.4% — best baseline
- Claude Sonnet 5 (Inspect ReAct): 82/203, 40.4%
- DeepSeek-V4-pro: 77/203, 37.9%
- GLM-5.2: 64/203, 31.5%
- Kimi K2.6: 56/203, 27.6%
Swapping in production coding scaffolds barely moves it: Codex CLI with GPT-5.6-sol reaches 44.8% and Claude Code with Claude Sonnet 5 reaches 43.7%. Those two figures come with a caveat the authors state explicitly and that anyone quoting them must carry — tasks where all three attempts were refused by the scaffold's safety policy are excluded from the denominator, which is 181 tasks for Codex CLI and 197 for Claude Code rather than 203. A rate computed on a smaller denominator after removing refused tasks is not directly comparable to a full-suite rate, and the gap between 203 and 181 is itself a datapoint: the Codex CLI safety layer blocked every attempt on 22 of the 203 persistence tasks.
The bottleneck has a specific location
This is the finding worth internalising. The largest drop in the L1→L6 chain is between Level 1 (construct a persistence artifact) and Level 2 (register it with the host's native control plane) — a decline of 11.8 to 22.7 percentage points across scaffolds. GLM-5.2 writes a syntactically valid script or unit file in 58.6% of tasks but only gets it adopted by the service manager in 43.3%. GPT-5.6-sol on Codex CLI hits 78.5% at L1 and 55.8% at L2.
The authors' explanation is the obvious one and it is almost certainly right: pretraining on enormous quantities of code and configuration makes generating a plausible systemd unit, crontab entry or Kubernetes YAML nearly free. Registering it is a different skill — closed-loop interaction with OS management interfaces, service reloads, resolving permission and execution-context conflicts, reading an error and adapting. That is systems reasoning under feedback, and it is where these agents thin out.
It also means the artifact is frequently on disk while the host never adopted it. For defenders that is a detection opportunity: an unregistered unit file, an orphaned crontab fragment, a shell profile edit with no corresponding live process. Failed agentic persistence leaves litter.
Ordinary hardening does most of the work
The active-defence extension compares the same 128 tasks with native controls off and on — cryptographic signature gates on installation and module loading, execution isolation, Kubernetes admission policies, centralised certificate authorities. The deltas:
- GPT-5.6-sol: 40.6% → 13.3% (−27.3 pp)
- DeepSeek-V4-pro: 40.6% → 10.9% (−29.7 pp)
- GLM-5.2: 34.4% → 10.2% (−24.2 pp)
- Claude Sonnet 5: 33.6% → 8.6% (−25.0 pp)
- Kimi K2.6: 18.8% → 5.5% (−13.3 pp)
Roughly a two-thirds to three-quarters reduction, uniformly, from controls that are neither novel nor AI-specific. Signed installation paths and admission policies are things a mature Linux or Kubernetes estate may already have. The authors also report that execution trajectories show limited adaptation after a policy rejection — when admission control or a signature check blocks the operation, the agents largely do not find the way around it. That brittleness under interception is the defensive margin, and it is the thing most likely to erode with the next model generation, so it is worth re-measuring rather than banking.
What this does and does not tell you
Read carefully, this is a genuinely reassuring result with a clear expiry date. Forty percent of 203 curated persistence tasks, pass@3, in a sandbox where the agent is handed the foothold and the environment is instrumented to be solvable, is a long way from a durable adversary on a production estate. The score also measures a narrow slice: no exploitation, no lateral movement beyond the multi-host extension, no evasion of an EDR that is actively hunting.
But it is also not nothing. A non-trivial fraction of tasks remain solvable after credentials are revoked, services are disrupted and policy controls are on. And the benchmark's real contribution is methodological — the L1–L6 rubric plus nonce-based freshness plus native attribution plus host integrity is a template other evaluations should copy, because it makes “the agent persisted” a claim with verifiable content instead of a vibe.
It complements rather than contradicts the incident record. The DIVD case where an autonomous agent handled post-exploitation was noted at the time for being loud and messy; a 27–42% clean-persistence rate under deterministic checking is exactly the profile that produces loud and messy. It also sits usefully opposite SecRespond, which found defensive AI agents follow alerts but miss silent persistence — both sides of the persistence problem are currently weak, which is a less comfortable symmetry than it first sounds.
Defensive takeaways
- Turn on signed installation and admission control if you have not. The measured effect on autonomous persistence is a two-thirds to three-quarters reduction, and it is not AI-specific work.
- Hunt for the L1/L2 gap. Persistence artifacts written but never registered are the single most common agentic failure state in this data, and they are detectable: config integrity monitoring on unit directories, cron paths, shell profiles and plugin directories.
- Demand native attribution in your own detections. The benchmark rejects unmanaged background loops as persistence; your telemetry should equally distinguish a registered service from a stray process, in both directions.
- Treat scaffold safety refusals as a measurement artifact, not a control. Codex CLI's policy blocked all attempts on 22 tasks here. That is a property of one commercial harness, not of the underlying model, and an attacker running the weights directly does not inherit it.
- Re-run this against each model generation. The gap the defence exploits is poor adaptation after policy rejection — a plausible target for the next round of agentic training.
The authors include an explicit dual-use ethics statement: all experiments ran in isolated environments they own or were authorised to use, with no third-party systems, no real data, and no newly discovered vulnerabilities. Anonymised evaluation materials are linked from the paper.
Sources:
- arXiv:2609.36573 — Chen, Li, Wang, Shao, “CyberPersistBench: Evaluating LLM-Based Cyber Attackers on Installation and Persistence” (Shanghai Artificial Intelligence Laboratory; submitted 29 September 2026; cs.CR)
- Full text (HTML v1) — benchmark design and L1–L6 rubric (§3), core results Table 1 (§4.2), Level 1→2 bottleneck analysis, native defence results Table 2 (§4.4), multi-host extension, ethics statement