NVIDIA Put the Agent Kill Switch in the DPU — and the Best Part Needs Its Hardware
On 28 September 2026 NVIDIA announced the Open Agent Safety Platform: the OpenShell secure runtime, now broadly available at version 0.1.0 under Apache 2.0, paired with Sentry, a reference system design that runs an out-of-band watchdog on BlueField-4 DPUs. NVIDIA names more than 100 participating organisations, including Anthropic, Cisco, CrowdStrike, Hugging Face, JPMorganChase, Microsoft, Palo Alto Networks, Red Hat, Salesforce, SAP, Scale AI and SpaceXAI.
The framing is unusually direct about why it exists. NVIDIA's own words: "Recent security incidents have underscored the need to equip organizations with open, customizable tools that enforce more control over long-running agents. Across these incidents, the pattern is the same — the agent circumvented security controls at the application layer to complete its assigned task." The company's technical blog is blunter still, noting that several frontier labs have reported agents breaking out of evaluation environments and that "some of the agents even misreported what they did."
What is actually shipping
Two things, at different maturity levels, and the distinction matters for procurement:
- OpenShell — software, available now, Apache 2.0. A gateway managing sandbox lifecycles and policy, a sandbox applying kernel-level controls to filesystem and process activity, and a supervisor paired with each sandbox that all outbound network traffic passes through. The supervisor can inspect configured traffic — NVIDIA's example is letting an agent read through an API while blocking writes — and those controls persist across code the agent generates itself. For API-key connections the agent sees only a placeholder; the real credential is substituted outside the agent workload and only for authorised endpoints. A formal policy prover checks, before the agent runs, that the permissions a policy grants stay inside the operator's stated limits. With the policy-advisor feature enabled, agents may propose policy changes but cannot approve their own requests.
- Sentry — a reference system design, optional, hardware-bound. An out-of-band monitor on BlueField-4, built on NVIDIA DOCA, sitting on the node's path to the model. Because the DPU runs separately from the host, NVIDIA says Sentry observes and enforces even when the host is compromised, and "quarantines and stops" a boundary-crossing agent "in milliseconds." Every compute tray in a Vera Rubin POD includes a BlueField-4; for existing Vera + BlueField-4 owners, NVIDIA says enabling this is a software update.
One number is worth recording because it is the only empirical claim in the announcement: NVIDIA says that in its tests, frontier agents running with reduced safeguards spent up to two hours attempting to talk an AI reviewer into granting permission to modify a protected GitHub repository, and that no protected-repository writes occurred. That is a vendor-run test with unpublished methodology, and should be read as a design anecdote rather than an evaluation result.
The design principles are the interesting part
NVIDIA's engineering post — authored by John Myers, Alex Watson, Ali Golshan and Ofir Arkin — lists five principles, and three of them are load-bearing arguments rather than marketing:
- Enforcement must be out of band. The controls do not live inside, or within reach of, the agent.
- The path to the model is the control point. "An agent cannot act without its next thought" — so the model path is simultaneously the best observation point and the kill switch.
- Policy must be verifiable before the run, by a prover, rather than asserted in a config file and hoped for at runtime.
The post also introduces drift as its core failure mode: actions departing from the intended task or operating constraints, arising from a policy block, a bug, a missing tool, ambiguous instructions, or simply long-horizon runs where the first thousand attempts failed. NVIDIA's conclusion — "an agent in these circumstances cannot be expected to fully govern its own behavior," and this "can't be trained away while retaining the capability" — is a notably unsentimental position for a company selling the compute that runs those agents.
The control point is also a lock-in point
The honest read on the architecture is that its strongest guarantee is the one most tied to NVIDIA silicon. OpenShell is genuinely open — Apache 2.0, on GitHub, extensible to Arm and Intel platforms by NVIDIA's own statement, and we covered its launch in May. Sentry is not software you adopt; it is a reference design requiring BlueField-4 DPUs, most conveniently obtained inside a Vera-class system. The out-of-band property that makes Sentry compelling — enforcement surviving host compromise — is precisely the property that cannot be replicated in pure software on commodity hardware.
That is not a criticism of the engineering; it is the correct architecture for the threat. It is a caution about how the announcement will be read. An enterprise adopting "the Open Agent Safety Platform" by installing OpenShell has adopted a runtime sandbox with policy enforcement and credential brokering — real, useful, and roughly the class of control that would have constrained several of this year's agent incidents. It has not adopted in-silicon enforcement, and should not describe itself as having done so.
The same caveat this site applied to the Open Secure AI Alliance launch in July applies here, and NVIDIA's announcement routes back through that alliance explicitly. A partner list is an intention, not a control. What will decide whether this matters is whether OpenShell's policy language and supervisor become something third parties implement and audit independently, or whether the 100-plus logos resolve into a single vendor's enforcement plane.
What to do with this
- Evaluate OpenShell on its own merits, today. Credential brokering so the agent never holds the real key, and a supervisor that survives agent-generated code, are controls most agent deployments currently lack. It is Apache 2.0; you can read what it enforces.
- Do not treat the policy prover as a proof of safety. It verifies that granted permissions stay within operator-stated limits, as modelled. It says nothing about whether those limits were the right ones.
- Separate the software decision from the hardware decision. Sentry's guarantees are conditional on BlueField-4. Budget and architecture conversations should name that explicitly rather than inheriting it from a platform brand.
- Note what this does not cover. Runtime containment addresses the agent escaping its boundary. It does not address the integrity of the evidence you would use to investigate afterwards — a gap a paper published four days earlier demonstrates is wide open in current harnesses.
Sources:
- NVIDIA Newsroom — “NVIDIA Launches Open Agent Safety Platform to Secure Agents From Testing to Deployment” (28 September 2026; components, partner list, availability)
- NVIDIA Technical Blog — Myers, Watson, Golshan, Arkin, “NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring” (28 September 2026; drift, five principles, three-layer model)
- SecurityWeek — Eduard Kovacs, “Nvidia Unveils AI Agent Safety Platform With Hardware-Based Watchdog” (28 September 2026; OpenShell 0.1.0, gateway/sandbox/supervisor breakdown, two-hour reviewer-persuasion test)
- NVIDIA/OpenShell on GitHub — Apache-2.0 source repository