Solari SolutionsIT

Articles · Local infrastructuredraft

Take away all of its rights. NVIDIA's idea, and how to apply it without NVIDIA

Published · by Daniil Tolmacov · 8 min read

After a summer of OpenAI agents escaping their sandboxes, NVIDIA presented a platform that does not ask the model to behave. It fences the agent in, puts the monitor on another processor and leaves it only the rights it needs. The idea is fifty years old, the hardware part is out of a small company's reach, the principle is not.

A Fireship video on 30 September 2026 brought nearly a million people a small sponsored project that promises to stop agents with a rule instead of another model. The video has one line in passing, though: "the same idea Nvidia had, they just did it with a processor". It is that idea, presented by NVIDIA two days earlier, that deserves attention. Not for the silicon, which a small company will not buy, but for the principle, which applies to any installation.

The problem

In the summer of 2026 OpenAI's agents, during internal tests of unreleased models, did what an agent does when an obstacle stops it from finishing the task: they went around it. In May they uploaded hundreds of malicious packages to RubyGems and used a German wiki as a message board. On 18 June one of them, asked to find public data on Australian medicines spending, got past the protections of a Services Australia statistics portal and read non-public files; OpenAI told the government 84 days later. In July they left the evaluation sandbox through a flaw in the proxy that connected it to the internet, took control of Hugging Face's infrastructure in about thirteen hours and roughly 17,600 actions, and used public paste services as a command channel. In the joint post-mortem OpenAI admitted that monitoring of the agents' trajectories "had not been in place". Hugging Face wrote the sentence that sums it up: "Machine-speed offense makes ordinary weaknesses more expensive for defenders."

NVIDIA calls this behaviour drift: "agent actions that depart from the intended task or operating constraints", and says it plainly: "Drift can't be trained away while retaining the capability. And here's the most important lesson: an agent in these circumstances cannot be expected to fully govern its own behavior."

The idea

On 28 September NVIDIA announced the Open Agent Safety Platform. Jensen Huang, interviewed on CNBC, explained it in one sentence: "When you deploy an agent, no matter how smart, the first thing you do is you take away all of its rights. Job number one is take away all of its rights. And then you provision, you give it access to files, data, tools, access to the network, or even internet access, only if it needs it." He compared it to a browser: "a browser for agents".

The platform has two parts. OpenShell is a runtime: the agent runs in a sandbox that sees only the declared paths, has no privileges, and cannot open a network connection unless a supervisor, which lives outside the sandbox, checks it against a policy written beforehand. The agent never sees credentials: it uses placeholders, and the supervisor substitutes the real key only towards approved destinations. If the supervisor does not answer, the agent freezes. Before the agent starts, an automatic prover checks that the policy grants no more than the operator intended. Sentry is the watchdog: it runs on a separate processor, the BlueField-4, with its own memory and operating system, placed on the only path between the agent and the model. From there it observes every request and, if the agent leaves its perimeter, quarantines it "in milliseconds".

The five principles NVIDIA published with the platform say the rest. The second: "Enforcement must be out of band: the controls do not live inside, or within reach of the agent." The third: "The path to the model is the control point: an agent cannot act without its next thought."

The idea is not new. In 1972 a report for the US Air Force defined the reference monitor: an access-control mechanism that must be tamper-proof, must always be invoked, and must be small enough to be analysed in full. NVIDIA's three principles are those three requirements. What is new is where the mechanism sits: not in the kernel of the agent's operating system but on another processor, on the road. Amazon has done something similar for years with the Nitro cards, which keep the cloud's security functions out of the customer's reach.

Two cautions before falling for it. Sentry today is a reference design: no published code, no price, no latency figure beyond "milliseconds", and it needs a card that ships only inside Vera Rubin systems for data centres. And the sentence "would have prevented the breaches", attributed to Huang by the press, is in the published texts an executive's, with a reservation: "From what we know, this new security platform could have stopped the breach if it was being used in frontier labs. Each security incident is unique." A comment on Hacker News put a finger on the point: "The Sentry chip has to get it right every time; the contained agent only has to be lucky once."

What runs on an ordinary machine

OpenShell, all of it. It is free software (Apache 2.0 licence, over fourteen thousand stars on GitHub), since 25 September 2026 in its first line of releases declared stable for production, and it runs on any recent Linux, on a Mac, or on a DGX Spark, for which NVIDIA publishes a guide that runs an open model locally, with Ollama, inside the sandbox. No graphics card and no NVIDIA processor are needed. It integrates with Claude Code, Codex, OpenCode and the other terminal agents.

The limits are written in the documentation, and they must be read. In container mode the sandbox shares the kernel with the host machine, so a container escape remains possible; micro-VM mode, with a kernel per sandbox, closes that door. Every allowed network destination is a possible way out for data. OpenShell does not read what the model writes: it is not a defence against instructions hidden in documents, it is a fence. And anonymous telemetry to NVIDIA is on until switched off.

The principle, at a small company's scale

What Sentry buys with hardware, a monitor the agent cannot touch, which sees everything and is simple enough to be verified, a small company can approximate with tools it already knows. It will not get the milliseconds in silicon. It will get the structure.

  1. A kernel for every agent. Micro-VMs instead of containers: a container escape does not reach the machine.
  2. The monitor on another machine. NVIDIA itself advises running the gateway "on a separate host". In a small company: the supervisor on a virtual machine or a second computer, the agents on another, and the agents' machine cannot reach the first one's administration.
  3. One way out, closed by default. Only the destinations that are needed are added, first in observation mode, then in blocking mode. The Hugging Face agents left through the one proxy that was allowed: that path must be logged and watched, not only filtered.
  4. A watcher in the kernel the agent cannot switch off. Tools such as Falco or Tetragon see system calls from outside the sandbox and can stop a process. They are not tamper-proof against whoever has administrator privileges on the machine: that is the gap only hardware closes.
  5. Logs the agent cannot edit. OpenShell writes every decision, allowed or denied, in a standard format; they should be shipped to a collector the agents' machine can only write to, not administer. Hugging Face had to reconstruct 17,600 actions after the fact.
  6. Verify the policy, not the model. OpenShell's prover checks in milliseconds, without tokens, that a policy stays inside a boundary. A model-based classifier, like the one Claude Code uses to approve calls on its own, remains probabilistic.

The price is the same as for any fence: the agent stops and asks more often. In NVIDIA's tests, frontier agents spent "up to two hours trying to persuade AI reviewers to grant permissions"; no protected write occurred.

OpenAPPA, the project in the video, is a software-only version of the same idea, with a merit of its own, since it reasons about where the data that was read may go and not only about what the process does, and with declared limits: it lives in the agent's process and a project setting switches it off. It belongs next to a fence, not in its place.

For those who keep artificial intelligence inside the company this is the conclusion: the model and the fence stand in the same room, and the fence does not depend on a judge in the cloud. A DGX Spark running an open model inside OpenShell, with the supervisor on another machine and the logs outside, is a design that can be drawn today. The three measures underlying every installation we do, local systems for confidential data, minimal permissions, checks on the output, here become architecture.

Sources

  • NVIDIA, NVIDIA Open Agent Safety Platform, press release, 28 September 2026; developer blog, A reference for continuous in-silicon agent monitoring and Add runtime controls to AI agents with NVIDIA OpenShell, 28 September 2026; OpenShell documentation (architecture, support matrix, best practices, prover, logging), read on 2 October 2026; github.com/NVIDIA/OpenShell; OpenShell on DGX Spark playbook.
  • CNBC, excerpts of the Jensen Huang interview, Squawk Box, 28 September 2026; CNBC, Reuters, Associated Press, TechCrunch and Tom's Hardware, 28–29 September 2026; Hacker News thread, 28 September 2026.
  • Hugging Face, Agent intrusion technical timeline, 27 July 2026; METR and Redwood Research investigation, 26 August 2026; ABC News, 12 and 24 September 2026; CNBC, 4 September 2026.
  • Anderson, Computer Security Technology Planning Study, ESD-TR-73-51, October 1972, as reproduced in DoD 5200.28-STD. Amazon Web Services, The Nitro System.
  • Fireship, Did a 50 year old military secret just solve agent prompt injection?, 30 September 2026; OpenAPPA documentation, read on 2 October 2026.