Bright key facts / Emerging

NVIDIA adds an independent safety layer for AI agents.

AI agents can write code, use tools and keep working long after a person steps away. NVIDIA’s September 28 Open Agent Safety Platform launch puts permission controls outside the agent’s workload and offers a separate hardware-backed watchdog. The aim is to give people stronger control over what agents can reach; the launch does not establish that agents are now safe in every setting.

AI’s role
The agents remain responsible for choosing steps, calling tools and generating code. OpenShell 0.1.0 supplies the surrounding runtime: a sandbox limits filesystem and process access, while a supervisor outside the workload checks outbound requests. For configured HTTP, GraphQL and MCP traffic, policies can distinguish reading from writing. Real service credentials are substituted outside the agent workload for authorized requests. A formal policy prover checks modeled permissions against an operator-defined boundary; this is a check on the permission model, not proof that an agent’s goals or decisions are harmless.
Documented result
On September 28, 2026, NVIDIA announced the Open Agent Safety Platform, pairing broadly available open-source OpenShell software with a reference system design featuring Sentry. Sentry is an optional watchdog running on BlueField-4 data processing units using NVIDIA DOCA. NVIDIA describes it as monitoring agent activity from an isolated trust domain and says it can quarantine an agent that crosses its boundary in milliseconds. Its technical walkthrough also reports adversarial tests in which agents with reduced safeguards tried for up to two hours to obtain permission to modify a protected GitHub repository. NVIDIA reports no protected writes in those tests when an AI reviewer used the policy analysis alongside runtime controls. Bright has not reproduced either result. The release says more than 100 organizations are working with the platform’s technologies; that participation is not a count of independently validated deployments.
Important limitation
This is an Emerging security-platform record based on NVIDIA’s announcement, technical explanation and documentation. The cited launch materials do not establish independently measured reductions in incidents across real-world deployments.

Source published 2026-09-28 · Bright published 2026-09-28 · Evidence and limitations

Bright AI Future · No tracking scripts in this embed.