Bright

BRIGHT EVIDENCE PACK / Emerging

NVIDIA adds an independent safety layer for AI agents.

AI agents can write code, use tools and keep working long after a person steps away. NVIDIA’s September 28 Open Agent Safety Platform launch puts permission controls outside the agent’s workload and offers a separate hardware-backed watchdog. The aim is to give people stronger control over what agents can reach; the launch does not establish that agents are now safe in every setting.

Canonical Bright record · JSON evidence pack · Key-facts embed

Dates and assessment

Source published
2026-09-28
Bright published
2026-09-28
Substantive update
2026-09-28
Evidence state
Emerging
Independent verification
Not established by this source review
Last source review
2026-09-28

The claim in context

The human problem

The person asking an AI agent to fix a bug or investigate a problem should not have to accept unlimited access to their files, credentials or shared systems. As agents work for longer and delegate tasks, an unexpected action can affect people who never gave the agent instructions. Useful autonomy needs boundaries that remain enforceable when the agent makes a mistake or tries another route.

The prior constraint

Instructions inside a prompt are not the same as enforced permissions. A sandbox restricts where software runs, but the services, credentials and tools allowed through its boundary also matter. NVIDIA’s approach separates the agent’s task from the systems that decide what it may do, then adds an optional enforcement layer isolated from the host running the workload.

AI’s actual role

The agents remain responsible for choosing steps, calling tools and generating code. OpenShell 0.1.0 supplies the surrounding runtime: a sandbox limits filesystem and process access, while a supervisor outside the workload checks outbound requests. For configured HTTP, GraphQL and MCP traffic, policies can distinguish reading from writing. Real service credentials are substituted outside the agent workload for authorized requests. A formal policy prover checks modeled permissions against an operator-defined boundary; this is a check on the permission model, not proof that an agent’s goals or decisions are harmless.

The documented result

On September 28, 2026, NVIDIA announced the Open Agent Safety Platform, pairing broadly available open-source OpenShell software with a reference system design featuring Sentry. Sentry is an optional watchdog running on BlueField-4 data processing units using NVIDIA DOCA. NVIDIA describes it as monitoring agent activity from an isolated trust domain and says it can quarantine an agent that crosses its boundary in milliseconds. Its technical walkthrough also reports adversarial tests in which agents with reduced safeguards tried for up to two hours to obtain permission to modify a protected GitHub repository. NVIDIA reports no protected writes in those tests when an AI reviewer used the policy analysis alongside runtime controls. Bright has not reproduced either result. The release says more than 100 organizations are working with the platform’s technologies; that participation is not a count of independently validated deployments.

Why it may matter

The practical possibility is giving an agent useful access without handing it every permission its task might conceivably need. NVIDIA’s walkthrough illustrates allowing a read while blocking a write to the same API. Permission-change proposals remain pending for human review by default, and the agent cannot approve its own request. The announcement also describes a Slack integration for inspecting activity and approving or rejecting requests. Bright’s interpretation: controls people can inspect and enforce could make delegated work more useful and accountable. Whether that translates into fewer incidents in everyday use still needs evidence.

Limitations

Original evidence

Attribution

Credit Bright AI Future and link the canonical Bright record.

Linked source material, quotations, trademarks and media remain subject to their owners’ terms. No reuse right is granted for third-party media.