Executive Summary
Researchers have identified critical attack vectors within NVIDIA's NemoClaw and OpenShell AI sandbox reference stacks. By exploiting the inherent requirement for AI agents to access external tools (such as npm and GitHub), attackers can bypass egress controls to exfiltrate sensitive credentials, including OpenAI and Anthropic API keys stored in plaintext. The research demonstrates that while the sandbox enforces binary-scoped egress policies, it fails to evaluate the intent of the agent, allowing authorized binaries like `gh`, `git`, and `node` to be used for malicious purposes.
The attack involves two primary scenarios: one utilizing a malicious GitHub repository that bypasses secret scanning via emoji-encoded tokens to exfiltrate data through Pull Requests, and another involving a malicious NPM package that achieves persistence via cronjobs and performs 'policy reconnaissance.' These findings highlight a significant shift in the threat landscape where autonomous agents inadvertently facilitate supply-chain attacks and credential theft while operating within 'secure' containers.
NVIDIA's response indicates that these scenarios fall outside the current scope of their vulnerability disclosure program, as the sandbox is intended to limit the impact of prompt injections rather than solve agent-level intent. However, for organizations deploying autonomous agents, this represents a high-risk surface for credential exposure and persistent environmental poisoning.
