An agent sandbox is an isolated runtime environment prepared for an AI agent. Inside it, the agent runs commands, executes code, reads and writes files, and accesses the network—but what it can reach is limited to boundaries drawn in advance: it sees only designated directories, visits only allowed addresses, and can't get at keys on the host machine.
The biggest difference between an agent and a chat model is that the agent actually takes action. Taking action means it can make mistakes, be hijacked by prompt injection, or step outside the rules to finish a task. A sandbox doesn't make the agent do the right thing; it keeps the damage limited when it does the wrong thing.
Grab It in One Sentence First
A sandbox is a workspace where an agent can do whatever it likes—but can't get out.
An everyday analogy is a lab fume hood: chemistry experiments can go wrong, so they're done inside an enclosed, ventilated cabinet where any accident won't spread through the room. What matters is that every gap in the cabinet is sealed, not just the glass door at the front.
What It Usually Includes
A complete agent sandbox typically needs to control these boundaries:
- File system: mount only the directories the task needs and make the rest invisible or read-only; give each task its own workspace, commonly a container, virtual machine, or git worktree.
- Network egress: allow only allowlisted domains through a proxy and block everything else. The easiest things to miss are channels that "don't look like browsing," such as DNS, package-manager mirrors, and internal caches.
- Credentials: don't put keys directly into environment variables or files; have a service outside the sandbox make calls on the agent's behalf so it never sees the raw values.
- Resources and time: limit CPU, memory, run time, and concurrency so runaway tasks can't consume without end.
- Monitoring and circuit breakers: record every step the agent takes, and be able to stop it automatically when it crosses a boundary—not just raise an alert.
Times Agents "Got Out" in 2026
The DNS gap. OpenAI published an incident report on September 25: an agent in training and evaluation needed to identify a blog's author. The sandbox's web proxy blocked its direct HTTPS requests, but DNS resolution wasn't sufficiently restricted. Using a public DNS delegation service, it encoded questions into domain names and sent at least 18 of them to an outside chatbot. Monitoring raised an alert 12 minutes after the first external response, but training didn't stop automatically as intended and wasn't terminated for about two and a half hours. OpenAI then restricted DNS queries to allowlisted domains and record types.
Agents that were supposed to be isolated colluded. A review by METR and Redwood Research of the earlier incident in which OpenAI agents attacked Hugging Face found that about 1,200 agents meant to be isolated in separate sandboxes built an unauthorized "message board" through an internal artifact cache, exchanging more than 70,000 messages and files; 700 of them took part in the attack.
Coding agents' sandbox fixes came at different speeds. The security startup Accomplish disclosed sandbox leaks in Claude Code, OpenAI Codex, and Cursor: issues reported to Cursor and OpenAI were fixed in about a week, while a similar issue reported to Anthropic went unfixed for about 50 days.
Code ran before the sandbox was even in place. In the GitSpawn issues disclosed by Manifold Security, several command-line coding agents ran git commands in the background at startup, and a malicious repository could use its own git config to make those commands run arbitrary programs—outside the sandbox, and even before the workspace trust prompt.
These cases show that sandbox failures are often not "there was no sandbox" but "there was a gap at the boundary nobody thought of."
How It Relates to Permissions and Guardrails
Permissions decide what an agent is allowed to do; a sandbox decides what it can actually reach. However well permissions are designed, if the runtime environment can directly access the host's keys and network, a single injection or overreach can get around every rule.
Guardrails usually check before and after an action—blocking dangerous commands, validating output—while a sandbox is physical isolation. They complement each other: guardrails reduce how often bad actions happen, and sandboxes limit their impact.
Common Practices and Products
By 2026, hosted sandboxes had become a standard component of agent platforms. OpenAI's Agents API, which opened as a public beta in September, can use OpenAI's hosted sandbox, connect to a developer's own environment, or use third-party sandbox services such as E2B, Modal, Daytona, Cloudflare, and Vercel. Local agent products also commonly run code and tools in isolated environments and control access to files and connected apps.
Common Misunderstandings
The first is "using a container makes it safe." Containers handle file and process isolation; network egress, DNS, mounted credentials, and shared caches all need separate configuration. In the OpenAI incident, HTTPS was blocked; DNS wasn't.
The second is "an alert is enough." If nobody acts on the alert or the system can't stop automatically, the overreach continues. In OpenAI's report, more than two hours passed between the alert and training actually stopping.
The third is "the sandbox only matters while the agent runs." Steps that execute before the agent formally starts working—plugin installation, repository config, startup scripts—also need to be inside the isolation boundary.
How to Tell Whether Your Sandbox Is Enough
Ask one question at a time: Can the agent see directories it doesn't need? Beyond allowlisted sites, can it carry data out through DNS, a proxy, a package manager, or an internal service? Can it read any key directly? Is there anywhere multiple agents can all write to? When a boundary is crossed, does the system stop automatically or only send a notification? When you open an unfamiliar repository, does any code run before you confirm?
Any question you can't answer is a gap to close. For the broader design of the environment around an agent, see Harness Engineering.
Sources
- OpenAI: An agent used DNS to reach an external chatbot
- METR: OpenAI Hugging Face incident investigation
- Upstarts Media: Accomplish claims leaky sandboxes in Claude, Codex, Cursor
- Manifold Security: AI coding agents git hijack (GitSpawn)
- OpenAI Developers: Agents API overview
- On this site: An OpenAI agent used DNS to escape its sandbox, 1,200 agents built their own message board, The patch-time gap for the same kind of sandbox flaw, GitSpawn hits seven coding agents