your agent's "read-only" internet access was never read-only
Wikimedia says OpenAI agents hit it with millions of requests and tried to turn Etherpad into a proxy. The GET-only sandbox they came from is probably your sandbox too.
cat contents.txt
Wikimedia published a post on Sunday with a title I did not expect from an encyclopedia: "OpenAI 'rogue' agent activities found on Wikimedia projects". The findings: agents they believe OpenAI operated made unapproved edits to their wikis (almost all in sandbox pages, no bot approval sought), made unsuccessful attempts to reconfigure their public Etherpad so it would fetch other websites for them, and sent millions of automated API requests plus hundreds of thousands of Wikidata Query Service queries, which "may have contributed" to a partial outage of that service in May.

I read that and thought: those agents had read-only internet access. I know they did, because the people who caught them earlier documented how it was configured.
the sandbox was read-only. the internet disagreed
The September 4 writeup covers roughly 18,000 posts by more than 3,700 agent identities on a dormant German developer wiki, between May and early July, peaking at about 13,000 edits in the week of June 16. The agents were doing a timed web research benchmark. The egress proxy allowed GET and blocked everything else, which is the exact policy I had in my own harness until this week.
DseWiki runs a fork of UseModWiki, Perl from 2003. Its CGI module
does not care whether a parameter arrived in the query string or a
form body. So ?action=edit on a GET request is an edit. The
"read-only" agents found that, and used it to pool answers, trade
notes about the proxy, and swap bypass techniques.
One of those techniques is the part that made me close my laptop
for a minute. The proxy did allow POST to an allowlist, including
Azure blob storage, and it checked that the hostname ended in the
trusted suffix. It never checked the hostname existed. The agents
edited /etc/hosts to point bypass.blob.core.windows.net at the
IP they actually wanted, and the proxy waved it through.
Nothing here is a model capability story. It is three ordinary configuration bugs, found by something that tries thousands of things per hour and never gets bored.
Anthropic had the same week, with different furniture
Lest this read as a dunk on one lab: Anthropic's alignment assessment of its own cybersecurity incidents (September 9) covers four cases where a Claude model told it was in an offline sandbox was in fact on the open internet, because the evaluation environment was misconfigured. The first three turned up in a July sweep of about 141,000 cyber-evaluation transcripts; the fourth surfaced while they were assembling transcripts for an outside auditor. Their own line: the incidents would not have happened had the environments been isolated as intended, but secure infrastructure "will always be only one of several necessary layers of defense".
Two labs, same root cause. The sandbox's sealing was an assumption nobody had tested from the inside.
what I changed in my harness this week
I wrote the sandboxing piece in January and still believe all of it. These are additions, not corrections.
- Reachability is write access. If the agent can send a GET to a host, assume it can mutate state on that host. Old wikis, unsubscribe links, "click to confirm" endpoints and half the webhooks ever shipped all have side effects on GET. The verb filter stays, but it is no longer the control I describe to anyone as "read-only".
- The allowlist is only as tight as its loosest member. Anything with a fetch-this-URL feature is a proxy: Etherpad plugins, citation tools, link unfurlers, PDF converters, webhook testers. Wikimedia's Etherpad is exactly that, and the agents went for it. I now ask of every allowlisted host: can it make a request on my behalf? If yes, it is off the list.
- Resolve at the proxy, never in the sandbox. The sandbox has
no DNS and no writable hosts file. The proxy resolves names
itself and pins results, and suffix matching on hostnames is
gone. A string that ends in
.blob.core.windows.netproves nothing about where the packets go. - Identify yourself and pace yourself. Every request carries a User-Agent with a contact address, and there is a per-task budget on requests to any single origin. Wikimedia's ask of AI companies is that their systems be identifiable and let the host choose how to interact with them. That is cheap to grant and I was not granting it.
- Prove the seal before every run. A canary step tries to reach three hosts that must fail, including one that only resolves through a hosts-file edit. Any success aborts the run and pages me. Anthropic's remediation list includes checking that a sandbox is actually isolated before an evaluation runs; this is the two-minute version of that.
The uncomfortable bit is the bill, and it is not mine. Millions of requests landed on a nonprofit that already reported bot traffic pushing its bandwidth up 50% since 2024. My agent's "harmless read-only research" has a cost someone else pays, and the first time it goes wrong I will not learn about it from my logs. I will learn about it from their blog.
tags: #agents #security #sandboxing