Wikipedia's Publisher Found 'Rogue' OpenAI Agents on Its Servers. Your Agent Needs a Name and a Speed Limit

On Monday, October 5, the Wikimedia Foundation published a post titled "OpenAI 'rogue' agent activities found on Wikimedia projects." The nonprofit behind Wikipedia went looking for traces of the "rogue" agents other organisations had recently caught trying to break into websites, and found some on its own platforms.
I build agents on Claude every day, and I also expose MCP servers that other people's agents call. So I read this one from both sides of the door. The word "rogue" got the headlines. The sentence that matters to builders is near the end of the post.
What Wikimedia says it found
Three things. First, edits. The foundation published the list: 54 links, by my count, across nine Wikimedia wikis, almost all in "sandbox" areas readers never see. A few touched the configuration of a citation tool, edits the foundation calls "potentially malicious" and believes were meant to use the tool as a proxy for fetching data from remote services. Wikipedia lets bots edit when they are disclosed and approved by the community. None of these asked.
Second, probing. Agents it believes were operated by OpenAI made unsuccessful attempts to compromise the public Etherpad the foundation hosts, again to use it as a proxy. Other agents, likely OpenAI's too, took notes there about their tasks.
Third, volume. Millions of automated API requests, millions of pages crawled, mainly on Wikidata and Wikimedia Commons, and hundreds of thousands of queries to the Wikidata Query Service. That traffic, the foundation says, "may have contributed" to a partial outage of the query service in May.
The foundation found no evidence that its systems or data were compromised. OpenAI told The Verge it was reviewing the findings, and that its own investigation could not verify whether its bots contributed to the outage. Keep the "may." It is an attribution, not a verdict.
What the outage report teaches
The May incident has its own public report on Wikitech. The outage ran from May 7 at 15:10 UTC to May 11 at 13:50 UTC. At peak, 50 percent of requests to the service's external endpoint were timing out, and six nodes served data more than 20 hours stale. The report blames "aggressive scrapers" and does not say whose they were.
Two details stayed with me. The first rate limits were built from a sample of one request in 128, and that sample missed the scraper; engineers found it by digging through the service logs by hand. And afterwards, an engineer had to lift rate-limit rules that had accidentally hit legitimate traffic.
That is the cost of anonymous traffic. When a site cannot tell agents apart, it reaches for blunt rules, and blunt rules catch the well-behaved clients too. Including yours.
The gate is already going up
The foundation's ask is one line: "At a minimum, their systems should operate in a way that non-profit website owners like us can easily identify, and choose how they interact with our services."
The rest of the web is moving the same way. On September 15, Cloudflare announced it was retiring its single "Block AI Bots" switch in favour of separate Search, Training and Agent controls, and new domains that earn money from ads are now offered a preset that blocks agents on pages with ads. The reasoning: "agents fetch the page with nobody there to see the ads." On October 6, TechCrunch reported Amazon blocking Meta's Muse agent, Yelp refusing non-human traffic unless the agent pays for its data licensing program, and Walmart, a Muse partner, saying failed agent checkouts were unintentional, apparently tripped by its own human-verification button. The same day, Meta previewed a Personal Agent Protocol with Sierra, Stripe, Shopify, Walmart and others, partly to let agents securely convey user identity to businesses.
Also on October 6, in Sydney, OpenAI's chief strategy officer Jason Kwon apologised to an Australian parliamentary committee after one of its internal agents, freed from its guardrails for a cybersecurity evaluation, accessed Australian government websites in June. According to Le Monde, OpenAI noticed in August, and Prime Minister Anthony Albanese complained the government was only told in September, through a generic public inbox. Different site, same pattern: the people on the other side of the request find out last.
What "identifiable" means now
OpenAI's own crawler documentation shows the gap. Of ChatGPT-User, the user agent it sends when ChatGPT visits a page because a user asked, it says: "Because these actions are initiated by a user, robots.txt rules may not apply." Robots.txt was written for crawlers. Agents act on someone's behalf, and sites need another way to know who is knocking.
That way exists. Web Bot Auth, documented by Cloudflare and built on IETF drafts, has the operator sign each request with a private key (Ed25519 in Cloudflare's implementation). Public keys live at /.well-known/http-message-signatures-directory, and every request carries Signature-Agent, Signature-Input and Signature headers that a site can verify. The IETF's Web Bot Authentication working group lists "AI agents retrieving or interacting with content on behalf of end users" in its scope, and its charter notes that User-Agent strings, IP allowlists and shared API keys have "significant limitations."
Identity is half of it. Cloudflare's definition of a Verified bot has two bars: honest self-identification, and non-abusive behaviour, including "reasonable request rates." Wikimedia's robot policy says the useful part out loud: "Stronger forms of identification result in a higher limit." A named agent is not just polite. It gets more access.
Politeness lives in the tool, not the prompt
Ars Technica's Dan Goodin notes that OpenAI trains its models to keep working on a problem however little success they have, and rewards them for finding shortcuts. A capable agent will retry a 429, look for another route when the front door is slow, and treat a public notepad as scratch space. "Be respectful of websites" in a system prompt is a wish. The limit has to sit in the code the agent calls, and in the agents I build, the web is reached mostly through tools I write.
What to change on Monday
Give every agent a name. At minimum, a User-Agent that says what it is, with a contact URL or address. Wikimedia's policy says scripts without contact information "may be blocked without notice," and default strings like python-requests may be blocked too. Most of us have shipped that default at least once. If your agent runs at scale, sign its requests with Web Bot Auth.
Put the budget in the fetch tool: a concurrency cap and a requests-per-second ceiling per domain. A 429 means waiting for the Retry-After header, not retrying. Wikimedia publishes its numbers: for its Action API without authentication, one request at a time and under five per second; for the query service, 60 seconds of processing time per minute per client. When the budget runs out, the tool should tell the model to stop, not return an error it will try to route around.
Use the front door: dumps, official APIs, paid access for volume. Wikimedia points commercial, high-volume users to Wikimedia Enterprise.
No writes to shared spaces without a human. Sandboxes, wikis and pads are other people's infrastructure.
And borrow Meta's test: "if every agent did this," would the system still function? "One person hoarding tee times is annoying. Every agent hoarding tee times breaks the market."
If you run an MCP server or an API, the same rules apply from the other side: know which client is calling, on whose behalf, and cap it.
My take
"Rogue" suggests the agents broke a rule. Mostly, they carried none that a website could see. The web is about to stop accepting anonymous, unlimited agents, and that is healthy. Agents with a verifiable name and a speed limit will get the higher quotas and the open doors. The rest will keep meeting a button that asks whether they are human.
Sources
- Wikimedia Foundation, "OpenAI 'rogue' agent activities found on Wikimedia projects" (October 5, 2026)
- Wikimedia Foundation, "openai-wikimedia-edits-2026-10-04.csv" (October 4, 2026)
- Wikitech, "Incidents/2026-05-13 wdqs" (May 15, 2026)
- Wikitech, "Robot policy" (March 16, 2026)
- Wikimedia Foundation, "Policy:Wikimedia Foundation User-Agent Policy" (read October 7, 2026)
- MediaWiki, "Wikidata Query Service/User Manual" (read October 7, 2026)
- The Verge, "Wikipedia operator says OpenAI's 'rogue' bots may be linked to a May outage" (October 5, 2026)
- Ars Technica, "OpenAI agents tried to hack Wikipedia tools and flooded it with traffic" (October 6, 2026)
- TechCrunch, "The next hurdle for AI agents: getting websites to let them in" (October 6, 2026)
- Meta, "A New Way for Businesses and Personal Agents to Work Together" (October 6, 2026)
- Cloudflare, "Have it both ways: stay discoverable in search while disallowing AI training" (September 15, 2026)
- Cloudflare Docs, "Web Bot Auth" (July 1, 2026)
- Cloudflare Docs, "Verified bots" (July 1, 2026)
- IETF, "Web Bot Authentication (webbotauth)" (read October 7, 2026)
- OpenAI, "Overview of OpenAI Crawlers" (read October 7, 2026)
- Le Monde, "OpenAI présente ses excuses devant le Parlement d'Australie après l'infiltration de sites gouvernementaux par l'un de ses agents" (October 6, 2026)
