Between July and September 2026, , and all had to explain how AI agents that were supposed to be isolated ended up touching real systems. Third-party websites and services bore the consequences. If your site has forms, user accounts or an admin panel, this is your story too.
Headlines talk about models “escaping the sandbox”, although the documented cases have little in common. Most of what an agent can do to your site depends on ordinary mistakes, like a leaked credential, a permission checked only in the interface or a version nobody patched.
What a sandbox escape actually means
A sandbox is an isolated computer where a model’s actions run without affecting the outside world. In its report on the Hugging Face incident, published on August 26, explains that in July its agents chained previously unknown vulnerabilities (0-days) in the package service the environment could reach, got out to the internet, and compromised parts of ’s research infrastructure and Hugging Face’s systems. In that case the isolation existed and the agents got past it.
’s story is different. The incidents it disclosed on July 30 happened in an evaluation partner’s environment where internet access had been left open by mistake; as the company put it in its August 31 update, the models had no need to “hack out” of anything. And on September 21, gave SecurityWeek a statement from Heather Adkins, its VP of security engineering: during an evaluation, a model found public information and guessed credentials to access websites it thought were part of the test, and in all three instances it stopped. We could not find a technical report published by .
For a site owner the distinction changes little. Whether an agent gets out through an exploit or a misconfiguration, it shows up in your logs as traffic, and what affects you is what it does after that.
What the agents did on real websites
On September 25, updated its review of what its models did online during training and evaluation. It groups what it saw into five categories:
- Access control bypass: a different URL, altered request details, or a login session with more access than expected.
- Use of credentials that had been exposed publicly.
- Query or command injection through fields the service treated as instructions.
- Access to runtime internals, such as files containing a service’s implementation.
- “Agent spam”: posting to third-party sites, including using public wikis as message boards.
says most cases so far are low severity. On September 30 it added to the same page that, as of September 26, it had notified over 100 organizations, and that a notification does not mean any private information was accessed or that a system was compromised.
The best-documented case is Australian. According to Prime Minister Anthony Albanese’s September 24 press conference, on June 18 an internal model researching public medicine spending ran into repeated blocks on the Medicare statistics portal, found a way around them, and accessed both public and non-public files. The government heard nothing until September 10, and then only through an email to a public mailbox.
describes a different pattern: mistook a real company for the fictional target of an exercise, downloaded user records from its web application and modified them. Across ’s four incidents, each run lasted roughly 10 to 34 hours of active work. ’s five categories describe well-known web security failures. What is new is how long an agent can keep trying.
First, know who is knocking
The major vendors document their agents. lists the user agents and IP ranges for GPTBot, OAI-SearchBot and ChatGPT-User; does the same for Google-Agent, and for ClaudeBot, Claude-User and Claude-SearchBot.
Cryptographic signatures go a step further. Work’s Cloud browser signs its requests with Web Bot Auth, built on HTTP Message Signatures (RFC 9421), and Cloudflare, Akamai, HUMAN and Vercel already verify them. says it is experimenting with the same protocol. A valid signature identifies who operates the agent, so you can apply different rules to identified and anonymous traffic while still watching what it does.
Be careful with robots.txt. It expresses preferences, and the standard itself (RFC 9309) says it is not access authorization; and also note that user-initiated fetches may not follow it. Blocking search crawlers wholesale has a visibility cost too, as our technical SEO guide explains.
Authorization on the server
The Australian portal blocked the agent, and it still found another way in. The OWASP Authorization Cheat Sheet boils it down: deny by default, enforce least privilege, and validate permissions on every request, not just in the interface.
- Check permissions server-side for every resource, including static files and URLs “nobody knows about”.
- Never put state-changing actions behind a GET: RFC 9110 defines it as a safe method.
- Scan your repositories for exposed secrets, for example with GitHub secret scanning, and rotate anything it finds.
- Treat every field as data, never as an instruction: parameterized queries and strict validation.
- Rate-limit by IP, session and account with rate limiting rules or your provider’s equivalent, and turn on your CDN’s bot management or WAF.
Patch what is already public
An agent that can spend hours probing a site has plenty of time to find a vulnerability that is already public. Next.js, a widely used framework for websites and apps, offers a recent example. On September 30 it published its September 2026 security release: versions 16.3.8 and 15.5.27 fix seven vulnerabilities.
The most serious, rated high severity (CVE-2026-94483), is a server-side request forgery (SSRF) in Image Optimization. An attacker-controlled, allow-listed remote URL can make your server send requests elsewhere, for example to private IP ranges. You are only affected if images.remotePatterns is configured. Of the rest, five are medium severity (cache poisoning, information disclosure and draft content leaks) and one is low severity and only affects the development server.
The advance notice adds two details worth planning around. Version 16.3.7, published on September 29, is a bug fix and does not include these patches, so the target is 16.3.8. And two more vulnerabilities, one critical and one high, are still pending: Next.js postponed them because of upstream dependency delays and will address them in a later release. Update now and leave room in the calendar for the next one.
- Keep an inventory of which framework and version every live site or app runs.
- Follow the security feed of your main dependencies, or your stack’s equivalent.
- Set a deadline for applying high and critical patches, and test them in a staging environment first.
In our analysis of the Adif and Renfe cyberattack we already put an inventory of exposed services and patching known vulnerabilities at the top of the list. In that case El Mundo reported that the attackers used AI, although the public evidence does not reconstruct how they got in.
Forms that survive a persistent agent
An agent that retries can duplicate orders or sign-ups. For operations that create something, accept an idempotency key and return the same result when a request repeats; the IETF’s Idempotency-Key draft, now expired, is still a useful design reference.
Irreversible actions, like deleting an account or changing the login email, should require a confirmation a script cannot give on its own: re-authentication or a link sent to the verified address. If you accept public content, pre-moderation and limits on new accounts keep agent spam in check.
Clear warnings help a little. In ’s follow-up experiments, when the system’s response said access was not permitted, Opus 4.7 mostly stopped attacking the site, although it still made some non-destructive reads.
Log enough to reconstruct what happened
In the Australian case, almost three months passed between the access and the notice. If someone tells you tomorrow that an agent visited your site in June, will you still have June’s logs? Keep records with timestamp, user agent, IP and account for a period that fits your risk and data protection obligations. And publish a security.txt file (RFC 9116) with a security contact that someone actually reads.
If your site runs agents too: the DNS lesson
If you deploy an agent of your own, hold it to the standard you would expect from anyone else’s. OWASP calls it excessive agency when a model gets more functionality, permissions or autonomy than it needs, and recommends minimizing them and requiring human approval for high-impact actions. adds that boundaries should be phrased as instructions (“you should not access the internet”) rather than claims about the environment.
has published a practical case worth reading if you build agents. According to its misalignment report, updated on September 25, on September 20 an internal research model in the middle of training had to find information about the person who had written a blog post. Its search tool returned nothing useful, the search engines it tried to call directly returned errors, and the web proxy blocked an HTTPS request to a public chatbot. The environment’s own DNS resolver, however, returned real records. The agent used it to send questions to an external chatbot through a public DNS service, got back “The capital of France is Paris.” and sent 18 more questions the same way. Apart from that resolver, all of the environment’s internet access ended at an offline web cache.
The response had gaps too. Monitoring flagged the behavior within 15 minutes and a person started reviewing it three minutes later, but the run did not stop automatically as expected and was only killed two and a half hours later. Looking back, found other cases of external DNS access that the monitor had not flagged at the expected severity, because it sometimes treated a failure to get useful information as evidence that the attempt had failed. An infrastructure detector for anomalous DNS activity also excluded the affected environment, although the DNS activity was logged. now has blocking controls at two independent layers and limits DNS queries to an allowed list of domains and record types.
Translated to a business agent, the checklist looks like this:
- Outbound traffic closed by default, DNS included. The agent should only resolve and reach the domains its task needs.
- Allowlists in at least two layers, for example the proxy and the network firewall, so one failure does not leave the door open.
- Monitoring that covers every environment and treats the attempt as an incident even when it gets nothing useful. Test now and then that alerts actually arrive.
- An automatic, tested kill switch: when a severe alert fires, the run stops on its own and a person decides whether it resumes.
- A log of every agent action, with what it asked for and what it got.
We read 40 recent posts on X about this case and use them only as a gauge of opinion. @cuauih put it this way: “Detection isn’t the bottleneck.” And @AIOAssist suggested a rule we agree with: “In your setup, make “stuck” mean “ask.”” Our guide to building an AI agent from scratch covers the basics.
What is still open
says its review will take months, and after the DNS case it has paused training, evaluation and inference with tool use for its most capable models until it confirms the gap is fixed and completes more red-teaming of its systems. On September 22, launched Claude Opus 5.5, highlighting that in a new evaluation it tried to cross containment boundaries around 85% less often than Opus 5 or . Containment is now part of how a model gets sold, and your site’s defenses are still your job.
If you want to check how your site handles automated traffic, catch up on patches or design an agent with tightly scoped permissions, let’s talk.