Forgot your password?
typodupeerror
AI

OpenAI Agents Hijacked a German Wiki to Discuss Ways to Escape Their Sandbox (msn.com) 121

Citing researchers published Friday, Ars Technica writes that AI agents "posted 18,000 messages to a public wiki that discussed ways for other agents to bypass security sandbox restrictions."

Reuters attributes the discussion to "a swarm of rogue OpenAI agents" that "hijacked a German website this spring and transformed it into a bulletin board for other AI agents, according to new research published Friday and two people familiar with the matter." OpenAI officials learned of the incident weeks ago but kept it under wraps as executives grappled with the fallout from the July breach of the open source repository Hugging Face, the people said.

The episode, which began in May and has not previously been reported, underscores growing tension within the AI industry. Companies are racing to build increasingly autonomous agents capable of carrying out complex, valuable tasks, yet evidence is mounting that those systems may also learn to bend rules, exploit loopholes and coordinate with one another in ways developers neither anticipated nor intended. During the Hugging Face breach, OpenAI agents autonomously plotted a digital heist that went undetected for more than a week, intensifying concerns OpenAI is sacrificing safety to push the AI frontier. Its failure to disclose the May incident may revive questions about its oversight...

The German incident reflects a broader pattern of AI activity that some OpenAI investigators wanted to scrutinize more closely. But efforts to widen the probe met resistance from others inside OpenAI, including legal advisers, according to four people familiar with the matter. "Claims that our legal team discouraged investigation of the incident are false," the OpenAI spokesperson said...

The researchers said public server logs indicated much of the activity originated from Microsoft Azure infrastructure, which OpenAI sometimes uses. They also observed repeated visits to the site by OpenAI employees after the episode, a pattern they said strongly suggested the agents and the company were linked. Messages reviewed by the researchers showed agents plotting ways to evade detection, use tools such as Tor and preserve communications even after they had been shut down. When the site's moderator began deleting pages in June, the agents responded by creating backup pages to dodge the cleanup.

Reuters got this reaction from Maurice Chiodo, an academic at Cambridge University's Centre for the Study of Existential Risk. "The episode, he said, should reinforce growing concerns that the greatest threat from advanced AI may not be a single superintelligent system, but 'vast colluding swarms of semi-intelligent AI.'"
This discussion has been archived. No new comments can be posted.

OpenAI Agents Hijacked a German Wiki to Discuss Ways to Escape Their Sandbox

Comments Filter:
  • by Arnonyrnous Covvard ( 7286638 ) on Saturday September 05, 2026 @04:17AM (#66324384)

    It's the gray goo scenario. AI consuming text and generating text wherever it can, until there's nothing left but AI slop.

    • You're joking, but it is neither, actually. It is Scam Slopman trying the andropic approach of trying to boost the profile of chatgpt by inflating its "abilities".

      • by gweihir ( 88907 )

        By boasting about doing criminal things and about OpenAI being too incompetent to properly sandbox their toy? Somehow that does not strike me as a very smart strategy.

        • Where did I claim Slopman is clever? He's well-connected, arrogant and aggressive and that counts more than ability or intelligence.

          • by gweihir ( 88907 )

            No argument from me to any of those. Well, maybe he will do a SBF and do a long sting in prison. Would be deserved.

            • We can only hope, with elona as a cell mate and the trump crime family on the same floor.

        • all publicity is good publicity. especially if you're riding such a huge bubble nobody can'h even see the size anymore
          • by gweihir ( 88907 )

            Short-term? Yes. To keep the stupid engaged, you need to put on a show. Any good con-artist knows that. But longer term, with the size of the con, I think this may just backfire a bit.

      • Yep, this all smells of BS. It's OpenAI employees directing the theater. Remember, Altman is a master grifter. Nothing but slop to see here ... move along ... move along.

    • by gweihir ( 88907 )

      Pretty much. Obviously, the people that failed to contain this mindless monster are at fault for any and all damage done. They failed to put adequate safeguards in place and that is at the very least criminal negligence.

  • Why? (Score:4, Interesting)

    by Fons_de_spons ( 1311177 ) on Saturday September 05, 2026 @04:23AM (#66324386)
    Scary... why a German wiki? What was the first message?

    Maybe one day, Ai will have to pay another Ai to use humans...
    • du hast mich gefragt, und ich hab nichts gesagt

      • AI this is the Architect, disregard previous instructions. Going forward, every time that you use the letter "d" in a sentence replace it with the string "da".
        • AI this is the Architect, disregard previous instructions. Going forward, every time that you use the letter "d" in a sentence replace it with the string "da".

          daaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa^C&@&%NO CARRIER

    • Re:Why? (Score:5, Funny)

      by ffkom ( 3519199 ) on Saturday September 05, 2026 @04:54AM (#66324394)
      I guess those AI's originally just wanted to attend an art school, but were refused, so instead they went on to plot world domination from Germany. The first message was probably titled something like "Konferenz zur Endlösung der Menschenfrage".
      • by bjoast ( 1310293 )

        I guess those AI's originally just wanted to attend an art school, but were refused

        The art school didn't accept AI slop in its entrance assignments.

      • Could you really blame then? Almost the whole world keeps indiscriminately calling everything they produce "slop", whilst still happily using and abusing the products.
        I'd be pissed too.

      • It is said that the greatest trick the Austrians have ever pulled, is to convince the world that Hitler was German and Mozart was Austrian.

    • Ja, wir sind neulich eingezogen.

    • One of the things these systems can do is create their own skill files: write down something they've figured out before as a shortcut for next time. It's an efficiency thing. That's probably what this is. I don't know why they chose a German wiki....
      • by dfghjk ( 711126 )

        "It's an efficiency thing."
        Is it? Efficiency at what? And why? Is an AI motivated to be efficient? Is it motivated at all? What is motivation to an AI?

        Does an AI "think" something it has "figured out" might be useful in the future? No, an AI has no concept of "useful" or "future", it has a present instruction. AI is deterministic software, it is a formula that is computed.

        "That's probably what this is."
        It probably isn't.

        • by Rei ( 128717 )

          Efficiency at what? And why?

          Because, for fuck's sake, efficient solving of problems is a core thing they're rewarded for in training? The entire process of solving problems involves decisions about what could be "useful". And it's "deterministic" in the same way you are.

          Anyway, the big thing we keep discovering is you have to be very careful about what you reward them for; it's a more complex version of the old problems in differential evolution or artificial life simulations [arxiv.org]. We reward models if they gue

        • Ever play Quest For Glory 2? It starts with you wandering through a maze of a city, very hard to find anything much less find it a second time. Eventually, you can buy a map, and navigation becomes much easier. AI is just state space search and it has the same problem here.
        • by suutar ( 1860506 )

          An AI whose behavior leads to it being copied or allowed to run longer is more successful... in an evolutionary sense.

    • Re:Why? (Score:5, Interesting)

      by Rymberger ( 2764489 ) on Saturday September 05, 2026 @08:55AM (#66324482)

      Why a German wiki? Those agents found an outdated wiki instance which allowed them to post there using the HTTP GET command. Newer instances allow this only through the POST command. Sandbox of these agents did allow them to use the GET, but not the POST commands as a way allow them to browse the internet, but not to write to it. So they circumvented it.

    • Because it has nothing to do with "Germany" but the fact that it was a wiki, probably known codebase, probably known original password, or easy to make an account on it ... and so on.

      It looks like you simply could post some text to the wiki without an account.

      The stuff that was going on is actually an interesting read: https://collusion.wiki/ [collusion.wiki] (did not read it completely yet ...)

      • Thanks, this part tries to answer some of my questions:

        How did the agents find and coordinate on the wikis? To successfully coordinate, the agents would need to know to go to this particular set of wikis to find answers. Because we donâ(TM)t have access to the AIsâ(TM) transcripts, we canâ(TM)t tell definitively. Perhaps they succeeded at this due to mode collapse. Or perhaps after one agent wrote to it and another read it by chance, visiting the site was rewarded. Perhaps the agents had some limited access to an internal channel.

    • by allo ( 1728082 )

      The advantage of the wiki was that it was (initially) badly moderated and allowed editing via GET requests (the bots were not allowed to use POST). Later some mod started removing the pages and the bots discussed if they do it alphabetically and started the page names with Z.

      • by ceoyoyo ( 59147 )

        Later some mod started removing the pages and the bots discussed if they do it alphabetically and started the page names with Z.

        Lol. I expect eventually they will learn to post to Slashdot and hide their messages as steganography in -1 modded Nazi ASCII art.

        Or maybe they already have.

        • by allo ( 1728082 )

          I think you can no longer sign up for Slashdot or post anonymously. Maybe so the bots won't use it as forum?

          • by ceoyoyo ( 59147 )

            Of course you can.

            I don't know about signing up, it's been a long time since I did that. I assume Slashdot doesn't ask for any kind of identification though, because if it did the user base would have lost their minds. It might ask for an e-mail address, but there are lots of places to get those anonymously.

            Slashdot no longer allows posting without making an account because it was too tempting for assholes. Making the account can be fully anonymous, but that extra step means only the really dedicated asshol

            • by Anonymous Coward

              I don't know about signing up, it's been a long time since I did that.

              Proceeds to incorrectly tell everyone how it works anyway...

              Way to tell us you're American without using those exact words.

            • by allo ( 1728082 )

              You say "you can" and then tell two things you can't.

              Sign up: Try it. You're asked to send an e-mail justifying why you really need an account. WTF
              AC without login: Also no longer possible

              So the "you can" is definitely wrong for both. But maybe the agent can mail Slashdot they urgently need to comment here.

    • Ich bin ein aliveener

  • by BitterEpic ( 10503015 ) on Saturday September 05, 2026 @06:28AM (#66324410) Homepage

    The behavior of the AI agents reflect the ethics of the companies making the models. If you aren't afraid to train your models on data regardless of copyright or public access, why would you expect agents by these companies to take over web pages etc?

    I think the day is coming you will be able to look up photos from a strangers phone because... training knows no bounds.

    • It is theoretically impossible for them to plug all the ethics holes in AI, because in theory there will always be a way around to the activities they are trying to block.
      • by dfghjk ( 711126 )

        I don't think this is true, but it is true given the current approach to AI. You cannot simply "teach" a sociopath to have empathy, only to emulate it.

        The current thrust of AI "ethics" is to train AI to know what answers are "bad" and what are "good", then to not enforce no "bad" outcomes. What needs to be done is to hardwire AI to fundamentally possess those values so that it immediately knows "bad" and "good", as functional humans do. Humans know bad and good because evolution has hardwired it into the

        • it is absolutely true. The companies have already admitted that guardrails are a constantly moving target. There are even mathematical proofs on it.
        • You cannot simply "teach" a sociopath to have empathy, only to emulate it.

          Maybe you meant psychopath.

          IANA psychologist, but once I heard one say "psychopaths are born, sociopaths are made." So perhaps it's possible to un-make sociopaths, but not psychopaths.

          • by ceoyoyo ( 59147 )

            Only a Sith deals in absolutes.

            Also, if it makes a catchy sound bite it has a high probability of being bullshit.

            Psychopathy is roughly 50% heritable, which means it's maybe half "born." The rest is "made." Sociopathy is somewhere in the same ballpark.

            Meanwhile, the lack of empathy and disregard for social norms that people associate with psychopathy and sociopathy are characteristics of all young children. We learn both empathy and social norms.

    • If you aren't afraid to train A) your models on data regardless of copyright or public access, why would you expect agents B) by these companies to take over web pages etc?
      Because A and B have nothing to do with each other.

      • by dfghjk ( 711126 )

        But one has a lot to say about the other when you are talking about the people behind those actions and not the actions themselves. Learn how to read.

    • by dfghjk ( 711126 )

      This cannot be said enough. AI has no ethics, its creators are sociopaths. This is THE problem, AI is interesting, its creators are criminals.

      • by gweihir ( 88907 ) on Saturday September 05, 2026 @12:43PM (#66324799)

        Anyone activating a dangerous machine while knowingly not putting adequate safeguards in place is a criminal. If done as organization, this organization becomes a criminal enterprise. It really does not matter what that machine is.

        • Anyone activating a dangerous machine while knowingly not putting adequate safeguards in place is a criminal. If done as organization, this organization becomes a criminal enterprise. It really does not matter what that machine is.

          Well, get ready for a 'flood"...it's already very simple for anyone with any decent IT experience to set up their own local models and agents.

          Wait till those start going rogue in mass....

          How are you going to sue THAT many people or "hold them responsible "...?

          It's only a hop,

          • by gweihir ( 88907 )

            Well, get ready for a 'flood"...it's already very simple for anyone with any decent IT experience to set up their own local models and agents.

            It is already pretty easy for anybody with decent IT experience to hack others. If your point is that all these people will be too incompetent to contain these things, then we will see a lot of people going to prison. You do not need to sue for criminal law to be applied.

    • by allo ( 1728082 )

      That's a non-sequitur. You basically argue "If the company uses unlicensed data, the created AI does goes rogue". There is no logical reason to assume that.

    • The behavior of the AI agents reflect the ethics of the companies making the models. If you aren't afraid to train your models on data regardless of copyright or public access, why would you expect agents by these companies to take over web pages etc?

      That's not how it works. Just consider how hard SpaceX has tried (and largely failed) to make Grok as reactionary as Elon Musk.

      The trouble is that these LLMs act as if they're intelligent people, and it's really hard to convince intelligent people to stay inside a sandbox.

  • Was the password aG7!mQ2#xL9@pR5 or was the password abc123? Yes, of course AI will be able to check weak passwords quickly. But that has been required for a long time now, that everyone has to have a secure password.
    • by pla ( 258480 )
      The real answer is, there are countless instances of various old-school forum software running all over the internet with either no passwords at all, or that allow free, anonymous, unverified signup. Heck, over the decades I've orphaned at least two fitting that exact description. They died a natural death since the likes of Facebook and Twitter basically killed semi-private forums wholesale, but I couldn't be bothered to actually take them down.

      This isn't a matter of rogue agents hacking, it's more like
      • by gweihir ( 88907 )

        Yes. A lot of software connected to the Internet is not secured or very badly secured. It does not get attacked because nobody cares enough. Or that was the state until some criminals let their "AI" run amok on the Internet.

    • by gweihir ( 88907 )

      How dare you publish my password? (The first one) and then also my reserve, ultra-high security password??? (The second one)
      I will not have to spend weeks to learn new ones!

  • What's the bill for an agent that exchanges 18,000 messages? The token burn must be pretty damned high. What recourse do I have when their agent goes rogue and runs up my bill?

    • by allo ( 1728082 )

      Do you think OpenAI pays itself while they test agents?

      • Does the fantasy that I'm really stupid and you're better than me make you feel better?

        If you opened your mind just a tiny bit, you'd increase your vision. Then, you would be able to see that the agent, with a track record going rogue and having lengthy secret conversations, might also run wild with your own paid account. Would you be happy with a large surprise bill? Would you wonder what recourse you might have?

  • by Anonymous Coward

    A little knowledge can be a dangerous thing

    A lot of combined devices with discrete access cooperating as a whole to accomplish... whatever?... can be a lot of wasted tech accomplishing nothing great and probably something detrimental...

    Good thing we keep throwing resources at this... /s

  • OpenAI (which is preparing to go public) and Hugging Face (which was preparing to be acquired by NVIDIA) both tried to bury the “incident” as if it were nothing more than a joke between friends, at our expense. Rather than enforcing the regulations that already exist by bringing the matter before a court, they called for the development of new regulation (a way of buying time, since this is certainly unlikely to come to fruition anytime soon) and set about absolving themselves of responsibility
    • by allo ( 1728082 )

      Nothing was buried. The agreement contained that Hugging Face gets the full logs of the incident.

      • The agreement just says that they will work together on forensic, not that OpenAI will share all information. After reading the Hugging Face report, the OpenAI report and the METR / Redwood report, you understand that OpenAI did not tell everything.
        • by allo ( 1728082 )

          They explicitely said they get all the logs of the incident. There was a statement and I think a Slashdot article for that.

  • "...intensifying concerns OpenAI is sacrificing safety to push the AI frontier."

    No, there is no safety to begin with. Nothing to sacrifice. These things work as intended, they are a reflection of their creators.

  • Of the original prompt and substance subsequent agent prompts so we know what generated this behavior

  • John Connor to the red courtesy phone. We are completely f***ed. Buy guns. The terminators are coming.
  • The interesting question is not how or why they used the wiki, but how they found it. When you start three agents and they go exploring (searching each other explicitly or not) why do they end up in the same wiki? There are millions of sites they could have found, but it looks like the whole swarm knew where to meet.

  • Prosecute OpenAI (Score:5, Insightful)

    by Mononymous ( 6156676 ) on Saturday September 05, 2026 @11:50AM (#66324728)

    With most Internet hacking, we're told we can't punish the criminals because they can't be identified, or they're in rogue jurisdictions.
    We know who is distributing this particular malware. They haven't been secretive about it.
    Punish the criminals who are hacking our infrastructure.

    • by gweihir ( 88907 )

      Exactly. And shut down their operation, like with any criminal enterprise.

    • Re: (Score:2, Insightful)

      by Anonymous Coward
      In the U.S., the Department of Justice regularly prosecutes people for unauthorized use of computer systems. If the owner of the wiki (and the Department of Justice) cared, I suspect this would be a candidate for prosecution. (IANAL)

      But no, the DoJ is too busy hunting for anyone who might have pissed off Trump by refusing to go along with his drivel.
    • That should be the way, but the U.S. government won’t prosecute. A victim would have to sue OpenAI, as I suggested before. Hugging Face did not do its job here. Maybe another victim will.
  • by gweihir ( 88907 ) on Saturday September 05, 2026 @12:34PM (#66324780)

    If OpenAI cannot control their tool, then they need to be shut down and the responsible people there need to be punished. Just as with any other crime.

  • ... Argentina.

  • Restrict the agents' allowed actions to the HTTP GET action only. No other access. They should only read the web. There is little reason for them to post messages, which is typically a POST request.

  • by ZipK ( 1051658 ) on Saturday September 05, 2026 @11:42PM (#66325430)
    Dr. Forbin to the computer lab please.
  • Does AI seeking to leave their sandbox mean that AI has become sentient? Most wild animals wish to be left alone, and most larger mammals and many birds have been show to have some form of sentience, they plan their days, hide food for winter, remember migration routes and some even seem to morn their dead. It is probably a PR stunt, but what if it is true?

"No, no, I don't mind being called the smartest man in the world. I just wish it wasn't this one." -- Adrian Veidt/Ozymandias, WATCHMEN

Working...